Groq On-Demand API
LLM 5.6/10Llama 3.1 8B $0.05/$0.08; GPT-OSS $0.075/$0.30; Llama 3.3 70B $0.59/$0.79 per 1M. LPU ultra-fast inference
Groq On-Demand API scores 5.6/10 for free value in LLM. The main constraint: Paid — Pay-as-you-go, from $0.05/M input. Llama 3.1 8B $0.05/$0.08; GPT-OSS $0.075/$0.30; Llama 3.3 70B $0.59/$0.79 per 1M. LPU ultra-fast inference.
Free value score
5.6 / 10Show the math
Key metrics
What you get
Paid — Pay-as-you-go, from $0.05/M input. Llama 3.1 8B $0.05/$0.08; GPT-OSS $0.075/$0.30; Llama 3.3 70B $0.59/$0.79 per 1M. LPU ultra-fast inference
13.3M input tokens per $1 on GPT-OSS; up to 20M input tokens/$1 on 8B — fastest cheap inference
Sources
-
"Input Token Price(Per Million Tokens). $0.075(13.3M / $1)* Output Token Price(Per Million Tokens). $0.30(3.33M / $1)*"
https://groq.com/pricing -
"the flagship Llama 3.3 70B Versatile is $0.59 input / ... per 1M"
https://www.aipricing.guru/groq-pricing/ -
"GPT-OSS 120B is present in the refreshed BenchLM Artificial Analysis table with score 23.8."
https://benchlm.ai/benchmarks/artificialAnalysis