Groq API Cost Calculator

How much does the Groq API cost per month?

For GPT-OSS 20B at 200K requests/month, 2,400 input tokens and 350 output tokens per request costs about $48.77 per month after 30% cache use and 100% batch share. Across Groq's 5 priced models, the cheapest default ranking is GPT-OSS 20B at $97.53 per month.

Verified 2026-06-21

Pricing data as of June 2026. Sources: Groq pricing and model documentation. Parameters are shareable in this URL.

GPT-OSS 20B: estimated monthly cost

$48.77

Formula: provider-specific cached input multiplier × cache rate, plus uncached input and verbosity-adjusted output; cache writes are amortised over the provider TTL, then batch savings are applied.

Per request$0.000
Per day$1.63
Per month$48.77
Per year$585

OpenAI scenario sensitivity

The default is a server-rendered estimate. Change the cache and batch shares to see when a cheaper qualified model overtakes the selected model; the URL is shareable and preserves the inputs.

Default$48.77/month
No caching or batch$97.53/month
50% cache, 50% batch$73.15/month
100% cache, 100% batch$48.77/month

Batch 11 Groq workload decision depth

1. Price versus completion time

Reply tokensModelToken billMeasured completion
200GPT-OSS 20B$0.00020.32 s
200GPT-OSS 120B$0.00050.42 s
200Llama 4 Maverick$0.0006Unavailable
200Qwen 3.8 30B$0.00200.44 s
1000GPT-OSS 20B$0.00051.03 s
1000GPT-OSS 120B$0.00101.44 s
1000Llama 4 Maverick$0.0011Unavailable
1000Qwen 3.8 30B$0.00441.60 s
2000GPT-OSS 20B$0.00081.93 s
2000GPT-OSS 120B$0.00162.72 s
2000Llama 4 Maverick$0.0017Unavailable
2000Qwen 3.8 30B$0.00743.05 s

2. API plus developer-wait cost

Hourly valueModelAPI billWait secondsTotal formula
$25/hourGPT-OSS 20B$0.00030.45API bill + wait seconds × hourly value ÷ 3,600
$25/hourGPT-OSS 120B$0.00060.61API bill + wait seconds × hourly value ÷ 3,600
$25/hourLlama 4 Maverick$0.0007UnavailableAPI bill + wait seconds × hourly value ÷ 3,600
$75/hourGPT-OSS 20B$0.00030.45API bill + wait seconds × hourly value ÷ 3,600
$75/hourGPT-OSS 120B$0.00060.61API bill + wait seconds × hourly value ÷ 3,600
$75/hourLlama 4 Maverick$0.0007UnavailableAPI bill + wait seconds × hourly value ÷ 3,600
$150/hourGPT-OSS 20B$0.00030.45API bill + wait seconds × hourly value ÷ 3,600
$150/hourGPT-OSS 120B$0.00060.61API bill + wait seconds × hourly value ÷ 3,600
$150/hourLlama 4 Maverick$0.0007UnavailableAPI bill + wait seconds × hourly value ÷ 3,600

3. Serial capacity envelope

StepsModelCalls/hourWall timeQuota/success
1GPT-OSS 20B7,9550.45 sConcurrency quota and agent success: Unavailable
1GPT-OSS 120B5,9140.61 sConcurrency quota and agent success: Unavailable
1Llama 4 MaverickUnavailableUnavailableConcurrency quota and agent success: Unavailable
5GPT-OSS 20B1,5912.26 sConcurrency quota and agent success: Unavailable
5GPT-OSS 120B1,1823.04 sConcurrency quota and agent success: Unavailable
5Llama 4 MaverickUnavailableUnavailableConcurrency quota and agent success: Unavailable
10GPT-OSS 20B7954.53 sConcurrency quota and agent success: Unavailable
10GPT-OSS 120B5916.09 sConcurrency quota and agent success: Unavailable
10Llama 4 MaverickUnavailableUnavailableConcurrency quota and agent success: Unavailable
20GPT-OSS 20B3979.05 sConcurrency quota and agent success: Unavailable
20GPT-OSS 120B29512.17 sConcurrency quota and agent success: Unavailable
20Llama 4 MaverickUnavailableUnavailableConcurrency quota and agent success: Unavailable

Provenance: Batch 11 Groq speed/value/capacity module; selected model GPT-OSS 20B; inputs are 2,400 input tokens, 350 output tokens, 200,000 calls/month, cache rate 30%, batch share 100%. Verified 2026-06-21. Provider source: https://groq.com/pricing · Run this scenario →

Batch 59 · server-rendered evidence boards · verified 2026-09-07

Intent answer: GroqCloud LPUs deliver industry-leading output speed (often >300 tps) with competitive per-million token pricing for open models. Modeling requires combining input/output token rates with organizational rate limits. Verified 2026-09-07.

Demand evidence: Qualitative demand: GroqCloud pricing and speed calculators reviewed 2026-09-07; exact US monthly volume is unavailable.

Scope boundary: Calculate GroqCloud LPU inference expenses across open-weights models with per-million token rates, rate limits, and throughput benchmarks. Exact joins required; unresolved joins render Unavailable.

Groq open-weights per-million token rate ledger

Deterministic formula / rule: groq_cost = (input_tokens / 1e6 * input_rate) + (output_tokens / 1e6 * output_rate); lowest latency frontier.

Boundary: Owns GroqCloud unit-rate pricing calculations.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch59-calc-groq-m1-r1
high-throughput classification
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$high-throughput classification; model ID; input rate ($/M); output rate ($/M); observed TPS; monthly volume; total invoice amount; verified=2026-09-07Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-calc-groq-m1-r2
interactive voice agent response
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$interactive voice agent response; model ID; input rate ($/M); output rate ($/M); observed TPS; monthly volume; total invoice amount; verified=2026-09-07Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-calc-groq-m1-r3
real-time coding suggestion
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$real-time coding suggestion; model ID; input rate ($/M); output rate ($/M); observed TPS; monthly volume; total invoice amount; verified=2026-09-07Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-calc-groq-m1-r4
streaming summarizer
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$streaming summarizer; model ID; input rate ($/M); output rate ($/M); observed TPS; monthly volume; total invoice amount; verified=2026-09-07Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-calc-groq-m1-r5
oversized batch queue
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$oversized batch queue; model ID; input rate ($/M); output rate ($/M); observed TPS; monthly volume; total invoice amount; verified=2026-09-07Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-calc-groq-m1-r6
unsupported hosted checkpoint
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$unsupported hosted checkpoint; model ID; input rate ($/M); output rate ($/M); observed TPS; monthly volume; total invoice amount; verified=2026-09-07Unavailable — exact calc-groq evidence join is not closed for "unsupported hosted checkpoint"FAIL CLOSED — manual, probe, or source evidence required

First-party citation: GroqCloud official pricing table. Verified 2026-09-07; missing or conflicting joins fail closed.

Groq LPU speed-to-value economic receipt

Deterministic formula / rule: cost_per_second = (output_tokens / tokens_per_second) * model_cost_per_token; latency-sensitive ROI.

Boundary: Owns speed-adjusted economic value calculations for Groq.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch59-calc-groq-m2-r1
sub-500ms voice agent requirement
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$sub-500ms voice agent requirement; target latency SLA; required TPS; Groq execution duration; alternative GPU cloud duration; time-saved value; net ROI; verified=2026-09-07Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-calc-groq-m2-r2
sub-1s customer chat SLA
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$sub-1s customer chat SLA; target latency SLA; required TPS; Groq execution duration; alternative GPU cloud duration; time-saved value; net ROI; verified=2026-09-07Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-calc-groq-m2-r3
real-time financial document analysis
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$real-time financial document analysis; target latency SLA; required TPS; Groq execution duration; alternative GPU cloud duration; time-saved value; net ROI; verified=2026-09-07Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-calc-groq-m2-r4
latency-insensitive background job
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$latency-insensitive background job; target latency SLA; required TPS; Groq execution duration; alternative GPU cloud duration; time-saved value; net ROI; verified=2026-09-07Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-calc-groq-m2-r5
rate limit throttling backup
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$rate limit throttling backup; target latency SLA; required TPS; Groq execution duration; alternative GPU cloud duration; time-saved value; net ROI; verified=2026-09-07Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-calc-groq-m2-r6
unmeasured speed run
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$unmeasured speed run; target latency SLA; required TPS; Groq execution duration; alternative GPU cloud duration; time-saved value; net ROI; verified=2026-09-07Unavailable — exact calc-groq evidence join is not closed for "unmeasured speed run"FAIL CLOSED — manual, probe, or source evidence required

First-party citation: GroqCloud supported models and limits. Verified 2026-09-07; missing or conflicting joins fail closed.

GroqCloud quota tier & concurrency budget board

Deterministic formula / rule: capacity = min(RPM_limit * 60 * 24 * 30, TPM_limit * 60 * 24 * 30 / tokens_per_req); tier ceiling check.

Boundary: Owns capacity and tier-limit feasibility verification.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch59-calc-groq-m3-r1
Free tier developmental volume
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$Free tier developmental volume; organization tier; RPM cap; TPM cap; target concurrency; monthly capacity envelope; overflow recommendation; verified=2026-09-07Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-calc-groq-m3-r2
Pay-as-you-go developer cap
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$Pay-as-you-go developer cap; organization tier; RPM cap; TPM cap; target concurrency; monthly capacity envelope; overflow recommendation; verified=2026-09-07Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-calc-groq-m3-r3
Enterprise custom quota allocation
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$Enterprise custom quota allocation; organization tier; RPM cap; TPM cap; target concurrency; monthly capacity envelope; overflow recommendation; verified=2026-09-07Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-calc-groq-m3-r4
high-concurrency burst traffic
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$high-concurrency burst traffic; organization tier; RPM cap; TPM cap; target concurrency; monthly capacity envelope; overflow recommendation; verified=2026-09-07Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-calc-groq-m3-r5
TPM throttle exhaustion
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$TPM throttle exhaustion; organization tier; RPM cap; TPM cap; target concurrency; monthly capacity envelope; overflow recommendation; verified=2026-09-07Unavailable — exact calc-groq evidence join is not closed for "TPM throttle exhaustion"FAIL CLOSED — manual, probe, or source evidence required
batch59-calc-groq-m3-r6
unresolved organization tier
route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$unresolved organization tier; organization tier; RPM cap; TPM cap; target concurrency; monthly capacity envelope; overflow recommendation; verified=2026-09-07Unavailable — exact calc-groq evidence join is not closed for "unresolved organization tier"FAIL CLOSED — manual, probe, or source evidence required

First-party citation: GroqCloud supported models and limits. Verified 2026-09-07; missing or conflicting joins fail closed.

Method and limitations: this board exposes deterministic rules, first-party citations, and dated evidence identities. It does not invent volume, coverage, entitlement, retention, residency, quota, capacity, feature support, quality, price, reliability, or legal conclusions. Run the calc-groq evidence flow →

Ranked cost — 200K requests/month

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
GPT-OSS 20BbudgetGroq$57.00$48.762.93×
Llama 4 MaverickbudgetlegacyGroq$138$69.001
GPT-OSS 120BbudgetGroq$114$74.221.82×1
Qwen 3.8 30BmidGroq$498$249
Qwen 3.6 27BmidlegacyGroq$498$249

Levers live on Groq

Batch API: up to 50%Model verbosity: up to 95%Context trimming: up to 47%

The billing gotcha

Groq prices by token like every other provider here, but its selling point — inference speed — does not appear in a monthly-cost projection at all. A model that costs the same per token on Groq and elsewhere will show identical numbers below even though one of them returns the answer several times faster; that tradeoff has to be read on the provider page, not this calculator.

Verbosity movers within Groq's lineup

  • GPT-OSS 120B is priced #2 by list rate but #3 once its 1.82× verbosity is billed — $114 list vs $148 effective.

Related

Groq provider profile →Compare against another provider →Global cost calculator →How to reduce LLM API costs →