Groq API Cost Calculator
How much does the Groq API cost per month?
For GPT-OSS 20B at 200K requests/month, 2,400 input tokens and 350 output tokens per request costs about $48.77 per month after 30% cache use and 100% batch share. Across Groq's 5 priced models, the cheapest default ranking is GPT-OSS 20B at $97.53 per month.
Pricing data as of June 2026. Sources: Groq pricing and model documentation. Parameters are shareable in this URL.
GPT-OSS 20B: estimated monthly cost
Formula: provider-specific cached input multiplier × cache rate, plus uncached input and verbosity-adjusted output; cache writes are amortised over the provider TTL, then batch savings are applied.
| Per request | $0.000 |
|---|---|
| Per day | $1.63 |
| Per month | $48.77 |
| Per year | $585 |
OpenAI scenario sensitivity
The default is a server-rendered estimate. Change the cache and batch shares to see when a cheaper qualified model overtakes the selected model; the URL is shareable and preserves the inputs.
| Default | $48.77/month |
|---|---|
| No caching or batch | $97.53/month |
| 50% cache, 50% batch | $73.15/month |
| 100% cache, 100% batch | $48.77/month |
Batch 11 Groq workload decision depth
1. Price versus completion time
| Reply tokens | Model | Token bill | Measured completion |
|---|---|---|---|
| 200 | GPT-OSS 20B | $0.0002 | 0.32 s |
| 200 | GPT-OSS 120B | $0.0005 | 0.42 s |
| 200 | Llama 4 Maverick | $0.0006 | Unavailable |
| 200 | Qwen 3.8 30B | $0.0020 | 0.44 s |
| 1000 | GPT-OSS 20B | $0.0005 | 1.03 s |
| 1000 | GPT-OSS 120B | $0.0010 | 1.44 s |
| 1000 | Llama 4 Maverick | $0.0011 | Unavailable |
| 1000 | Qwen 3.8 30B | $0.0044 | 1.60 s |
| 2000 | GPT-OSS 20B | $0.0008 | 1.93 s |
| 2000 | GPT-OSS 120B | $0.0016 | 2.72 s |
| 2000 | Llama 4 Maverick | $0.0017 | Unavailable |
| 2000 | Qwen 3.8 30B | $0.0074 | 3.05 s |
2. API plus developer-wait cost
| Hourly value | Model | API bill | Wait seconds | Total formula |
|---|---|---|---|---|
| $25/hour | GPT-OSS 20B | $0.0003 | 0.45 | API bill + wait seconds × hourly value ÷ 3,600 |
| $25/hour | GPT-OSS 120B | $0.0006 | 0.61 | API bill + wait seconds × hourly value ÷ 3,600 |
| $25/hour | Llama 4 Maverick | $0.0007 | Unavailable | API bill + wait seconds × hourly value ÷ 3,600 |
| $75/hour | GPT-OSS 20B | $0.0003 | 0.45 | API bill + wait seconds × hourly value ÷ 3,600 |
| $75/hour | GPT-OSS 120B | $0.0006 | 0.61 | API bill + wait seconds × hourly value ÷ 3,600 |
| $75/hour | Llama 4 Maverick | $0.0007 | Unavailable | API bill + wait seconds × hourly value ÷ 3,600 |
| $150/hour | GPT-OSS 20B | $0.0003 | 0.45 | API bill + wait seconds × hourly value ÷ 3,600 |
| $150/hour | GPT-OSS 120B | $0.0006 | 0.61 | API bill + wait seconds × hourly value ÷ 3,600 |
| $150/hour | Llama 4 Maverick | $0.0007 | Unavailable | API bill + wait seconds × hourly value ÷ 3,600 |
3. Serial capacity envelope
| Steps | Model | Calls/hour | Wall time | Quota/success |
|---|---|---|---|---|
| 1 | GPT-OSS 20B | 7,955 | 0.45 s | Concurrency quota and agent success: Unavailable |
| 1 | GPT-OSS 120B | 5,914 | 0.61 s | Concurrency quota and agent success: Unavailable |
| 1 | Llama 4 Maverick | Unavailable | Unavailable | Concurrency quota and agent success: Unavailable |
| 5 | GPT-OSS 20B | 1,591 | 2.26 s | Concurrency quota and agent success: Unavailable |
| 5 | GPT-OSS 120B | 1,182 | 3.04 s | Concurrency quota and agent success: Unavailable |
| 5 | Llama 4 Maverick | Unavailable | Unavailable | Concurrency quota and agent success: Unavailable |
| 10 | GPT-OSS 20B | 795 | 4.53 s | Concurrency quota and agent success: Unavailable |
| 10 | GPT-OSS 120B | 591 | 6.09 s | Concurrency quota and agent success: Unavailable |
| 10 | Llama 4 Maverick | Unavailable | Unavailable | Concurrency quota and agent success: Unavailable |
| 20 | GPT-OSS 20B | 397 | 9.05 s | Concurrency quota and agent success: Unavailable |
| 20 | GPT-OSS 120B | 295 | 12.17 s | Concurrency quota and agent success: Unavailable |
| 20 | Llama 4 Maverick | Unavailable | Unavailable | Concurrency quota and agent success: Unavailable |
Provenance: Batch 11 Groq speed/value/capacity module; selected model GPT-OSS 20B; inputs are 2,400 input tokens, 350 output tokens, 200,000 calls/month, cache rate 30%, batch share 100%. Verified 2026-06-21. Provider source: https://groq.com/pricing · Run this scenario →
Batch 59 · server-rendered evidence boards · verified 2026-09-07
Intent answer: GroqCloud LPUs deliver industry-leading output speed (often >300 tps) with competitive per-million token pricing for open models. Modeling requires combining input/output token rates with organizational rate limits. Verified 2026-09-07.
Demand evidence: Qualitative demand: GroqCloud pricing and speed calculators reviewed 2026-09-07; exact US monthly volume is unavailable.
Scope boundary: Calculate GroqCloud LPU inference expenses across open-weights models with per-million token rates, rate limits, and throughput benchmarks. Exact joins required; unresolved joins render Unavailable.
Groq open-weights per-million token rate ledger
Deterministic formula / rule: groq_cost = (input_tokens / 1e6 * input_rate) + (output_tokens / 1e6 * output_rate); lowest latency frontier.
Boundary: Owns GroqCloud unit-rate pricing calculations.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch59-calc-groq-m1-r1high-throughput classification | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$high-throughput classification; model ID; input rate ($/M); output rate ($/M); observed TPS; monthly volume; total invoice amount; verified=2026-09-07 | Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receipt | UNTESTED — module 1 rule is reproducible but no production observation is claimed |
batch59-calc-groq-m1-r2interactive voice agent response | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$interactive voice agent response; model ID; input rate ($/M); output rate ($/M); observed TPS; monthly volume; total invoice amount; verified=2026-09-07 | Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receipt | UNTESTED — module 1 rule is reproducible but no production observation is claimed |
batch59-calc-groq-m1-r3real-time coding suggestion | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$real-time coding suggestion; model ID; input rate ($/M); output rate ($/M); observed TPS; monthly volume; total invoice amount; verified=2026-09-07 | Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receipt | UNTESTED — module 1 rule is reproducible but no production observation is claimed |
batch59-calc-groq-m1-r4streaming summarizer | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$streaming summarizer; model ID; input rate ($/M); output rate ($/M); observed TPS; monthly volume; total invoice amount; verified=2026-09-07 | Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receipt | UNTESTED — module 1 rule is reproducible but no production observation is claimed |
batch59-calc-groq-m1-r5oversized batch queue | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$oversized batch queue; model ID; input rate ($/M); output rate ($/M); observed TPS; monthly volume; total invoice amount; verified=2026-09-07 | Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receipt | UNTESTED — module 1 rule is reproducible but no production observation is claimed |
batch59-calc-groq-m1-r6unsupported hosted checkpoint | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$unsupported hosted checkpoint; model ID; input rate ($/M); output rate ($/M); observed TPS; monthly volume; total invoice amount; verified=2026-09-07 | Unavailable — exact calc-groq evidence join is not closed for "unsupported hosted checkpoint" | FAIL CLOSED — manual, probe, or source evidence required |
First-party citation: GroqCloud official pricing table. Verified 2026-09-07; missing or conflicting joins fail closed.
Groq LPU speed-to-value economic receipt
Deterministic formula / rule: cost_per_second = (output_tokens / tokens_per_second) * model_cost_per_token; latency-sensitive ROI.
Boundary: Owns speed-adjusted economic value calculations for Groq.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch59-calc-groq-m2-r1sub-500ms voice agent requirement | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$sub-500ms voice agent requirement; target latency SLA; required TPS; Groq execution duration; alternative GPU cloud duration; time-saved value; net ROI; verified=2026-09-07 | Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch59-calc-groq-m2-r2sub-1s customer chat SLA | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$sub-1s customer chat SLA; target latency SLA; required TPS; Groq execution duration; alternative GPU cloud duration; time-saved value; net ROI; verified=2026-09-07 | Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch59-calc-groq-m2-r3real-time financial document analysis | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$real-time financial document analysis; target latency SLA; required TPS; Groq execution duration; alternative GPU cloud duration; time-saved value; net ROI; verified=2026-09-07 | Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch59-calc-groq-m2-r4latency-insensitive background job | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$latency-insensitive background job; target latency SLA; required TPS; Groq execution duration; alternative GPU cloud duration; time-saved value; net ROI; verified=2026-09-07 | Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch59-calc-groq-m2-r5rate limit throttling backup | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$rate limit throttling backup; target latency SLA; required TPS; Groq execution duration; alternative GPU cloud duration; time-saved value; net ROI; verified=2026-09-07 | Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch59-calc-groq-m2-r6unmeasured speed run | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$unmeasured speed run; target latency SLA; required TPS; Groq execution duration; alternative GPU cloud duration; time-saved value; net ROI; verified=2026-09-07 | Unavailable — exact calc-groq evidence join is not closed for "unmeasured speed run" | FAIL CLOSED — manual, probe, or source evidence required |
First-party citation: GroqCloud supported models and limits. Verified 2026-09-07; missing or conflicting joins fail closed.
GroqCloud quota tier & concurrency budget board
Deterministic formula / rule: capacity = min(RPM_limit * 60 * 24 * 30, TPM_limit * 60 * 24 * 30 / tokens_per_req); tier ceiling check.
Boundary: Owns capacity and tier-limit feasibility verification.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch59-calc-groq-m3-r1Free tier developmental volume | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$Free tier developmental volume; organization tier; RPM cap; TPM cap; target concurrency; monthly capacity envelope; overflow recommendation; verified=2026-09-07 | Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch59-calc-groq-m3-r2Pay-as-you-go developer cap | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$Pay-as-you-go developer cap; organization tier; RPM cap; TPM cap; target concurrency; monthly capacity envelope; overflow recommendation; verified=2026-09-07 | Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch59-calc-groq-m3-r3Enterprise custom quota allocation | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$Enterprise custom quota allocation; organization tier; RPM cap; TPM cap; target concurrency; monthly capacity envelope; overflow recommendation; verified=2026-09-07 | Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch59-calc-groq-m3-r4high-concurrency burst traffic | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$high-concurrency burst traffic; organization tier; RPM cap; TPM cap; target concurrency; monthly capacity envelope; overflow recommendation; verified=2026-09-07 | Unavailable — frozen calc-groq fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch59-calc-groq-m3-r5TPM throttle exhaustion | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$TPM throttle exhaustion; organization tier; RPM cap; TPM cap; target concurrency; monthly capacity envelope; overflow recommendation; verified=2026-09-07 | Unavailable — exact calc-groq evidence join is not closed for "TPM throttle exhaustion" | FAIL CLOSED — manual, probe, or source evidence required |
batch59-calc-groq-m3-r6unresolved organization tier | route=$/llm-cost-calculator/groq; owner=$calc-groq; scenario=$unresolved organization tier; organization tier; RPM cap; TPM cap; target concurrency; monthly capacity envelope; overflow recommendation; verified=2026-09-07 | Unavailable — exact calc-groq evidence join is not closed for "unresolved organization tier" | FAIL CLOSED — manual, probe, or source evidence required |
First-party citation: GroqCloud supported models and limits. Verified 2026-09-07; missing or conflicting joins fail closed.
Method and limitations: this board exposes deterministic rules, first-party citations, and dated evidence identities. It does not invent volume, coverage, entitlement, retention, residency, quota, capacity, feature support, quality, price, reliability, or legal conclusions. Run the calc-groq evidence flow →
Ranked cost — 200K requests/month
| Model | Provider | List monthly | Effective monthly | Verbosity | Rank Δ |
|---|---|---|---|---|---|
| GPT-OSS 20Bbudget | Groq | $57.00 | $48.76 | 2.93× | — |
| Llama 4 Maverickbudgetlegacy | Groq | $138 | $69.00 | — | ▲1 |
| GPT-OSS 120Bbudget | Groq | $114 | $74.22 | 1.82× | ▼1 |
| Qwen 3.8 30Bmid | Groq | $498 | $249 | — | — |
| Qwen 3.6 27Bmidlegacy | Groq | $498 | $249 | — | — |
Levers live on Groq
The billing gotcha
Groq prices by token like every other provider here, but its selling point — inference speed — does not appear in a monthly-cost projection at all. A model that costs the same per token on Groq and elsewhere will show identical numbers below even though one of them returns the answer several times faster; that tradeoff has to be read on the provider page, not this calculator.
Verbosity movers within Groq's lineup
- GPT-OSS 120B is priced #2 by list rate but #3 once its 1.82× verbosity is billed — $114 list vs $148 effective.
