GPT-OSS 20B on Groq API Pricing: Blazing Speed at Budget Rates
Explore GPT-OSS 20B on Groq API pricing ($0.07/M input, $0.30/M output), ultra-high token generation speed on Groq LPUs, and budget cost efficiency.
Full specs, context window and API limits →How much does GPT-OSS 20B cost per million tokens?
GPT-OSS 20B on Groq costs $0.07 per million input tokens and $0.30 per million output tokens ($0.1275/M blended at 3:1). Delivers lightning-fast inference on Groq LPUs for classification and code generation. Verified 2026-09-08.
How much does GPT-OSS 20B cost per 1,000 requests?
Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 2.93× verbosity factor.
| Request shape | Input tokens | Output tokens | Cost / 1,000 requests |
|---|---|---|---|
| Short | 100 | 50 | $0.0515 |
| Medium | 1,000 | 500 | $0.5145 |
| Long | 4,000 | 2,000 | $2.0580 |
Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.
Three model-specific pricing decisions
Groq-hosted 20B economics stay separate from open-weight specifications and the 120B page. Every bill below uses text-token units.
1. Fixed workload bills and reasoning-output sensitivity
| Workload | Input / output | 100K requests | Expansion treatment |
|---|---|---|---|
| Classification | 300 / 30 | $3.15 | Text tokens |
| Coding agent | 4,000 / 1,000 | $60.00 | Text tokens |
| Reasoning ×2 | 4,000 / 1,000 | $90.00 | Output doubled explicitly |
2. 20B → 120B cost-per-accepted-run threshold
| Fixed input / output | GPT-OSS 20B | GPT-OSS 120B (Cerebras) | Narrow decision boundary |
|---|---|---|---|
| Coding run · 4,000 / 1,000 | $60.00 | $215.00 | 15% accepted-result uplift required |
| Agent run · 8,000 / 1,200 | $96.00 | $370.00 | 20% accepted-result uplift required |
Formula: requests × (input tokens × input $/M + output tokens × output $/M) ÷ 1,000,000. The uplift is a planning threshold, not a measured quality claim.
3. Price / TTFT / throughput frontier
| Host/model | Listed input / output per M | Speed evidence | Capacity boundary |
|---|---|---|---|
| GPT-OSS 20B | $0.07 / $0.30 | 1120 tokens/sec; TTFT 140 ms; 5 measured samples | Rate limit and concurrency unavailable |
| GPT-OSS 120B (Cerebras) | $0.35 / $0.75 | 2450 tokens/sec; TTFT 90 ms; 5 measured samples | Do not infer queue capacity |
Verified 2026-04-06. Luna is the data owner. “Unavailable” means no compatible dated evidence was found; it is never treated as zero. First-party price source · Run this Batch 5 scenario.
All three Batch 5 decisions are server-rendered for GPT-OSS 20B; fixed inputs, formulas, dated sources, speed sample state, and unavailable mechanics are visible.
Batch 67 · exact model pricing decision contributions · verified 2026-09-07
Exact model boundary: Groq gpt-oss-20b (slug gpt-oss-20b). First-party provider pricing and API documentation remain fact owners.
LPU high-throughput token pricing and monthly spend matrix
Frozen Batch 67 scenario board. Formula / deterministic rule: monthly_spend = calls * ((in_tokens * 0.07 + out_tokens * 0.30) / 1M) Boundary: Owns Groq LPU token expenditure calculations for GPT-OSS 20B.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch67-gpt-oss-20b-m1-r150K low-latency conversational replies | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=50K low-latency conversational replies; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 50K low-latency conversational replies is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-20b-m1-r2200K code generation tasks | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=200K code generation tasks; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 200K code generation tasks is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-20b-m1-r31M automated customer triage passes | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=1M automated customer triage passes; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 1M automated customer triage passes is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-20b-m1-r4high-concurrency request surge | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=high-concurrency request surge; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — high-concurrency request surge is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-20b-m1-r5batch offline ingestion queue | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=batch offline ingestion queue; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — batch offline ingestion queue is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-20b-m1-r6unresolved billing currency | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=unresolved billing currency; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — unresolved billing currency has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: Groq Cloud official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.
Inference turnaround latency and streaming UX SLA audit
Frozen Batch 67 scenario board. Formula / deterministic rule: turnaround = ttft + (tokens_out / tps); Groq LPU ultra-high generation throughput Boundary: Owns turnaround time benchmarks and user experience responsiveness.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch67-gpt-oss-20b-m2-r1real-time conversational streaming (<80ms) | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=real-time conversational streaming (<80ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — real-time conversational streaming (<80ms) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-20b-m2-r2automated code autocomplete (<150ms) | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=automated code autocomplete (<150ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — automated code autocomplete (<150ms) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-20b-m2-r3instant document summary (<300ms) | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=instant document summary (<300ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — instant document summary (<300ms) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-20b-m2-r4long-form synthetic data generation | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=long-form synthetic data generation; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — long-form synthetic data generation is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-20b-m2-r5network contention latency buffer | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=network contention latency buffer; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — network contention latency buffer is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-20b-m2-r6unmeasured speed fixture | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=unmeasured speed fixture; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — unmeasured speed fixture is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
First-party provenance: Groq Cloud documentation; verification date 2026-09-07. Missing or conflicting joins fail closed.
Model tier escalation and multi-model routing model
Frozen Batch 67 scenario board. Formula / deterministic rule: blended_cost = 20b_volume * 20b_cost + 120b_volume * 120b_cost Boundary: Owns two-tier architectural routing between fast 20B triage and deep 120B generation.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch67-gpt-oss-20b-m3-r1100% GPT-OSS 20B baseline ($0.07/$0.30) | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=100% GPT-OSS 20B baseline ($0.07/$0.30); escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 100% GPT-OSS 20B baseline ($0.07/$0.30) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-20b-m3-r290% 20B triage / 10% 120B escalation | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=90% 20B triage / 10% 120B escalation; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 90% 20B triage / 10% 120B escalation is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-20b-m3-r380% 20B triage / 20% 120B escalation | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=80% 20B triage / 20% 120B escalation; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 80% 20B triage / 20% 120B escalation is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-20b-m3-r450% 20B triage / 50% 120B escalation | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=50% 20B triage / 50% 120B escalation; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 50% 20B triage / 50% 120B escalation is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-20b-m3-r5100% direct GPT-OSS 120B execution | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=100% direct GPT-OSS 120B execution; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 100% direct GPT-OSS 120B execution is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-20b-m3-r6unresolved confidence threshold trigger | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=unresolved confidence threshold trigger; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — unresolved confidence threshold trigger has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: Groq Cloud official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.
Method and limitations: formulas are deterministic; observed and assumed inputs are labeled; no missing provider, host, account, region, realm, alias, snapshot, revision, weight, artifact, control, tool, modality, workload, rate-period, timestamp, or result is transferred. Run the gpt-oss-20b Batch 67 scenario →
GPT-OSS 20B on Groq API Pricing: Blazing Speed at Budget Rates
GPT-OSS 20B on Groq costs $0.07 per million input tokens and $0.30 per million output tokens ($0.1275/M blended at 3:1). Delivers lightning-fast inference on Groq LPUs for classification and code generation. Verified 2026-09-08.
LPU high-throughput token pricing and monthly spend matrix
Frozen Batch 75 scenario board. Formula / deterministic rule: monthly_spend = calls * ((in_tokens * 0.07 + out_tokens * 0.30) / 1M) Boundary: Owns Groq LPU token expenditure calculations for GPT-OSS 20B.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch75-gpt-oss-20b-m1-r150K low-latency conversational replies | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=50K low-latency conversational replies; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 50K low-latency conversational replies is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-20b-m1-r2200K code generation tasks | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=200K code generation tasks; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 200K code generation tasks is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-20b-m1-r31M automated customer triage passes | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=1M automated customer triage passes; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 1M automated customer triage passes is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-20b-m1-r4high-concurrency request surge | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=high-concurrency request surge; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — high-concurrency request surge is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-20b-m1-r5batch offline ingestion queue | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=batch offline ingestion queue; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — batch offline ingestion queue is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-20b-m1-r6unresolved billing currency | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=unresolved billing currency; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — unresolved billing currency has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: Groq official API pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.
Inference turnaround latency and streaming UX SLA audit
Frozen Batch 75 scenario board. Formula / deterministic rule: turnaround = ttft + (tokens_out / tps); Groq LPU ultra-high generation throughput Boundary: Owns turnaround time benchmarks and user experience responsiveness.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch75-gpt-oss-20b-m2-r1real-time conversational streaming (<80ms) | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=real-time conversational streaming (<80ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — real-time conversational streaming (<80ms) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-20b-m2-r2automated code autocomplete (<150ms) | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=automated code autocomplete (<150ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — automated code autocomplete (<150ms) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-20b-m2-r3instant document summary (<300ms) | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=instant document summary (<300ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — instant document summary (<300ms) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-20b-m2-r4long-form synthetic data generation | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=long-form synthetic data generation; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — long-form synthetic data generation is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-20b-m2-r5network contention latency buffer | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=network contention latency buffer; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — network contention latency buffer is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-20b-m2-r6unmeasured speed fixture | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=unmeasured speed fixture; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — unmeasured speed fixture is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
First-party provenance: Groq developer documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Model tier escalation and multi-model routing model
Frozen Batch 75 scenario board. Formula / deterministic rule: blended_cost = 20b_volume * 20b_cost + 120b_volume * 120b_cost Boundary: Owns two-tier architectural routing between fast 20B triage and deep 120B generation.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch75-gpt-oss-20b-m3-r1100% GPT-OSS 20B baseline ($0.07/$0.30) | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=100% GPT-OSS 20B baseline ($0.07/$0.30); escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 100% GPT-OSS 20B baseline ($0.07/$0.30) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-20b-m3-r290% 20B triage / 10% 120B escalation | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=90% 20B triage / 10% 120B escalation; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 90% 20B triage / 10% 120B escalation is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-20b-m3-r380% 20B triage / 20% 120B escalation | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=80% 20B triage / 20% 120B escalation; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 80% 20B triage / 20% 120B escalation is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-20b-m3-r450% 20B triage / 50% 120B escalation | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=50% 20B triage / 50% 120B escalation; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 50% 20B triage / 50% 120B escalation is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-20b-m3-r5100% direct GPT-OSS 120B execution | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=100% direct GPT-OSS 120B execution; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 100% direct GPT-OSS 120B execution is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-20b-m3-r6unresolved confidence threshold trigger | model=gpt-oss-20b; slug=gpt-oss-20b; provider=Groq; scenario=unresolved confidence threshold trigger; escalation mix; total monthly queries; blended spend; cost savings vs pure 120B; accuracy retention; routing recommendation; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — unresolved confidence threshold trigger has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: Groq official API pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.
How fast is GPT-OSS 20B?
How much does GPT-OSS 20B cost at scale?
| Tokens / month | Est. cost (blended 3:1) |
|---|---|
| 100,000 | $0.01 |
| 1,000,000 | $0.13 |
| 10,000,000 | $1.31 |
| 100,000,000 | $13.12 |
How does GPT-OSS 20B compare with other models?
What is GPT-OSS 20B best for?
What should you explore next for GPT-OSS 20B?
What are common questions about GPT-OSS 20B?
Is GPT-OSS 20B cheaper than Muse Spark 1.3 Contributor?
GPT-OSS 20B costs $0.13/M blended tokens, Muse Spark 1.3 Contributor costs $0.13/M — Muse Spark 1.3 Contributor is cheaper.
How much does 1 million tokens cost with GPT-OSS 20B?
At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.13. Pure input costs $0.07/M; pure output costs $0.30/M.
What does GPT-OSS 20B cost at high volume?
At 100 million blended tokens a month, GPT-OSS 20B costs approximately $13.12. See the cost-at-scale table below for other volumes.
