GPT-OSS 120B on Groq API Pricing: Frontier Open Weights at LPU Speeds
Explore GPT-OSS 120B on Groq API pricing ($0.15/M input, $0.60/M output), ultra-high LPU streaming tokens/sec, and high-throughput cost efficiency.
Full specs, context window and API limits →How much does GPT-OSS 120B cost per million tokens?
GPT-OSS 120B on Groq costs $0.15 per million input tokens and $0.60 per million output tokens ($0.2625/M blended at 3:1). Delivers unmatched inference speeds on Groq LPU hardware. Verified 2026-09-08.
How much does GPT-OSS 120B cost per 1,000 requests?
Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 1.82× verbosity factor.
| Request shape | Input tokens | Output tokens | Cost / 1,000 requests |
|---|---|---|---|
| Short | 100 | 50 | $0.0696 |
| Medium | 1,000 | 500 | $0.6960 |
| Long | 4,000 | 2,000 | $2.7840 |
Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.
Three model-specific pricing decisions
This is the Groq-hosted GPT-OSS 120B owner. Cerebras rates are a separately dated hosting input; model specifications and self-hosting claims remain linked elsewhere.
1. Groq versus Cerebras hosting bill and throughput frontier
| Host | Input / output per M | 100K / 2K / 500 bill | Throughput evidence |
|---|---|---|---|
| Groq | $0.15 / $0.60 | $60.00 | 780 tokens/sec |
| Cerebras | $0.35 / $0.75 | $107.50 | 2450 tokens/sec |
2. 120B → 20B accepted-result uplift threshold
| Fixed workload (input / output) | GPT-OSS 120B | GPT-OSS 20B | Decision boundary |
|---|---|---|---|
| Agent loop · 8,000 / 1,200 | $192.00 | $96.00 | 15% accepted-result uplift required to justify the higher bill |
| Coding · 4,000 / 1,000 | $120.00 | $60.00 | 20% accepted-result uplift required to justify the higher bill |
Formula: 100,000 × (input tokens × input $/M + output tokens × output $/M) ÷ 1,000,000. Uplift is a decision threshold, not a measured quality claim.
3. Hosted API versus self-hosting ledger
| Option | Sourced calculation | Unknowns kept unavailable | Decision |
|---|---|---|---|
| Groq API | $192.00 | None for listed token bill | Use when utilization is variable |
| Self-hosted | API spend: not applicable | Hardware, utilization, region, operations | Supply user inputs before comparing |
| Cerebras API | $370.00 | Provider-specific throughput/SLA gaps | Compare only dated host rows |
Verified 2026-04-06. Luna is the data owner for this server-rendered module. “Unavailable” means no compatible dated evidence was found; it is never treated as zero. First-party price source · Run this exact scenario.
All three decisions are server-rendered for GPT-OSS 120B; fixed inputs, formulas, dated sources, and unavailable states are intentionally visible.
Batch 67 · exact model pricing decision contributions · verified 2026-09-07
Exact model boundary: Groq gpt-oss-120b (slug gpt-oss-120b). First-party provider pricing and API documentation remain fact owners.
LPU high-throughput token pricing and monthly spend matrix
Frozen Batch 67 scenario board. Formula / deterministic rule: monthly_spend = calls * ((in_tokens * 0.15 + out_tokens * 0.60) / 1M) Boundary: Owns Groq LPU token expenditure calculations.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch67-gpt-oss-120b-m1-r125K low-latency conversational replies | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=25K low-latency conversational replies; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 25K low-latency conversational replies is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-120b-m1-r2100K code generation tasks | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=100K code generation tasks; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 100K code generation tasks is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-120b-m1-r3500K automated customer triage passes | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=500K automated customer triage passes; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 500K automated customer triage passes is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-120b-m1-r4high-concurrency request surge | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=high-concurrency request surge; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — high-concurrency request surge is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-120b-m1-r5batch offline ingestion queue | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=batch offline ingestion queue; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — batch offline ingestion queue is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-120b-m1-r6unresolved billing currency | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=unresolved billing currency; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — unresolved billing currency has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: Groq Cloud official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.
Inference turnaround latency and streaming UX SLA audit
Frozen Batch 67 scenario board. Formula / deterministic rule: turnaround = ttft + (tokens_out / tps); Groq LPU ultra-high generation throughput Boundary: Owns turnaround time benchmarks and user experience responsiveness.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch67-gpt-oss-120b-m2-r1real-time conversational streaming (<100ms) | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=real-time conversational streaming (<100ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — real-time conversational streaming (<100ms) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-120b-m2-r2automated code autocomplete (<200ms) | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=automated code autocomplete (<200ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — automated code autocomplete (<200ms) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-120b-m2-r3instant document summary (<500ms) | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=instant document summary (<500ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — instant document summary (<500ms) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-120b-m2-r4long-form synthetic data generation | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=long-form synthetic data generation; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — long-form synthetic data generation is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-120b-m2-r5network contention latency buffer | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=network contention latency buffer; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — network contention latency buffer is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-120b-m2-r6unmeasured speed fixture | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=unmeasured speed fixture; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — unmeasured speed fixture is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
First-party provenance: Groq Cloud documentation; verification date 2026-09-07. Missing or conflicting joins fail closed.
Self-hosted vLLM vs Grok Cloud managed LPU financial break-even
Frozen Batch 67 scenario board. Formula / deterministic rule: cloud_cost = tokens * 0.2625 / 1M; self_hosted = gpu_instances * 730 * hourly_rate Boundary: Owns infrastructure commitment and cloud vs on-premise trade-offs.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch67-gpt-oss-120b-m3-r110M tokens/month (Cloud optimal) | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=10M tokens/month (Cloud optimal); monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 10M tokens/month (Cloud optimal) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-120b-m3-r250M tokens/month (Cloud optimal) | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=50M tokens/month (Cloud optimal); monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 50M tokens/month (Cloud optimal) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-120b-m3-r3250M tokens/month (Cloud optimal) | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=250M tokens/month (Cloud optimal); monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 250M tokens/month (Cloud optimal) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-120b-m3-r41B tokens/month transition boundary | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=1B tokens/month transition boundary; monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — 1B tokens/month transition boundary is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-120b-m3-r5idle dedicated cluster waste | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=idle dedicated cluster waste; monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — idle dedicated cluster waste is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch67-gpt-oss-120b-m3-r6unsupported GPU cluster configuration | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=unsupported GPU cluster configuration; monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicit | Unavailable — unsupported GPU cluster configuration has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: Groq Cloud official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.
Method and limitations: formulas are deterministic; observed and assumed inputs are labeled; no missing provider, host, account, region, realm, alias, snapshot, revision, weight, artifact, control, tool, modality, workload, rate-period, timestamp, or result is transferred. Run the gpt-oss-120b Batch 67 scenario →
GPT-OSS 120B on Groq API Pricing: Frontier Open Weights at LPU Speeds
GPT-OSS 120B on Groq costs $0.15 per million input tokens and $0.60 per million output tokens ($0.2625/M blended at 3:1). Delivers unmatched inference speeds on Groq LPU hardware. Verified 2026-09-08.
LPU high-throughput token pricing and monthly spend matrix
Frozen Batch 75 scenario board. Formula / deterministic rule: monthly_spend = calls * ((in_tokens * 0.15 + out_tokens * 0.60) / 1M) Boundary: Owns Groq LPU token expenditure calculations.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch75-gpt-oss-120b-m1-r125K low-latency conversational replies | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=25K low-latency conversational replies; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 25K low-latency conversational replies is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-120b-m1-r2100K code generation tasks | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=100K code generation tasks; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 100K code generation tasks is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-120b-m1-r3500K automated customer triage passes | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=500K automated customer triage passes; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 500K automated customer triage passes is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-120b-m1-r4high-concurrency request surge | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=high-concurrency request surge; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — high-concurrency request surge is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-120b-m1-r5batch offline ingestion queue | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=batch offline ingestion queue; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — batch offline ingestion queue is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-120b-m1-r6unresolved billing currency | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=unresolved billing currency; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — unresolved billing currency has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: Groq official API pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.
Inference turnaround latency and streaming UX SLA audit
Frozen Batch 75 scenario board. Formula / deterministic rule: turnaround = ttft + (tokens_out / tps); Groq LPU ultra-high generation throughput Boundary: Owns turnaround time benchmarks and user experience responsiveness.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch75-gpt-oss-120b-m2-r1real-time conversational streaming (<100ms) | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=real-time conversational streaming (<100ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — real-time conversational streaming (<100ms) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-120b-m2-r2automated code autocomplete (<200ms) | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=automated code autocomplete (<200ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — automated code autocomplete (<200ms) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-120b-m2-r3instant document summary (<500ms) | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=instant document summary (<500ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — instant document summary (<500ms) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-120b-m2-r4long-form synthetic data generation | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=long-form synthetic data generation; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — long-form synthetic data generation is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-120b-m2-r5network contention latency buffer | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=network contention latency buffer; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — network contention latency buffer is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-120b-m2-r6unmeasured speed fixture | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=unmeasured speed fixture; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — unmeasured speed fixture is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
First-party provenance: Groq developer documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Self-hosted vLLM vs Groq Cloud managed LPU financial break-even
Frozen Batch 75 scenario board. Formula / deterministic rule: cloud_cost = tokens * 0.2625 / 1M; self_hosted = gpu_instances * 730 * hourly_rate Boundary: Owns infrastructure commitment and cloud vs on-premise trade-offs.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch75-gpt-oss-120b-m3-r110M tokens/month (Cloud optimal) | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=10M tokens/month (Cloud optimal); monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 10M tokens/month (Cloud optimal) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-120b-m3-r250M tokens/month (Cloud optimal) | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=50M tokens/month (Cloud optimal); monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 50M tokens/month (Cloud optimal) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-120b-m3-r3250M tokens/month (Cloud optimal) | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=250M tokens/month (Cloud optimal); monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 250M tokens/month (Cloud optimal) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-120b-m3-r41B tokens/month transition boundary | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=1B tokens/month transition boundary; monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 1B tokens/month transition boundary is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-120b-m3-r5idle dedicated cluster waste | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=idle dedicated cluster waste; monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — idle dedicated cluster waste is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch75-gpt-oss-120b-m3-r6unsupported GPU cluster configuration | model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=unsupported GPU cluster configuration; monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — unsupported GPU cluster configuration has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: Groq official API pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.
How fast is GPT-OSS 120B?
How much does GPT-OSS 120B cost at scale?
| Tokens / month | Est. cost (blended 3:1) |
|---|---|
| 100,000 | $0.03 |
| 1,000,000 | $0.26 |
| 10,000,000 | $2.62 |
| 100,000,000 | $26.25 |
How does GPT-OSS 120B compare with other models?
What is GPT-OSS 120B best for?
What should you explore next for GPT-OSS 120B?
What are common questions about GPT-OSS 120B?
Is GPT-OSS 120B cheaper than GPT-4o Mini?
GPT-OSS 120B costs $0.26/M blended tokens, GPT-4o Mini costs $0.26/M — GPT-4o Mini is cheaper.
How much does 1 million tokens cost with GPT-OSS 120B?
At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.26. Pure input costs $0.15/M; pure output costs $0.60/M.
What does GPT-OSS 120B cost at high volume?
At 100 million blended tokens a month, GPT-OSS 120B costs approximately $26.25. See the cost-at-scale table below for other volumes.
