← Back to all pricing

GPT-OSS 120B on Groq API Pricing: Frontier Open Weights at LPU Speeds

Explore GPT-OSS 120B on Groq API pricing ($0.15/M input, $0.60/M output), ultra-high LPU streaming tokens/sec, and high-throughput cost efficiency.

Full specs, context window and API limits →

How much does GPT-OSS 120B cost per million tokens?

GPT-OSS 120B on Groq costs $0.15 per million input tokens and $0.60 per million output tokens ($0.2625/M blended at 3:1). Delivers unmatched inference speeds on Groq LPU hardware. Verified 2026-09-08.

Verified 2026-09-07 source
Input
$0.15/M
Output
$0.60/M
Blended
$0.26/M
Provider
Verified 2026-04-06source

How much does GPT-OSS 120B cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 1.82× verbosity factor.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.0696
Medium1,000500$0.6960
Long4,0002,000$2.7840

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.

Three model-specific pricing decisions

This is the Groq-hosted GPT-OSS 120B owner. Cerebras rates are a separately dated hosting input; model specifications and self-hosting claims remain linked elsewhere.

1. Groq versus Cerebras hosting bill and throughput frontier

HostInput / output per M100K / 2K / 500 billThroughput evidence
Groq$0.15 / $0.60$60.00780 tokens/sec
Cerebras$0.35 / $0.75$107.502450 tokens/sec

2. 120B → 20B accepted-result uplift threshold

Fixed workload (input / output)GPT-OSS 120BGPT-OSS 20BDecision boundary
Agent loop · 8,000 / 1,200$192.00$96.0015% accepted-result uplift required to justify the higher bill
Coding · 4,000 / 1,000$120.00$60.0020% accepted-result uplift required to justify the higher bill

Formula: 100,000 × (input tokens × input $/M + output tokens × output $/M) ÷ 1,000,000. Uplift is a decision threshold, not a measured quality claim.

3. Hosted API versus self-hosting ledger

OptionSourced calculationUnknowns kept unavailableDecision
Groq API$192.00None for listed token billUse when utilization is variable
Self-hostedAPI spend: not applicableHardware, utilization, region, operationsSupply user inputs before comparing
Cerebras API$370.00Provider-specific throughput/SLA gapsCompare only dated host rows

Verified 2026-04-06. Luna is the data owner for this server-rendered module. “Unavailable” means no compatible dated evidence was found; it is never treated as zero. First-party price source · Run this exact scenario.

All three decisions are server-rendered for GPT-OSS 120B; fixed inputs, formulas, dated sources, and unavailable states are intentionally visible.

Batch 67 · exact model pricing decision contributions · verified 2026-09-07

Exact model boundary: Groq gpt-oss-120b (slug gpt-oss-120b). First-party provider pricing and API documentation remain fact owners.

LPU high-throughput token pricing and monthly spend matrix

Frozen Batch 67 scenario board. Formula / deterministic rule: monthly_spend = calls * ((in_tokens * 0.15 + out_tokens * 0.60) / 1M) Boundary: Owns Groq LPU token expenditure calculations.

Frozen scenario / field IDModel, identity, provider, and evidence fieldsResultState
batch67-gpt-oss-120b-m1-r1
25K low-latency conversational replies
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=25K low-latency conversational replies; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 25K low-latency conversational replies is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch67-gpt-oss-120b-m1-r2
100K code generation tasks
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=100K code generation tasks; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 100K code generation tasks is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch67-gpt-oss-120b-m1-r3
500K automated customer triage passes
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=500K automated customer triage passes; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 500K automated customer triage passes is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch67-gpt-oss-120b-m1-r4
high-concurrency request surge
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=high-concurrency request surge; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — high-concurrency request surge is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch67-gpt-oss-120b-m1-r5
batch offline ingestion queue
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=batch offline ingestion queue; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — batch offline ingestion queue is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch67-gpt-oss-120b-m1-r6
unresolved billing currency
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=unresolved billing currency; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — unresolved billing currency has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: Groq Cloud official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.

Inference turnaround latency and streaming UX SLA audit

Frozen Batch 67 scenario board. Formula / deterministic rule: turnaround = ttft + (tokens_out / tps); Groq LPU ultra-high generation throughput Boundary: Owns turnaround time benchmarks and user experience responsiveness.

Frozen scenario / field IDModel, identity, provider, and evidence fieldsResultState
batch67-gpt-oss-120b-m2-r1
real-time conversational streaming (<100ms)
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=real-time conversational streaming (<100ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — real-time conversational streaming (<100ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch67-gpt-oss-120b-m2-r2
automated code autocomplete (<200ms)
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=automated code autocomplete (<200ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — automated code autocomplete (<200ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch67-gpt-oss-120b-m2-r3
instant document summary (<500ms)
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=instant document summary (<500ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — instant document summary (<500ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch67-gpt-oss-120b-m2-r4
long-form synthetic data generation
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=long-form synthetic data generation; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — long-form synthetic data generation is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch67-gpt-oss-120b-m2-r5
network contention latency buffer
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=network contention latency buffer; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — network contention latency buffer is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch67-gpt-oss-120b-m2-r6
unmeasured speed fixture
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=unmeasured speed fixture; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — unmeasured speed fixture is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict

First-party provenance: Groq Cloud documentation; verification date 2026-09-07. Missing or conflicting joins fail closed.

Self-hosted vLLM vs Grok Cloud managed LPU financial break-even

Frozen Batch 67 scenario board. Formula / deterministic rule: cloud_cost = tokens * 0.2625 / 1M; self_hosted = gpu_instances * 730 * hourly_rate Boundary: Owns infrastructure commitment and cloud vs on-premise trade-offs.

Frozen scenario / field IDModel, identity, provider, and evidence fieldsResultState
batch67-gpt-oss-120b-m3-r1
10M tokens/month (Cloud optimal)
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=10M tokens/month (Cloud optimal); monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 10M tokens/month (Cloud optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch67-gpt-oss-120b-m3-r2
50M tokens/month (Cloud optimal)
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=50M tokens/month (Cloud optimal); monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 50M tokens/month (Cloud optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch67-gpt-oss-120b-m3-r3
250M tokens/month (Cloud optimal)
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=250M tokens/month (Cloud optimal); monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 250M tokens/month (Cloud optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch67-gpt-oss-120b-m3-r4
1B tokens/month transition boundary
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=1B tokens/month transition boundary; monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — 1B tokens/month transition boundary is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch67-gpt-oss-120b-m3-r5
idle dedicated cluster waste
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=idle dedicated cluster waste; monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — idle dedicated cluster waste is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch67-gpt-oss-120b-m3-r6
unsupported GPU cluster configuration
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=unsupported GPU cluster configuration; monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-07; measurement versus assumption=explicitUnavailable — unsupported GPU cluster configuration has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: Groq Cloud official pricing; verification date 2026-09-07. Missing or conflicting joins fail closed.

Method and limitations: formulas are deterministic; observed and assumed inputs are labeled; no missing provider, host, account, region, realm, alias, snapshot, revision, weight, artifact, control, tool, modality, workload, rate-period, timestamp, or result is transferred. Run the gpt-oss-120b Batch 67 scenario →

Batch 75 Verified Model Pricing IntelligenceModel owner: gpt-oss-120b (Groq)Audit date: 2026-09-08

GPT-OSS 120B on Groq API Pricing: Frontier Open Weights at LPU Speeds

GPT-OSS 120B on Groq costs $0.15 per million input tokens and $0.60 per million output tokens ($0.2625/M blended at 3:1). Delivers unmatched inference speeds on Groq LPU hardware. Verified 2026-09-08.

LPU high-throughput token pricing and monthly spend matrix

Frozen Batch 75 scenario board. Formula / deterministic rule: monthly_spend = calls * ((in_tokens * 0.15 + out_tokens * 0.60) / 1M) Boundary: Owns Groq LPU token expenditure calculations.

Frozen scenario / field IDModel, identity, provider, and evidence fieldsResultState
batch75-gpt-oss-120b-m1-r1
25K low-latency conversational replies
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=25K low-latency conversational replies; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 25K low-latency conversational replies is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch75-gpt-oss-120b-m1-r2
100K code generation tasks
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=100K code generation tasks; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 100K code generation tasks is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch75-gpt-oss-120b-m1-r3
500K automated customer triage passes
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=500K automated customer triage passes; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 500K automated customer triage passes is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch75-gpt-oss-120b-m1-r4
high-concurrency request surge
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=high-concurrency request surge; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — high-concurrency request surge is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch75-gpt-oss-120b-m1-r5
batch offline ingestion queue
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=batch offline ingestion queue; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — batch offline ingestion queue is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch75-gpt-oss-120b-m1-r6
unresolved billing currency
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=unresolved billing currency; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — unresolved billing currency has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: Groq official API pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.

Inference turnaround latency and streaming UX SLA audit

Frozen Batch 75 scenario board. Formula / deterministic rule: turnaround = ttft + (tokens_out / tps); Groq LPU ultra-high generation throughput Boundary: Owns turnaround time benchmarks and user experience responsiveness.

Frozen scenario / field IDModel, identity, provider, and evidence fieldsResultState
batch75-gpt-oss-120b-m2-r1
real-time conversational streaming (<100ms)
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=real-time conversational streaming (<100ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — real-time conversational streaming (<100ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch75-gpt-oss-120b-m2-r2
automated code autocomplete (<200ms)
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=automated code autocomplete (<200ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — automated code autocomplete (<200ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch75-gpt-oss-120b-m2-r3
instant document summary (<500ms)
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=instant document summary (<500ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — instant document summary (<500ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch75-gpt-oss-120b-m2-r4
long-form synthetic data generation
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=long-form synthetic data generation; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — long-form synthetic data generation is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch75-gpt-oss-120b-m2-r5
network contention latency buffer
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=network contention latency buffer; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — network contention latency buffer is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch75-gpt-oss-120b-m2-r6
unmeasured speed fixture
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=unmeasured speed fixture; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — unmeasured speed fixture is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict

First-party provenance: Groq developer documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Self-hosted vLLM vs Groq Cloud managed LPU financial break-even

Frozen Batch 75 scenario board. Formula / deterministic rule: cloud_cost = tokens * 0.2625 / 1M; self_hosted = gpu_instances * 730 * hourly_rate Boundary: Owns infrastructure commitment and cloud vs on-premise trade-offs.

Frozen scenario / field IDModel, identity, provider, and evidence fieldsResultState
batch75-gpt-oss-120b-m3-r1
10M tokens/month (Cloud optimal)
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=10M tokens/month (Cloud optimal); monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 10M tokens/month (Cloud optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch75-gpt-oss-120b-m3-r2
50M tokens/month (Cloud optimal)
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=50M tokens/month (Cloud optimal); monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 50M tokens/month (Cloud optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch75-gpt-oss-120b-m3-r3
250M tokens/month (Cloud optimal)
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=250M tokens/month (Cloud optimal); monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 250M tokens/month (Cloud optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch75-gpt-oss-120b-m3-r4
1B tokens/month transition boundary
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=1B tokens/month transition boundary; monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 1B tokens/month transition boundary is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch75-gpt-oss-120b-m3-r5
idle dedicated cluster waste
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=idle dedicated cluster waste; monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — idle dedicated cluster waste is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch75-gpt-oss-120b-m3-r6
unsupported GPU cluster configuration
model=gpt-oss-120b; slug=gpt-oss-120b; provider=Groq; scenario=unsupported GPU cluster configuration; monthly token volume; Groq Cloud spend; 8x H100 GPU hosting cost; cost differential; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — unsupported GPU cluster configuration has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: Groq official API pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.

How fast is GPT-OSS 120B?

Tokens / sec
780
TTFT
160 ms
Rank
#4 of 31
$ / M ÷ t/s
$0.0003
Measured with 5 runs on a fixed prompt — see the full methodology.

How much does GPT-OSS 120B cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.03
1,000,000$0.26
10,000,000$2.62
100,000,000$26.25

How does GPT-OSS 120B compare with other models?

GPT-OSS 20B$0.13/MLlama 4 Maverick$0.30/MQwen 3.8 30B$1.20/MQwen 3.6 27B$1.20/MGPT-4o Mini$0.26/MGrok-3 Mini$0.26/MMistral Small 3.1$0.26/M
See all Groq models →

What is GPT-OSS 120B best for?

#10 for Structured Data Extraction#11 for Writing & Content#11 for Chatbots & Support
Looking for a cheaper option?
GPT-OSS 20B is 50% cheaper — a drop-in migration. See all 8 alternatives to GPT-OSS 120B

What are common questions about GPT-OSS 120B?

Is GPT-OSS 120B cheaper than GPT-4o Mini?

GPT-OSS 120B costs $0.26/M blended tokens, GPT-4o Mini costs $0.26/M — GPT-4o Mini is cheaper.

How much does 1 million tokens cost with GPT-OSS 120B?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.26. Pure input costs $0.15/M; pure output costs $0.60/M.

What does GPT-OSS 120B cost at high volume?

At 100 million blended tokens a month, GPT-OSS 120B costs approximately $26.25. See the cost-at-scale table below for other volumes.

Try GPT-OSS 120B for free

Run real prompts against GPT-OSS 120B and every other model on this page in one workspace.

Try GPT-OSS 120B Free