Cerebras GPT-OSS 120B API Pricing: Extreme Speed on Wafer Scale
Analyze Cerebras GPT-OSS 120B API pricing ($0.35/M input, $0.75/M output), extreme wafer-scale inference speed (2,000+ tps), and high-throughput cost efficiency.
Full specs, context window and API limits →How much does GPT-OSS 120B (Cerebras) cost per million tokens?
Cerebras GPT-OSS 120B costs $0.35 per million input tokens and $0.75 per million output tokens ($0.45/M blended at 3:1). Provides ultra-fast open-weights inference on the Cerebras CS-3 system. Verified 2026-09-08.
How much does GPT-OSS 120B (Cerebras) cost per 1,000 requests?
Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 2.32× verbosity factor.
| Request shape | Input tokens | Output tokens | Cost / 1,000 requests |
|---|---|---|---|
| Short | 100 | 50 | $0.1220 |
| Medium | 1,000 | 500 | $1.2200 |
| Long | 4,000 | 2,000 | $4.8800 |
Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.
Three model-specific pricing decisions
This owner is the Cerebras-hosted 120B delivery surface. The API bill is joined to the controlled speed sample; quota, hardware, and SLA claims remain unavailable.
1. Fixed agent bill joined to completion time
| Agent shape | 100K bill | TTFT / throughput sample | Estimated completion time |
|---|---|---|---|
| 8,000 input / 1,200 output | $370.00 | 2450 tokens/sec; TTFT 90 ms; 5 measured samples | 0.49 sec output-only estimate |
| 32,000 / 4,000 long agent | $1420.00 | 2450 tokens/sec; TTFT 90 ms; 5 measured samples | 1.63 sec output-only estimate |
Formula: output tokens ÷ measured tokens/sec; this excludes queueing and network time.
2. Request-volume and output-length capacity table
| Monthly requests | Output each | Token bill | Queue/concurrency boundary |
|---|---|---|---|
| 100K | 500 | $107.50 | Concurrency unavailable |
| 1M | 500 | $1075.00 | Concurrency unavailable |
| 1M | 2,000 | $2200.00 | Output length is the sourced sensitivity |
3. SDK adoption and quota evidence ledger
| Evidence | Dated result | Spend calculation | Missing boundary |
|---|---|---|---|
| Replay-backed SDK examples | API docs / examples | $107.50 | SLA and account quota unavailable |
| Account quota | Unavailable | $107.50 | Cerebras quota terms not in price record |
| Self-hosting | Unavailable | API spend is not transferable | Hardware and utilization unavailable |
Verified 2026-06-14. Luna is the data owner. “Unavailable” means no compatible dated evidence was found; it is never treated as zero. First-party price source · Run this Batch 5 scenario.
All three Batch 5 decisions are server-rendered for GPT-OSS 120B (Cerebras); fixed inputs, formulas, dated sources, speed sample state, and unavailable mechanics are visible.
Batch 68 · exact model pricing decision contributions · verified 2026-09-08
Exact model boundary: Cerebras cerebras-gpt-oss-120b (slug cerebras-gpt-oss-120b). First-party provider pricing and API documentation remain fact owners.
Wafer-scale token pricing and high-throughput monthly spend matrix
Frozen Batch 68 scenario board. Formula / deterministic rule: monthly_spend = calls * ((in_tokens * 0.35 + out_tokens * 0.75) / 1M) Boundary: Owns Cerebras CS-3 hardware acceleration token tariff modeling.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch68-cerebras-gpt-oss-120b-m1-r150K instant customer support queries | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=50K instant customer support queries; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 50K instant customer support queries is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-gpt-oss-120b-m1-r2200K real-time document summarizations | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=200K real-time document summarizations; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 200K real-time document summarizations is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-gpt-oss-120b-m1-r31M high-concurrency event extraction calls | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=1M high-concurrency event extraction calls; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 1M high-concurrency event extraction calls is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-gpt-oss-120b-m1-r4batch offline ingestion queue | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=batch offline ingestion queue; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — batch offline ingestion queue is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-gpt-oss-120b-m1-r5enterprise dedicated CS-3 provisioning | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=enterprise dedicated CS-3 provisioning; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — enterprise dedicated CS-3 provisioning is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-gpt-oss-120b-m1-r6unresolved billing currency | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=unresolved billing currency; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — unresolved billing currency has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: Cerebras official pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.
Ultra-high token throughput SLA and generation turnaround audit
Frozen Batch 68 scenario board. Formula / deterministic rule: turnaround = ttft + (tokens_out / tps); Cerebras Wafer-Scale Engine 2,000+ tps Boundary: Owns turnaround time benchmarks and user experience responsiveness.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch68-cerebras-gpt-oss-120b-m2-r1sub-50ms conversational streaming response | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=sub-50ms conversational streaming response; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — sub-50ms conversational streaming response is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-gpt-oss-120b-m2-r2rapid code autocomplete (<100ms) | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=rapid code autocomplete (<100ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — rapid code autocomplete (<100ms) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-gpt-oss-120b-m2-r3instant multi-page legal summary (<250ms) | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=instant multi-page legal summary (<250ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — instant multi-page legal summary (<250ms) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-gpt-oss-120b-m2-r4high-concurrency request surge | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=high-concurrency request surge; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — high-concurrency request surge is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-gpt-oss-120b-m2-r5network transit buffer | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=network transit buffer; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — network transit buffer is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-gpt-oss-120b-m2-r6unmeasured speed fixture | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=unmeasured speed fixture; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — unmeasured speed fixture is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
First-party provenance: Cerebras documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Hardware acceleration cost efficiency and self-hosting break-even
Frozen Batch 68 scenario board. Formula / deterministic rule: managed_cost = tokens * blended_rate; cluster_cost = (8 * H100_hourly + power) * 730 Boundary: Owns cloud wafer-scale API versus private on-prem GPU cluster economics.
| Frozen scenario / field ID | Model, identity, provider, and evidence fields | Result | State |
|---|---|---|---|
batch68-cerebras-gpt-oss-120b-m3-r1low intermittent workload (<10M tokens/mo) | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=low intermittent workload (<10M tokens/mo); monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — low intermittent workload (<10M tokens/mo) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-gpt-oss-120b-m3-r250M monthly tokens (API optimal) | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=50M monthly tokens (API optimal); monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 50M monthly tokens (API optimal) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-gpt-oss-120b-m3-r3250M high-throughput enterprise scale | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=250M high-throughput enterprise scale; monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 250M high-throughput enterprise scale is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-gpt-oss-120b-m3-r41B+ constant saturation (cluster threshold) | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=1B+ constant saturation (cluster threshold); monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — 1B+ constant saturation (cluster threshold) is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-gpt-oss-120b-m3-r5uncommitted hardware idle hours | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=uncommitted hardware idle hours; monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — uncommitted hardware idle hours is a frozen fixture pending exact identity, configuration, and denominator joins. | UNTESTED — assumption cannot establish a verdict |
batch68-cerebras-gpt-oss-120b-m3-r6unresolved data center energy tariff | model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=unresolved data center energy tariff; monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicit | Unavailable — unresolved data center energy tariff has no matched, dated bilateral observation. | FAIL CLOSED — manual, probe, or source evidence required |
First-party provenance: Cerebras official pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.
Method and limitations: formulas are deterministic; observed and assumed inputs are labeled; no missing provider, host, account, region, realm, alias, snapshot, revision, weight, artifact, control, tool, modality, workload, rate-period, timestamp, or result is transferred. Run the cerebras-gpt-oss-120b Batch 68 scenario →
How fast is GPT-OSS 120B (Cerebras)?
How much does GPT-OSS 120B (Cerebras) cost at scale?
| Tokens / month | Est. cost (blended 3:1) |
|---|---|
| 100,000 | $0.04 |
| 1,000,000 | $0.45 |
| 10,000,000 | $4.50 |
| 100,000,000 | $45.00 |
How does GPT-OSS 120B (Cerebras) compare with other models?
What is GPT-OSS 120B (Cerebras) best for?
What should you explore next for GPT-OSS 120B (Cerebras)?
What are common questions about GPT-OSS 120B (Cerebras)?
Is GPT-OSS 120B (Cerebras) cheaper than Codestral?
GPT-OSS 120B (Cerebras) costs $0.45/M blended tokens, Codestral costs $0.45/M — Codestral is cheaper.
How much does 1 million tokens cost with GPT-OSS 120B (Cerebras)?
At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.45. Pure input costs $0.35/M; pure output costs $0.75/M.
What does GPT-OSS 120B (Cerebras) cost at high volume?
At 100 million blended tokens a month, GPT-OSS 120B (Cerebras) costs approximately $45.00. See the cost-at-scale table below for other volumes.
