← Back to all pricing

Cerebras GPT-OSS 120B API Pricing: Extreme Speed on Wafer Scale

Analyze Cerebras GPT-OSS 120B API pricing ($0.35/M input, $0.75/M output), extreme wafer-scale inference speed (2,000+ tps), and high-throughput cost efficiency.

Full specs, context window and API limits →

How much does GPT-OSS 120B (Cerebras) cost per million tokens?

Cerebras GPT-OSS 120B costs $0.35 per million input tokens and $0.75 per million output tokens ($0.45/M blended at 3:1). Provides ultra-fast open-weights inference on the Cerebras CS-3 system. Verified 2026-09-08.

Verified 2026-09-08 source
Input
$0.35/M
Output
$0.75/M
Blended
$0.45/M
Provider
Verified 2026-06-14source

How much does GPT-OSS 120B (Cerebras) cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 2.32× verbosity factor.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.1220
Medium1,000500$1.2200
Long4,0002,000$4.8800

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.

Three model-specific pricing decisions

This owner is the Cerebras-hosted 120B delivery surface. The API bill is joined to the controlled speed sample; quota, hardware, and SLA claims remain unavailable.

1. Fixed agent bill joined to completion time

Agent shape100K billTTFT / throughput sampleEstimated completion time
8,000 input / 1,200 output$370.002450 tokens/sec; TTFT 90 ms; 5 measured samples0.49 sec output-only estimate
32,000 / 4,000 long agent$1420.002450 tokens/sec; TTFT 90 ms; 5 measured samples1.63 sec output-only estimate

Formula: output tokens ÷ measured tokens/sec; this excludes queueing and network time.

2. Request-volume and output-length capacity table

Monthly requestsOutput eachToken billQueue/concurrency boundary
100K500$107.50Concurrency unavailable
1M500$1075.00Concurrency unavailable
1M2,000$2200.00Output length is the sourced sensitivity

3. SDK adoption and quota evidence ledger

EvidenceDated resultSpend calculationMissing boundary
Replay-backed SDK examplesAPI docs / examples$107.50SLA and account quota unavailable
Account quotaUnavailable$107.50Cerebras quota terms not in price record
Self-hostingUnavailableAPI spend is not transferableHardware and utilization unavailable

Verified 2026-06-14. Luna is the data owner. “Unavailable” means no compatible dated evidence was found; it is never treated as zero. First-party price source · Run this Batch 5 scenario.

All three Batch 5 decisions are server-rendered for GPT-OSS 120B (Cerebras); fixed inputs, formulas, dated sources, speed sample state, and unavailable mechanics are visible.

Batch 68 · exact model pricing decision contributions · verified 2026-09-08

Exact model boundary: Cerebras cerebras-gpt-oss-120b (slug cerebras-gpt-oss-120b). First-party provider pricing and API documentation remain fact owners.

Wafer-scale token pricing and high-throughput monthly spend matrix

Frozen Batch 68 scenario board. Formula / deterministic rule: monthly_spend = calls * ((in_tokens * 0.35 + out_tokens * 0.75) / 1M) Boundary: Owns Cerebras CS-3 hardware acceleration token tariff modeling.

Frozen scenario / field IDModel, identity, provider, and evidence fieldsResultState
batch68-cerebras-gpt-oss-120b-m1-r1
50K instant customer support queries
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=50K instant customer support queries; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 50K instant customer support queries is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m1-r2
200K real-time document summarizations
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=200K real-time document summarizations; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 200K real-time document summarizations is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m1-r3
1M high-concurrency event extraction calls
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=1M high-concurrency event extraction calls; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 1M high-concurrency event extraction calls is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m1-r4
batch offline ingestion queue
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=batch offline ingestion queue; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — batch offline ingestion queue is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m1-r5
enterprise dedicated CS-3 provisioning
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=enterprise dedicated CS-3 provisioning; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — enterprise dedicated CS-3 provisioning is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m1-r6
unresolved billing currency
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=unresolved billing currency; workload; prompt tokens; completion tokens; monthly spend; speed advantage; cost verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — unresolved billing currency has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: Cerebras official pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.

Ultra-high token throughput SLA and generation turnaround audit

Frozen Batch 68 scenario board. Formula / deterministic rule: turnaround = ttft + (tokens_out / tps); Cerebras Wafer-Scale Engine 2,000+ tps Boundary: Owns turnaround time benchmarks and user experience responsiveness.

Frozen scenario / field IDModel, identity, provider, and evidence fieldsResultState
batch68-cerebras-gpt-oss-120b-m2-r1
sub-50ms conversational streaming response
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=sub-50ms conversational streaming response; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — sub-50ms conversational streaming response is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m2-r2
rapid code autocomplete (<100ms)
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=rapid code autocomplete (<100ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — rapid code autocomplete (<100ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m2-r3
instant multi-page legal summary (<250ms)
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=instant multi-page legal summary (<250ms); task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — instant multi-page legal summary (<250ms) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m2-r4
high-concurrency request surge
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=high-concurrency request surge; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — high-concurrency request surge is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m2-r5
network transit buffer
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=network transit buffer; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — network transit buffer is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m2-r6
unmeasured speed fixture
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=unmeasured speed fixture; task; response SLA; TTFT; generation speed; SLA compliance; responsiveness champion; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — unmeasured speed fixture is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict

First-party provenance: Cerebras documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Hardware acceleration cost efficiency and self-hosting break-even

Frozen Batch 68 scenario board. Formula / deterministic rule: managed_cost = tokens * blended_rate; cluster_cost = (8 * H100_hourly + power) * 730 Boundary: Owns cloud wafer-scale API versus private on-prem GPU cluster economics.

Frozen scenario / field IDModel, identity, provider, and evidence fieldsResultState
batch68-cerebras-gpt-oss-120b-m3-r1
low intermittent workload (<10M tokens/mo)
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=low intermittent workload (<10M tokens/mo); monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — low intermittent workload (<10M tokens/mo) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m3-r2
50M monthly tokens (API optimal)
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=50M monthly tokens (API optimal); monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 50M monthly tokens (API optimal) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m3-r3
250M high-throughput enterprise scale
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=250M high-throughput enterprise scale; monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 250M high-throughput enterprise scale is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m3-r4
1B+ constant saturation (cluster threshold)
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=1B+ constant saturation (cluster threshold); monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — 1B+ constant saturation (cluster threshold) is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m3-r5
uncommitted hardware idle hours
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=uncommitted hardware idle hours; monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — uncommitted hardware idle hours is a frozen fixture pending exact identity, configuration, and denominator joins.UNTESTED — assumption cannot establish a verdict
batch68-cerebras-gpt-oss-120b-m3-r6
unresolved data center energy tariff
model=cerebras-gpt-oss-120b; slug=cerebras-gpt-oss-120b; provider=Cerebras; scenario=unresolved data center energy tariff; monthly token volume; managed Cerebras spend; self-hosted 8x H100 cost; financial break-even; deployment verdict; prompt/config/input/output/result/cache/checkpoint/artifact hashes=required; evidence=2026-09-08; measurement versus assumption=explicitUnavailable — unresolved data center energy tariff has no matched, dated bilateral observation.FAIL CLOSED — manual, probe, or source evidence required

First-party provenance: Cerebras official pricing; verification date 2026-09-08. Missing or conflicting joins fail closed.

Method and limitations: formulas are deterministic; observed and assumed inputs are labeled; no missing provider, host, account, region, realm, alias, snapshot, revision, weight, artifact, control, tool, modality, workload, rate-period, timestamp, or result is transferred. Run the cerebras-gpt-oss-120b Batch 68 scenario →

How fast is GPT-OSS 120B (Cerebras)?

Tokens / sec
2450
TTFT
90 ms
Rank
#1 of 31
$ / M ÷ t/s
$0.0002
Measured with 5 runs on a fixed prompt — see the full methodology.

How much does GPT-OSS 120B (Cerebras) cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.04
1,000,000$0.45
10,000,000$4.50
100,000,000$45.00

How does GPT-OSS 120B (Cerebras) compare with other models?

GLM 4.7 (Cerebras)$2.38/MCodestral$0.45/MGPT-5.4 Nano$0.46/MGemini 3.1 Flash Lite$0.56/M
See all Cerebras models →

What is GPT-OSS 120B (Cerebras) best for?

#3 for Writing & Content#3 for Agents & Tool Use#3 for Chatbots & Support
Looking for a cheaper option?
GPT-OSS 20B is 70.8% cheaper — a config migration. See all 8 alternatives to GPT-OSS 120B (Cerebras)

What are common questions about GPT-OSS 120B (Cerebras)?

Is GPT-OSS 120B (Cerebras) cheaper than Codestral?

GPT-OSS 120B (Cerebras) costs $0.45/M blended tokens, Codestral costs $0.45/M — Codestral is cheaper.

How much does 1 million tokens cost with GPT-OSS 120B (Cerebras)?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.45. Pure input costs $0.35/M; pure output costs $0.75/M.

What does GPT-OSS 120B (Cerebras) cost at high volume?

At 100 million blended tokens a month, GPT-OSS 120B (Cerebras) costs approximately $45.00. See the cost-at-scale table below for other volumes.

Try GPT-OSS 120B (Cerebras) for free

Run real prompts against GPT-OSS 120B (Cerebras) and every other model on this page in one workspace.

Try GPT-OSS 120B (Cerebras) Free