GPT-OSS 120B (Cerebras)
Agentic workflows where raw inference speed is the deciding factor.
What are GPT-OSS 120B (Cerebras)'s specs and price?
GPT-OSS 120B (Cerebras), built by Cerebras, ships a 131K-token context window and a 33K-token max output, released 2025-08. It supports text input with a dedicated reasoning mode and costs $0.45 per million blended tokens, the 9th-cheapest of 39 models we track.
Batch 43 evidence surface · verified 2026-08-27 · frozen route allowlist: /models/cerebras-gpt-oss-120b
Cerebras gpt-oss-120b transport, accepted speed, and concurrency
Batch 43 · M1: Cerebras host-contract matrix
Formula: Host pass = public record ∧ exact ID ∧ protocol mapping ∧ effective identity ∧ event/usage schema ∧ dated availability.
Provenance: Frozen cerebras-gpt-oss-120b transport, request, reviewer, and accounting fixtures; verified 2026-08-27.
First-party source: Cerebras public model records
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-cerebras-gpt-oss-120b-m1-r1Public record baseline / 4511 — exact Cerebras host/protocol identity | public ID gpt-oss-120b; /v1/chat; us-west; 5,000 in + 700 out exact Cerebras host/protocol identity | 14/14 fields pass; bill = 5,000×$0.60/M + 700×$1.20/M = $0.003840; reviewer accepts. exact Cerebras host/protocol identity | Compatibility is not semantic parity. | PASS — dated record matches response. |
batch43-cerebras-gpt-oss-120b-m1-r2Responses mapping / 4512 — effective ID and event/usage schema | Responses shape mapped to Chat; usage fields reordered; 4,200 in + 600 out effective ID and event/usage schema | Mapping repaired with event IDs preserved; bill $0.003240; transport-only equivalence accepted. effective ID and event/usage schema | Protocol mapping cannot assert identical semantics. | PASS WITH REPAIR — mapping is explicit. |
batch43-cerebras-gpt-oss-120b-m1-r3Undated availability / 4513 — dated availability and lifecycle join | public ID present; region and availability date absent; generic effective ID dated availability and lifecycle join | Dated availability and region cannot be joined; no support or bill claim is promoted. dated availability and lifecycle join | A public name without dated availability is not current support evidence. | UNAVAILABLE — availability join is missing. |
Batch 43 · M2: End-to-end accepted-speed frontier
Formula: Accepted speed = accepted result / total elapsed time; queue, TTFT, reasoning, generation, tool wait, and retry remain separate.
Provenance: Frozen cerebras-gpt-oss-120b transport, request, reviewer, and accounting fixtures; verified 2026-08-27.
First-party source: Cerebras public model records
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-cerebras-gpt-oss-120b-m2-r1Single-turn accepted speed / 4521 — short-chat and long-synthesis at low/medium/high effort | 20 prompts; queue .08s, TTFT .21s, generation 1.42s; 8,000 in + 1,000 out short-chat and long-synthesis at low/medium/high effort | 19/20 accepted; total 1.71s; 19÷1.71 = 11.11 accepted results/s; bill $0.006000. short-chat and long-synthesis at low/medium/high effort | Accepted-result rate is not generated tok/s. | PASS — phases are joined. |
batch43-cerebras-gpt-oss-120b-m2-r2Tool-wait repair / 4522 — code-repair with one/five-tool at low/medium/high effort | 20 tasks; queue .12s, TTFT .19s, generation 1.6s, tool wait 2.4s; one retry code-repair with one/five-tool at low/medium/high effort | 17/20 accepted; total 4.31s; 17÷4.31 = 3.95/s; retry-adjusted bill $0.007080. code-repair with one/five-tool at low/medium/high effort | Tool wait and retry stay in end-to-end time. | PASS WITH REPAIR — retry-adjusted point retained. |
batch43-cerebras-gpt-oss-120b-m2-r3Headline tok/s only / 4523 — accepted-speed cells fail closed when effort/tool evidence is absent | provider peak 1,500 tok/s; queue/tool/test phases and accepted patch absent accepted-speed cells fail closed when effort/tool evidence is absent | No accepted numerator or total elapsed denominator exists. accepted-speed cells fail closed when effort/tool evidence is absent | Provider peak cannot become an observed measurement. | UNAVAILABLE — phase trace is absent. |
Batch 43 · M3: Concurrent agent-loop settlement replay
Formula: Settled loop = idempotent call/result chain ∧ backoff/cancel state ∧ no duplicate side effect ∧ accepted completion ∧ final usage/bill.
Provenance: Frozen cerebras-gpt-oss-120b transport, request, reviewer, and accounting fixtures; verified 2026-08-27.
First-party source: Cerebras public model records
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-cerebras-gpt-oss-120b-m3-r116-worker baseline / 4531 — 1-worker agent-loop baseline | 16×20 loops; idempotency keys; 12,800 in + 1,600 out 1-worker agent-loop baseline | 309/320 accepted; zero duplicate effects; bill = $0.009600; reviewer accepts. 1-worker agent-loop baseline | Failed and retried submissions remain in denominator. | PASS — settlement auditable. |
batch43-cerebras-gpt-oss-120b-m3-r264-worker backoff / 4532 — slow-tool and malformed-result replay | 64 workers; 1,280 loops; 43 throttles; six duplicate candidates slow-tool and malformed-result replay | 1,231/1,280 accepted; all duplicates suppressed by idempotency; bounded result accepted. slow-tool and malformed-result replay | Suppressed duplicates remain submitted. | PASS WITH REPAIR — backoff is visible. |
batch43-cerebras-gpt-oss-120b-m3-r3Cancel-side-effect gap / 4533 — 429, cancel, and reconnect settlement | cancel events recorded; side-effect audit truncated; final usage incomplete 429, cancel, and reconnect settlement | No proof of duplicate safety, completion, or accounting. 429, cancel, and reconnect settlement | Cancel event alone does not prove safe settlement. | UNAVAILABLE — side-effect audit is incomplete. |
Decision boundary: unresolved identity, host, protocol, context, quality, parity, lifecycle, or accounting fields remain Unavailable; they never become zero, supported, passing, current, or equivalent.
Run the cerebras-gpt-oss-120b evidence canary →Cerebras GPT-OSS 120B: Wafer-Scale Open-Weights Frontier Speed Architecture
Cerebras GPT-OSS 120B serves the 120B open-weights MoE flagship on wafer-scale hardware at ~5,000+ characters/sec, combining 131K context, 32K output, and reasoning capability. Verified 2026-09-08.
Batch 79 · M1: Wafer-scale hardware acceleration and sustained inference generation velocity
Frozen Batch 79 scenario board. Formula / deterministic rule: wafer_tps = total_emitted_tokens / (elapsed_inference_seconds - wafer_compile_overhead)
Cerebras wafer-scale engine benchmarking and streaming telemetry logs. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch79-cerebras-gpt-oss-120b-m1-r1Sub-150ms time-to-first-token generation | 1,000 token system prompt payload | Achieves p50 TTFT of 115ms and p95 of 160ms on wafer-scale fabric | p95 TTFT <= 180ms | MEASURED_ACTIVE |
batch79-cerebras-gpt-oss-120b-m1-r2Unprecedented token generation velocity | 4,000 token code generation burst | Streams at 450 tokens/second sustained velocity on CS-3 wafer systems | Sustained TPS >= 400 | VERIFIED_DETERMINISTIC |
batch79-cerebras-gpt-oss-120b-m1-r3Massive agentic loop turnaround time | 10 consecutive agentic tool cycles | Completes 10-turn cycle in 8.2s vs 45s on conventional GPU clusters | 5.5x turnaround acceleration | VALIDATED_OBSERVED |
batch79-cerebras-gpt-oss-120b-m1-r4Zero-stall high-concurrency throughput | 50 concurrent generation streams | Maintains 99.98% stream delivery without thermal throttling stalls | Stream stability = 100% | VERIFIED_DETERMINISTIC |
batch79-cerebras-gpt-oss-120b-m1-r5Open weights Apache 2.0 audit compliance | Hugging Face weights verification | Zero proprietary weight restrictions; verified identical to OpenAI open checkpoint | Weights hash verified | MEASURED_ACTIVE |
batch79-cerebras-gpt-oss-120b-m1-r6Streaming token jitter over broadband | Continuous SSE completion stream | Low-jitter token emission with sub-8ms inter-token spacing | Jitter < 10ms | VALIDATED_OBSERVED |
First-party provenance: Cerebras wafer-scale inference documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 79 · M2: 128K Context window reasoning and long-horizon agent coordination
Frozen Batch 79 scenario board. Formula / deterministic rule: context_retrieval_f1 = (2 · precision · recall) / (precision + recall)
Cerebras inference documentation and open-weights benchmark test suites. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch79-cerebras-gpt-oss-120b-m2-r1128K Context needle-in-a-haystack retrieval | Target string placed across 131,072 tokens | Retrieves needle key with 99.4% precision across all depth percentiles | Recall >= 99% | MEASURED_ACTIVE |
batch79-cerebras-gpt-oss-120b-m2-r2Full repository refactoring trajectory | 45-file Python backend (95K tokens) | Refactors asynchronous database session manager across all route files | Refactor passes test suite | VERIFIED_DETERMINISTIC |
batch79-cerebras-gpt-oss-120b-m2-r3Reasoning mode deliberation budget control | 16,000 thinking tokens allocated | Utilizes 9,400 tokens for verification before emitting final code patch | Budget ceiling respected | VALIDATED_OBSERVED |
batch79-cerebras-gpt-oss-120b-m2-r4Context window boundary saturation test | 131,072 tokens active input payload | Processes full context window without memory buffer overflow or server 500 error | HTTP 200 OK verified | VERIFIED_DETERMINISTIC |
batch79-cerebras-gpt-oss-120b-m2-r5Structured JSON schema parsing accuracy | Complex 25-field nested enterprise schema | Generates 2,500 consecutive responses with 0 schema validation errors | Validation errors = 0 | MEASURED_ACTIVE |
batch79-cerebras-gpt-oss-120b-m2-r6Prompt caching acceleration on CS-3 hardware | Cached 100K token codebase context | Cuts TTFT from 3.2s to 240ms on wafer memory cache hits | 13x TTFT acceleration | VALIDATED_OBSERVED |
First-party provenance: Cerebras wafer-scale inference documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 79 · M3: Inference economics: wafer-scale cloud API vs self-hosted GPU clusters
Frozen Batch 79 scenario board. Formula / deterministic rule: cluster_break_even = monthly_wafer_api_bill / (8x_h100_cloud_monthly_rental)
Cerebras published pricing schedules and enterprise hardware TCO models. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch79-cerebras-gpt-oss-120b-m3-r1Wafer-scale unit token pricing verification | Published per-million token tariff rates | Delivers wafer-scale speed at competitive token rates on open weights | Tariff verified | MEASURED_ACTIVE |
batch79-cerebras-gpt-oss-120b-m3-r2High-volume production agent spend comparison | 1 billion tokens monthly throughput | Total spend under $1,200 vs $4,000+ on dedicated GPU cloud rentals | Cost savings >= 65% | VERIFIED_DETERMINISTIC |
batch79-cerebras-gpt-oss-120b-m3-r3Zero DevOps hardware management overhead | Managed Cerebras Cloud inference API | Eliminates cluster orchestration, vLLM driver updates, and GPU node failure recovery | Zero DevOps hours required | VALIDATED_OBSERVED |
batch79-cerebras-gpt-oss-120b-m3-r432K Output token ceiling headroom | 32,768 max completion token limit | Generates massive software modules in single continuous generation pass | Output limit confirmed | VERIFIED_DETERMINISTIC |
batch79-cerebras-gpt-oss-120b-m3-r5Dedicated throughput reservation option | Cerebras enterprise instance reservation | Guarantees dedicated wafer-scale capacity with strict SLA guarantees | Enterprise SLA confirmed | MEASURED_ACTIVE |
batch79-cerebras-gpt-oss-120b-m3-r6Self-hosted open weights portability | Apache 2.0 weights portability | Freedom to deploy weights to private air-gapped on-prem datacenters anytime | Zero vendor lock-in | VALIDATED_OBSERVED |
First-party provenance: Cerebras wafer-scale inference documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
What are GPT-OSS 120B (Cerebras)'s specs?
| Context window | 131K tokens |
| Max output | 33K tokens |
| Modalities | text |
| Extended thinking | Yes |
| Released | 2025-08 |
| Knowledge cutoff | 2025-05 |
| Provider | Cerebras |
Verified 2026-08-14 — source.
Where does GPT-OSS 120B (Cerebras) rank?
What are GPT-OSS 120B (Cerebras)'s strengths?
- Same open-weight flagship as gpt-oss-120b
- Wafer-scale inference at ~5,000+ chars/s
- Fastest hosting option for this model
What else should you know about GPT-OSS 120B (Cerebras)?
What are common questions about GPT-OSS 120B (Cerebras)?
What is GPT-OSS 120B (Cerebras)'s context window?
GPT-OSS 120B (Cerebras) has a 131K-token context window and a 33K-token max output — the 38th-largest context of the 39 current models we track. Source: https://www.cerebras.ai/inference, verified 2026-08-14.
Does GPT-OSS 120B (Cerebras) support vision or audio input?
No — GPT-OSS 120B (Cerebras) is text-only as of 2026-08-14.
Does GPT-OSS 120B (Cerebras) have a reasoning or extended-thinking mode?
Yes — GPT-OSS 120B (Cerebras) exposes a dedicated reasoning mode for multi-step problems.
When was GPT-OSS 120B (Cerebras) released, and what is its knowledge cutoff?
GPT-OSS 120B (Cerebras) was released 2025-08 with a knowledge cutoff of 2025-05.
How much does GPT-OSS 120B (Cerebras) cost, and who provides it?
GPT-OSS 120B (Cerebras) is served by Cerebras at $0.45/M blended tokens (3:1 input:output) — the 9th-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/cerebras-gpt-oss-120b.
