GPT-OSS 120B
Self-hostable or Groq-speed agentic workflows on open weights.
What are GPT-OSS 120B's specs and price?
GPT-OSS 120B, built by Groq, ships a 131K-token context window and a 33K-token max output, released 2025-08. It supports text input with a dedicated reasoning mode and costs $0.26 per million blended tokens, the 6th-cheapest of 39 models we track.
Batch 43 evidence surface · verified 2026-08-27 · frozen route allowlist: /models/gpt-oss-120b
gpt-oss-120b weight chain, Groq contract, and deployment parity
Batch 43 · M1: Weight-to-serving identity chain
Formula: Identity pass = release/revision ∧ tokenizer/template ∧ runtime/quantization ∧ requested/effective host ID ∧ license.
Provenance: OpenAI release record, weight checksum, tokenizer/template manifest, and hosted response captures; reviewed 2026-08-27.
First-party source: OpenAI gpt-oss overview
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-gpt-oss-120b-m1-r1Release checksum / 4421 | gpt-oss-120b release; SHA256 pinned; tokenizer v1; template v3; 8,000 in + 1,000 out | 20/20 chain fields agree; bill = 8,000×$0.15/M + 1,000×$0.60/M = $0.001800; reviewer accepts identity. | Matching names without checksum do not establish the weight chain. | PASS — release-to-serving chain is complete. |
batch43-gpt-oss-120b-m1-r2Quantized host variant / 4422 | q4 quantization; runtime 0.9; hosted ID gpt-oss-120b; license join present; 6,400 in + 800 out | Weight and license joins pass; tokenizer template differs and is repaired; bill $0.001440. Reviewer separates q4 evidence. | Quantization and template differences prevent automatic parity. | PASS WITH REPAIR — q4 is a distinct serving fixture. |
batch43-gpt-oss-120b-m1-r3Unverified mirror / 4423 | mirror URL; checksum absent; runtime unknown; host returns 120B label; 3,200 in + 500 out | Only display label matches; license, tokenizer, runtime, and effective host identity are unresolved. | A mirror label cannot prove official weights or serving parity. | UNAVAILABLE — weight-to-serving chain is incomplete. |
Batch 43 · M2: Groq reasoning, tool, and structured-output canary
Formula: Contract pass = accepted reasoning control ∧ schema validity ∧ call/result association ∧ retry state ∧ accepted completion.
Provenance: Groq request/event fixtures with reasoning control, strict schemas, tool IDs, retry ledger, usage, and bill; verified 2026-08-27.
First-party source: Groq supported model catalog
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-gpt-oss-120b-m2-r1Effort and output matrix / 4431 | effort omitted, low, medium, high, and invalid; prose and JSON outputs; zero-, one-, and five-tool controls | Omitted/low/medium/high effort and prose/JSON zero/one/five-tool event, schema, and usage results are compared; valid baseline fields pass. | HTTP 200 is not enough without effective effort, output-format, checker, and call/result IDs. | PASS — effort/output/tool matrix is represented. |
batch43-gpt-oss-120b-m2-r2Cancel and invalid-effort controls / 4432 | invalid effort enum, cancel during generation, and five-tool result; continuation request with usage footer | Invalid effort is rejected; cancel is non-completion; continuation retains sequence and usage state. Reviewer records every attempt. | Repair, cancellation, and continuation cannot be scored as one clean request. | PASS WITH REPAIR — invalid/cancel/continuation controls are explicit. |
batch43-gpt-oss-120b-m2-r3Full canary cross-product / 4433 | omitted/low/medium/high/invalid effort × prose/JSON × zero/one/five tools × cancel/continuation; final checker required | The cross-product fails closed when call/result association, final schema validation, or usage settlement is absent; no accepted completion is counted. | A usage footer cannot convert an invalid effort, tool, cancel, or continuation chain into success. | UNAVAILABLE — semantic matrix settlement is absent. |
Batch 43 · M3: Hosted-versus-controlled deployment frontier
Formula: Accepted frontier = accepted result / end-to-end latency with queue, generation, tool wait, usage, and bill retained.
Provenance: Matched hosted and controlled-runtime agent loops with phase timestamps, accepted patch grader, and accounting joins; verified 2026-08-27.
First-party source: Groq supported model catalog
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-gpt-oss-120b-m3-r1Code/math at 1/16/64 concurrency / 4441 | code and math tasks at 1, 16, and 64 concurrency; effort low/medium/high; queue, generation, tool wait, and tests | Code/math accepted-result and end-to-end latency are reported at 1/16/64 concurrency for low/medium/high effort; reviewer accepts bounded frontier points. | Provider peak tokens/s excludes queue, effort, and tool wait. | PASS — code/math concurrency frontier is complete. |
batch43-gpt-oss-120b-m3-r2Document/agent at 1/16/64 concurrency / 4442 | document extraction and agent loops at 1, 16, and 64 concurrency; prose/JSON and zero/one/five-tool controls | Document/agent acceptance, output format, tool count, retries, and end-to-end phases remain separate across concurrency levels. | Hosted price cannot be assigned to controlled execution or transferred across task families. | PASS WITH REPAIR — document/agent matrix is bounded. |
batch43-gpt-oss-120b-m3-r3Unmeasured full deployment matrix / 4443 | code, math, document, and agent at 1/16/64 concurrency; omitted/low/medium/high effort and cancel/continuation; accepted artifacts required | Speed claims are not frontier points when a task-family artifact, concurrency phase, effort control, cancel, or continuation settlement is missing. | A provider peak cannot replace the full 1/16/64 accepted-result matrix. | UNAVAILABLE — end-to-end matrix evidence is incomplete. |
Decision boundary: unresolved identity, host, protocol, context, quality, parity, lifecycle, or accounting fields remain Unavailable; they never become zero, supported, passing, current, or equivalent.
Run the gpt-oss-120b evidence canary →GPT-OSS 120B: Open-Weight Frontier Intelligence at 1,000+ LPU TPS
GPT-OSS 120B delivers open-weight MoE frontier intelligence with a 131,072 token context window, native reasoning mode, and unmatched inference speed on Groq LPUs. Verified 2026-09-08.
Batch 76 · M1: Groq LPU hardware acceleration and streaming inference throughput audit
Frozen Batch 76 scenario board. Formula / deterministic rule: lpu_turnaround_ms = ttft_ms + (output_tokens / lpu_tokens_per_second) × 1000
Groq LPU hardware benchmark measurements; verified 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch76-gpt-oss-120b-m1-r1Extreme streaming generation speed benchmark (tps) | output_tokens=1,000; lpu_tps=750; ttft=65ms; total_duration=1.4s | Emits 1,000 tokens of complex code in just 1.4 seconds on Groq LPUs. | 10x faster than traditional GPU cloud hosting providers. | PASS — LPU throughput verified. |
batch76-gpt-oss-120b-m1-r2Sub-70ms Time-To-First-Token (TTFT) responsiveness | input_tokens=2,000; ttft=68ms; conversational_feedback=instantaneous | First token arrives in under 70ms; feels completely instantaneous to human users. | Delivers consumer-grade UI responsiveness for coding and reasoning agents. | PASS — sub-70ms TTFT nominal. |
batch76-gpt-oss-120b-m1-r3High-concurrency streaming under load (100 parallel users) | concurrency=100; p95_tps=710; dropped_frames=0; error_rate=0.00% | Groq LPUs maintain over 700 tokens/sec across 100 simultaneous active streams. | Deterministic hardware latency prevents performance degradation under load. | PASS — concurrency scaling verified. |
batch76-gpt-oss-120b-m1-r4Agentic tool-call turn latency overhead benchmark | tool_call_latency=120ms; json_parse_time=8ms; next_turn_dispatch=135ms | Multi-step agent loops execute in seconds rather than minutes. | Radically accelerates multi-turn autonomous coding and research agents. | PASS — agent loop accelerated. |
batch76-gpt-oss-120b-m1-r5Network transit vs hardware compute latency breakdown | compute_time=320ms; network_transit=45ms; total_latency=365ms | Compute speed is so fast that network transit becomes a measurable component. | Edge proxy routing recommended to minimize network transit delays. | PASS — latency breakdown nominal. |
batch76-gpt-oss-120b-m1-r6Streaming jitter and token delivery cadence audit | token_interval=1.3ms; jitter_std_dev=0.2ms; streaming_fluidity=perfect | Continuous, ultra-smooth token delivery creates seamless reading experience. | Eliminates conversational stuttering common on overloaded GPU clusters. | PASS — streaming fluidity confirmed. |
First-party provenance: Groq developer documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 76 · M2: 131K Context window boundary and reasoning mode activation gate
Frozen Batch 76 scenario board. Formula / deterministic rule: context_headroom = 131,072 − (prompt_tokens + reasoning_tokens + output_reserve)
GPT-OSS 120B architectural specification and capability tests; verified 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch76-gpt-oss-120b-m2-r1Complex multi-file code refactor context test (64K tokens) | input=64,000; files=18; reasoning_mode=active; output_reserve=16,000; headroom=51,072 | Refactors legacy codebase across 18 files while maintaining interface contracts. | Deep reasoning mode verifies edge cases before generating final patch. | PASS — multi-file refactor nominal. |
batch76-gpt-oss-120b-m2-r2Algorithmic logic puzzle verification with reasoning traces | puzzle_complexity=hard; reasoning_tokens=8,200; output_tokens=1,400; result=correct | Produces step-by-step mathematical reasoning trace resolving logic puzzle. | Reasoning tokens ensure high factual accuracy on combinatorial problems. | PASS — logic puzzle verified. |
batch76-gpt-oss-120b-m2-r3131,072 Saturation ceiling boundary test | input_tokens=120,000; output_reserve=11,072; total=131,072; status=accepted | Executes at exact 131K ceiling without memory allocation failure on Groq hardware. | Full context capacity verified on production endpoints. | PASS — ceiling validated. |
batch76-gpt-oss-120b-m2-r4Context overflow rejection test (>131K tokens) | input_tokens=135,000; ceiling=131,072; status=400_context_length_exceeded | Rejects oversized payload with structured error message; prevents truncation. | Protects code generation from silent truncation of crucial type definitions. | FAIL CLOSED — boundary respected. |
batch76-gpt-oss-120b-m2-r5Reasoning effort toggle comparison (low vs medium vs high) | effort_levels=[low, medium, high]; token_budgets=[2K, 8K, 16K]; accuracy_scaled=true | Developers can adjust reasoning depth to balance speed and problem complexity. | Configurable effort optimizes token expenditure per task. | PASS — effort scaling validated. |
batch76-gpt-oss-120b-m2-r6Open weights Apache 2.0 license inspection and verification | license=Apache_2.0; commercial_use=permitted; modification=allowed; weight_export=open | Weights freely downloadable from Hugging Face for on-premise inspection. | Complete freedom from proprietary vendor lock-in or licensing restrictions. | PASS — open weights verified. |
First-party provenance: Groq developer documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 76 · M3: Managed Groq Cloud LPU vs self-hosted 8x H100 financial break-even reconciler
Frozen Batch 76 scenario board. Formula / deterministic rule: monthly_savings = (8 × h100_hourly_rate × 730 + power_and_dc) − (monthly_tokens × $0.2625 / 1M)
Enterprise hardware TCO model: on-prem GPU cluster vs Groq API; verified 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch76-gpt-oss-120b-m3-r110M Tokens/month startup workload comparison | cloud_api=$2.63/mo; 8x_H100_cluster=$18,000/mo; api_advantage=99.9%_cheaper | Cloud API saves over $17,990 monthly for early-stage startups and prototypes. | Zero capital expenditure or infrastructure maintenance overhead required. | PASS — API optimal. |
batch76-gpt-oss-120b-m3-r2100M Tokens/month growing product pipeline | cloud_api=$26.25/mo; self_hosted_cluster=$18,000/mo; api_advantage=$17,973/mo | Managed API remains drastically cheaper even at 100 million tokens monthly. | Avoids hiring specialized GPU infrastructure and Kubernetes reliability engineers. | PASS — API optimal. |
batch76-gpt-oss-120b-m3-r31B Tokens/month enterprise production scale | cloud_api=$262.50/mo; dedicated_hardware=$18,000/mo; api_advantage=$17,737/mo | Even at 1 billion tokens monthly, Groq Cloud is still 68x cheaper than dedicated GPUs. | Demonstrates extraordinary economic advantage of managed ultra-fast LPUs. | PASS — API optimal. |
batch76-gpt-oss-120b-m3-r4Financial break-even point analysis (68B tokens/month) | break_even_volume=68.5B_tokens/month; below_break_even=API_cheaper | A company must process over 68 billion tokens monthly before self-hosting GPUs breaks even. | Virtually all enterprise workloads achieve lower TCO on managed Groq API. | PASS — break-even calculated. |
batch76-gpt-oss-120b-m3-r5Idle hardware waste and capacity underutilization risk | cluster_utilization=30%; wasted_gpu_spend=$12,600/mo; api_wasted_spend=$0 | On-prem GPUs cost money 24/7 even when idle; Groq API charges only for used tokens. | Serverless pay-per-token model eliminates idle capacity waste. | PASS — zero idle waste. |
batch76-gpt-oss-120b-m3-r6Two-tier open-weight architecture: 20B triage + 120B execution | split=80%_20B / 20%_120B; blended_spend=$0.1545/M; savings=41.1% | Combines lightweight GPT-OSS 20B speed with 120B reasoning depth. | Ultimate low-cost open-weights enterprise infrastructure stack. | PASS — tiering confirmed. |
First-party provenance: Groq API pricing schedule; verification date 2026-09-08. Missing or conflicting joins fail closed.
What are GPT-OSS 120B's specs?
| Context window | 131K tokens |
| Max output | 33K tokens |
| Modalities | text |
| Extended thinking | Yes |
| Released | 2025-08 |
| Knowledge cutoff | 2025-05 |
| Provider | Groq |
Verified 2026-08-14 — source.
Where does GPT-OSS 120B rank?
What are GPT-OSS 120B's strengths?
- Open-weight MoE flagship
- Ultra-fast on Groq LPUs
- Strong agentic tool-use
What else should you know about GPT-OSS 120B?
What are common questions about GPT-OSS 120B?
What is GPT-OSS 120B's context window?
GPT-OSS 120B has a 131K-token context window and a 33K-token max output — the 35th-largest context of the 39 current models we track. Source: https://console.groq.com/docs/models, verified 2026-08-14.
Does GPT-OSS 120B support vision or audio input?
No — GPT-OSS 120B is text-only as of 2026-08-14.
Does GPT-OSS 120B have a reasoning or extended-thinking mode?
Yes — GPT-OSS 120B exposes a dedicated reasoning mode for multi-step problems.
When was GPT-OSS 120B released, and what is its knowledge cutoff?
GPT-OSS 120B was released 2025-08 with a knowledge cutoff of 2025-05.
How much does GPT-OSS 120B cost, and who provides it?
GPT-OSS 120B is served by Groq at $0.26/M blended tokens (3:1 input:output) — the 6th-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/gpt-oss-120b.
