GPT-OSS 20B
Cost-sensitive agentic tasks that don’t need the full 120B model.
What are GPT-OSS 20B's specs and price?
GPT-OSS 20B, built by Groq, ships a 131K-token context window and a 33K-token max output, released 2025-08. It supports text input with a dedicated reasoning mode and costs $0.13 per million blended tokens, the 4th-cheapest of 39 models we track.
Batch 43 evidence surface · verified 2026-08-27 · frozen route allowlist: /models/gpt-oss-20b
gpt-oss-20b constrained deployment and bounded escalation
Batch 43 · M1: Compact weight-and-runtime identity ledger
Formula: Identity pass = exact weights/revision ∧ tokenizer/template ∧ quantization/runtime ∧ host ID ∧ effective response; 120B evidence cannot transfer.
Provenance: OpenAI release and runtime manifests joined to 20B request captures; exact checksums and effective IDs reviewed 2026-08-27.
First-party source: OpenAI gpt-oss overview
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-gpt-oss-20b-m1-r1Laptop/workstation/hosted admission / 4451 | laptop, workstation, and hosted admission profiles; 20B checksum; tokenizer v1; template v3; q4 runtime | Laptop, workstation, and hosted identity/admission fields are joined separately; reviewer accepts only the exact 20B profile. | 120B checksum or performance cannot fill a laptop, workstation, or hosted 20B field. | PASS — deployment admission is exact. |
batch43-gpt-oss-20b-m1-r2Runtime template drift / 4452 | 20B weights exact; template v2 runtime; host ID correct; 3,800 in + 600 out | Weights and host pass; template mismatch causes 2/10 prompt failures; bill $0.000930; reviewer isolates runtime variant. | Same weights do not establish same prompt contract. | PASS WITH REPAIR — template drift remains visible. |
batch43-gpt-oss-20b-m1-r3Unnamed quantization / 4453 | 20B display label; quantization and checksum absent; effective response truncated | Exact serving identity and accounting cannot be joined. | Parameter count and label cannot identify a runtime. | UNAVAILABLE — runtime identity is missing. |
Batch 43 · M2: Constrained tool/schema contract matrix
Formula: Accepted = control ∧ strict schema ∧ call/result IDs ∧ repair state ∧ final checker ∧ usage/bill.
Provenance: 20B strict-schema/tool fixtures with checker output, repair attempts, and usage joins; verified 2026-08-27.
First-party source: OpenAI gpt-oss overview
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-gpt-oss-20b-m2-r1Low/medium/high extraction and code-repair / 4461 | low, medium, and high controls across extraction and code-repair prompts; strict schema; 1 call | Extraction and code-repair reliability is reported per low/medium/high control; schema, call/result, and final-checker fields pass where present. | All required fields must be checked after the tool result. | PASS — extraction/code-repair matrix is explicit. |
batch43-gpt-oss-20b-m2-r2Planning and arithmetic reliability / 4462 | low/medium/high planning and arithmetic prompts; invalid enum then deterministic repair; two attempts | Planning and arithmetic reliability is reported by control; first invalid attempt fails and second repair is retained with its cost. | Retry cost and original failure remain in the fixture; controls cannot be blended. | PASS WITH REPAIR — reliability matrix is disclosed. |
batch43-gpt-oss-20b-m2-r3Unsupported control settlement / 4463 | extraction/code-repair/planning/arithmetic at low/medium/high; unsupported nested field; tool result and final checker required | Request acceptance is not semantic reliability; missing final checker or control-specific output leaves the relevant matrix cell unavailable. | Unsupported fields and absent control evidence fail closed. | UNAVAILABLE — full reliability settlement is absent. |
Batch 43 · M3: Bounded escalation state machine
Formula: Escalation = classified failure → deterministic repair/retry/tool correction/20B handoff/human review, preserving identity, charge, and acceptance.
Provenance: Agent-loop state traces with failure classes, transitions, handoff IDs, and reviewer decisions; verified 2026-08-27.
First-party source: OpenAI gpt-oss overview
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-gpt-oss-20b-m3-r1Explicit retry and tool-correction path / 4471 | extraction/code-repair/planning/arithmetic failure S1; explicit retry, schema repair, and tool-correction transitions | Transition S1→retry/tool-correction→accepted preserves original failure, control level, and charge; reviewer accepts bounded state machine. | A retry or tool correction must preserve the original failure state. | PASS — escalation is auditable. |
batch43-gpt-oss-20b-m3-r2Hosted handoff after retry/tool failure / 4472 | laptop/workstation retry exhausted after tool timeout; hosted handoff ID linked; 20B usage separate | Accepted results before/after handoff are reported separately; retry and tool-correction costs remain visible. | A handoff cannot be counted as a 20B completion or hide a retry. | PASS WITH REPAIR — handoff boundary is explicit. |
batch43-gpt-oss-20b-m3-r3Human-review orphan / 4473 | failure class present; reviewer ticket missing; final actor and bill unclear | Transition path stops at escalation; acceptance, actor identity, and accounting are unresolved. | An escalation label is not a human decision. | UNAVAILABLE — terminal review state is missing. |
Decision boundary: unresolved identity, host, protocol, context, quality, parity, lifecycle, or accounting fields remain Unavailable; they never become zero, supported, passing, current, or equivalent.
Run the gpt-oss-20b evidence canary →GPT-OSS 20B: Ultra-Fast Budget Open-Weights Intelligence at $0.07/M
GPT-OSS 20B offers compact open-weight MoE efficiency with a 131,072 token context window, optional reasoning mode, and rock-bottom $0.07/M token rates on Groq LPUs. Verified 2026-09-08.
Batch 76 · M1: Rock-bottom token pricing ($0.07/$0.30) and high-volume classification spend
Frozen Batch 76 scenario board. Formula / deterministic rule: monthly_spend = calls × ((in_tokens × $0.07 + out_tokens × $0.30) / 1M)
Groq official pricing schedule and high-volume test fixtures; verified 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch76-gpt-oss-20b-m1-r1100K Intent routing classifications per day (3M/month) | volume=3,000,000; avg_in=200; avg_out=20; monthly_spend=$13.80 | Processes 3 million user intent classifications monthly for less than $14 total spend. | Ultra-low cost unlocks unlimited intent parsing without budget constraints. | PASS — classification nominal. |
batch76-gpt-oss-20b-m1-r2500K Customer support ticket triage passes | volume=500,000; in_tokens=500; out_tokens=50; monthly_spend=$25.00 | Automated ticket tagging, urgency detection, and department assignment for $25. | Replaces expensive proprietary models on routine categorization tasks. | PASS — ticket triage verified. |
batch76-gpt-oss-20b-m1-r310M Serverless IoT telemetry anomaly detection events | volume=10,000,000; in_tokens=100; out_tokens=10; monthly_spend=$100.00 | Processes 10 million real-time sensor events monthly on a $100 budget. | Brings conversational intelligence to high-frequency IoT telemetry streams. | PASS — IoT telemetry nominal. |
batch76-gpt-oss-20b-m1-r4Batch offline processing queue for content indexing | batch_tokens=50,000,000; batch_cost=$6.38; turnaround=45_minutes | Processes 50 million tokens of offline documentation indexing for just $6.38. | Extreme cost efficiency for bulk vector database enrichment. | PASS — batch economy confirmed. |
batch76-gpt-oss-20b-m1-r5Cost comparison vs GPT-4o-mini ($0.1275 vs $0.375 blended) | gpt_oss_blended=$0.1275/M; gpt_4o_mini=$0.375/M; savings=66.0% | Delivers 66% lower token expenditure than OpenAI’s budget flagship. | Substantial cost savings for high-volume consumer applications. | PASS — cost advantage verified. |
batch76-gpt-oss-20b-m1-r6Open weights Apache 2.0 commercial licensing verification | license=Apache_2.0; commercial_use=approved; redistribution=allowed | Complete ownership of weights without risk of provider deprecation or price hikes. | Ensures perpetual business continuity for enterprise deployments. | PASS — licensing verified. |
First-party provenance: Groq API pricing schedule; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 76 · M2: Sub-50ms Time-To-First-Token (TTFT) and streaming UX latency audit
Frozen Batch 76 scenario board. Formula / deterministic rule: total_latency_ms = ttft_ms + (output_tokens / tokens_per_second) × 1000
Groq LPU hardware speed benchmarks for compact models; verified 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch76-gpt-oss-20b-m2-r1Real-time conversational streaming latency (<50ms) | input=500; output=100; ttft=42ms; tps=920; duration=150ms | First token arrives in 42ms; entire 100-token completion finishes in 150ms. | Fastest conversational streaming available in the AI industry today. | PASS — sub-50ms TTFT confirmed. |
batch76-gpt-oss-20b-m2-r2Instant search query normalization and semantic expansion | input=50; output=25; ttft=35ms; tps=950; duration=61ms | Expands and corrects user search queries in 61ms total execution time. | Can run synchronously inside search autocomplete dropdowns without lag. | PASS — search autocomplete nominal. |
batch76-gpt-oss-20b-m2-r3High-frequency code autocomplete inline completion | input=1,500; output=40; ttft=48ms; tps=900; duration=92ms | Delivers IDE inline code suggestions in under 100ms total latency. | Matches native IDE typing speed for seamless developer flow. | PASS — code autocomplete nominal. |
batch76-gpt-oss-20b-m2-r4High-concurrency streaming under peak load (250 simultaneous) | concurrency=250; p95_ttft=55ms; dropped_packets=0; error_rate=0.00% | Maintains sub-60ms P95 latency even under 250 concurrent user streams. | Groq LPUs provide deterministic real-time performance at scale. | PASS — concurrency scaling verified. |
batch76-gpt-oss-20b-m2-r5Network transit vs hardware compute latency breakdown | compute_latency=45ms; network_transit=35ms; client_perceived=80ms | Hardware execution is so rapid that network transit represents 44% of delay. | Edge hosting recommended to maximize real-time benefits. | PASS — latency breakdown verified. |
batch76-gpt-oss-20b-m2-r6Streaming token jitter and visual smoothness audit | token_interval=1.1ms; jitter_std_dev=0.15ms; streaming_fluidity=flawless | Emits tokens with clockwork precision; zero pauses or rendering hitches. | Provides the most fluid streaming visual experience possible. | PASS — streaming smoothness validated. |
First-party provenance: Groq developer documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 76 · M3: Two-tier architectural routing: 20B triage vs 120B deep execution
Frozen Batch 76 scenario board. Formula / deterministic rule: savings_pct = 1 − ((20b_share × 20b_cost + 120b_share × 120b_cost) / 120b_cost)
Two-tier architectural routing model across open-weight model family; verified 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch76-gpt-oss-20b-m3-r185% 20B triage / 15% 120B escalation routing mix | volume=1M; 20b_share=850K; 120b_share=150K; blended_spend=$148; savings=43.6% | Saves 43.6% compared to executing 100% of queries on GPT-OSS 120B. | Maintains high answer quality while capturing budget tier economics. | PASS — routing balance optimal. |
batch76-gpt-oss-20b-m3-r295% 20B triage / 5% 120B escalation customer support pipeline | volume=10M; 20b_share=9.5M; 120b_share=500K; blended_spend=$1,343; savings=48.8% | Saves over $1,280 monthly on 10 million support conversations. | Routine FAQs resolved instantly by 20B; complex disputes escalate to 120B. | PASS — support routing nominal. |
batch76-gpt-oss-20b-m3-r3Confidence score threshold routing decision boundary | confidence_metric=entropy; threshold=0.25; low_entropy=20B; high_entropy=120B | Ambiguous or mathematically complex queries automatically route to 120B. | Deterministic confidence metrics prevent low-quality answers from reaching users. | PASS — confidence routing validated. |
batch76-gpt-oss-20b-m3-r4On-premise edge device deployment (MacBook / RTX 4090) | vram_required=14GB; quantization=4_bit; edge_tps=65; hardware_cost=$1,800 | Compact 20B parameter size fits comfortably in consumer VRAM for local execution. | Enables 100% private offline edge AI without any cloud API dependency. | PASS — edge deployment verified. |
batch76-gpt-oss-20b-m3-r5Failover redundancy during upstream internet disconnections | internet_status=offline; local_20b=active; core_functions_retained=true | Local 20B instance maintains essential device operation when cloud connectivity drops. | Provides critical resilience for industrial and automotive embedded systems. | PASS — offline resilience nominal. |
batch76-gpt-oss-20b-m3-r6Full-lifecycle TCO comparison vs proprietary cloud APIs | annual_tokens=12B; 20b_spend=$1,530; proprietary_spend=$15,000+; savings=89.8% | Enterprise saves over $13,000 annually per 12 billion tokens processed. | Proves that open-weight small models deliver unbeatable enterprise ROI. | PASS — TCO superiority confirmed. |
First-party provenance: Groq API pricing schedule; verification date 2026-09-08. Missing or conflicting joins fail closed.
What are GPT-OSS 20B's specs?
| Context window | 131K tokens |
| Max output | 33K tokens |
| Modalities | text |
| Extended thinking | Yes |
| Released | 2025-08 |
| Knowledge cutoff | 2025-05 |
| Provider | Groq |
Verified 2026-08-14 — source.
Where does GPT-OSS 20B rank?
What are GPT-OSS 20B's strengths?
- Compact open-weight model
- Very cheap on Groq
- Still supports reasoning mode
What else should you know about GPT-OSS 20B?
What are common questions about GPT-OSS 20B?
What is GPT-OSS 20B's context window?
GPT-OSS 20B has a 131K-token context window and a 33K-token max output — the 36th-largest context of the 39 current models we track. Source: https://console.groq.com/docs/models, verified 2026-08-14.
Does GPT-OSS 20B support vision or audio input?
No — GPT-OSS 20B is text-only as of 2026-08-14.
Does GPT-OSS 20B have a reasoning or extended-thinking mode?
Yes — GPT-OSS 20B exposes a dedicated reasoning mode for multi-step problems.
When was GPT-OSS 20B released, and what is its knowledge cutoff?
GPT-OSS 20B was released 2025-08 with a knowledge cutoff of 2025-05.
How much does GPT-OSS 20B cost, and who provides it?
GPT-OSS 20B is served by Groq at $0.13/M blended tokens (3:1 input:output) — the 4th-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/gpt-oss-20b.
