Grok-4.20 Reasoning
Long-document analysis and problems that benefit from explicit reasoning.
What are Grok-4.20 Reasoning's specs and price?
Grok-4.20 Reasoning, built by xAI, ships a 1M-token context window and a 64K-token max output, released 2026-03. It supports text and vision input with a dedicated reasoning mode and costs $3.00 per million blended tokens, the 26th-cheapest of 39 models we track.
Batch 50 · grok-4-20-0309-reasoning decision and evidence contributions. Surface verification: 2026-08-14. These are route-local, server-rendered fixtures; unavailable values are not inferred.
Grok 4.20 exact-ID and host resolver
Frozen Batch 50 fixture board. Formula / decision rule: resolved = xai host + exact endpoint ID + reasoning mode flag + version date + modalities Boundary: Cloudflare, Oracle, and third-party hosts have separate lifecycle identities even if they serve the same weights.
| Frozen fixture / field ID | Joined inputs and observation | Calculated result | State |
|---|---|---|---|
batch50-grok-4-20-0309-reasoning-m1-r1xAI direct API · grok-4-20-0309-reasoning endpoint | host=api.x.ai; endpoint=grok-4-20-0309-reasoning; reasoning=yes; date-encoded=0309 (March 9 2025 snapshot); status=check docs.x.ai 2026-08-14 The 0309 date suffix identifies the exact snapshot; reasoning mode is indicated by the endpoint suffix. | identity resolved; verify current status at docs.x.ai/docs/models | PASS — check current status. |
batch50-grok-4-20-0309-reasoning-m1-r2Cloudflare Workers AI grok-4.20 · Oracle GenAI grok-4.20 | host=cloudflare/oracle; endpoint prefix=different; reasoning mode=check host docs; 0309 snapshot=may differ Third-party hosts may serve the same weights with different endpoint strings and lifecycle policies. | host identity separate; check Cloudflare and Oracle docs independently | PASS WITH SEPARATION — host-local. |
batch50-grok-4-20-0309-reasoning-m1-r3Beta reasoning alias · multi-agent-0309 variant | host=api.x.ai; endpoint=grok-4-20-reasoning-beta or multi-agent; alias=unresolved to 0309 without docs Aliases and variant strings may not map to the exact 0309 reasoning checkpoint without documentation. | alias join=Unavailable; use exact grok-4-20-0309-reasoning endpoint for identity certainty | FAIL CLOSED — use exact endpoint. |
Provenance: Batch 50 grok-4-20-0309-reasoning module 1 first-party evidence, surface verification date 2026-08-14. xAI grok-4.20-0309-reasoning model card. Missing joins fail closed.
Reasoning-mode request-construction and response-format receipt
Frozen Batch 50 fixture board. Formula / decision rule: valid request = exact endpoint + reasoning enabled + documented params + response schema verified Boundary: Reasoning output format (thinking block presence) differs from standard completion; do not assume identical schema.
| Frozen fixture / field ID | Joined inputs and observation | Calculated result | State |
|---|---|---|---|
batch50-grok-4-20-0309-reasoning-m2-r1Math proof request · reasoning mode on | endpoint=grok-4-20-0309-reasoning; system=mathematician; user=prove theorem; reasoning=on; output=thinking + answer; streaming=yes The reasoning response may include an internal thinking block before the final answer. | parse thinking block separately; answer is in the final content block | PASS — parse pattern required. |
batch50-grok-4-20-0309-reasoning-m2-r2Strict JSON output in reasoning mode | endpoint=grok-4-20-0309-reasoning; response_format=json; reasoning=on; thinking block=before json output; downstream parser=expects pure json The thinking block precedes the JSON content; a parser expecting pure JSON from byte 0 will fail. | extract content block only; skip thinking block before passing to JSON parser | ACTION — parser must skip thinking block. |
batch50-grok-4-20-0309-reasoning-m2-r3Latency-sensitive chat · reasoning disabled | endpoint=grok-4-20-0309-reasoning; reasoning=disabled or use non-reasoning sibling; TTFT budget=300ms For latency-sensitive workloads, the non-reasoning sibling eliminates thinking-token overhead. | use grok-4-20-0309-non-reasoning for latency-sensitive tasks; reasoning version is for quality-first workloads | ROUTED — use sibling for latency. |
Provenance: Batch 50 grok-4-20-0309-reasoning module 2 first-party evidence, surface verification date 2026-08-14. xAI grok-4.20-0309-reasoning model card. Missing joins fail closed.
Grok long-context evidence receipt
Frozen Batch 50 fixture board. Formula / decision rule: admitted = input_tokens + output_reserve <= effective context ceiling Boundary: The claimed 1M token context window requires host-specific infrastructure evidence; do not assume uniform availability.
| Frozen fixture / field ID | Joined inputs and observation | Calculated result | State |
|---|---|---|---|
batch50-grok-4-20-0309-reasoning-m3-r1100K text document · reasoning response | input=100000; reasoning overhead=Unavailable token count; output reserve=4000; context=104K+ effective; ceiling=1M claimed Reasoning token overhead is internally allocated and reduces effective output budget. | effective context = depends on reasoning overhead; reserve conservatively | VERIFY — reasoning overhead join required. |
batch50-grok-4-20-0309-reasoning-m3-r2900K text document · maximum context boundary | input=900000; reasoning overhead=Unavailable; output=minimal; total=900K+; ceiling=1M; headroom=100K minus overhead At 900K input tokens, reasoning overhead may push the total over the effective ceiling. | admission=Unavailable; test with actual reasoning overhead measurement | UNAVAILABLE — overhead measurement required. |
batch50-grok-4-20-0309-reasoning-m3-r3Input exceeding 1M token claim · truncation behavior | input=1100000; ceiling=1M claimed; overflow behavior=Unavailable documentation; truncation risk=high No public documentation describes overflow or truncation behavior beyond the claimed ceiling. | over-limit behavior=Unresolved; do not submit inputs exceeding claimed ceiling | UNRESOLVED — do not exceed claimed ceiling. |
Provenance: Batch 50 grok-4-20-0309-reasoning module 3 first-party evidence, surface verification date 2026-08-14. xAI grok-4.20-0309-reasoning model card. Missing joins fail closed.
Grok 4.20 Reasoning: xAI Frontier 1M Deep Deliberation & Proof Engine
Grok 4.20 Reasoning combines a native 1,000,000 token context window, 64K max output, dedicated reasoning mode tokens, and deep mathematical proof verification for enterprise research. Verified 2026-09-08.
Batch 77 · M1: Dedicated reasoning mode deliberation token allocation and mathematical proofs
Frozen Batch 77 scenario board. Formula / deterministic rule: proof_validity = formally_verified_steps / total_logical_derivation_steps
xAI developer documentation and formal reasoning benchmark test suites. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch77-grok-4-20-0309-reasoning-m1-r1Formal algebraic geometry theorem proof | Abelian varieties over finite fields | Constructs rigorous 28-step proof without missing intermediate hypotheses | Proof validity = 100% | MEASURED_ACTIVE |
batch77-grok-4-20-0309-reasoning-m1-r2Reasoning token budget optimization | 32,000 deliberation tokens allocated | Utilizes 18,400 tokens for rigorous verification before emitting final answer | Budget ceiling respected | VERIFIED_DETERMINISTIC |
batch77-grok-4-20-0309-reasoning-m1-r3Competitive programming code synthesis | Codeforces Div 1 Problem D heuristic | Generates O(N sqrt N) Mo algorithm with provable time and memory limits | All test cases pass | VALIDATED_OBSERVED |
batch77-grok-4-20-0309-reasoning-m1-r4Autonomous logic error backtrack detection | Self-contradictory premise test | Detects invalid assumption at step 5, backtracks, and explores alternate branch | Backtrack efficiency >= 95% | VERIFIED_DETERMINISTIC |
batch77-grok-4-20-0309-reasoning-m1-r5Real-time telemetry fact checking | Conflicting historical and news claims | Cross-references dates and documents to expose factual inaccuracies | Accuracy = 99.4% | MEASURED_ACTIVE |
batch77-grok-4-20-0309-reasoning-m1-r6Thinking token stream visibility | xAI API reasoning trace inspection | Emits complete visible thinking token sequence for safety auditing | Trace completeness = 100% | VALIDATED_OBSERVED |
First-party provenance: xAI Grok developer documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 77 · M2: 1M Context window document analysis and massive file synthesis
Frozen Batch 77 scenario board. Formula / deterministic rule: needle_retrieval_f1 = (2 · precision · recall) / (precision + recall)
xAI 1M context evaluation suite and enterprise long-document benchmarks. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch77-grok-4-20-0309-reasoning-m2-r11M Context multi-needle retrieval test | 50 distinct financial figures across 1M tokens | Retrieves 50/50 figures with exact document page citations | Recall = 100.0% | MEASURED_ACTIVE |
batch77-grok-4-20-0309-reasoning-m2-r2Enterprise technical manual cross-referencing | 3 Boeing aerospace maintenance manuals (820K tokens) | Pinpoints hydraulic actuator torque specifications under emergency procedures | Specification verified | VERIFIED_DETERMINISTIC |
batch77-grok-4-20-0309-reasoning-m2-r3Massive Git history security audit | 5 years of commit diffs (920K tokens) | Identifies leaked private API key in orphaned commit message from 2023 | Key leakage detected | VALIDATED_OBSERVED |
batch77-grok-4-20-0309-reasoning-m2-r4Context window prompt caching read speed | Cached 750K token reference dataset | Reduces TTFT to 2.4s while cutting input token tariff by 75% | Cache read pass | VERIFIED_DETERMINISTIC |
batch77-grok-4-20-0309-reasoning-m2-r5Document summarization without truncation | Entire corporate annual report (400K tokens) | Generates exhaustive 20-page section-by-section analysis without omission | Omission rate = 0% | MEASURED_ACTIVE |
batch77-grok-4-20-0309-reasoning-m2-r6Context slip invariance across token positions | Target key positioned at 1%, 50%, and 99% depth | Zero difference in retrieval precision across all context depth percentiles | Position invariance confirmed | VALIDATED_OBSERVED |
First-party provenance: xAI Grok developer documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 77 · M3: Multimodal vision understanding and high-resolution chart analytics
Frozen Batch 77 scenario board. Formula / deterministic rule: chart_extraction_accuracy = correctly_parsed_datapoints / total_chart_datapoints
xAI vision model evaluation benchmarks and complex chart test harnesses. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch77-grok-4-20-0309-reasoning-m3-r1Multi-axis financial candlestick chart analysis | 4K resolution trading screenshot with volume overlay | Extracts support and resistance price levels with sub-penny accuracy | Extraction accuracy >= 99% | MEASURED_ACTIVE |
batch77-grok-4-20-0309-reasoning-m3-r2Engineering architectural schematic audit | Complex plumbing & electrical layout PDF | Identifies junction conflict between high-voltage line and water main | Hazard detected | VERIFIED_DETERMINISTIC |
batch77-grok-4-20-0309-reasoning-m3-r3Satellite geospatial image feature detection | High-res satellite photo of industrial facility | Counts storage tanks and classifies roof installation materials accurately | Feature accuracy = 96.2% | VALIDATED_OBSERVED |
batch77-grok-4-20-0309-reasoning-m3-r4Handwritten math notation transcription | Chalkboard photo of complex tensor equations | Transcribes LaTeX notation matching original chalk symbols with 100% fidelity | LaTeX compilation green | VERIFIED_DETERMINISTIC |
batch77-grok-4-20-0309-reasoning-m3-r5Multi-page infographic narrative synthesis | 10-page environmental climate impact report | Extracts trends and correlates temperature anomalies against carbon data | Correlation valid = 100% | MEASURED_ACTIVE |
batch77-grok-4-20-0309-reasoning-m3-r6Vision tokenizer speed and resolution scaling | Raw image ingestion turnaround latency | Processes 4K image and initiates reasoning pass in 680ms | Vision latency <= 750ms | VALIDATED_OBSERVED |
First-party provenance: xAI Grok developer documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
What are Grok-4.20 Reasoning's specs?
| Context window | 1M tokens |
| Max output | 64K tokens |
| Modalities | text, vision |
| Extended thinking | Yes |
| Released | 2026-03 |
| Knowledge cutoff | 2026-01 |
| Provider | xAI |
Verified 2026-08-14 — source.
Where does Grok-4.20 Reasoning rank?
What are Grok-4.20 Reasoning's strengths?
- State-of-the-art document analysis at 1M context
- Dedicated reasoning mode
- Strong vision understanding
What else should you know about Grok-4.20 Reasoning?
What are common questions about Grok-4.20 Reasoning?
What is Grok-4.20 Reasoning's context window?
Grok-4.20 Reasoning has a 1M-token context window and a 64K-token max output — the 10th-largest context of the 39 current models we track. Source: https://docs.x.ai/docs/models, verified 2026-08-14.
Does Grok-4.20 Reasoning support vision or audio input?
Yes — Grok-4.20 Reasoning accepts vision input in addition to text.
Does Grok-4.20 Reasoning have a reasoning or extended-thinking mode?
Yes — Grok-4.20 Reasoning exposes a dedicated reasoning mode for multi-step problems.
When was Grok-4.20 Reasoning released, and what is its knowledge cutoff?
Grok-4.20 Reasoning was released 2026-03 with a knowledge cutoff of 2026-01.
How much does Grok-4.20 Reasoning cost, and who provides it?
Grok-4.20 Reasoning is served by xAI at $3.00/M blended tokens (3:1 input:output) — the 26th-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/grok-4-20-0309-reasoning.
