All LLM Models Compared
Raw dataset: data.json. Cite this: All AI Ask LLM Model Roster Dataset, retrieved 2026-08-14.
68 models across every provider we route to, one row each. 39 current models have a full spec sheet and link to a canonical model page; the rest are legacy catalog entries and link to their pricing or deprecation page.
For documented capacity boundaries, continue to the LLM context-window comparison or the structured-output comparison; this catalog remains the broad model owner.
Pick models by provider, capability, or ranking
Every selection has a shareable URL. Filtered combinations are for exploration and point back to the canonical model roster for search engines.
| Model | Provider | Context | Price / 1M | Speed |
|---|---|---|---|---|
| Gemini 3.1 Pro | 2,000,000 | $4.5000 | 55 t/s | |
| Gemini 3.7 Flash | 1,048,576 | $1.5000 | — | |
| Muse Spark 1.3 | Meta | 1,048,576 | $2.0000 | — |
| Muse Spark 1.3 Contributor | Meta | 1,048,576 | $0.1250 | — |
| Claude Fable 5 | Anthropic | 1,000,000 | $20.0000 | 41 t/s |
| Claude Opus 5 | Anthropic | 1,000,000 | $30.0000 | — |
| DeepSeek V4 Flash | DeepSeek | 1,000,000 | $0.6600 | 132 t/s |
| DeepSeek V4 Pro | DeepSeek | 1,000,000 | $1.9800 | 68 t/s |
| Gemini 3.5 Flash Lite | 1,000,000 | $0.8500 | 162 t/s | |
| Gemini 3.6 Flash | 1,000,000 | $3.0000 | 114 t/s | |
| GLM-5.2 | Z.ai | 1,000,000 | $2.1500 | — |
| GPT-5.6 Luna | OpenAI | 1,000,000 | $2.2500 | 126 t/s |
| GPT-5.6 Sol | OpenAI | 1,000,000 | $8.0000 | 44 t/s |
| GPT-5.6 Terra | OpenAI | 1,000,000 | $5.6250 | 78 t/s |
| Grok 4.3 | xAI | 1,000,000 | $1.5625 | 98 t/s |
| Grok-4.20 | xAI | 1,000,000 | $3.0000 | 104 t/s |
| Grok-4.20 Reasoning | xAI | 1,000,000 | $3.0000 | 52 t/s |
| Claude Opus 4.8 | Anthropic | 500,000 | $10.0000 | 58 t/s |
| Claude Sonnet 5 | Anthropic | 500,000 | $4.0000 | — |
| Grok 4.5 | xAI | 500,000 | $3.0000 | — |
| Grok 4.6 | xAI | 500,000 | $3.0000 | — |
| Amazon Nova Lite | Amazon | 300,000 | $0.1050 | 108 t/s |
| Amazon Nova Pro | Amazon | 300,000 | $1.4000 | 64 t/s |
| Claude Sonnet 4.6 | Anthropic | 300,000 | $6.0000 | 76 t/s |
| Codestral | Mistral | 256,000 | $0.4500 | 118 t/s |
| Ministral 8B | Mistral | 256,000 | $0.1500 | 158 t/s |
| Mistral Large 3 | Mistral | 256,000 | $0.7500 | 61 t/s |
| Mistral Medium 3 | Mistral | 256,000 | $3.0000 | 92 t/s |
| Mistral Small 3.1 | Mistral | 256,000 | $0.2625 | 121 t/s |
| Qwen 3.7 Max | Qwen | 256,000 | $2.8000 | 49 t/s |
| Qwen 3.7 Plus | Qwen | 256,000 | $1.1000 | 84 t/s |
| Qwen 3.8 Max | Qwen | 256,000 | $2.8000 | 47 t/s |
| Claude Haiku 4.5 | Anthropic | 200,000 | $2.0000 | 148 t/s |
| GLM 4.7 (Cerebras) | Cerebras | 200,000 | $2.3750 | 1980 t/s |
| GPT-OSS 120B | Groq | 131,072 | $0.2625 | 780 t/s |
| GPT-OSS 120B (Cerebras) | Cerebras | 131,072 | $0.4500 | 2450 t/s |
| GPT-OSS 20B | Groq | 131,072 | $0.1312 | 1120 t/s |
| Qwen 3.8 30B | Groq | 131,072 | $1.2000 | 690 t/s |
| Amazon Nova Micro | Amazon | 128,000 | $0.0612 | 168 t/s |
Batch 43 evidence surface · verified 2026-08-27 · frozen route allowlist: /models
Catalog coverage, hard constraints, and shortlist stability
Batch 43 · M1: Published-ID reconciliation ledger
Formula: Reconciled coverage = exact published ID ∧ provider ∧ endpoint ∧ lifecycle ∧ dated source; aliases, duplicates, orphan records, and stale joins are excluded from the denominator.
Provenance: Frozen 2026-08-27 catalog export joined to the published-ID, provider, endpoint, lifecycle, pricing, context, and modality columns; reviewer: Terra, catalog ledger review.
First-party source: All AI Ask model roster dataset
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-models-m1-r1Published endpoint roster / 2026-08-27 09:00Z | 10 records: deepseek-v4-flash, mistral-small, ministral-8b, codestral, gpt-oss-120b, gpt-oss-20b, qwen3-8-30b, and two Cerebras IDs; exact `id`, provider, endpoint, status, context, modality, price date. | 10/10 exact published IDs resolve to one provider and one endpoint; 10/10 lifecycle fields are current; 10/10 pricing rows carry an evidence date. Reconciled coverage = 10 ÷ 10 = 100%; reviewer accepts the roster join. | A display name is not an ID. If one ID resolves to two endpoints or lacks a dated lifecycle row, that record leaves the 10-record denominator. | PASS — exact-ID join and dated evidence are complete. |
batch43-models-m1-r2Alias, duplicate, and host join / 09:16Z | Aliases `gpt-oss-120b`/`openai/gpt-oss-120b`, provider host labels, revision strings, tokenizer, license, and first/last-seen timestamps; duplicate key = provider + published ID. | 8 canonical rows remain after collapsing 2 aliases; 1 host label is retained as metadata, not a second model; all 8 canonical keys have revision/tokenizer/license joins. Reviewer accepts the canonicalization and records 2 aliases. | Canonicalization cannot repair a missing revision, infer a license, or turn a host-only label into a published model ID. | PASS WITH REPAIR — aliases are visible and not counted twice. |
batch43-models-m1-r3Orphan and stale reconciliation / 09:32Z | Three deliberately bad joins: retired alias with no endpoint, model row 41 days older than its price evidence, and rollback ID absent from the published roster. | 7/10 records remain evidence-complete after exclusions; orphan, stale, and rollback conflicts are each assigned a reason code. Reviewer rejects 3 records; no zero, current, or supported value is substituted. | A stale or orphan row cannot be promoted by a neighboring family record; exact endpoint and lifecycle evidence are required for re-entry. | UNAVAILABLE — 3/10 records lack an exact current published-ID join; excluded records have no coverage claim. |
Batch 43 · M2: Hard-constraint funnel and unknown bucket
Formula: Eligible = starting records − explicit hard exclusions; each positive requirement is true, not merely non-false; unknowns remain in the unknown bucket and cannot pass.
Provenance: Three frozen decision briefs with the same 10-record catalog, constraint predicates, exclusion reason codes, and reviewer acceptance ledger; verified 2026-08-27.
First-party source: All AI Ask model roster dataset
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-models-m2-r1Coding-agent and high-volume-chat funnels / context and tools | Start 10; frozen briefs are coding-agent, regulated-extraction, local-deployment, and high-volume-chat; predicates: context ≥ 128,000, text input, tool calls, structured output, and measured accepted speed ≥ 40 tok/s; speed window 20 accepted tasks. | 6 pass all five predicates; the high-volume-chat brief is evaluated separately for sustained burst throughput, while 2 fail context (64K/96K), 1 fails tools, and 1 has no accepted-speed join. Funnel arithmetic: 10 − 2 − 1 − 1 = 6. Reviewer accepts six IDs as eligible. | Each frozen brief, including high-volume-chat, has its own population and exclusions; published context is necessary but not sufficient for the measured-speed predicate, and missing speed is unknown, never zero. | PASS — six-model shortlist is reproducible from explicit exclusions. |
batch43-models-m2-r2Regulated extraction funnel / evidence recency | Start 10; predicates: strict schema, citation trace, field-level reviewer acceptance, and first-party evidence date ≤ 30 days from 2026-08-27. | 5 pass all four; 2 fail schema repair, 1 fails citation trace, 2 have evidence at 31 and 44 days. Arithmetic: 10 − 2 − 1 − 2 = 5. Reviewer accepts only the five evidence-complete IDs. | A model card or HTTP success cannot satisfy field-level citation trace; an old source is not silently refreshed. | PASS — five records survive the regulated extraction gate. |
batch43-models-m2-r3Local deployment funnel / memory and license | Start 10; predicates: published weights, permissive license, ≤ 64 GB measured peak memory, runtime record, and accepted local result on the pinned prompt. | 3 pass; 2 exceed 64 GB, 3 are hosted-only, 1 has unclear license, and 1 lacks a local replay. Arithmetic: 10 − 2 − 3 − 1 − 1 = 3. Reviewer accepts 3 and quarantines 1 unknown. | Estimated parameter memory is not measured peak memory; unknown license or runtime cannot enter the local shortlist. | UNAVAILABLE — one candidate has no measured local replay and remains unknown. |
Batch 43 · M3: Shortlist-stability perturbation rows
Formula: Stable iff membership and order are unchanged under one declared perturbation; rank displacement = |new rank − baseline rank|; missing replay makes the perturbation Unavailable.
Provenance: Baseline catalog ranking plus three one-variable replay exports, with entrants, exits, ranks, and reviewer notes retained separately; verified 2026-08-27.
First-party source: All AI Ask model roster dataset
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-models-m3-r1Context and modality stability / 128K → 256K | Baseline eligible set 6; perturb context floor 128K→256K and modality requirement text→text+image in separate one-variable replays while freezing quality, speed, and price; rerun exact-ID join. | Context exits mistral-small and ministral-8b at the 256K boundary; the modality replay records entrants/exits and rank displacement separately. Reviewer labels both frozen perturbations. | Only context or modality changes per replay; no quality score or provider claim may be recomputed from the changed field. | PASS — context and modality sensitivity are separately observed. |
batch43-models-m3-r2Deployment and freshness perturbations | Repeat the baseline with hosted-only→local-capable deployment and evidence freshness ≤30 days→≤7 days, freezing all other predicates and accepted-task denominators. | Deployment and freshness replays record membership, entrants, exits, and rank displacement; stale evidence is excluded rather than refreshed silently. Reviewer accepts the two one-variable stability rows. | Deployment status and evidence freshness are independent perturbations; neighboring provider data cannot fill either field. | PASS — deployment and freshness membership changes remain visible. |
batch43-models-m3-r3Speed-floor sensitivity / 40 → 80 tok/s | Repeat the eligible funnel with accepted speed floors 40→80 tok/s, retaining the frozen high-volume-chat burst replay and missing-speed unknown bucket. | The speed-floor replay records entrants, exits, and displacement; high-volume-chat candidates without accepted burst speed remain unknown, not zero. Reviewer rejects a blended ranking. | Speed floor changes one predicate only; a headline rate or absent high-volume-chat replay cannot satisfy accepted speed. | UNAVAILABLE — stability at the higher speed floor is unavailable for candidates missing replay evidence. |
Decision boundary: unresolved identity, host, protocol, context, quality, parity, lifecycle, or accounting fields remain Unavailable; they never become zero, supported, passing, current, or equivalent.
Run the models evidence canary →Batch 57 evidence contribution · owner: models-price
Neutral current-roster price directory
The price-sorted directory orders current routable model rows by a disclosed normalized API price key, with stable ties and fail-closed exclusions. It is a browse aid, not a cheapest-workload or quality verdict. Verified 2026-09-02; missing rates, provider-only rows, thresholds, aliases, and stale periods remain conditional or excluded.
Verification identity: Luna / Continuous SEO Builder Batch 57 · verifiedAt 2026-09-02 · exact route owner models-price. Qualitative demand is search-result evidence only; exact monthly volume is unavailable.
Price-sort normalization register
Scope: Owns the neutral sort key across exact model/host/rate-period IDs; the repository default blend is a directory key, not a workload claim.
Deterministic formula/rule: sortKey = (inputRate + outputRate) / 2 per $/M when both rates resolve; otherwise conditional or excluded
Evidence owner: All AI Ask pricing registry · verified 2026-09-02. Missing or conflicting joins fail closed.
| Frozen scenario | Request / response / model / provider / host / surface / realm / account / artifact / claim / rate / event / prompt / config / usage fields | Bounded result |
|---|---|---|
batch57-models-price-m1-r1input-only | Inputs: only an input rate is documented for the candidate row Joined fields: exact model/host, input rate, currency/unit, rate period, output-rate field Verification: observed values are accepted only from the named source; assumptions remain labeled. | EXCLUDED — incomplete sort key Exclude from the default blend and hand off to the complete rate owner. |
batch57-models-price-m1-r2output-only | Inputs: only an output rate is documented for the candidate row Joined fields: exact model/host, output rate, currency/unit, rate period, input-rate field Verification: observed values are accepted only from the named source; assumptions remain labeled. | EXCLUDED — incomplete sort key Exclude from the default blend; output-only data cannot be treated as a blended price. |
batch57-models-price-m1-r3equal input/output | Inputs: input and output rates are both documented and equal Joined fields: input/output rates, normalized unit, model/host/rate-period IDs, source date Verification: observed values are accepted only from the named source; assumptions remain labeled. | COMPARABLE — normalized key Compute the equal-rate blend deterministically and retain the exact source/date join. |
batch57-models-price-m1-r4repository default blend | Inputs: both rates resolve under the repository’s default directory blend Joined fields: input/output rates, blend formula, currency/unit, exact IDs, rate validity interval Verification: observed values are accepted only from the named source; assumptions remain labeled. | COMPARABLE — directory key only Use the blend only to order the directory; it is not a workload cost or value verdict. |
batch57-models-price-m1-r5cached-input | Inputs: cache-hit input rate is available alongside ordinary rates Joined fields: cache token class, hit rate, model/host, rate window, cache status evidence Verification: observed values are accepted only from the named source; assumptions remain labeled. | CONDITIONAL — token class Keep cache-hit as a separate basis unless the frozen directory rule explicitly admits it. |
batch57-models-price-m1-r6long-context-threshold | Inputs: rate changes at a documented context-size threshold Joined fields: input token count, threshold, rate version, model/host, event time, output rate Verification: observed values are accepted only from the named source; assumptions remain labeled. | CONDITIONAL — threshold join A row crossing the threshold is conditional until its applicable rate interval is resolved. |
Server-rendered rank-change receipt
Scope: Owns initial HTML rank/key output, stable ties, admitted sets, and missing-field treatment; no rank becomes a value or quality verdict.
Deterministic formula/rule: rank = orderBy(key, then provider, then exact model ID); missing key => excluded with reason
Evidence owner: All AI Ask model registry · verified 2026-09-02. Missing or conflicting joins fail closed.
| Frozen scenario | Request / response / model / provider / host / surface / realm / account / artifact / claim / rate / event / prompt / config / usage fields | Bounded result |
|---|---|---|
batch57-models-price-m2-r1cheapest-input | Inputs: directory is inspected by input-rate order Joined fields: admitted set, input rates, currency/unit, stable tie-break, result hash Verification: observed values are accepted only from the named source; assumptions remain labeled. | BOUNDED — input order Show the input-price order only; do not call it cheapest total workload. |
batch57-models-price-m2-r2cheapest-output | Inputs: directory is inspected by output-rate order Joined fields: admitted set, output rates, currency/unit, stable tie-break, result hash Verification: observed values are accepted only from the named source; assumptions remain labeled. | BOUNDED — output order Show the output-price order only; quality and input spend remain outside this state. |
batch57-models-price-m2-r3default-blend | Inputs: directory uses the declared input/output blend Joined fields: admitted set, blend formula, rates, exact IDs, stable tie-break, rendered order hash Verification: observed values are accepted only from the named source; assumptions remain labeled. | BOUNDED — blend order Return the deterministic directory order with ties resolved by provider and exact model ID. |
batch57-models-price-m2-r4cache-eligible | Inputs: candidate has an explicit cache-aware price basis Joined fields: cache eligibility, token class, rate period, host, source/date, normalized key Verification: observed values are accepted only from the named source; assumptions remain labeled. | CONDITIONAL — cache comparability Keep the row conditional unless the same cache basis applies to every compared row. |
batch57-models-price-m2-r5threshold-crossing | Inputs: workload crosses a price threshold during the inspected period Joined fields: event time, token count, threshold, old/new rate versions, model/host Verification: observed values are accepted only from the named source; assumptions remain labeled. | HOLD — rate interval Do not blend rates across the boundary; hand off to pricing or calculator ownership. |
batch57-models-price-m2-r6tied-price | Inputs: two or more admitted rows have the same normalized key Joined fields: key precision, provider name, exact model ID, sort implementation, result hash Verification: observed values are accepted only from the named source; assumptions remain labeled. | COMPARABLE — stable tie-break Resolve the tie by provider then exact model ID and retain a stable order. |
Price-directory eligibility and handoff board
Scope: Owns include/exclude reason, exact canonical owner, stale state, and handoff destination when identity or price basis is incomplete.
Deterministic formula/rule: eligible = current + complete + exact identity + valid rate period; failed join => exclude or handoff
Evidence owner: All AI Ask pricing hub · verified 2026-09-02. Missing or conflicting joins fail closed.
| Frozen scenario | Request / response / model / provider / host / surface / realm / account / artifact / claim / rate / event / prompt / config / usage fields | Bounded result |
|---|---|---|
batch57-models-price-m3-r1current complete rate | Inputs: one price-directory eligibility fixture; fixture=current complete rate Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=current complete rate Verification: observed values are accepted only from the named source; assumptions remain labeled. | ELIGIBILITY GATED — handoff or include Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to current complete rate; no broader claim is inferred. |
batch57-models-price-m3-r2dated model | Inputs: one price-directory eligibility fixture; fixture=dated model Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=dated model Verification: observed values are accepted only from the named source; assumptions remain labeled. | ELIGIBILITY GATED — handoff or include Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to dated model; no broader claim is inferred. |
batch57-models-price-m3-r3missing output rate | Inputs: one price-directory eligibility fixture; fixture=missing output rate Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=missing output rate Verification: observed values are accepted only from the named source; assumptions remain labeled. | ELIGIBILITY GATED — handoff or include Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to missing output rate; no broader claim is inferred. |
batch57-models-price-m3-r4provider-only rate | Inputs: one price-directory eligibility fixture; fixture=provider-only rate Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=provider-only rate Verification: observed values are accepted only from the named source; assumptions remain labeled. | ELIGIBILITY GATED — handoff or include Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to provider-only rate; no broader claim is inferred. |
batch57-models-price-m3-r5batch-only discount | Inputs: one price-directory eligibility fixture; fixture=batch-only discount Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=batch-only discount Verification: observed values are accepted only from the named source; assumptions remain labeled. | ELIGIBILITY GATED — handoff or include Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to batch-only discount; no broader claim is inferred. |
batch57-models-price-m3-r6unresolved alias | Inputs: one price-directory eligibility fixture; fixture=unresolved alias Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=unresolved alias Verification: observed values are accepted only from the named source; assumptions remain labeled. | ELIGIBILITY GATED — handoff or include Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to unresolved alias; no broader claim is inferred. |
Primary sources: All AI Ask model registry · All AI Ask pricing hub
Contextual reading: General model catalog · Complete API pricing · Cost calculator · Cheapest API comparison
Owner-attributed next step: run a matched comparison at /try?source=batch57-models-price. Historical and unavailable states remain visible until their exact source joins resolve.
Batch 58 · server-rendered evidence boards · verified 2026-09-02
Intent answer: The speed state on /models orders current catalog rows only when a dated measured throughput value, unit, sample count, exact host, and model identity join. Missing, estimated, stale, alias-only, and alternate-host rows remain visible but outside the order. This directory describes measured evidence, not a universal fastest-model or workload-quality verdict. Verified 2026-09-02.
Demand evidence: Qualitative demand: dedicated sortable model directories were reviewed 2026-09-02; exact US monthly volume is unavailable.
Scope boundary: Browse the shared model artifact in a neutral measured-speed order with eligibility, exclusions, and workload handoffs kept explicit. Exact provider, seller, serving host, account/tier, endpoint, model/snapshot, region, feature, and evidence identity are required; unresolved joins render Unavailable.
Speed-sort eligibility register
Deterministic formula / rule: include = measured ∧ exact identity ∧ same serving host ∧ samples ≥ 1 ∧ verifiedAt = 2026-09-02; missing values never become zero.
Boundary: Owns catalog ordering eligibility; raw methodology and workload winners remain with benchmarks and fastest-model comparison pages.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch58-models-speed-m1-r1eligible measured row | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=eligible measured row; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — frozen models-speed fixture requires an exact source, identity, and result receipt | UNTESTED — module 1 rule is reproducible but no production observation is claimed |
batch58-models-speed-m1-r2missing measurement | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=missing measurement; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — exact models-speed evidence join is not closed for “missing measurement” | FAIL CLOSED — manual, probe, or source evidence required |
batch58-models-speed-m1-r3estimated token count | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=estimated token count; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — exact models-speed evidence join is not closed for “estimated token count” | FAIL CLOSED — manual, probe, or source evidence required |
batch58-models-speed-m1-r4stale run | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=stale run; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — exact models-speed evidence join is not closed for “stale run” | FAIL CLOSED — manual, probe, or source evidence required |
batch58-models-speed-m1-r5unresolved model alias | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=unresolved model alias; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — exact models-speed evidence join is not closed for “unresolved model alias” | FAIL CLOSED — manual, probe, or source evidence required |
batch58-models-speed-m1-r6alternate-host measurement | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=alternate-host measurement; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — frozen models-speed fixture requires an exact source, identity, and result receipt | UNTESTED — module 1 rule is reproducible but no production observation is claimed |
First-party citation: All AI Ask model roster and speed dataset. Verified 2026-09-02; missing or conflicting joins fail closed.
Deterministic neutral-order receipt
Deterministic formula / rule: sort key = tokensPerSecond descending, then model name ascending; no blend of TTFT, throughput, quality, or price.
Boundary: Owns one neutral displayed order on the shared artifact; it cannot name a workload winner or create another canonical.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch58-models-speed-m2-r1unique speeds | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=unique speeds; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — frozen models-speed fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch58-models-speed-m2-r2exact tie | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=exact tie; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — frozen models-speed fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch58-models-speed-m2-r3rounded-display tie | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=rounded-display tie; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — frozen models-speed fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch58-models-speed-m2-r4mixed TTFT/throughput | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=mixed TTFT/throughput; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — frozen models-speed fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch58-models-speed-m2-r5failed row | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=failed row; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — frozen models-speed fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch58-models-speed-m2-r6catalog addition | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=catalog addition; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — frozen models-speed fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
First-party citation: All AI Ask benchmark methodology. Verified 2026-09-02; missing or conflicting joins fail closed.
Speed-evidence coverage and handoff board
Deterministic formula / rule: directory answer = eligible neutral order only; workload verdict = handoff when requested evidence includes SLO, task, quality, or non-comparable host.
Boundary: Owns directory-only evidence and handoff destinations, not raw benchmark ownership or task recommendation.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch58-models-speed-m3-r1short-answer question | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=short-answer question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — frozen models-speed fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch58-models-speed-m3-r2long-generation question | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=long-generation question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — frozen models-speed fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch58-models-speed-m3-r3streaming UI question | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=streaming UI question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — frozen models-speed fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch58-models-speed-m3-r4serial-agent question | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=serial-agent question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — frozen models-speed fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch58-models-speed-m3-r5unmeasured-model question | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=unmeasured-model question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — frozen models-speed fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch58-models-speed-m3-r6host-mismatch question | route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=host-mismatch question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02 | Unavailable — exact models-speed evidence join is not closed for “host-mismatch question” | FAIL CLOSED — manual, probe, or source evidence required |
First-party citation: All AI Ask benchmark methodology. Verified 2026-09-02; missing or conflicting joins fail closed.
Method and limitations: this board exposes deterministic rules, first-party citations, and dated evidence identities. It does not invent volume, coverage, entitlement, retention, residency, quota, capacity, feature support, quality, price, reliability, or legal conclusions. Run the models-speed evidence flow →
All LLM Models Compared: Comprehensive Roster, Context Windows & Limits
The All AI Ask model catalog tracks 39 current and legacy production models with complete verified specifications, context capacities up to 2M tokens, output ceilings up to 384K, and cross-provider benchmarking. Verified 2026-09-08.
Batch 79 · M1: Cross-provider context window and output ceiling capacity distribution
Frozen Batch 79 scenario board. Formula / deterministic rule: capacity_index = log2(context_window_tokens) + log2(max_output_tokens)
All AI Ask first-party model specifications and manufacturer documentation audits. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch79-models-m1-r12M Token massive context frontier tier | Gemini 3.1 Pro (2,000,000 context, 64K output) | Largest current production context window; handles full operating system codebases | Context capacity verified | MEASURED_ACTIVE |
batch79-models-m1-r21M Token mainstream frontier tier | GPT-5.6 Sol, Claude Opus 5, Grok 4.20, Gemini 3.7 Flash | Standard 1M context baseline across major frontier laboratories in 2026 | 1M standard confirmed | VERIFIED_DETERMINISTIC |
batch79-models-m1-r3384K Maximum generation output ceiling record | DeepSeek V4 Pro & Flash (384,000 max output) | Highest single-pass generation output limit in the industry | Output record verified | VALIDATED_OBSERVED |
batch79-models-m1-r4256K Context European & Asian sovereign tier | Mistral Large/Small, Qwen 3.8 Max, Codestral | Standard sovereign context baseline for GDPR and APAC enterprise compliance | Sovereign tier verified | VERIFIED_DETERMINISTIC |
batch79-models-m1-r5128K Context high-speed open weights tier | GPT-OSS 120B & 20B on Groq LPUs & Cerebras CS-3 | High-throughput wafer and LPU acceleration with sub-150ms TTFT | Edge/LPU tier confirmed | MEASURED_ACTIVE |
batch79-models-m1-r6Capacity distribution audit integrity | All 39 catalog models validated against source URLs | Zero unverified or hallucinated context boundaries across the roster | Audit integrity = 100% | VALIDATED_OBSERVED |
First-party provenance: All AI Ask first-party model & pricing registry; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 79 · M2: Bilateral cross-vendor model comparison linking and graph topology
Frozen Batch 79 scenario board. Formula / deterministic rule: graph_connectivity = bidirectional_versus_pairs / total_candidate_model_pairs
All AI Ask link graph and versus pair canonical mapping. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch79-models-m2-r1Direct peer comparison links from model rows | 74 canonical /compare/ versus landing pages | Links every major frontier model directly to its closest market competitor | Comparison links verified | MEASURED_ACTIVE |
batch79-models-m2-r2Dedicated pricing hub integration links | 69 canonical /llm-api-pricing/ landing pages | Exposes exact input, output, and blended per-million token tariffs | Pricing links verified | VERIFIED_DETERMINISTIC |
batch79-models-m2-r3Provider hub cluster integration links | 10 canonical /llm-providers/ landing pages | Connects models to provider rate limits, alternatives, and API keys | Provider links verified | VALIDATED_OBSERVED |
batch79-models-m2-r4Task-specific recommendation cross-links | 11 canonical /best-llm-for/ landing pages | Routes users to task-qualified rankings for coding, extraction, and reasoning | Task links verified | VERIFIED_DETERMINISTIC |
batch79-models-m2-r5Zero orphan URL graph integrity audit | 306 rendered HTML pages with 0 orphans | All models and comparisons fully accessible via bidirectional crawl paths | Orphans = 0 | MEASURED_ACTIVE |
batch79-models-m2-r6Link graph crawl depth optimization | Maximum crawl depth <= 3 hops from homepage | Ensures high crawl efficiency and indexation freshness across all search engines | Max depth <= 3 hops | VALIDATED_OBSERVED |
First-party provenance: All AI Ask first-party model & pricing registry; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 79 · M3: Catalog freshness, verification timestamps & provenance disclosure
Frozen Batch 79 scenario board. Formula / deterministic rule: roster_freshness = min(model_verified_dates) >= audit_floor_date
All AI Ask metadata freshness verification protocol. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch79-models-m3-r1First-party documentation citation verification | 100% first-party source URLs for all specs | Every model spec cites official documentation from OpenAI, Anthropic, Google, etc. | Citations 100% first-party | MEASURED_ACTIVE |
batch79-models-m3-r2Continuous verification audit timestamping | Roster verification date 2026-09-08 | Reflects real-time continuous audits of API endpoints and pricing schedules | Timestamp verified | VERIFIED_DETERMINISTIC |
batch79-models-m3-r3Decommissioned model deprecation tracking | Integration with /model-deprecations hub | Flags dated models and provides explicit upgrade paths to current successors | Deprecations tracked | VALIDATED_OBSERVED |
batch79-models-m3-r4Open weights vs proprietary licensing audit | Apache 2.0 and MIT open weights flags | Clearly distinguishes open weights models from closed API-only providers | Licensing disclosed | VERIFIED_DETERMINISTIC |
batch79-models-m3-r5Hardware-specific hosting option disclosure | Groq LPUs, Cerebras CS-3, Alibaba Studio | Identifies specialized hosting platforms providing distinctive latency or data residency | Hosting disclosed | MEASURED_ACTIVE |
batch79-models-m3-r6Machine-readable dataset parity verification | models/data.json synchronized with HTML view | Zero divergence between client-facing table and programmatic JSON distribution | Data parity = 100% | VALIDATED_OBSERVED |
First-party provenance: All AI Ask first-party model & pricing registry; verification date 2026-09-08. Missing or conflicting joins fail closed.
| Model | Provider | Context | Max output | Modalities | Reasoning | $/M blended | tok/s | Status |
|---|---|---|---|---|---|---|---|---|
| Gemini 3.1 Pro | 2M | 64K | text, vision, audio | Yes | $4.50 | 55 | current | |
| Muse Spark 1.3 Contributor | Meta | 1.0M | 128K | text | Yes | $0.13 | — | current |
| Gemini 3.7 Flash | 1.0M | 66K | text, vision, audio | Yes | $1.50 | — | current | |
| Muse Spark 1.3 | Meta | 1.0M | 128K | text | Yes | $2.00 | — | current |
| Gemini 2.5 Flash Lite | 1M | 8K | text, vision | No | $0.18 | — | legacy | |
| DeepSeek V4 Flash | DeepSeek | 1M | 384K | text | No | $0.66 | 132 | current |
| Gemini 3.5 Flash Lite | 1M | 64K | text, vision | Yes | $0.85 | 162 | current | |
| Gemini 2.5 Flash | 1M | 8K | text, vision, audio | No | $0.85 | — | legacy | |
| Grok 4.3 | xAI | 1M | 64K | text, vision | Yes | $1.56 | 98 | current |
| DeepSeek V4 Pro | DeepSeek | 1M | 384K | text | Yes | $1.98 | 68 | current |
| GLM-5.2 | Z.ai | 1M | 64K | text | Yes | $2.15 | — | current |
| GPT-5.6 Luna | OpenAI | 1M | 64K | text, vision | No | $2.25 | 126 | current |
| Grok-3 | xAI | 1M | 8K | text, vision | No | $2.50 | — | legacy |
| Grok-4.20 Reasoning | xAI | 1M | 64K | text, vision | Yes | $3.00 | 52 | current |
| Grok-4.20 | xAI | 1M | 32K | text, vision | No | $3.00 | 104 | current |
| Gemini 3.6 Flash | 1M | 64K | text, vision, audio | Yes | $3.00 | 114 | current | |
| GPT-5.6 Terra | OpenAI | 1M | 128K | text, vision | Yes | $5.63 | 78 | current |
| GPT-5.6 Sol | OpenAI | 1M | 128K | text, vision | Yes | $8.00 | 44 | current |
| Claude Fable 5 | Anthropic | 1M | 128K | text, vision | Yes | $20.00 | 41 | current |
| Claude Opus 5 | Anthropic | 1M | 128K | text, vision | Yes | $30.00 | — | current |
| Grok 4.6 | xAI | 500K | 64K | text, vision | Yes | $3.00 | — | current |
| Grok 4.5 | xAI | 500K | 64K | text, vision | Yes | $3.00 | — | current |
| Claude Sonnet 5 | Anthropic | 500K | 64K | text, vision | Yes | $4.00 | — | current |
| Claude Opus 4.8 | Anthropic | 500K | 64K | text, vision | Yes | $10.00 | 58 | current |
| Amazon Nova Lite | Amazon | 300K | 16K | text, vision, audio | No | $0.11 | 108 | current |
| Amazon Nova Pro | Amazon | 300K | 33K | text, vision, audio | No | $1.40 | 64 | current |
| Claude Sonnet 4.6 | Anthropic | 300K | 64K | text, vision | Yes | $6.00 | 76 | current |
| Ministral 8B | Mistral | 256K | 33K | text, vision | No | $0.15 | 158 | current |
| Mistral Small 3.1 | Mistral | 256K | 33K | text, vision | Yes | $0.26 | 121 | current |
| Codestral | Mistral | 256K | 33K | text | No | $0.45 | 118 | current |
| Mistral Large 3 | Mistral | 256K | 33K | text, vision | No | $0.75 | 61 | current |
| Qwen 3.7 Plus | Qwen | 256K | 33K | text, vision | No | $1.10 | 84 | current |
| Qwen 3.8 Max | Qwen | 256K | 33K | text, vision | Yes | $2.80 | 47 | current |
| Qwen 3.7 Max | Qwen | 256K | 33K | text, vision | Yes | $2.80 | 49 | current |
| Mistral Medium 3 | Mistral | 256K | 33K | text, vision | No | $3.00 | 92 | current |
| Claude Haiku 4.5 | Anthropic | 200K | 32K | text, vision | No | $2.00 | 148 | current |
| GLM 4.7 (Cerebras) | Cerebras | 200K | 33K | text | Yes | $2.38 | 1980 | current |
| Claude Sonnet 4.5 | Anthropic | 200K | 32K | text, vision | Yes | $6.00 | — | legacy |
| Claude Sonnet 4 | Anthropic | 200K | 32K | text, vision | No | $6.00 | — | legacy |
| Claude Opus 4 | Anthropic | 200K | 4K | text, vision | No | $30.00 | — | legacy |
| GPT-OSS 20B | Groq | 131K | 33K | text | Yes | $0.13 | 1120 | current |
| GPT-OSS 120B | Groq | 131K | 33K | text | Yes | $0.26 | 780 | current |
| GPT-OSS 120B (Cerebras) | Cerebras | 131K | 33K | text | Yes | $0.45 | 2450 | current |
| Qwen 3.8 30B | Groq | 131K | 33K | text, vision | Yes | $1.20 | 690 | current |
| Amazon Nova Micro | Amazon | 128K | 8K | text | No | $0.06 | 168 | current |
| GPT-4o Mini | OpenAI | 128K | 16K | text, vision | No | $0.26 | — | legacy |
| GLM-5.1 | Z.ai | 128K | 8K | text | No | $1.00 | — | legacy |
| GPT-4o | OpenAI | 128K | 16K | text, vision | No | $4.38 | — | legacy |
| GPT-4 Turbo | OpenAI | 128K | 4K | text, vision | No | $15.00 | — | legacy |
| GPT-5 Nano | OpenAI | — | — | — | — | $0.14 | — | legacy |
| Grok-3 Mini | xAI | — | — | — | — | $0.26 | — | legacy |
| Llama 4 Maverick | Groq | — | — | — | — | $0.30 | — | legacy |
| GPT-5.4 Nano | OpenAI | — | — | — | — | $0.46 | — | legacy |
| Gemini 3.1 Flash Lite | — | — | — | — | $0.56 | — | legacy | |
| GPT-5 Mini | OpenAI | — | — | — | — | $0.69 | — | legacy |
| Qwen 3.6 27B | Groq | — | — | — | — | $1.20 | — | legacy |
| GPT-5.4 Mini | OpenAI | — | — | — | — | $1.69 | — | legacy |
| Gemini 3.1 Flash | — | — | — | — | $1.69 | — | legacy | |
| o3-Mini | OpenAI | — | — | — | — | $1.93 | — | legacy |
| Gemini 3.5 Flash | — | — | — | — | $3.38 | — | legacy | |
| GPT-5 | OpenAI | — | — | — | — | $3.44 | — | legacy |
| GPT-4.1 | OpenAI | — | — | — | — | $3.50 | — | legacy |
| GPT-5.4 | OpenAI | — | — | — | — | $5.63 | — | legacy |
| Claude Opus 4.7 | Anthropic | — | — | — | — | $10.00 | — | legacy |
| Claude Opus 4.6 | Anthropic | — | — | — | — | $10.00 | — | legacy |
| Claude Opus 4.5 | Anthropic | — | — | — | — | $10.00 | — | legacy |
| Claude Opus 4.1 | Anthropic | — | — | — | — | $30.00 | — | legacy |
| GPT-5.4 Pro | OpenAI | — | — | — | — | $67.50 | — | legacy |
39 of 68 models have full specs (legacy catalog entries deliberately have none); 31 have measured speed. Data verified 2026-08-14.
Which LLMs have the largest context windows?
Gemini 3.1 Pro has the largest documented context window among current models in the All AI Ask roster, at 2M tokens. Gemini 3.7 Flash follows at 1.0M tokens. This ranking covers live models only, uses each provider’s published specification, and is useful when a workload must fit a long document, codebase, or conversation in one request.
Largest context window LLMs
Current models ranked by the maximum context window documented in their spec sheet. Canonical /models view.
| Rank | Model | Provider | Value |
|---|---|---|---|
| 1 | Gemini 3.1 Pro | 2M tokens | |
| 2 | Gemini 3.7 Flash | 1.0M tokens | |
| 3 | Muse Spark 1.3 | Meta | 1.0M tokens |
| 4 | Muse Spark 1.3 Contributor | Meta | 1.0M tokens |
| 5 | Claude Fable 5 | Anthropic | 1M tokens |
| 6 | Claude Opus 5 | Anthropic | 1M tokens |
| 7 | DeepSeek V4 Flash | DeepSeek | 1M tokens |
| 8 | DeepSeek V4 Pro | DeepSeek | 1M tokens |
| 9 | Gemini 3.5 Flash Lite | 1M tokens | |
| 10 | Gemini 3.6 Flash | 1M tokens | |
| 11 | GLM-5.2 | Z.ai | 1M tokens |
| 12 | GPT-5.6 Luna | OpenAI | 1M tokens |
| 13 | GPT-5.6 Sol | OpenAI | 1M tokens |
| 14 | GPT-5.6 Terra | OpenAI | 1M tokens |
| 15 | Grok 4.3 | xAI | 1M tokens |
| 16 | Grok-4.20 | xAI | 1M tokens |
| 17 | Grok-4.20 Reasoning | xAI | 1M tokens |
| 18 | Claude Opus 4.8 | Anthropic | 500K tokens |
| 19 | Claude Sonnet 5 | Anthropic | 500K tokens |
| 20 | Grok 4.5 | xAI | 500K tokens |
| 21 | Grok 4.6 | xAI | 500K tokens |
| 22 | Amazon Nova Lite | Amazon | 300K tokens |
| 23 | Amazon Nova Pro | Amazon | 300K tokens |
| 24 | Claude Sonnet 4.6 | Anthropic | 300K tokens |
| 25 | Codestral | Mistral | 256K tokens |
| 26 | Ministral 8B | Mistral | 256K tokens |
| 27 | Mistral Large 3 | Mistral | 256K tokens |
| 28 | Mistral Medium 3 | Mistral | 256K tokens |
| 29 | Mistral Small 3.1 | Mistral | 256K tokens |
| 30 | Qwen 3.7 Max | Qwen | 256K tokens |
| 31 | Qwen 3.7 Plus | Qwen | 256K tokens |
| 32 | Qwen 3.8 Max | Qwen | 256K tokens |
| 33 | Claude Haiku 4.5 | Anthropic | 200K tokens |
| 34 | GLM 4.7 (Cerebras) | Cerebras | 200K tokens |
| 35 | GPT-OSS 120B | Groq | 131K tokens |
| 36 | GPT-OSS 120B (Cerebras) | Cerebras | 131K tokens |
| 37 | GPT-OSS 20B | Groq | 131K tokens |
| 38 | Qwen 3.8 30B | Groq | 131K tokens |
| 39 | Amazon Nova Micro | Amazon | 128K tokens |
Which LLMs generate the most output tokens?
DeepSeek V4 Flash has the largest documented maximum output among current models in the All AI Ask roster, at 384K tokens. DeepSeek V4 Pro follows at 384K tokens. Maximum output is a generation limit, not a promise that every response will use that many tokens; compare it separately from context capacity, price, latency, and task quality.
LLMs with the largest max output
Current models ranked by their documented maximum output-token allowance. Canonical /models view.
| Rank | Model | Provider | Value |
|---|---|---|---|
| 1 | DeepSeek V4 Flash | DeepSeek | 384K tokens |
| 2 | DeepSeek V4 Pro | DeepSeek | 384K tokens |
| 3 | Claude Fable 5 | Anthropic | 128K tokens |
| 4 | Claude Opus 5 | Anthropic | 128K tokens |
| 5 | GPT-5.6 Sol | OpenAI | 128K tokens |
| 6 | GPT-5.6 Terra | OpenAI | 128K tokens |
| 7 | Muse Spark 1.3 | Meta | 128K tokens |
| 8 | Muse Spark 1.3 Contributor | Meta | 128K tokens |
| 9 | Gemini 3.7 Flash | 66K tokens | |
| 10 | Claude Opus 4.8 | Anthropic | 64K tokens |
| 11 | Claude Sonnet 4.6 | Anthropic | 64K tokens |
| 12 | Claude Sonnet 5 | Anthropic | 64K tokens |
| 13 | Gemini 3.1 Pro | 64K tokens | |
| 14 | Gemini 3.5 Flash Lite | 64K tokens | |
| 15 | Gemini 3.6 Flash | 64K tokens | |
| 16 | GLM-5.2 | Z.ai | 64K tokens |
| 17 | GPT-5.6 Luna | OpenAI | 64K tokens |
| 18 | Grok 4.3 | xAI | 64K tokens |
| 19 | Grok 4.5 | xAI | 64K tokens |
| 20 | Grok 4.6 | xAI | 64K tokens |
| 21 | Grok-4.20 Reasoning | xAI | 64K tokens |
| 22 | Amazon Nova Pro | Amazon | 33K tokens |
| 23 | Codestral | Mistral | 33K tokens |
| 24 | GLM 4.7 (Cerebras) | Cerebras | 33K tokens |
| 25 | GPT-OSS 120B | Groq | 33K tokens |
| 26 | GPT-OSS 120B (Cerebras) | Cerebras | 33K tokens |
| 27 | GPT-OSS 20B | Groq | 33K tokens |
| 28 | Ministral 8B | Mistral | 33K tokens |
| 29 | Mistral Large 3 | Mistral | 33K tokens |
| 30 | Mistral Medium 3 | Mistral | 33K tokens |
| 31 | Mistral Small 3.1 | Mistral | 33K tokens |
| 32 | Qwen 3.7 Max | Qwen | 33K tokens |
| 33 | Qwen 3.7 Plus | Qwen | 33K tokens |
| 34 | Qwen 3.8 30B | Groq | 33K tokens |
| 35 | Qwen 3.8 Max | Qwen | 33K tokens |
| 36 | Claude Haiku 4.5 | Anthropic | 32K tokens |
| 37 | Grok-4.20 | xAI | 32K tokens |
| 38 | Amazon Nova Lite | Amazon | 16K tokens |
| 39 | Amazon Nova Micro | Amazon | 8K tokens |
Which LLMs have the newest knowledge cutoff?
Claude Opus 5 has the newest disclosed knowledge cutoff among current models in the All AI Ask roster, listed as 2026-05. Claude Sonnet 5 follows at 2026-05. Models without a published cutoff are omitted rather than treated as current. A newer cutoff can reduce stale answers, but retrieval, source quality, and the prompt still determine whether a response is up to date.
LLMs with the newest knowledge cutoff
Current models with a disclosed cutoff, ranked from newest to oldest; undisclosed cutoffs are omitted. Canonical /models view.
| Rank | Model | Provider | Value |
|---|---|---|---|
| 1 | Claude Opus 5 | Anthropic | 2026-05 |
| 2 | Claude Sonnet 5 | Anthropic | 2026-05 |
| 3 | Gemini 3.6 Flash | 2026-04 | |
| 4 | Grok 4.3 | xAI | 2026-04 |
| 5 | GPT-5.6 Luna | OpenAI | 2026-03 |
| 6 | GPT-5.6 Sol | OpenAI | 2026-03 |
| 7 | GPT-5.6 Terra | OpenAI | 2026-03 |
| 8 | Qwen 3.8 Max | Qwen | 2026-03 |
| 9 | Claude Fable 5 | Anthropic | 2026-02 |
| 10 | DeepSeek V4 Flash | DeepSeek | 2026-02 |
| 11 | DeepSeek V4 Pro | DeepSeek | 2026-02 |
| 12 | Gemini 3.5 Flash Lite | 2026-02 | |
| 13 | GLM-5.2 | Z.ai | 2026-02 |
| 14 | Grok 4.5 | xAI | 2026-02 |
| 15 | Qwen 3.8 30B | Groq | 2026-02 |
| 16 | Claude Opus 4.8 | Anthropic | 2026-01 |
| 17 | Grok-4.20 | xAI | 2026-01 |
| 18 | Grok-4.20 Reasoning | xAI | 2026-01 |
| 19 | Qwen 3.7 Max | Qwen | 2026-01 |
| 20 | Qwen 3.7 Plus | Qwen | 2026-01 |
| 21 | Claude Sonnet 4.6 | Anthropic | 2025-12 |
| 22 | Mistral Medium 3 | Mistral | 2025-12 |
| 23 | Mistral Small 3.1 | Mistral | 2025-12 |
| 24 | Gemini 3.1 Pro | 2025-11 | |
| 25 | GLM 4.7 (Cerebras) | Cerebras | 2025-10 |
| 26 | Mistral Large 3 | Mistral | 2025-10 |
| 27 | Claude Haiku 4.5 | Anthropic | 2025-08 |
| 28 | Ministral 8B | Mistral | 2025-07 |
| 29 | Codestral | Mistral | 2025-06 |
| 30 | GPT-OSS 120B | Groq | 2025-05 |
| 31 | GPT-OSS 120B (Cerebras) | Cerebras | 2025-05 |
| 32 | GPT-OSS 20B | Groq | 2025-05 |
| 33 | Amazon Nova Lite | Amazon | 2024-10 |
| 34 | Amazon Nova Micro | Amazon | 2024-10 |
| 35 | Amazon Nova Pro | Amazon | 2024-10 |
Try any model for free
Every model in this table, one workspace, one API key.
Try It Free