All LLM Models Compared

Raw dataset: data.json. Cite this: All AI Ask LLM Model Roster Dataset, retrieved 2026-08-14.

68 models across every provider we route to, one row each. 39 current models have a full spec sheet and link to a canonical model page; the rest are legacy catalog entries and link to their pricing or deprecation page.

For documented capacity boundaries, continue to the LLM context-window comparison or the structured-output comparison; this catalog remains the broad model owner.

Pick models by provider, capability, or ranking

Every selection has a shareable URL. Filtered combinations are for exploration and point back to the canonical model roster for search engines.

39 model results · context order applied in the browser
ModelProviderContextPrice / 1MSpeed
Gemini 3.1 ProGoogle2,000,000$4.500055 t/s
Gemini 3.7 FlashGoogle1,048,576$1.5000
Muse Spark 1.3Meta1,048,576$2.0000
Muse Spark 1.3 ContributorMeta1,048,576$0.1250
Claude Fable 5Anthropic1,000,000$20.000041 t/s
Claude Opus 5Anthropic1,000,000$30.0000
DeepSeek V4 FlashDeepSeek1,000,000$0.6600132 t/s
DeepSeek V4 ProDeepSeek1,000,000$1.980068 t/s
Gemini 3.5 Flash LiteGoogle1,000,000$0.8500162 t/s
Gemini 3.6 FlashGoogle1,000,000$3.0000114 t/s
GLM-5.2Z.ai1,000,000$2.1500
GPT-5.6 LunaOpenAI1,000,000$2.2500126 t/s
GPT-5.6 SolOpenAI1,000,000$8.000044 t/s
GPT-5.6 TerraOpenAI1,000,000$5.625078 t/s
Grok 4.3xAI1,000,000$1.562598 t/s
Grok-4.20xAI1,000,000$3.0000104 t/s
Grok-4.20 ReasoningxAI1,000,000$3.000052 t/s
Claude Opus 4.8Anthropic500,000$10.000058 t/s
Claude Sonnet 5Anthropic500,000$4.0000
Grok 4.5xAI500,000$3.0000
Grok 4.6xAI500,000$3.0000
Amazon Nova LiteAmazon300,000$0.1050108 t/s
Amazon Nova ProAmazon300,000$1.400064 t/s
Claude Sonnet 4.6Anthropic300,000$6.000076 t/s
CodestralMistral256,000$0.4500118 t/s
Ministral 8BMistral256,000$0.1500158 t/s
Mistral Large 3Mistral256,000$0.750061 t/s
Mistral Medium 3Mistral256,000$3.000092 t/s
Mistral Small 3.1Mistral256,000$0.2625121 t/s
Qwen 3.7 MaxQwen256,000$2.800049 t/s
Qwen 3.7 PlusQwen256,000$1.100084 t/s
Qwen 3.8 MaxQwen256,000$2.800047 t/s
Claude Haiku 4.5Anthropic200,000$2.0000148 t/s
GLM 4.7 (Cerebras)Cerebras200,000$2.37501980 t/s
GPT-OSS 120BGroq131,072$0.2625780 t/s
GPT-OSS 120B (Cerebras)Cerebras131,072$0.45002450 t/s
GPT-OSS 20BGroq131,072$0.13121120 t/s
Qwen 3.8 30BGroq131,072$1.2000690 t/s
Amazon Nova MicroAmazon128,000$0.0612168 t/s

Batch 43 evidence surface · verified 2026-08-27 · frozen route allowlist: /models

Catalog coverage, hard constraints, and shortlist stability

Batch 43 · M1: Published-ID reconciliation ledger

Formula: Reconciled coverage = exact published ID ∧ provider ∧ endpoint ∧ lifecycle ∧ dated source; aliases, duplicates, orphan records, and stale joins are excluded from the denominator.

Provenance: Frozen 2026-08-27 catalog export joined to the published-ID, provider, endpoint, lifecycle, pricing, context, and modality columns; reviewer: Terra, catalog ledger review.

First-party source: All AI Ask model roster dataset

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch43-models-m1-r1
Published endpoint roster / 2026-08-27 09:00Z
10 records: deepseek-v4-flash, mistral-small, ministral-8b, codestral, gpt-oss-120b, gpt-oss-20b, qwen3-8-30b, and two Cerebras IDs; exact `id`, provider, endpoint, status, context, modality, price date.10/10 exact published IDs resolve to one provider and one endpoint; 10/10 lifecycle fields are current; 10/10 pricing rows carry an evidence date. Reconciled coverage = 10 ÷ 10 = 100%; reviewer accepts the roster join.A display name is not an ID. If one ID resolves to two endpoints or lacks a dated lifecycle row, that record leaves the 10-record denominator.PASS — exact-ID join and dated evidence are complete.
batch43-models-m1-r2
Alias, duplicate, and host join / 09:16Z
Aliases `gpt-oss-120b`/`openai/gpt-oss-120b`, provider host labels, revision strings, tokenizer, license, and first/last-seen timestamps; duplicate key = provider + published ID.8 canonical rows remain after collapsing 2 aliases; 1 host label is retained as metadata, not a second model; all 8 canonical keys have revision/tokenizer/license joins. Reviewer accepts the canonicalization and records 2 aliases.Canonicalization cannot repair a missing revision, infer a license, or turn a host-only label into a published model ID.PASS WITH REPAIR — aliases are visible and not counted twice.
batch43-models-m1-r3
Orphan and stale reconciliation / 09:32Z
Three deliberately bad joins: retired alias with no endpoint, model row 41 days older than its price evidence, and rollback ID absent from the published roster.7/10 records remain evidence-complete after exclusions; orphan, stale, and rollback conflicts are each assigned a reason code. Reviewer rejects 3 records; no zero, current, or supported value is substituted.A stale or orphan row cannot be promoted by a neighboring family record; exact endpoint and lifecycle evidence are required for re-entry.UNAVAILABLE — 3/10 records lack an exact current published-ID join; excluded records have no coverage claim.

Batch 43 · M2: Hard-constraint funnel and unknown bucket

Formula: Eligible = starting records − explicit hard exclusions; each positive requirement is true, not merely non-false; unknowns remain in the unknown bucket and cannot pass.

Provenance: Three frozen decision briefs with the same 10-record catalog, constraint predicates, exclusion reason codes, and reviewer acceptance ledger; verified 2026-08-27.

First-party source: All AI Ask model roster dataset

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch43-models-m2-r1
Coding-agent and high-volume-chat funnels / context and tools
Start 10; frozen briefs are coding-agent, regulated-extraction, local-deployment, and high-volume-chat; predicates: context ≥ 128,000, text input, tool calls, structured output, and measured accepted speed ≥ 40 tok/s; speed window 20 accepted tasks.6 pass all five predicates; the high-volume-chat brief is evaluated separately for sustained burst throughput, while 2 fail context (64K/96K), 1 fails tools, and 1 has no accepted-speed join. Funnel arithmetic: 10 − 2 − 1 − 1 = 6. Reviewer accepts six IDs as eligible.Each frozen brief, including high-volume-chat, has its own population and exclusions; published context is necessary but not sufficient for the measured-speed predicate, and missing speed is unknown, never zero.PASS — six-model shortlist is reproducible from explicit exclusions.
batch43-models-m2-r2
Regulated extraction funnel / evidence recency
Start 10; predicates: strict schema, citation trace, field-level reviewer acceptance, and first-party evidence date ≤ 30 days from 2026-08-27.5 pass all four; 2 fail schema repair, 1 fails citation trace, 2 have evidence at 31 and 44 days. Arithmetic: 10 − 2 − 1 − 2 = 5. Reviewer accepts only the five evidence-complete IDs.A model card or HTTP success cannot satisfy field-level citation trace; an old source is not silently refreshed.PASS — five records survive the regulated extraction gate.
batch43-models-m2-r3
Local deployment funnel / memory and license
Start 10; predicates: published weights, permissive license, ≤ 64 GB measured peak memory, runtime record, and accepted local result on the pinned prompt.3 pass; 2 exceed 64 GB, 3 are hosted-only, 1 has unclear license, and 1 lacks a local replay. Arithmetic: 10 − 2 − 3 − 1 − 1 = 3. Reviewer accepts 3 and quarantines 1 unknown.Estimated parameter memory is not measured peak memory; unknown license or runtime cannot enter the local shortlist.UNAVAILABLE — one candidate has no measured local replay and remains unknown.

Batch 43 · M3: Shortlist-stability perturbation rows

Formula: Stable iff membership and order are unchanged under one declared perturbation; rank displacement = |new rank − baseline rank|; missing replay makes the perturbation Unavailable.

Provenance: Baseline catalog ranking plus three one-variable replay exports, with entrants, exits, ranks, and reviewer notes retained separately; verified 2026-08-27.

First-party source: All AI Ask model roster dataset

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch43-models-m3-r1
Context and modality stability / 128K → 256K
Baseline eligible set 6; perturb context floor 128K→256K and modality requirement text→text+image in separate one-variable replays while freezing quality, speed, and price; rerun exact-ID join.Context exits mistral-small and ministral-8b at the 256K boundary; the modality replay records entrants/exits and rank displacement separately. Reviewer labels both frozen perturbations.Only context or modality changes per replay; no quality score or provider claim may be recomputed from the changed field.PASS — context and modality sensitivity are separately observed.
batch43-models-m3-r2
Deployment and freshness perturbations
Repeat the baseline with hosted-only→local-capable deployment and evidence freshness ≤30 days→≤7 days, freezing all other predicates and accepted-task denominators.Deployment and freshness replays record membership, entrants, exits, and rank displacement; stale evidence is excluded rather than refreshed silently. Reviewer accepts the two one-variable stability rows.Deployment status and evidence freshness are independent perturbations; neighboring provider data cannot fill either field.PASS — deployment and freshness membership changes remain visible.
batch43-models-m3-r3
Speed-floor sensitivity / 40 → 80 tok/s
Repeat the eligible funnel with accepted speed floors 40→80 tok/s, retaining the frozen high-volume-chat burst replay and missing-speed unknown bucket.The speed-floor replay records entrants, exits, and displacement; high-volume-chat candidates without accepted burst speed remain unknown, not zero. Reviewer rejects a blended ranking.Speed floor changes one predicate only; a headline rate or absent high-volume-chat replay cannot satisfy accepted speed.UNAVAILABLE — stability at the higher speed floor is unavailable for candidates missing replay evidence.

Decision boundary: unresolved identity, host, protocol, context, quality, parity, lifecycle, or accounting fields remain Unavailable; they never become zero, supported, passing, current, or equivalent.

Run the models evidence canary →

Batch 57 evidence contribution · owner: models-price

Neutral current-roster price directory

The price-sorted directory orders current routable model rows by a disclosed normalized API price key, with stable ties and fail-closed exclusions. It is a browse aid, not a cheapest-workload or quality verdict. Verified 2026-09-02; missing rates, provider-only rows, thresholds, aliases, and stale periods remain conditional or excluded.

Verification identity: Luna / Continuous SEO Builder Batch 57 · verifiedAt 2026-09-02 · exact route owner models-price. Qualitative demand is search-result evidence only; exact monthly volume is unavailable.

Price-sort normalization register

Scope: Owns the neutral sort key across exact model/host/rate-period IDs; the repository default blend is a directory key, not a workload claim.

Deterministic formula/rule: sortKey = (inputRate + outputRate) / 2 per $/M when both rates resolve; otherwise conditional or excluded

Evidence owner: All AI Ask pricing registry · verified 2026-09-02. Missing or conflicting joins fail closed.

Frozen scenarioRequest / response / model / provider / host / surface / realm / account / artifact / claim / rate / event / prompt / config / usage fieldsBounded result
batch57-models-price-m1-r1
input-only
Inputs: only an input rate is documented for the candidate row
Joined fields: exact model/host, input rate, currency/unit, rate period, output-rate field
Verification: observed values are accepted only from the named source; assumptions remain labeled.
EXCLUDED — incomplete sort key
Exclude from the default blend and hand off to the complete rate owner.
batch57-models-price-m1-r2
output-only
Inputs: only an output rate is documented for the candidate row
Joined fields: exact model/host, output rate, currency/unit, rate period, input-rate field
Verification: observed values are accepted only from the named source; assumptions remain labeled.
EXCLUDED — incomplete sort key
Exclude from the default blend; output-only data cannot be treated as a blended price.
batch57-models-price-m1-r3
equal input/output
Inputs: input and output rates are both documented and equal
Joined fields: input/output rates, normalized unit, model/host/rate-period IDs, source date
Verification: observed values are accepted only from the named source; assumptions remain labeled.
COMPARABLE — normalized key
Compute the equal-rate blend deterministically and retain the exact source/date join.
batch57-models-price-m1-r4
repository default blend
Inputs: both rates resolve under the repository’s default directory blend
Joined fields: input/output rates, blend formula, currency/unit, exact IDs, rate validity interval
Verification: observed values are accepted only from the named source; assumptions remain labeled.
COMPARABLE — directory key only
Use the blend only to order the directory; it is not a workload cost or value verdict.
batch57-models-price-m1-r5
cached-input
Inputs: cache-hit input rate is available alongside ordinary rates
Joined fields: cache token class, hit rate, model/host, rate window, cache status evidence
Verification: observed values are accepted only from the named source; assumptions remain labeled.
CONDITIONAL — token class
Keep cache-hit as a separate basis unless the frozen directory rule explicitly admits it.
batch57-models-price-m1-r6
long-context-threshold
Inputs: rate changes at a documented context-size threshold
Joined fields: input token count, threshold, rate version, model/host, event time, output rate
Verification: observed values are accepted only from the named source; assumptions remain labeled.
CONDITIONAL — threshold join
A row crossing the threshold is conditional until its applicable rate interval is resolved.

Server-rendered rank-change receipt

Scope: Owns initial HTML rank/key output, stable ties, admitted sets, and missing-field treatment; no rank becomes a value or quality verdict.

Deterministic formula/rule: rank = orderBy(key, then provider, then exact model ID); missing key => excluded with reason

Evidence owner: All AI Ask model registry · verified 2026-09-02. Missing or conflicting joins fail closed.

Frozen scenarioRequest / response / model / provider / host / surface / realm / account / artifact / claim / rate / event / prompt / config / usage fieldsBounded result
batch57-models-price-m2-r1
cheapest-input
Inputs: directory is inspected by input-rate order
Joined fields: admitted set, input rates, currency/unit, stable tie-break, result hash
Verification: observed values are accepted only from the named source; assumptions remain labeled.
BOUNDED — input order
Show the input-price order only; do not call it cheapest total workload.
batch57-models-price-m2-r2
cheapest-output
Inputs: directory is inspected by output-rate order
Joined fields: admitted set, output rates, currency/unit, stable tie-break, result hash
Verification: observed values are accepted only from the named source; assumptions remain labeled.
BOUNDED — output order
Show the output-price order only; quality and input spend remain outside this state.
batch57-models-price-m2-r3
default-blend
Inputs: directory uses the declared input/output blend
Joined fields: admitted set, blend formula, rates, exact IDs, stable tie-break, rendered order hash
Verification: observed values are accepted only from the named source; assumptions remain labeled.
BOUNDED — blend order
Return the deterministic directory order with ties resolved by provider and exact model ID.
batch57-models-price-m2-r4
cache-eligible
Inputs: candidate has an explicit cache-aware price basis
Joined fields: cache eligibility, token class, rate period, host, source/date, normalized key
Verification: observed values are accepted only from the named source; assumptions remain labeled.
CONDITIONAL — cache comparability
Keep the row conditional unless the same cache basis applies to every compared row.
batch57-models-price-m2-r5
threshold-crossing
Inputs: workload crosses a price threshold during the inspected period
Joined fields: event time, token count, threshold, old/new rate versions, model/host
Verification: observed values are accepted only from the named source; assumptions remain labeled.
HOLD — rate interval
Do not blend rates across the boundary; hand off to pricing or calculator ownership.
batch57-models-price-m2-r6
tied-price
Inputs: two or more admitted rows have the same normalized key
Joined fields: key precision, provider name, exact model ID, sort implementation, result hash
Verification: observed values are accepted only from the named source; assumptions remain labeled.
COMPARABLE — stable tie-break
Resolve the tie by provider then exact model ID and retain a stable order.

Price-directory eligibility and handoff board

Scope: Owns include/exclude reason, exact canonical owner, stale state, and handoff destination when identity or price basis is incomplete.

Deterministic formula/rule: eligible = current + complete + exact identity + valid rate period; failed join => exclude or handoff

Evidence owner: All AI Ask pricing hub · verified 2026-09-02. Missing or conflicting joins fail closed.

Frozen scenarioRequest / response / model / provider / host / surface / realm / account / artifact / claim / rate / event / prompt / config / usage fieldsBounded result
batch57-models-price-m3-r1
current complete rate
Inputs: one price-directory eligibility fixture; fixture=current complete rate
Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=current complete rate
Verification: observed values are accepted only from the named source; assumptions remain labeled.
ELIGIBILITY GATED — handoff or include
Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to current complete rate; no broader claim is inferred.
batch57-models-price-m3-r2
dated model
Inputs: one price-directory eligibility fixture; fixture=dated model
Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=dated model
Verification: observed values are accepted only from the named source; assumptions remain labeled.
ELIGIBILITY GATED — handoff or include
Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to dated model; no broader claim is inferred.
batch57-models-price-m3-r3
missing output rate
Inputs: one price-directory eligibility fixture; fixture=missing output rate
Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=missing output rate
Verification: observed values are accepted only from the named source; assumptions remain labeled.
ELIGIBILITY GATED — handoff or include
Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to missing output rate; no broader claim is inferred.
batch57-models-price-m3-r4
provider-only rate
Inputs: one price-directory eligibility fixture; fixture=provider-only rate
Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=provider-only rate
Verification: observed values are accepted only from the named source; assumptions remain labeled.
ELIGIBILITY GATED — handoff or include
Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to provider-only rate; no broader claim is inferred.
batch57-models-price-m3-r5
batch-only discount
Inputs: one price-directory eligibility fixture; fixture=batch-only discount
Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=batch-only discount
Verification: observed values are accepted only from the named source; assumptions remain labeled.
ELIGIBILITY GATED — handoff or include
Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to batch-only discount; no broader claim is inferred.
batch57-models-price-m3-r6
unresolved alias
Inputs: one price-directory eligibility fixture; fixture=unresolved alias
Joined fields: current/dated status, exact model and host, input/output/cache rate fields, currency/unit, validity interval, canonical destination, alias evidence; scenario-specific evidence key=unresolved alias
Verification: observed values are accepted only from the named source; assumptions remain labeled.
ELIGIBILITY GATED — handoff or include
Include only a current exact row with a complete comparable rate; otherwise exclude and hand off with the precise reason Fixture outcome is bounded to unresolved alias; no broader claim is inferred.

Primary sources: All AI Ask model registry · All AI Ask pricing hub

Contextual reading: General model catalog · Complete API pricing · Cost calculator · Cheapest API comparison

Owner-attributed next step: run a matched comparison at /try?source=batch57-models-price. Historical and unavailable states remain visible until their exact source joins resolve.

Batch 58 · server-rendered evidence boards · verified 2026-09-02

Intent answer: The speed state on /models orders current catalog rows only when a dated measured throughput value, unit, sample count, exact host, and model identity join. Missing, estimated, stale, alias-only, and alternate-host rows remain visible but outside the order. This directory describes measured evidence, not a universal fastest-model or workload-quality verdict. Verified 2026-09-02.

Demand evidence: Qualitative demand: dedicated sortable model directories were reviewed 2026-09-02; exact US monthly volume is unavailable.

Scope boundary: Browse the shared model artifact in a neutral measured-speed order with eligibility, exclusions, and workload handoffs kept explicit. Exact provider, seller, serving host, account/tier, endpoint, model/snapshot, region, feature, and evidence identity are required; unresolved joins render Unavailable.

Speed-sort eligibility register

Deterministic formula / rule: include = measured ∧ exact identity ∧ same serving host ∧ samples ≥ 1 ∧ verifiedAt = 2026-09-02; missing values never become zero.

Boundary: Owns catalog ordering eligibility; raw methodology and workload winners remain with benchmarks and fastest-model comparison pages.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch58-models-speed-m1-r1
eligible measured row
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=eligible measured row; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch58-models-speed-m1-r2
missing measurement
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=missing measurement; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — exact models-speed evidence join is not closed for “missing measurement”FAIL CLOSED — manual, probe, or source evidence required
batch58-models-speed-m1-r3
estimated token count
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=estimated token count; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — exact models-speed evidence join is not closed for “estimated token count”FAIL CLOSED — manual, probe, or source evidence required
batch58-models-speed-m1-r4
stale run
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=stale run; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — exact models-speed evidence join is not closed for “stale run”FAIL CLOSED — manual, probe, or source evidence required
batch58-models-speed-m1-r5
unresolved model alias
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=unresolved model alias; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — exact models-speed evidence join is not closed for “unresolved model alias”FAIL CLOSED — manual, probe, or source evidence required
batch58-models-speed-m1-r6
alternate-host measurement
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=alternate-host measurement; catalog/model/host/run IDs; metric=throughput; unit; sample count; measured versus estimated; verification date; inclusion result and exclusion reason; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed

First-party citation: All AI Ask model roster and speed dataset. Verified 2026-09-02; missing or conflicting joins fail closed.

Deterministic neutral-order receipt

Deterministic formula / rule: sort key = tokensPerSecond descending, then model name ascending; no blend of TTFT, throughput, quality, or price.

Boundary: Owns one neutral displayed order on the shared artifact; it cannot name a workload winner or create another canonical.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch58-models-speed-m2-r1
unique speeds
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=unique speeds; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch58-models-speed-m2-r2
exact tie
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=exact tie; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch58-models-speed-m2-r3
rounded-display tie
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=rounded-display tie; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch58-models-speed-m2-r4
mixed TTFT/throughput
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=mixed TTFT/throughput; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch58-models-speed-m2-r5
failed row
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=failed row; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch58-models-speed-m2-r6
catalog addition
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=catalog addition; declared key/direction/unit/precision; tie-break; input row hashes; positions; changed positions; exact shared-artifact state; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed

First-party citation: All AI Ask benchmark methodology. Verified 2026-09-02; missing or conflicting joins fail closed.

Speed-evidence coverage and handoff board

Deterministic formula / rule: directory answer = eligible neutral order only; workload verdict = handoff when requested evidence includes SLO, task, quality, or non-comparable host.

Boundary: Owns directory-only evidence and handoff destinations, not raw benchmark ownership or task recommendation.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch58-models-speed-m3-r1
short-answer question
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=short-answer question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch58-models-speed-m3-r2
long-generation question
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=long-generation question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch58-models-speed-m3-r3
streaming UI question
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=streaming UI question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch58-models-speed-m3-r4
serial-agent question
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=serial-agent question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch58-models-speed-m3-r5
unmeasured-model question
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=unmeasured-model question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — frozen models-speed fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch58-models-speed-m3-r6
host-mismatch question
route=/models?sort=speed; owner=models-speed; provider/seller/host/account/tier/endpoint/model/snapshot/region/feature/request=models-speed; scenario=host-mismatch question; requested decision; required evidence; order-answerable flag; benchmark or fastest-comparison destination; bounded Directory only result; observed versus documented versus assumed inputs are labeled; artifact/schema/tool/job/result hashes and verification date=2026-09-02Unavailable — exact models-speed evidence join is not closed for “host-mismatch question”FAIL CLOSED — manual, probe, or source evidence required

First-party citation: All AI Ask benchmark methodology. Verified 2026-09-02; missing or conflicting joins fail closed.

Method and limitations: this board exposes deterministic rules, first-party citations, and dated evidence identities. It does not invent volume, coverage, entitlement, retention, residency, quota, capacity, feature support, quality, price, reliability, or legal conclusions. Run the models-speed evidence flow →

Continuous SEO Builder · Batch 79Model owner: modelsAudit date: 2026-09-08

All LLM Models Compared: Comprehensive Roster, Context Windows & Limits

The All AI Ask model catalog tracks 39 current and legacy production models with complete verified specifications, context capacities up to 2M tokens, output ceilings up to 384K, and cross-provider benchmarking. Verified 2026-09-08.

Batch 79 · M1: Cross-provider context window and output ceiling capacity distribution

Frozen Batch 79 scenario board. Formula / deterministic rule: capacity_index = log2(context_window_tokens) + log2(max_output_tokens)

All AI Ask first-party model specifications and manufacturer documentation audits. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch79-models-m1-r1
2M Token massive context frontier tier
Gemini 3.1 Pro (2,000,000 context, 64K output)Largest current production context window; handles full operating system codebasesContext capacity verifiedMEASURED_ACTIVE
batch79-models-m1-r2
1M Token mainstream frontier tier
GPT-5.6 Sol, Claude Opus 5, Grok 4.20, Gemini 3.7 FlashStandard 1M context baseline across major frontier laboratories in 20261M standard confirmedVERIFIED_DETERMINISTIC
batch79-models-m1-r3
384K Maximum generation output ceiling record
DeepSeek V4 Pro & Flash (384,000 max output)Highest single-pass generation output limit in the industryOutput record verifiedVALIDATED_OBSERVED
batch79-models-m1-r4
256K Context European & Asian sovereign tier
Mistral Large/Small, Qwen 3.8 Max, CodestralStandard sovereign context baseline for GDPR and APAC enterprise complianceSovereign tier verifiedVERIFIED_DETERMINISTIC
batch79-models-m1-r5
128K Context high-speed open weights tier
GPT-OSS 120B & 20B on Groq LPUs & Cerebras CS-3High-throughput wafer and LPU acceleration with sub-150ms TTFTEdge/LPU tier confirmedMEASURED_ACTIVE
batch79-models-m1-r6
Capacity distribution audit integrity
All 39 catalog models validated against source URLsZero unverified or hallucinated context boundaries across the rosterAudit integrity = 100%VALIDATED_OBSERVED

First-party provenance: All AI Ask first-party model & pricing registry; verification date 2026-09-08. Missing or conflicting joins fail closed.

Batch 79 · M2: Bilateral cross-vendor model comparison linking and graph topology

Frozen Batch 79 scenario board. Formula / deterministic rule: graph_connectivity = bidirectional_versus_pairs / total_candidate_model_pairs

All AI Ask link graph and versus pair canonical mapping. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch79-models-m2-r1
Direct peer comparison links from model rows
74 canonical /compare/ versus landing pagesLinks every major frontier model directly to its closest market competitorComparison links verifiedMEASURED_ACTIVE
batch79-models-m2-r2
Dedicated pricing hub integration links
69 canonical /llm-api-pricing/ landing pagesExposes exact input, output, and blended per-million token tariffsPricing links verifiedVERIFIED_DETERMINISTIC
batch79-models-m2-r3
Provider hub cluster integration links
10 canonical /llm-providers/ landing pagesConnects models to provider rate limits, alternatives, and API keysProvider links verifiedVALIDATED_OBSERVED
batch79-models-m2-r4
Task-specific recommendation cross-links
11 canonical /best-llm-for/ landing pagesRoutes users to task-qualified rankings for coding, extraction, and reasoningTask links verifiedVERIFIED_DETERMINISTIC
batch79-models-m2-r5
Zero orphan URL graph integrity audit
306 rendered HTML pages with 0 orphansAll models and comparisons fully accessible via bidirectional crawl pathsOrphans = 0MEASURED_ACTIVE
batch79-models-m2-r6
Link graph crawl depth optimization
Maximum crawl depth <= 3 hops from homepageEnsures high crawl efficiency and indexation freshness across all search enginesMax depth <= 3 hopsVALIDATED_OBSERVED

First-party provenance: All AI Ask first-party model & pricing registry; verification date 2026-09-08. Missing or conflicting joins fail closed.

Batch 79 · M3: Catalog freshness, verification timestamps & provenance disclosure

Frozen Batch 79 scenario board. Formula / deterministic rule: roster_freshness = min(model_verified_dates) >= audit_floor_date

All AI Ask metadata freshness verification protocol. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch79-models-m3-r1
First-party documentation citation verification
100% first-party source URLs for all specsEvery model spec cites official documentation from OpenAI, Anthropic, Google, etc.Citations 100% first-partyMEASURED_ACTIVE
batch79-models-m3-r2
Continuous verification audit timestamping
Roster verification date 2026-09-08Reflects real-time continuous audits of API endpoints and pricing schedulesTimestamp verifiedVERIFIED_DETERMINISTIC
batch79-models-m3-r3
Decommissioned model deprecation tracking
Integration with /model-deprecations hubFlags dated models and provides explicit upgrade paths to current successorsDeprecations trackedVALIDATED_OBSERVED
batch79-models-m3-r4
Open weights vs proprietary licensing audit
Apache 2.0 and MIT open weights flagsClearly distinguishes open weights models from closed API-only providersLicensing disclosedVERIFIED_DETERMINISTIC
batch79-models-m3-r5
Hardware-specific hosting option disclosure
Groq LPUs, Cerebras CS-3, Alibaba StudioIdentifies specialized hosting platforms providing distinctive latency or data residencyHosting disclosedMEASURED_ACTIVE
batch79-models-m3-r6
Machine-readable dataset parity verification
models/data.json synchronized with HTML viewZero divergence between client-facing table and programmatic JSON distributionData parity = 100%VALIDATED_OBSERVED

First-party provenance: All AI Ask first-party model & pricing registry; verification date 2026-09-08. Missing or conflicting joins fail closed.

Explore all 39 model spec sheets
Largest context
Gemini 3.1 Pro
2M tokens
Cheapest current model
Amazon Nova Micro
$0.06/M blended
Fastest measured
GPT-OSS 120B (Cerebras)
2450 tok/s
Sort by:ContextPriceSpeedName
ModelProviderContextMax outputModalitiesReasoning$/M blendedtok/sStatus
Gemini 3.1 ProGoogle2M64Ktext, vision, audioYes$4.5055current
Muse Spark 1.3 ContributorMeta1.0M128KtextYes$0.13current
Gemini 3.7 FlashGoogle1.0M66Ktext, vision, audioYes$1.50current
Muse Spark 1.3Meta1.0M128KtextYes$2.00current
Gemini 2.5 Flash LiteGoogle1M8Ktext, visionNo$0.18legacy
DeepSeek V4 FlashDeepSeek1M384KtextNo$0.66132current
Gemini 3.5 Flash LiteGoogle1M64Ktext, visionYes$0.85162current
Gemini 2.5 FlashGoogle1M8Ktext, vision, audioNo$0.85legacy
Grok 4.3xAI1M64Ktext, visionYes$1.5698current
DeepSeek V4 ProDeepSeek1M384KtextYes$1.9868current
GLM-5.2Z.ai1M64KtextYes$2.15current
GPT-5.6 LunaOpenAI1M64Ktext, visionNo$2.25126current
Grok-3xAI1M8Ktext, visionNo$2.50legacy
Grok-4.20 ReasoningxAI1M64Ktext, visionYes$3.0052current
Grok-4.20xAI1M32Ktext, visionNo$3.00104current
Gemini 3.6 FlashGoogle1M64Ktext, vision, audioYes$3.00114current
GPT-5.6 TerraOpenAI1M128Ktext, visionYes$5.6378current
GPT-5.6 SolOpenAI1M128Ktext, visionYes$8.0044current
Claude Fable 5Anthropic1M128Ktext, visionYes$20.0041current
Claude Opus 5Anthropic1M128Ktext, visionYes$30.00current
Grok 4.6xAI500K64Ktext, visionYes$3.00current
Grok 4.5xAI500K64Ktext, visionYes$3.00current
Claude Sonnet 5Anthropic500K64Ktext, visionYes$4.00current
Claude Opus 4.8Anthropic500K64Ktext, visionYes$10.0058current
Amazon Nova LiteAmazon300K16Ktext, vision, audioNo$0.11108current
Amazon Nova ProAmazon300K33Ktext, vision, audioNo$1.4064current
Claude Sonnet 4.6Anthropic300K64Ktext, visionYes$6.0076current
Ministral 8BMistral256K33Ktext, visionNo$0.15158current
Mistral Small 3.1Mistral256K33Ktext, visionYes$0.26121current
CodestralMistral256K33KtextNo$0.45118current
Mistral Large 3Mistral256K33Ktext, visionNo$0.7561current
Qwen 3.7 PlusQwen256K33Ktext, visionNo$1.1084current
Qwen 3.8 MaxQwen256K33Ktext, visionYes$2.8047current
Qwen 3.7 MaxQwen256K33Ktext, visionYes$2.8049current
Mistral Medium 3Mistral256K33Ktext, visionNo$3.0092current
Claude Haiku 4.5Anthropic200K32Ktext, visionNo$2.00148current
GLM 4.7 (Cerebras)Cerebras200K33KtextYes$2.381980current
Claude Sonnet 4.5Anthropic200K32Ktext, visionYes$6.00legacy
Claude Sonnet 4Anthropic200K32Ktext, visionNo$6.00legacy
Claude Opus 4Anthropic200K4Ktext, visionNo$30.00legacy
GPT-OSS 20BGroq131K33KtextYes$0.131120current
GPT-OSS 120BGroq131K33KtextYes$0.26780current
GPT-OSS 120B (Cerebras)Cerebras131K33KtextYes$0.452450current
Qwen 3.8 30BGroq131K33Ktext, visionYes$1.20690current
Amazon Nova MicroAmazon128K8KtextNo$0.06168current
GPT-4o MiniOpenAI128K16Ktext, visionNo$0.26legacy
GLM-5.1Z.ai128K8KtextNo$1.00legacy
GPT-4oOpenAI128K16Ktext, visionNo$4.38legacy
GPT-4 TurboOpenAI128K4Ktext, visionNo$15.00legacy
GPT-5 NanoOpenAI$0.14legacy
Grok-3 MinixAI$0.26legacy
Llama 4 MaverickGroq$0.30legacy
GPT-5.4 NanoOpenAI$0.46legacy
Gemini 3.1 Flash LiteGoogle$0.56legacy
GPT-5 MiniOpenAI$0.69legacy
Qwen 3.6 27BGroq$1.20legacy
GPT-5.4 MiniOpenAI$1.69legacy
Gemini 3.1 FlashGoogle$1.69legacy
o3-MiniOpenAI$1.93legacy
Gemini 3.5 FlashGoogle$3.38legacy
GPT-5OpenAI$3.44legacy
GPT-4.1OpenAI$3.50legacy
GPT-5.4OpenAI$5.63legacy
Claude Opus 4.7Anthropic$10.00legacy
Claude Opus 4.6Anthropic$10.00legacy
Claude Opus 4.5Anthropic$10.00legacy
Claude Opus 4.1Anthropic$30.00legacy
GPT-5.4 ProOpenAI$67.50legacy

39 of 68 models have full specs (legacy catalog entries deliberately have none); 31 have measured speed. Data verified 2026-08-14.

Which LLMs have the largest context windows?

Gemini 3.1 Pro has the largest documented context window among current models in the All AI Ask roster, at 2M tokens. Gemini 3.7 Flash follows at 1.0M tokens. This ranking covers live models only, uses each provider’s published specification, and is useful when a workload must fit a long document, codebase, or conversation in one request.

Verified 2026-08-14 source

Largest context window LLMs

Current models ranked by the maximum context window documented in their spec sheet. Canonical /models view.

RankModelProviderValue
1Gemini 3.1 ProGoogle2M tokens
2Gemini 3.7 FlashGoogle1.0M tokens
3Muse Spark 1.3Meta1.0M tokens
4Muse Spark 1.3 ContributorMeta1.0M tokens
5Claude Fable 5Anthropic1M tokens
6Claude Opus 5Anthropic1M tokens
7DeepSeek V4 FlashDeepSeek1M tokens
8DeepSeek V4 ProDeepSeek1M tokens
9Gemini 3.5 Flash LiteGoogle1M tokens
10Gemini 3.6 FlashGoogle1M tokens
11GLM-5.2Z.ai1M tokens
12GPT-5.6 LunaOpenAI1M tokens
13GPT-5.6 SolOpenAI1M tokens
14GPT-5.6 TerraOpenAI1M tokens
15Grok 4.3xAI1M tokens
16Grok-4.20xAI1M tokens
17Grok-4.20 ReasoningxAI1M tokens
18Claude Opus 4.8Anthropic500K tokens
19Claude Sonnet 5Anthropic500K tokens
20Grok 4.5xAI500K tokens
21Grok 4.6xAI500K tokens
22Amazon Nova LiteAmazon300K tokens
23Amazon Nova ProAmazon300K tokens
24Claude Sonnet 4.6Anthropic300K tokens
25CodestralMistral256K tokens
26Ministral 8BMistral256K tokens
27Mistral Large 3Mistral256K tokens
28Mistral Medium 3Mistral256K tokens
29Mistral Small 3.1Mistral256K tokens
30Qwen 3.7 MaxQwen256K tokens
31Qwen 3.7 PlusQwen256K tokens
32Qwen 3.8 MaxQwen256K tokens
33Claude Haiku 4.5Anthropic200K tokens
34GLM 4.7 (Cerebras)Cerebras200K tokens
35GPT-OSS 120BGroq131K tokens
36GPT-OSS 120B (Cerebras)Cerebras131K tokens
37GPT-OSS 20BGroq131K tokens
38Qwen 3.8 30BGroq131K tokens
39Amazon Nova MicroAmazon128K tokens

Which LLMs generate the most output tokens?

DeepSeek V4 Flash has the largest documented maximum output among current models in the All AI Ask roster, at 384K tokens. DeepSeek V4 Pro follows at 384K tokens. Maximum output is a generation limit, not a promise that every response will use that many tokens; compare it separately from context capacity, price, latency, and task quality.

Verified 2026-08-14 source

LLMs with the largest max output

Current models ranked by their documented maximum output-token allowance. Canonical /models view.

RankModelProviderValue
1DeepSeek V4 FlashDeepSeek384K tokens
2DeepSeek V4 ProDeepSeek384K tokens
3Claude Fable 5Anthropic128K tokens
4Claude Opus 5Anthropic128K tokens
5GPT-5.6 SolOpenAI128K tokens
6GPT-5.6 TerraOpenAI128K tokens
7Muse Spark 1.3Meta128K tokens
8Muse Spark 1.3 ContributorMeta128K tokens
9Gemini 3.7 FlashGoogle66K tokens
10Claude Opus 4.8Anthropic64K tokens
11Claude Sonnet 4.6Anthropic64K tokens
12Claude Sonnet 5Anthropic64K tokens
13Gemini 3.1 ProGoogle64K tokens
14Gemini 3.5 Flash LiteGoogle64K tokens
15Gemini 3.6 FlashGoogle64K tokens
16GLM-5.2Z.ai64K tokens
17GPT-5.6 LunaOpenAI64K tokens
18Grok 4.3xAI64K tokens
19Grok 4.5xAI64K tokens
20Grok 4.6xAI64K tokens
21Grok-4.20 ReasoningxAI64K tokens
22Amazon Nova ProAmazon33K tokens
23CodestralMistral33K tokens
24GLM 4.7 (Cerebras)Cerebras33K tokens
25GPT-OSS 120BGroq33K tokens
26GPT-OSS 120B (Cerebras)Cerebras33K tokens
27GPT-OSS 20BGroq33K tokens
28Ministral 8BMistral33K tokens
29Mistral Large 3Mistral33K tokens
30Mistral Medium 3Mistral33K tokens
31Mistral Small 3.1Mistral33K tokens
32Qwen 3.7 MaxQwen33K tokens
33Qwen 3.7 PlusQwen33K tokens
34Qwen 3.8 30BGroq33K tokens
35Qwen 3.8 MaxQwen33K tokens
36Claude Haiku 4.5Anthropic32K tokens
37Grok-4.20xAI32K tokens
38Amazon Nova LiteAmazon16K tokens
39Amazon Nova MicroAmazon8K tokens

Which LLMs have the newest knowledge cutoff?

Claude Opus 5 has the newest disclosed knowledge cutoff among current models in the All AI Ask roster, listed as 2026-05. Claude Sonnet 5 follows at 2026-05. Models without a published cutoff are omitted rather than treated as current. A newer cutoff can reduce stale answers, but retrieval, source quality, and the prompt still determine whether a response is up to date.

Verified 2026-08-14 source

LLMs with the newest knowledge cutoff

Current models with a disclosed cutoff, ranked from newest to oldest; undisclosed cutoffs are omitted. Canonical /models view.

RankModelProviderValue
1Claude Opus 5Anthropic2026-05
2Claude Sonnet 5Anthropic2026-05
3Gemini 3.6 FlashGoogle2026-04
4Grok 4.3xAI2026-04
5GPT-5.6 LunaOpenAI2026-03
6GPT-5.6 SolOpenAI2026-03
7GPT-5.6 TerraOpenAI2026-03
8Qwen 3.8 MaxQwen2026-03
9Claude Fable 5Anthropic2026-02
10DeepSeek V4 FlashDeepSeek2026-02
11DeepSeek V4 ProDeepSeek2026-02
12Gemini 3.5 Flash LiteGoogle2026-02
13GLM-5.2Z.ai2026-02
14Grok 4.5xAI2026-02
15Qwen 3.8 30BGroq2026-02
16Claude Opus 4.8Anthropic2026-01
17Grok-4.20xAI2026-01
18Grok-4.20 ReasoningxAI2026-01
19Qwen 3.7 MaxQwen2026-01
20Qwen 3.7 PlusQwen2026-01
21Claude Sonnet 4.6Anthropic2025-12
22Mistral Medium 3Mistral2025-12
23Mistral Small 3.1Mistral2025-12
24Gemini 3.1 ProGoogle2025-11
25GLM 4.7 (Cerebras)Cerebras2025-10
26Mistral Large 3Mistral2025-10
27Claude Haiku 4.5Anthropic2025-08
28Ministral 8BMistral2025-07
29CodestralMistral2025-06
30GPT-OSS 120BGroq2025-05
31GPT-OSS 120B (Cerebras)Cerebras2025-05
32GPT-OSS 20BGroq2025-05
33Amazon Nova LiteAmazon2024-10
34Amazon Nova MicroAmazon2024-10
35Amazon Nova ProAmazon2024-10

Try any model for free

Every model in this table, one workspace, one API key.

Try It Free