Live results · Jun 21, 2026 · Audit verified 2026-09-01

Premium AI Model Tests — The Flagships on Genuinely Hard Tasks

Our cheap-model tests answer “what's the best value LLM?” This page asks the opposite question: when you reach for the most capable models money can buy, what do the extra dollars and seconds actually buy you? We took the four high-end flagships — GPT-5.4 Pro, Claude Opus 4.6, Gemini 3.1 Pro and GLM 5.2 in Max reasoning mode — and sent each the same deliberately hard prompts through the live All AI Ask API. Every number below is real, and both tasks were verified programmatically, not just eyeballed.

2
Hard tasks
4
Flagship models
8
Live API runs
41×
Price spread

The tasks

Complex Reasoning

Solving a Constraint Logic Puzzle

A five-house constraint-satisfaction puzzle with five interlocking clues and a single valid solution. The model must reason through both cases, eliminate the dead end, and report the exact arrangement and count.

Best value: Claude Opus 4.8 (100/100)
View full results →
〈/〉
Advanced Coding

Median of Two Sorted Arrays in O(log n)

A classic hard algorithm: compute the median of two sorted lists in O(log(min(m,n))) time. A merge is explicitly disallowed, so the model must implement the tricky binary-search partition correctly — including empty-list and even/odd edge cases — and return code only.

Best value: GLM 5.2 (Max) (100/100)
View full results →

Overall leaderboard

Averaged across both hard tasks. Accuracy is graded against each task's published criteria and cross-checked programmatically. Ties on accuracy break toward the cheaper model.

#ModelAvg accuracyAvg speedList priceTotal cost
🥇
Claude Opus 4.8Anthropic
10099.1 t/s$25/M$0.036505
2
GPT-5.4 ProOpenAI
10012.2 t/s$180/M$0.2565
3
GLM 5.2 (Max)Z.ai
99.559.4 t/s$4.4/M$0.02237
4
Gemini 3.1 ProGoogle
98.520.5 t/s$12/M$0.010744

Batch 51 · premium-model-tests evidence contributions. Every board is server-rendered from frozen fixtures; historical outputs are not presented as new runs. Verification date: 2026-09-01.

Deterministic test receipts

These values are read from the exported test records. Prompt hashes identify the immutable test input; raw-output hashes identify each verbatim model response. Version and validation fields are recorded metadata, not a new run or a re-grade.

Test / modelPrompt SHA-256Raw output SHA-256GraderSolver / harnessValidation flags
constraint-logic-puzzle
gpt-5.4-pro
119b0025774a31bd409ae9f1f387f96eaea0c8eb3f793dc0186496fecc3df336beea172063b8f2d3e67aa379ad19d2339907b45b4525c9897ca741db8b2e46b5claude-agent-grader-v1constraint-solver-v1-120-permutationsrecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent
constraint-logic-puzzle
claude-opus-4-8
119b0025774a31bd409ae9f1f387f96eaea0c8eb3f793dc0186496fecc3df336cac6fe506285c705adbd84cb5a5d0e1560dbde69fe6ce791e76167c1ebd38fe5claude-agent-grader-v1constraint-solver-v1-120-permutationsrecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected
constraint-logic-puzzle
gemini-3.1-pro
119b0025774a31bd409ae9f1f387f96eaea0c8eb3f793dc0186496fecc3df336598a199b193a2c86c727861f860be1aafa3c6b395b594c7a2e95c1863f4051ddclaude-agent-grader-v1constraint-solver-v1-120-permutationsrecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent
constraint-logic-puzzle
glm-5.2
119b0025774a31bd409ae9f1f387f96eaea0c8eb3f793dc0186496fecc3df3362dfceaafb7a800a36299028ff1928a564e4158c4228e5d77bf4c32a676288310claude-agent-grader-v1constraint-solver-v1-120-permutationsrecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected
median-two-sorted-arrays
gpt-5.4-pro
abc8c17072e616feae256681c64aa944e31b145375338fb5269d4eb831708a57a2aefdcfc47376d961f5c43a54d8fe0f7226dd2a8e0334b1127cbd755f2f3945claude-agent-grader-v1median-harness-v1-seed-2026-5000-casesrecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
median-two-sorted-arrays
claude-opus-4-8
abc8c17072e616feae256681c64aa944e31b145375338fb5269d4eb831708a572ba7b9c31f2a871ff6037dee27a26cfa4aba1d36f417e2b01ac65473b802a3a6claude-agent-grader-v1median-harness-v1-seed-2026-5000-casesrecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
median-two-sorted-arrays
gemini-3.1-pro
abc8c17072e616feae256681c64aa944e31b145375338fb5269d4eb831708a57c87151f5b92f95381d9b0bf0ec89b041834f164eda008b3f0a2baa885fc37617claude-agent-grader-v1median-harness-v1-seed-2026-5000-casesrecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
median-two-sorted-arrays
glm-5.2
abc8c17072e616feae256681c64aa944e31b145375338fb5269d4eb831708a571e909c8417f150d1aac0eb4b7d4c3219ed4678847a6c0d18a24005722a5b6ae3claude-agent-grader-v1median-harness-v1-seed-2026-5000-casesrecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present

Premium cohort identity and eligibility receipt

Frozen Batch 51 fixture board. Formula / decision rule: historical cohort = exact ID + host + snapshot + available run + comparable prompt; later catalog entries do not append Boundary: The four-model 2026-06-21 cohort is immutable; a renamed or newer model cannot be silently added.

Frozen fixture / field IDIdentity keysDeterministic ruleOutput / bounded stateValidation
batch51-premium-model-tests-m1-r1
four included exact IDs
suite=premium; run=2026-06-21; provider/host/model IDs=exact; snapshots=joined; two tests=availableall four exact IDs and both test joins are presenthistorical inclusion=Included; cohort size=4INCLUDED — frozen cohort
batch51-premium-model-tests-m1-r2
renamed alias
suite=premium; run=2026-06-21; requested ID=alias; resolved ID=not proven equal; host=joinedalias rename requires exact ID equivalence evidenceinclusion=Unavailable; do not merge by display nameUNAVAILABLE — alias join
batch51-premium-model-tests-m1-r3
newer untested flagship
suite=premium; run=2026-06-21; model=post-run flagship; run availability=noneno historical run means no cohort membershipexclude from historical four-model cohortEXCLUDED — untested
batch51-premium-model-tests-m1-r4
failed endpoint
suite=premium; run=2026-06-21; exact model/host=joined; endpoint status=failed; raw hash=presentfailed endpoint has no valid comparable resultexclude from complete cohort; preserve failure receiptEXCLUDED — failed endpoint
batch51-premium-model-tests-m1-r5
alternate host
suite=premium; run=2026-06-21; model display name=same; host=alternate; snapshot=Unknownhost is an identity key, not a display labelcomparability=No; do not transfer runEXCLUDED — host mismatch
batch51-premium-model-tests-m1-r6
unresolved model-version fixtures
suite=premium; run=2026-06-21; model ID=ambiguous; snapshot=Unknown; prompt hash=joinedversion ambiguity blocks selectioninclusion=Unavailable; require version joinUNAVAILABLE — version identity

Provenance: Batch 51 premium-model-tests module 1; historical run 2026-06-21; audit verification 2026-09-01. First-party premium-model suite evidence (JSON). Missing or conflicting joins fail closed.

Two-test consistency and hard-failure matrix

Frozen Batch 51 fixture board. Formula / decision rule: complete = valid solver/harness + rubric + identity joins for both tests; hard failure blocks overall rank Boundary: One missing or invalid required test prevents an overall premium-suite rank.

Frozen fixture / field IDIdentity keysDeterministic ruleOutput / bounded stateValidation
batch51-premium-model-tests-m2-r1
both-pass
suite=premium; model/host=exact; logic solver=pass; median harness=pass; rubric=joined; run=2026-06-21both required tests pass their verification gatescomplete-case=Yes; cross-test consistency=PassPASS — both tests
batch51-premium-model-tests-m2-r2
logic-pass/code-fail
suite=premium; exact test/model join; logic solver=pass; median harness=fail; raw hashes=joinedany hard failure blocks suite acceptancehard-failure=code; overall rank=UnavailableBLOCKED — code failure
batch51-premium-model-tests-m2-r3
code-pass/logic-fail
suite=premium; exact test/model join; logic solver=fail; median harness=pass; raw hashes=joinedany hard failure blocks suite acceptancehard-failure=logic; overall rank=UnavailableBLOCKED — logic failure
batch51-premium-model-tests-m2-r4
instruction-leak
suite=premium; exact model/run; test output contains reasoning wrapper where code-only required; grader=joinedcode-only hard failure is preserved separately from correctnesscomplete-case=No; no overall rankBLOCKED — instruction leak
batch51-premium-model-tests-m2-r5
incomplete run
suite=premium; exact model/run; one test status=missing; prompt hash=present; metrics=partialmissing test evidence breaks two-test denominatorcomplete-case=No; exclusion reason=incomplete runBLOCKED — incomplete
batch51-premium-model-tests-m2-r6
missing-harness fixtures
suite=premium; exact model/run; logic solver version=present; median harness version=missingverification version is decision-criticaloverall rank=Unavailable; require harness joinUNAVAILABLE — harness identity

Provenance: Batch 51 premium-model-tests module 2; historical run 2026-06-21; audit verification 2026-09-01. First-party premium-model suite evidence (JSON). Missing or conflicting joins fail closed.

Premium observed-frontier sensitivity board

Frozen Batch 51 fixture board. Formula / decision rule: frontier(policy) = eligible recorded metrics after declared hard gates/normalization; missing price => Unavailable Boundary: Two frozen prompts cannot establish a universal frontier-model verdict.

Frozen fixture / field IDIdentity keysDeterministic ruleOutput / bounded stateValidation
batch51-premium-model-tests-m3-r1
accuracy-first
suite=premium; policy=accuracy-first; recorded rubric scores=joined; eligible models=complete casesselect by recorded accuracy then declared tie rulewinner/tie=Unavailable until all complete scores joinUNAVAILABLE — score join
batch51-premium-model-tests-m3-r2
both-tests-pass
suite=premium; policy=both-tests-pass; solver/harness gates=joined; eligible models=passersfilter to models passing both exact verifierseligible set=Unavailable without complete receiptUNAVAILABLE — eligibility
batch51-premium-model-tests-m3-r3
lowest run cost among passers
suite=premium; policy=lowest recorded run cost; cost=historical only; pass gates=joinedchoose lowest recorded cost among exact passerswinner=Unavailable if any cost field is missing; current tariff excludedUNAVAILABLE — cost field
batch51-premium-model-tests-m3-r4
lowest latency among passers
suite=premium; policy=lowest recorded latency; latency=joined; pass gates=joinedchoose lowest recorded latency among exact passerswinner/tie=Unavailable without joined latency vectorUNAVAILABLE — latency field
batch51-premium-model-tests-m3-r5
balanced normalized score
suite=premium; policy=balanced; normalization=declared per test; complete cases=joinednormalize only recorded test metrics under declared policyrank movement=Unavailable; no cross-test imputationUNAVAILABLE — normalization
batch51-premium-model-tests-m3-r6
no-price-field policies
suite=premium; policy=quality/latency only; price=missing; exact run IDs=joinedquality/latency policy may proceed only if its own fields joinprice-dependent result=Unavailable; two-prompt boundary remainsBOUNDED — no price

Provenance: Batch 51 premium-model-tests module 3; historical run 2026-06-21; audit verification 2026-09-01. First-party premium-model suite evidence (JSON). Missing or conflicting joins fail closed.

Run the premium-model-tests Batch 51 evidence scenario →

How we tested

  • The cohort: the four high-end flagships — GPT-5.4 Pro, Claude Opus 4.6, Gemini 3.1 Pro, and GLM 5.2 run in Max reasoning mode.
  • Identical prompts: each model received the same prompt through the live API; GLM 5.2 used reasoningEffort: "max", the others their default flagship settings.
  • Real metrics: latency, token counts, and cost come straight from the API for each run. (GPT-5.4 Pro's token count is estimated from output length — the API under-reports it for that model.)
  • Verified accuracy: the logic puzzle was checked by exhaustive search and every code answer was executed against a 5,000-case correctness harness — so the grades aren't guesswork.

Pit the flagships against your own hard prompt

Send one prompt to every premium model at once and watch the speed, cost, and quality side by side.

Try the live playground free →