Live results · Jun 16, 2026 · Audit verified 2026-09-01

Cheap AI Model Tests — Everything Under $3 / Million Tokens

Picking a budget LLM usually means guessing. So we stopped guessing. We took every model priced under $3 per million output tokens that you can reach from a free trial, and sent each one the exact same prompts through the live All AI Ask API. Every number below — speed, cost, and the model outputs themselves — comes from real API calls, not marketing decks.

3
Tasks tested
20
Models
9
Providers
60
Live API runs

The three tests

{ }
Code Generation

Writing a Code Snippet

A focused coding task: produce a correct, efficient, 0-indexed iterative Fibonacci function in Python — and nothing but the code.

Winner: GPT-5.4 Nano (100/100)
View full results →
Copywriting

Writing a Short Paragraph

A plain-English writing task: explain what an API is to a non-technical small-business owner in 3–4 jargon-free sentences using one analogy.

Winner: GPT-5.4 Nano (97/100)
View full results →
[ ]
Structured Data

Extracting Structured Data

A structured-output task: read one sentence and return strict JSON with a string name, numeric price, and boolean stock flag — no markdown, no prose.

Winner: GPT-5.4 Nano (100/100)
View full results →

Overall leaderboard

Averaged across all 3 tasks. Accuracy is graded by the agent against each task's published criteria.

#ModelAvg accuracyAvg speedTotal cost
🥇
GPT-5.4 NanoOpenAI
9957.4 t/s$0.000274
2
Muse Spark 1.3 ContributorMeta
99195.2 t/s$0.000547
3
Gemini 3.1 Flash LiteGoogle
9782.6 t/s$0.000339
4
CodestralMistral
96.392.4 t/s$0.000227
5
Mistral Medium 3Mistral
9635.7 t/s$0.000505
6
Llama 3.1 8BGroq
95.3274.2 t/s$0.000029
7
Mistral Small 3.1Mistral
95.377.3 t/s$0.000161
8
Llama 3.3 70BGroq
95181.9 t/s$0.000324
9
Amazon Nova MicroAmazon
91.3102.6 t/s$0.000029
10
Amazon Nova LiteAmazon
90.3110.1 t/s$0.00006
11
Ministral 8BMistral
9060 t/s$0.000061
12
Llama 4 ScoutGroq
89.7194.8 t/s$0.000101
13
DeepSeek V4 ProDeepSeek
82.780.5 t/s$0.000684
14
DeepSeek V4 FlashDeepSeek
8266.4 t/s$0.000173
15
Grok 4.3xAI
80.728.9 t/s$0.000984
16
GPT-OSS 120BGroq
80.7322.3 t/s$0.000309
17
GPT-OSS 20BGroq
80502.9 t/s$0.000205
18
GPT-OSS 120B (Cerebras)Cerebras
77.3490.4 t/s$0.000534
19
GLM 4.7 (Cerebras)Cerebras
75.7531.4 t/s$0.00511
20
Qwen 3 32BGroq
73.7351.5 t/s$0.001475

Batch 51 · cheap-model-tests evidence contributions. Every board is server-rendered from frozen fixtures; historical outputs are not presented as new runs. Verification date: 2026-09-01.

Deterministic test receipts

These values are read from the exported test records. Prompt hashes identify the immutable test input; raw-output hashes identify each verbatim model response. Version and validation fields are recorded metadata, not a new run or a re-grade.

Test / modelPrompt SHA-256Raw output SHA-256GraderSolver / harnessValidation flags
code-snippet-fibonacci
gpt-5.4-nano
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1a57327357e981f19bce34cfb0c23297cc50ded8b330f716c24f57ccfa80bd655claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
code-snippet-fibonacci
gemini-3.1-flash-lite
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1412402373211028756fd80bb8d4dcb36eccff6f30f720459c6d8b76ce0baf11dclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
code-snippet-fibonacci
grok-4.3
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1cb0caeb9645e1bae32af74ffcd103c8181c0a774deae8d40a2b64051c67ea769claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present
code-snippet-fibonacci
llama-4-scout
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c14f153c75da4a20c233cecae0c093c4443c29afbcddbfea4cfbff72aa26a10986claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
code-snippet-fibonacci
gpt-oss-120b
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c19bfe8306c29b98bdce17c9ed9ba9c4cf1049a28bab74a0b5100ff111d73bd7f3claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present
code-snippet-fibonacci
gpt-oss-20b
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c10677c7b2d9fb1de2c1a8337499c71c8635ae1c8b03730c41a502f4e6d50f3b9cclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present
code-snippet-fibonacci
llama-3.3-70b
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1d25a76102a8f210d19b5e4a03fd019813491443e11f3afebd03780c35e682936claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
code-snippet-fibonacci
llama-3.1-8b
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1116ae97fa49c8db2ecc0a76bf31769aef837266ce264e4bc80ff2a30b6543f3eclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
code-snippet-fibonacci
qwen3-32b
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1dde27b8a27cec3bbf94b0170dccb6a9b6a685bf200e497a6cbbc5099627535cbclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present
code-snippet-fibonacci
deepseek-v4-flash
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1115b101e08808e34b5e082f1075d331d63d5f542516c47b4f7f472b7d06c37d9claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present
code-snippet-fibonacci
deepseek-v4-pro
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1c0e8b6d5774ef5d5da18b4b9c9de94b1f71289f1d9ca1cd66ecabfaae344605cclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present
code-snippet-fibonacci
mistral-medium
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1cd7a80fb1cd78f7ad1f83292b9c855a52c54fefd44aba9ffbd079cc96811eb4bclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
code-snippet-fibonacci
mistral-small
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1cd7a80fb1cd78f7ad1f83292b9c855a52c54fefd44aba9ffbd079cc96811eb4bclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
code-snippet-fibonacci
ministral-8b
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1cd7a80fb1cd78f7ad1f83292b9c855a52c54fefd44aba9ffbd079cc96811eb4bclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
code-snippet-fibonacci
codestral
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1cd7a80fb1cd78f7ad1f83292b9c855a52c54fefd44aba9ffbd079cc96811eb4bclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
code-snippet-fibonacci
cerebras-gpt-oss-120b
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1a1d43cb0157cd3d04bde4fb07d1e3152e86036c295e9832872c4513e58674bb1claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present
code-snippet-fibonacci
cerebras-glm-4.7
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c15b09a6ecae0b0c4ebeaf0fed639ffa7d38aa30e8baea93c2c861070b29a3c3aaclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present
code-snippet-fibonacci
nova-lite
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c13e1f8b3777e1bd772c4a101c70e97e395ed1f59eca5dfe703ebc3a2d80427572claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
code-snippet-fibonacci
nova-micro
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c11306c9cd67e804a40617e532eae72b59247ae26c25cadafdbdfbaf2e2ed6783eclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
code-snippet-fibonacci
muse-spark-1.3-contributor
d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1d25a76102a8f210d19b5e4a03fd019813491443e11f3afebd03780c35e682936claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present
short-paragraph-api-explainer
gpt-5.4-nano
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b210474e92d1d5e7b227186ddd9bafd7b85eca3b0f6ea97f96802ba00be291eacclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent
short-paragraph-api-explainer
gemini-3.1-flash-lite
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b49218b08d18f2a0dd5dba68b6668566ad86dc1fcbeb1723e451e3645e2d4ca04claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent
short-paragraph-api-explainer
grok-4.3
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b48fae9524b713506b1f6b09554eb82950abde0a4984f6347a9c78919ce41721dclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected
short-paragraph-api-explainer
llama-4-scout
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b83228ff5ce5ac7c9ecabc30dfa237090b8ec9b6dc72f77267518fe8c40d0b637claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent
short-paragraph-api-explainer
gpt-oss-120b
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2bf17d529b32e8318c8308ad703494589a37779ea74472aba86838a028a9e928c6claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected
short-paragraph-api-explainer
gpt-oss-20b
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b4411063b4c273b0afc9a3772750ec6fa696aefb9b36c114e92f132ea1caf2379claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected
short-paragraph-api-explainer
llama-3.3-70b
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b2e09f63a045968a36d9c6b84339e4d114315caf0517d2cb6c462d52d8af4a4bfclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent
short-paragraph-api-explainer
llama-3.1-8b
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2bd7f22bbe7f7691ff725dda79cf560f6728b09e1e3fa1de0cbd4540b1cc76c9e4claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent
short-paragraph-api-explainer
qwen3-32b
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b22026849898d090052075fe07250eabb2e732df791f4d4feba28072ec64d0017claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected
short-paragraph-api-explainer
deepseek-v4-flash
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b906b44b06b7b4069da988941d6585b42a0d79893daca850355c5daa16f558629claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected
short-paragraph-api-explainer
deepseek-v4-pro
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2bdb24a6153d2e5708fa525382b3d86de16d49280e6c9abfda1f4d3c3208366b2cclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected
short-paragraph-api-explainer
mistral-medium
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b02e692e523998e3ad049e53da39b7f2e2d8d39261850a745bb83535aee12b616claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent
short-paragraph-api-explainer
mistral-small
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b72d56f1bba9212bd52f3d46654a59e63a777afb136bed3c17717f12dccd49e18claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent
short-paragraph-api-explainer
ministral-8b
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2bf5d10afe812657d8e9bef6d747896de45fed08dc2ccadc4a667b22a530436797claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent
short-paragraph-api-explainer
codestral
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b180e79ca95622fde6cfafe1cf08986e522f2cba0ae1055bd55c29d5fc05e99efclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent
short-paragraph-api-explainer
cerebras-gpt-oss-120b
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b12f4e4bb9fd46c077d70df1647cdcd5c3e19eceb402b921cf7bb4354cf43f646claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected
short-paragraph-api-explainer
cerebras-glm-4.7
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b0e1b58f056a721da872756c06e5bc79a89891c741c3cb008cebfbf43f3c2f69fclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected
short-paragraph-api-explainer
nova-lite
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b69266104fafcd72851076486e9ce9dfb2affe28154c6e264d87600aaa2b4852cclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent
short-paragraph-api-explainer
nova-micro
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2bb1ad898bfd5c053c271216637e95673746042e17a95f394499333025e786e95bclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent
short-paragraph-api-explainer
muse-spark-1.3-contributor
06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b9d356e2dad28656f0b2f9d97826f8c5cd2037802fe95e69214bb21929dbacd90claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent
structured-data-extraction
gpt-5.4-nano
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2ebe74112168afc03369b78be17b508792711cb92a84e9c97e48379182dac8206claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
gemini-3.1-flash-lite
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f23826fa83b2b396c1d710eb5fa27351deab82b5354f768a9743ac558b33a096feclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
grok-4.3
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2023d73fe1562655765d39c7d7a1306a35666d00426c2d6accfb71cc38ceec5b5claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present
structured-data-extraction
llama-4-scout
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f26755f288b404d7655920b1a485c24cf630a3035286e1fb3f6b6178cd6d6a4761claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
gpt-oss-120b
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f22e3929c7452e4f92273386be003daae6fb8e13e8ea98598cda2684cdd132afb5claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present
structured-data-extraction
gpt-oss-20b
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f25c48aa19881b50afbbe4df83e7044701da97b6cc8c7715095aa149ae86710a4dclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present
structured-data-extraction
llama-3.3-70b
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397fclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
llama-3.1-8b
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397fclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
qwen3-32b
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2cb6eb2d20f562af88a59cbdd63c0b0468efda7ef969d378038ddef1d65697701claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present
structured-data-extraction
deepseek-v4-flash
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f25b97436361cd389f75d422e6939e7ff9afca4be0a5504c8d9979ab4f59315454claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present
structured-data-extraction
deepseek-v4-pro
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2d7ceebe74e6c44e5be3735489299cbe9e45cae2554eb1d88198d02bcc3dc9adfclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present
structured-data-extraction
mistral-medium
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397fclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
mistral-small
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397fclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
ministral-8b
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f26755f288b404d7655920b1a485c24cf630a3035286e1fb3f6b6178cd6d6a4761claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
codestral
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397fclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
cerebras-gpt-oss-120b
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2ba31906fe16dbdf92c8823917e422918f4484f82a3ddb724c37831f9ba6f0937claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present
structured-data-extraction
cerebras-glm-4.7
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2404c792a7901808747ed616173a4b57c2d04e7587705ffb6dd4866234417d85aclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present
structured-data-extraction
nova-lite
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f26755f288b404d7655920b1a485c24cf630a3035286e1fb3f6b6178cd6d6a4761claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
nova-micro
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f23826fa83b2b396c1d710eb5fa27351deab82b5354f768a9743ac558b33a096feclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
muse-spark-1.3-contributor
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397fclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present

Frozen cohort eligibility and drift ledger

Frozen Batch 51 fixture board. Formula / decision rule: frozen eligible = run-date price < $3/M + reachable at run + lifecycle eligible; current catalog never rewrites history Boundary: Historical inclusion is evaluated at the 2026-06-16 run and remains separate from today’s catalog.

Frozen fixture / field IDIdentity keysDeterministic ruleOutput / bounded stateValidation
batch51-cheap-model-tests-m1-r1
included-below-cap
suite=cheap; run=2026-06-16; exact model/host=joined; output price at run=< $3/M; reachability=pass; lifecycle=activeall frozen eligibility joins are truefrozen eligibility=Included; current drift checked separatelyINCLUDED — historical receipt
batch51-cheap-model-tests-m1-r2
exactly-at-cap
suite=cheap; run=2026-06-16; exact model/host=joined; output price=$3/M; reachability=pass; cap rule=strict belowstrict “under $3” excludes equalityfrozen eligibility=Excluded; no current price substitutionEXCLUDED — boundary
batch51-cheap-model-tests-m1-r3
above-cap
suite=cheap; run=2026-06-16; exact model/host=joined; output price>$3/M; reachability=passprice at run fails the strict capfrozen eligibility=ExcludedEXCLUDED — price rule
batch51-cheap-model-tests-m1-r4
missing-price
suite=cheap; run=2026-06-16; exact model/host=joined; output price=Unavailable; reachability=passmissing price cannot be assumed below capfrozen eligibility=Unavailable; do not infer inclusionUNAVAILABLE — price join
batch51-cheap-model-tests-m1-r5
unreachable-at-run
suite=cheap; run=2026-06-16; exact model/host=joined; output price=< $3/M; reachability=failreachability is a required eligibility fieldfrozen eligibility=Excluded from completed cohortEXCLUDED — unreachable
batch51-cheap-model-tests-m1-r6
now-retired fixtures
suite=cheap; run=2026-06-16; exact model/host=joined; run lifecycle=eligible; current lifecycle=retiredcurrent lifecycle drift does not rewrite the run-date statehistorical eligibility preserved; current drift=RetiredDRIFT — history preserved

Provenance: Batch 51 cheap-model-tests module 1; historical run 2026-06-16; audit verification 2026-09-01. First-party cheap-model suite evidence (JSON). Missing or conflicting joins fail closed.

Suite completeness and aggregation audit

Frozen Batch 51 fixture board. Formula / decision rule: complete case = six identity/status/metric joins for all three tasks; macro average denominator excludes invalid joins Boundary: No aggregate is emitted when model identity or a required task field is missing.

Frozen fixture / field IDIdentity keysDeterministic ruleOutput / bounded stateValidation
batch51-cheap-model-tests-m2-r1
all-three-complete
suite=cheap; model/host=exact; tasks=fibonacci, API explainer, extraction; prompt hashes=joined; run=2026-06-16; status=passthree valid task rows and all required metrics joinmacro denominator=3; complete-case eligibility=YesELIGIBLE — complete case
batch51-cheap-model-tests-m2-r2
one-missing-task
suite=cheap; model/host=exact; fibonacci=pass; API explainer=missing; extraction=pass; prompt hash=missinga missing task breaks three-task completenessaggregate=blocked; denominator not imputedBLOCKED — missing task
batch51-cheap-model-tests-m2-r3
provider error
suite=cheap; model/host=exact; task=extraction; status=provider_error; run timestamp=joinedprovider error is a failed evidence row, not a scoreaggregate=blocked for this model; exclusion reason=provider errorEXCLUDED — provider error
batch51-cheap-model-tests-m2-r4
zero-token anomaly
suite=cheap; model/host=exact; task=fibonacci; status=success; output tokens=0; raw hash=joinedsuccess with zero output tokens fails metric sanitycomplete-case eligibility=No; anomaly retainedBLOCKED — zero-token anomaly
batch51-cheap-model-tests-m2-r5
duplicate model ID
suite=cheap; model ID=duplicate; hosts=host-a/host-b; task rows=conflictingexact model and host pair must be uniqueaggregate=blocked; identity conflict preservedBLOCKED — duplicate identity
batch51-cheap-model-tests-m2-r6
stale-result fixtures
suite=cheap; model/host=exact; run=2026-06-16; audit=2026-09-01; lifecycle/price current join=stalestaleness labels the evidence; it does not alter recorded metricscomplete-case may remain Yes; current drift=staleDATED — stale historical result

Provenance: Batch 51 cheap-model-tests module 2; historical run 2026-06-16; audit verification 2026-09-01. First-party cheap-model suite evidence (JSON). Missing or conflicting joins fail closed.

Budget-suite rank sensitivity board

Frozen Batch 51 fixture board. Formula / decision rule: score(policy) = declared task weighting over normalized recorded inputs; absent metrics => Unavailable, never imputed Boundary: Every output covers three frozen prompts only and is not a general model-quality verdict.

Frozen fixture / field IDIdentity keysDeterministic ruleOutput / bounded stateValidation
batch51-cheap-model-tests-m3-r1
equal task weights
suite=cheap; policy=(1/3,1/3,1/3); eligible complete cases=joined; inputs=recorded accuracy onlymacro-average over three task scoreswinner/tie=Unavailable until joined complete-case scores are recalculatedUNAVAILABLE — score join
batch51-cheap-model-tests-m3-r2
coding-heavy
suite=cheap; policy=coding-heavy; fibonacci weight declared; complete cases=joinedapply only declared coding-heavy weights to complete casesrank movement=Unavailable without a joined score vectorUNAVAILABLE — score vector
batch51-cheap-model-tests-m3-r3
writing-heavy
suite=cheap; policy=writing-heavy; API-explainer weight declared; complete cases=joinedapply writing-heavy policy; no cross-task substitutionwinner/tie=Unavailable; decision boundary=policy declaration requiredUNAVAILABLE — score vector
batch51-cheap-model-tests-m3-r4
extraction-heavy
suite=cheap; policy=extraction-heavy; extraction weight declared; complete cases=joinedapply extraction-heavy policy to frozen extraction evidencerank movement=Unavailable absent normalized inputsUNAVAILABLE — score vector
batch51-cheap-model-tests-m3-r5
accuracy-floor-first
suite=cheap; policy=accuracy floor; floor=declared; eligible cases=joinedfilter below-floor cases before secondary rankingwinner=Unavailable until every floor and score field joinsUNAVAILABLE — floor join
batch51-cheap-model-tests-m3-r6
lowest-observed-cost policies
suite=cheap; policy=lowest recorded run cost; cost field=historical; missing cost=presentmissing cost cannot win a cost policywinner/tie=Unavailable; current tariff is out of scopeUNAVAILABLE — cost field

Provenance: Batch 51 cheap-model-tests module 3; historical run 2026-06-16; audit verification 2026-09-01. First-party cheap-model suite evidence (JSON). Missing or conflicting joins fail closed.

Run the cheap-model-tests Batch 51 evidence scenario →

How we tested

  • The cohort: every model under $3 / million output tokens reachable from a trial account (20 models, 9 providers).
  • Identical prompts: each model received the same prompt with default settings — one model per request.
  • Real metrics: latency, token counts, and cost are returned directly by the API for each run.
  • Accuracy: graded 0–100 by the agent against each task's published criteria — the full output for every model is shown so you can check the grading yourself.

Run your own prompt across all of them

Send one prompt to every cheap model at once and watch the speed, cost, and quality side by side.

Try the live playground free →