Cheap AI Model Tests — Everything Under $3 / Million Tokens
Picking a budget LLM usually means guessing. So we stopped guessing. We took every model priced under $3 per million output tokens that you can reach from a free trial, and sent each one the exact same prompts through the live All AI Ask API. Every number below — speed, cost, and the model outputs themselves — comes from real API calls, not marketing decks.
The three tests
Writing a Code Snippet
A focused coding task: produce a correct, efficient, 0-indexed iterative Fibonacci function in Python — and nothing but the code.
Writing a Short Paragraph
A plain-English writing task: explain what an API is to a non-technical small-business owner in 3–4 jargon-free sentences using one analogy.
Extracting Structured Data
A structured-output task: read one sentence and return strict JSON with a string name, numeric price, and boolean stock flag — no markdown, no prose.
Overall leaderboard
Averaged across all 3 tasks. Accuracy is graded by the agent against each task's published criteria.
| # | Model | Avg accuracy | Avg speed | Total cost |
|---|---|---|---|---|
| 🥇 | GPT-5.4 NanoOpenAI | 99 | 57.4 t/s | $0.000274 |
| 2 | Muse Spark 1.3 ContributorMeta | 99 | 195.2 t/s | $0.000547 |
| 3 | Gemini 3.1 Flash LiteGoogle | 97 | 82.6 t/s | $0.000339 |
| 4 | CodestralMistral | 96.3 | 92.4 t/s | $0.000227 |
| 5 | Mistral Medium 3Mistral | 96 | 35.7 t/s | $0.000505 |
| 6 | Llama 3.1 8BGroq | 95.3 | 274.2 t/s | $0.000029 |
| 7 | Mistral Small 3.1Mistral | 95.3 | 77.3 t/s | $0.000161 |
| 8 | Llama 3.3 70BGroq | 95 | 181.9 t/s | $0.000324 |
| 9 | Amazon Nova MicroAmazon | 91.3 | 102.6 t/s | $0.000029 |
| 10 | Amazon Nova LiteAmazon | 90.3 | 110.1 t/s | $0.00006 |
| 11 | Ministral 8BMistral | 90 | 60 t/s | $0.000061 |
| 12 | Llama 4 ScoutGroq | 89.7 | 194.8 t/s | $0.000101 |
| 13 | DeepSeek V4 ProDeepSeek | 82.7 | 80.5 t/s | $0.000684 |
| 14 | DeepSeek V4 FlashDeepSeek | 82 | 66.4 t/s | $0.000173 |
| 15 | Grok 4.3xAI | 80.7 | 28.9 t/s | $0.000984 |
| 16 | GPT-OSS 120BGroq | 80.7 | 322.3 t/s | $0.000309 |
| 17 | GPT-OSS 20BGroq | 80 | 502.9 t/s | $0.000205 |
| 18 | GPT-OSS 120B (Cerebras)Cerebras | 77.3 | 490.4 t/s | $0.000534 |
| 19 | GLM 4.7 (Cerebras)Cerebras | 75.7 | 531.4 t/s | $0.00511 |
| 20 | Qwen 3 32BGroq | 73.7 | 351.5 t/s | $0.001475 |
Batch 51 · cheap-model-tests evidence contributions. Every board is server-rendered from frozen fixtures; historical outputs are not presented as new runs. Verification date: 2026-09-01.
Deterministic test receipts
These values are read from the exported test records. Prompt hashes identify the immutable test input; raw-output hashes identify each verbatim model response. Version and validation fields are recorded metadata, not a new run or a re-grade.
| Test / model | Prompt SHA-256 | Raw output SHA-256 | Grader | Solver / harness | Validation flags |
|---|---|---|---|---|---|
code-snippet-fibonaccigpt-5.4-nano | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | a57327357e981f19bce34cfb0c23297cc50ded8b330f716c24f57ccfa80bd655 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccigemini-3.1-flash-lite | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 412402373211028756fd80bb8d4dcb36eccff6f30f720459c6d8b76ce0baf11d | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccigrok-4.3 | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | cb0caeb9645e1bae32af74ffcd103c8181c0a774deae8d40a2b64051c67ea769 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present |
code-snippet-fibonaccillama-4-scout | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 4f153c75da4a20c233cecae0c093c4443c29afbcddbfea4cfbff72aa26a10986 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccigpt-oss-120b | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 9bfe8306c29b98bdce17c9ed9ba9c4cf1049a28bab74a0b5100ff111d73bd7f3 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present |
code-snippet-fibonaccigpt-oss-20b | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 0677c7b2d9fb1de2c1a8337499c71c8635ae1c8b03730c41a502f4e6d50f3b9c | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present |
code-snippet-fibonaccillama-3.3-70b | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | d25a76102a8f210d19b5e4a03fd019813491443e11f3afebd03780c35e682936 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccillama-3.1-8b | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 116ae97fa49c8db2ecc0a76bf31769aef837266ce264e4bc80ff2a30b6543f3e | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonacciqwen3-32b | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | dde27b8a27cec3bbf94b0170dccb6a9b6a685bf200e497a6cbbc5099627535cb | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present |
code-snippet-fibonaccideepseek-v4-flash | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 115b101e08808e34b5e082f1075d331d63d5f542516c47b4f7f472b7d06c37d9 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present |
code-snippet-fibonaccideepseek-v4-pro | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | c0e8b6d5774ef5d5da18b4b9c9de94b1f71289f1d9ca1cd66ecabfaae344605c | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present |
code-snippet-fibonaccimistral-medium | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | cd7a80fb1cd78f7ad1f83292b9c855a52c54fefd44aba9ffbd079cc96811eb4b | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccimistral-small | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | cd7a80fb1cd78f7ad1f83292b9c855a52c54fefd44aba9ffbd079cc96811eb4b | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonacciministral-8b | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | cd7a80fb1cd78f7ad1f83292b9c855a52c54fefd44aba9ffbd079cc96811eb4b | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccicodestral | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | cd7a80fb1cd78f7ad1f83292b9c855a52c54fefd44aba9ffbd079cc96811eb4b | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccicerebras-gpt-oss-120b | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | a1d43cb0157cd3d04bde4fb07d1e3152e86036c295e9832872c4513e58674bb1 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present |
code-snippet-fibonaccicerebras-glm-4.7 | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 5b09a6ecae0b0c4ebeaf0fed639ffa7d38aa30e8baea93c2c861070b29a3c3aa | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, code-fence=present |
code-snippet-fibonaccinova-lite | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 3e1f8b3777e1bd772c4a101c70e97e395ed1f59eca5dfe703ebc3a2d80427572 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccinova-micro | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | 1306c9cd67e804a40617e532eae72b59247ae26c25cadafdbdfbaf2e2ed6783e | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
code-snippet-fibonaccimuse-spark-1.3-contributor | d7291bb1d553582240d939b4f9968735f29004f81c9c6136d1f60876605339c1 | d25a76102a8f210d19b5e4a03fd019813491443e11f3afebd03780c35e682936 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, code-fence=present |
short-paragraph-api-explainergpt-5.4-nano | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | 210474e92d1d5e7b227186ddd9bafd7b85eca3b0f6ea97f96802ba00be291eac | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent |
short-paragraph-api-explainergemini-3.1-flash-lite | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | 49218b08d18f2a0dd5dba68b6668566ad86dc1fcbeb1723e451e3645e2d4ca04 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent |
short-paragraph-api-explainergrok-4.3 | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | 48fae9524b713506b1f6b09554eb82950abde0a4984f6347a9c78919ce41721d | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected |
short-paragraph-api-explainerllama-4-scout | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | 83228ff5ce5ac7c9ecabc30dfa237090b8ec9b6dc72f77267518fe8c40d0b637 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent |
short-paragraph-api-explainergpt-oss-120b | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | f17d529b32e8318c8308ad703494589a37779ea74472aba86838a028a9e928c6 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected |
short-paragraph-api-explainergpt-oss-20b | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | 4411063b4c273b0afc9a3772750ec6fa696aefb9b36c114e92f132ea1caf2379 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected |
short-paragraph-api-explainerllama-3.3-70b | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | 2e09f63a045968a36d9c6b84339e4d114315caf0517d2cb6c462d52d8af4a4bf | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent |
short-paragraph-api-explainerllama-3.1-8b | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | d7f22bbe7f7691ff725dda79cf560f6728b09e1e3fa1de0cbd4540b1cc76c9e4 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent |
short-paragraph-api-explainerqwen3-32b | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | 22026849898d090052075fe07250eabb2e732df791f4d4feba28072ec64d0017 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected |
short-paragraph-api-explainerdeepseek-v4-flash | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | 906b44b06b7b4069da988941d6585b42a0d79893daca850355c5daa16f558629 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected |
short-paragraph-api-explainerdeepseek-v4-pro | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | db24a6153d2e5708fa525382b3d86de16d49280e6c9abfda1f4d3c3208366b2c | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected |
short-paragraph-api-explainermistral-medium | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | 02e692e523998e3ad049e53da39b7f2e2d8d39261850a745bb83535aee12b616 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent |
short-paragraph-api-explainermistral-small | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | 72d56f1bba9212bd52f3d46654a59e63a777afb136bed3c17717f12dccd49e18 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent |
short-paragraph-api-explainerministral-8b | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | f5d10afe812657d8e9bef6d747896de45fed08dc2ccadc4a667b22a530436797 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent |
short-paragraph-api-explainercodestral | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | 180e79ca95622fde6cfafe1cf08986e522f2cba0ae1055bd55c29d5fc05e99ef | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent |
short-paragraph-api-explainercerebras-gpt-oss-120b | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | 12f4e4bb9fd46c077d70df1647cdcd5c3e19eceb402b921cf7bb4354cf43f646 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected |
short-paragraph-api-explainercerebras-glm-4.7 | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | 0e1b58f056a721da872756c06e5bc79a89891c741c3cb008cebfbf43f3c2f69f | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected |
short-paragraph-api-explainernova-lite | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | 69266104fafcd72851076486e9ce9dfb2affe28154c6e264d87600aaa2b4852c | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent |
short-paragraph-api-explainernova-micro | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | b1ad898bfd5c053c271216637e95673746042e17a95f394499333025e786e95b | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent |
short-paragraph-api-explainermuse-spark-1.3-contributor | 06bc5a20094984a8cf1ca8fc4ab7e6b83d36d0297d4d7a02d522f5b9433a9e2b | 9d356e2dad28656f0b2f9d97826f8c5cd2037802fe95e69214bb21929dbacd90 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent |
structured-data-extractiongpt-5.4-nano | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | ebe74112168afc03369b78be17b508792711cb92a84e9c97e48379182dac8206 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractiongemini-3.1-flash-lite | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 3826fa83b2b396c1d710eb5fa27351deab82b5354f768a9743ac558b33a096fe | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractiongrok-4.3 | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 023d73fe1562655765d39c7d7a1306a35666d00426c2d6accfb71cc38ceec5b5 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present |
structured-data-extractionllama-4-scout | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 6755f288b404d7655920b1a485c24cf630a3035286e1fb3f6b6178cd6d6a4761 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractiongpt-oss-120b | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 2e3929c7452e4f92273386be003daae6fb8e13e8ea98598cda2684cdd132afb5 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present |
structured-data-extractiongpt-oss-20b | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 5c48aa19881b50afbbe4df83e7044701da97b6cc8c7715095aa149ae86710a4d | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present |
structured-data-extractionllama-3.3-70b | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397f | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractionllama-3.1-8b | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397f | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractionqwen3-32b | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | cb6eb2d20f562af88a59cbdd63c0b0468efda7ef969d378038ddef1d65697701 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present |
structured-data-extractiondeepseek-v4-flash | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 5b97436361cd389f75d422e6939e7ff9afca4be0a5504c8d9979ab4f59315454 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present |
structured-data-extractiondeepseek-v4-pro | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | d7ceebe74e6c44e5be3735489299cbe9e45cae2554eb1d88198d02bcc3dc9adf | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present |
structured-data-extractionmistral-medium | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397f | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractionmistral-small | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397f | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractionministral-8b | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 6755f288b404d7655920b1a485c24cf630a3035286e1fb3f6b6178cd6d6a4761 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractioncodestral | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397f | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractioncerebras-gpt-oss-120b | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | ba31906fe16dbdf92c8823917e422918f4484f82a3ddb724c37831f9ba6f0937 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present |
structured-data-extractioncerebras-glm-4.7 | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 404c792a7901808747ed616173a4b57c2d04e7587705ffb6dd4866234417d85a | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present |
structured-data-extractionnova-lite | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 6755f288b404d7655920b1a485c24cf630a3035286e1fb3f6b6178cd6d6a4761 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractionnova-micro | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 3826fa83b2b396c1d710eb5fa27351deab82b5354f768a9743ac558b33a096fe | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractionmuse-spark-1.3-contributor | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397f | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
Frozen cohort eligibility and drift ledger
Frozen Batch 51 fixture board. Formula / decision rule: frozen eligible = run-date price < $3/M + reachable at run + lifecycle eligible; current catalog never rewrites history Boundary: Historical inclusion is evaluated at the 2026-06-16 run and remains separate from today’s catalog.
| Frozen fixture / field ID | Identity keys | Deterministic rule | Output / bounded state | Validation |
|---|---|---|---|---|
batch51-cheap-model-tests-m1-r1included-below-cap | suite=cheap; run=2026-06-16; exact model/host=joined; output price at run=< $3/M; reachability=pass; lifecycle=active | all frozen eligibility joins are true | frozen eligibility=Included; current drift checked separately | INCLUDED — historical receipt |
batch51-cheap-model-tests-m1-r2exactly-at-cap | suite=cheap; run=2026-06-16; exact model/host=joined; output price=$3/M; reachability=pass; cap rule=strict below | strict “under $3” excludes equality | frozen eligibility=Excluded; no current price substitution | EXCLUDED — boundary |
batch51-cheap-model-tests-m1-r3above-cap | suite=cheap; run=2026-06-16; exact model/host=joined; output price>$3/M; reachability=pass | price at run fails the strict cap | frozen eligibility=Excluded | EXCLUDED — price rule |
batch51-cheap-model-tests-m1-r4missing-price | suite=cheap; run=2026-06-16; exact model/host=joined; output price=Unavailable; reachability=pass | missing price cannot be assumed below cap | frozen eligibility=Unavailable; do not infer inclusion | UNAVAILABLE — price join |
batch51-cheap-model-tests-m1-r5unreachable-at-run | suite=cheap; run=2026-06-16; exact model/host=joined; output price=< $3/M; reachability=fail | reachability is a required eligibility field | frozen eligibility=Excluded from completed cohort | EXCLUDED — unreachable |
batch51-cheap-model-tests-m1-r6now-retired fixtures | suite=cheap; run=2026-06-16; exact model/host=joined; run lifecycle=eligible; current lifecycle=retired | current lifecycle drift does not rewrite the run-date state | historical eligibility preserved; current drift=Retired | DRIFT — history preserved |
Provenance: Batch 51 cheap-model-tests module 1; historical run 2026-06-16; audit verification 2026-09-01. First-party cheap-model suite evidence (JSON). Missing or conflicting joins fail closed.
Suite completeness and aggregation audit
Frozen Batch 51 fixture board. Formula / decision rule: complete case = six identity/status/metric joins for all three tasks; macro average denominator excludes invalid joins Boundary: No aggregate is emitted when model identity or a required task field is missing.
| Frozen fixture / field ID | Identity keys | Deterministic rule | Output / bounded state | Validation |
|---|---|---|---|---|
batch51-cheap-model-tests-m2-r1all-three-complete | suite=cheap; model/host=exact; tasks=fibonacci, API explainer, extraction; prompt hashes=joined; run=2026-06-16; status=pass | three valid task rows and all required metrics join | macro denominator=3; complete-case eligibility=Yes | ELIGIBLE — complete case |
batch51-cheap-model-tests-m2-r2one-missing-task | suite=cheap; model/host=exact; fibonacci=pass; API explainer=missing; extraction=pass; prompt hash=missing | a missing task breaks three-task completeness | aggregate=blocked; denominator not imputed | BLOCKED — missing task |
batch51-cheap-model-tests-m2-r3provider error | suite=cheap; model/host=exact; task=extraction; status=provider_error; run timestamp=joined | provider error is a failed evidence row, not a score | aggregate=blocked for this model; exclusion reason=provider error | EXCLUDED — provider error |
batch51-cheap-model-tests-m2-r4zero-token anomaly | suite=cheap; model/host=exact; task=fibonacci; status=success; output tokens=0; raw hash=joined | success with zero output tokens fails metric sanity | complete-case eligibility=No; anomaly retained | BLOCKED — zero-token anomaly |
batch51-cheap-model-tests-m2-r5duplicate model ID | suite=cheap; model ID=duplicate; hosts=host-a/host-b; task rows=conflicting | exact model and host pair must be unique | aggregate=blocked; identity conflict preserved | BLOCKED — duplicate identity |
batch51-cheap-model-tests-m2-r6stale-result fixtures | suite=cheap; model/host=exact; run=2026-06-16; audit=2026-09-01; lifecycle/price current join=stale | staleness labels the evidence; it does not alter recorded metrics | complete-case may remain Yes; current drift=stale | DATED — stale historical result |
Provenance: Batch 51 cheap-model-tests module 2; historical run 2026-06-16; audit verification 2026-09-01. First-party cheap-model suite evidence (JSON). Missing or conflicting joins fail closed.
Budget-suite rank sensitivity board
Frozen Batch 51 fixture board. Formula / decision rule: score(policy) = declared task weighting over normalized recorded inputs; absent metrics => Unavailable, never imputed Boundary: Every output covers three frozen prompts only and is not a general model-quality verdict.
| Frozen fixture / field ID | Identity keys | Deterministic rule | Output / bounded state | Validation |
|---|---|---|---|---|
batch51-cheap-model-tests-m3-r1equal task weights | suite=cheap; policy=(1/3,1/3,1/3); eligible complete cases=joined; inputs=recorded accuracy only | macro-average over three task scores | winner/tie=Unavailable until joined complete-case scores are recalculated | UNAVAILABLE — score join |
batch51-cheap-model-tests-m3-r2coding-heavy | suite=cheap; policy=coding-heavy; fibonacci weight declared; complete cases=joined | apply only declared coding-heavy weights to complete cases | rank movement=Unavailable without a joined score vector | UNAVAILABLE — score vector |
batch51-cheap-model-tests-m3-r3writing-heavy | suite=cheap; policy=writing-heavy; API-explainer weight declared; complete cases=joined | apply writing-heavy policy; no cross-task substitution | winner/tie=Unavailable; decision boundary=policy declaration required | UNAVAILABLE — score vector |
batch51-cheap-model-tests-m3-r4extraction-heavy | suite=cheap; policy=extraction-heavy; extraction weight declared; complete cases=joined | apply extraction-heavy policy to frozen extraction evidence | rank movement=Unavailable absent normalized inputs | UNAVAILABLE — score vector |
batch51-cheap-model-tests-m3-r5accuracy-floor-first | suite=cheap; policy=accuracy floor; floor=declared; eligible cases=joined | filter below-floor cases before secondary ranking | winner=Unavailable until every floor and score field joins | UNAVAILABLE — floor join |
batch51-cheap-model-tests-m3-r6lowest-observed-cost policies | suite=cheap; policy=lowest recorded run cost; cost field=historical; missing cost=present | missing cost cannot win a cost policy | winner/tie=Unavailable; current tariff is out of scope | UNAVAILABLE — cost field |
Provenance: Batch 51 cheap-model-tests module 3; historical run 2026-06-16; audit verification 2026-09-01. First-party cheap-model suite evidence (JSON). Missing or conflicting joins fail closed.
How we tested
- The cohort: every model under $3 / million output tokens reachable from a trial account (20 models, 9 providers).
- Identical prompts: each model received the same prompt with default settings — one model per request.
- Real metrics: latency, token counts, and cost are returned directly by the API for each run.
- Accuracy: graded 0–100 by the agent against each task's published criteria — the full output for every model is shown so you can check the grading yourself.
