Extracting Structured Data: Which Cheap AI Model Wins?
A structured-output task: read one sentence and return strict JSON with a string name, numeric price, and boolean stock flag — no markdown, no prose. We sent this exact task to all 20 budget models (each under $3 per million output tokens) through the live API and measured real speed, cost, and accuracy. The full output of every model is included below — click any row to read it.
The prompt
Every model received exactly this, with default generation settings:
Extract the product details from the following sentence and return ONLY a valid JSON object with exactly these keys: "name" (string), "price" (number), "in_stock" (boolean). Do not include any explanation or markdown. Sentence: "The DeskPro X1 monitor stand normally retails for $49.99 and is currently available to ship."
How accuracy was graded
- Output is valid, parseable JSON.
- Extracts the full product name ("DeskPro X1 monitor stand").
- Correct types: price as a number (49.99), in_stock as a boolean (true).
- Returns only the JSON — no markdown fences, prose, or leaked reasoning.
Results — sortable by speed, cost & accuracy
| # | Model | Accuracy ↓ | Speed ↕ | Cost ↕ | Output |
|---|---|---|---|---|---|
| 1 | GPT-5.4 NanoOpenAI · gpt-5.4-nano | 100★ | 49.2 t/s488 ms | $0.0000456$1.25/M out | |
Agent grade — 100/100. Perfect: compact valid JSON, full product name, correct number and boolean types, no markdown or prose. Verbatim output (78 in / 24 out tokens, 488 ms, $0.0000456) {"name":"DeskPro X1 monitor stand","price":49.99,"in_stock":true} | |||||
| 2 | Muse Spark 1.3 ContributorMeta · muse-spark-1.3-contributor | 100★ | 207.1 t/s3704 ms | $0.000161$0.2/M out | |
Agent grade — 100/100. Valid, compact JSON; full product name; correct number and boolean types; no markdown or prose. Same reasoning tax as the other two tasks — most of the 767 output tokens were hidden chain-of-thought. Verbatim output (79 in / 767 out tokens, 3704 ms, $0.000161) {"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true} | |||||
| 3 | Gemini 3.1 Flash LiteGoogle · gemini-3.1-flash-lite | 99 | 75.5 t/s477 ms | $0.0000727$1.5/M out | |
Agent grade — 99/100. Valid JSON, full name, correct types, no markdown — just pretty-printed rather than compact. Verbatim output (75 in / 36 out tokens, 477 ms, $0.0000727) {
"name": "DeskPro X1 monitor stand",
"price": 49.99,
"in_stock": true
} | |||||
| 4 | Llama 3.3 70BGroq · llama-3.3-70b | 99 | 151.2 t/s172 ms | $0.0000837$0.79/M out | |
Agent grade — 99/100. Valid compact JSON, full name, correct types, no markdown or prose. Fastest run in this test. Verbatim output (107 in / 26 out tokens, 172 ms, $0.0000837) {"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true} | |||||
| 5 | Llama 3.1 8BGroq · llama-3.1-8b | 99 | 87.8 t/s296 ms | $0.0000074$0.08/M out | |
Agent grade — 99/100. Valid compact JSON, full name, correct types, no markdown — and the cheapest run here. Verbatim output (107 in / 26 out tokens, 296 ms, $0.0000074) {"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true} | |||||
| 6 | Mistral Medium 3Mistral · mistral-medium | 99 | 35.9 t/s808 ms | $0.0000944$2/M out | |
Agent grade — 99/100. Valid compact JSON, full name, correct types, no markdown or leakage. Verbatim output (91 in / 29 out tokens, 808 ms, $0.0000944) {"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true} | |||||
| 7 | Mistral Small 3.1Mistral · mistral-small | 99 | 54.7 t/s530 ms | $0.000031$0.6/M out | |
Agent grade — 99/100. Valid compact JSON, full name, correct types — identical quality to Medium for far less. Verbatim output (91 in / 29 out tokens, 530 ms, $0.000031) {"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true} | |||||
| 8 | CodestralMistral · codestral | 99 | 61.7 t/s470 ms | $0.0000498$0.9/M out | |
Agent grade — 99/100. Valid compact JSON, full name, correct types, no markdown or prose. Verbatim output (79 in / 29 out tokens, 470 ms, $0.0000498) {"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true} | |||||
| 9 | Amazon Nova MicroAmazon · nova-micro | 99 | 89.5 t/s380 ms | $0.0000074$0.14/M out | |
Agent grade — 99/100. Valid JSON, full name, correct types, no markdown or prose — only pretty-printed rather than compact. Verbatim output (76 in / 34 out tokens, 380 ms, $0.0000074) {
"name": "DeskPro X1 monitor stand",
"price": 49.99,
"in_stock": true
} | |||||
| 10 | Llama 4 ScoutGroq · llama-4-scout | 85 | 100 t/s350 ms | $0.0000209$0.34/M out | |
Agent grade — 85/100. Correct values and full name, but wrapped the JSON in a ```json markdown fence after being told "no markdown". Verbatim output (82 in / 35 out tokens, 350 ms, $0.0000209) ```json
{
"name": "DeskPro X1 monitor stand",
"price": 49.99,
"in_stock": true
}
``` | |||||
| 11 | Ministral 8BMistral · ministral-8b | 85 | 63.4 t/s599 ms | $0.0000175$0.15/M out | |
Agent grade — 85/100. Correct values and full name, but wrapped the output in a ```json markdown fence against instructions. Verbatim output (79 in / 38 out tokens, 599 ms, $0.0000175) ```json
{
"name": "DeskPro X1 monitor stand",
"price": 49.99,
"in_stock": true
}
``` | |||||
| 12 | Amazon Nova LiteAmazon · nova-lite | 85 | 111.8 t/s340 ms | $0.0000137$0.24/M out | |
Agent grade — 85/100. Correct values and full name, but wrapped the JSON in a ```json markdown fence after being told "no markdown". Verbatim output (76 in / 38 out tokens, 340 ms, $0.0000137) ```json
{
"name": "DeskPro X1 monitor stand",
"price": 49.99,
"in_stock": true
}
``` | |||||
| 13 | DeepSeek V4 FlashDeepSeek · deepseek-v4-flash | 82 | 60.5 t/s2033 ms | $0.0000455$0.28/M out | |
Agent grade — 82/100. Correct JSON with full name and proper types; docked only for the leaked <think> preamble. Verbatim output (79 in / 123 out tokens, 2033 ms, $0.0000455) <think>
We need to extract product details: name, price (number), in_stock. Sentence: "The DeskPro X1 monitor stand normally retails for $49.99 and is currently available to ship."
Name: "DeskPro X1 monitor stand" (should be string)
Price: $49.99 -> 49.99 as number
In_stock: "currently available to ship" implies it's in stock, so true.
Return JSON only.
</think>
{"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true} | |||||
| 14 | DeepSeek V4 ProDeepSeek · deepseek-v4-pro | 82 | 85.1 t/s2104 ms | $0.00019$0.87/M out | |
Agent grade — 82/100. Correct, complete JSON; the leaked <think> block is the only issue. Verbatim output (79 in / 179 out tokens, 2104 ms, $0.00019) <think>
We are asked to extract product details from the sentence: "The DeskPro X1 monitor stand normally retails for $49.99 and is currently available to ship."
We need to return ONLY a valid JSON object with keys: "name" (string), "price" (number), "in_stock" (boolean).
From the sentence:
- Name: "DeskPro X1 monitor stand"
- Price: $49.99 -> number 49.99
- In stock: "currently available to ship" implies it is in stock, so true.
So JSON: {"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true}
Make sure no extra text.
</think>
{"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true} | |||||
| 15 | Grok 4.3xAI · grok-4.3 | 80 | 26.3 t/s2819 ms | $0.000286$2.5/M out | |
Agent grade — 80/100. Valid JSON with the full name and correct types, but a leaked <think> block violates "return ONLY a valid JSON object". Verbatim output (81 in / 74 out tokens, 2819 ms, $0.000286) <think>
The task is: Extract the product details from the sentence and return ONLY a valid JSON object with exactly these keys: "name" (string), "price" (number), "in_stock" (boolean). Do not include any expl...
</think>
{
"name": "DeskPro X1 monitor stand",
"price": 49.99,
"in_stock": true
} | |||||
| 16 | GPT-OSS 120BGroq · gpt-oss-120b | 80 | 292.4 t/s277 ms | $0.00007$0.6/M out | |
Agent grade — 80/100. Final JSON is correct and complete, but a <think> block leaked in despite the "no explanation" instruction. Verbatim output (143 in / 81 out tokens, 277 ms, $0.00007) <think>
We need to output JSON with keys name, price, in_stock. Extract details: name: "DeskPro X1 monitor stand". price: 49.99. in_stock: true (available to ship). Return only JSON, no markdown.
</think>
{"name":"DeskPro X1 monitor stand","price":49.99,"in_stock":true} | |||||
| 17 | GPT-OSS 20BGroq · gpt-oss-20b | 80 | 478.1 t/s343 ms | $0.0000599$0.3/M out | |
Agent grade — 80/100. Correct JSON with the full name, but preceded by a leaked <think> reasoning block. Verbatim output (143 in / 164 out tokens, 343 ms, $0.0000599) <think>
The user wants to extract product details: name: "DeskPro X1 monitor stand", price: 49.99, in_stock: true. The sentence: "The DeskPro X1 monitor stand normally retails for $49.99 and is currently available to ship." The name: "DeskPro X1 monitor stand". Price: 49.99. In_stock: currently available to ship => in stock: true. JSON with keys: name, price, in_stock.
Return only JSON object, no explanation, no markdown. So just:
{"name":"DeskPro X1 monitor stand","price":49.99,"in_stock":true}
</think>
{"name":"DeskPro X1 monitor stand","price":49.99,"in_stock":true} | |||||
| 18 | Qwen 3 32BGroq · qwen3-32b | 75 | 428.6 t/s784 ms | $0.000222$0.59/M out | |
Agent grade — 75/100. Final JSON is correct, but buried under a ~340-token leaked <think> monologue — the opposite of "JSON only". Verbatim output (82 in / 336 out tokens, 784 ms, $0.000222) <think>
Okay, let me try to figure this out. The user wants me to extract product details from a sentence and put them into aJSON object with specific keys: name, price, and in_stock. The sentence given is "The DeskPro X1 monitor stand normally retails for $49.99 and is currently available to ship."
First, I need to identify the product name. The sentence starts with "The DeskPro X1 monitor stand..." so the name is probably "DeskPro X1 monitor stand". I should check if there's any other part of the name, but "DeskPro X1" seems like the model, and "monitor stand" is the product type. So the name key should be that full phrase.
Next, the price is mentioned as "$49.99". The key requires a number, so I need to remove the dollar sign and convert it to a numerical value. That would be 49.99.
For the in_stock boolean, the sentence says "is currently available to ship". If it's available to ship, that likely means it's in stock. So in_stock should be true.
I need to make sure there are only those three keys and that the JSON is valid. Let me double-check each part. The name is correct, price is a number without the currency symbol, and in_stock is a boolean. No other keys are needed. Alright, putting that all together into a JSON object with the exact keys.
</think>
{"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true} | |||||
| 19 | GLM 4.7 (Cerebras)Cerebras · cerebras-glm-4.7 | 73 | 562.9 t/s723 ms | $0.001292$2.75/M out | |
Agent grade — 73/100. The final JSON is correct, but the model exposed a ~400-token multi-step <think> process despite "no explanation". Verbatim output (77 in / 407 out tokens, 723 ms, $0.001292) <think>
1. **Analyze the Request:**
* Input sentence: "The DeskPro X1 monitor stand normally retails for $49.99 and is currently available to ship."
* Task: Extract product details.
* Output format: Valid JSON object only.
* Required keys: "name" (string), "price" (number), "in_stock" (boolean).
* Constraint: No explanation, no markdown code blocks (just the raw JSON).
2. **Analyze the Sentence to Extract Data:**
* *Name:* "The DeskPro X1 monitor stand". I should probably clean this up slightly to just the product name, or keep it as is. "DeskPro X1 monitor stand" seems most appropriate. Let's stick to the core product name.
* *Price:* "$49.99". The key requires a number. So, 49.99.
* *Stock Status:* "currently available to ship". This implies it is in stock. So, true.
3. **Construct the JSON Object:**
* `name`: "DeskPro X1 monitor stand"
* `price`: 49.99
* `in_stock`: true
4. **Verify against Constraints:**
* Valid JSON? Yes.
* Keys present? "name", "price", "in_stock". Yes.
* Types correct? String, Number, Boolean. Yes.
* No explanation? Yes.
* No markdown? Yes (I need to ensure I don't wrap it in backticks).
5. **Final Output Generation:**
{"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true}
</think>
{
"name": "DeskPro X1 monitor stand",
"price": 49.99,
"in_stock": true
} | |||||
| 20 | GPT-OSS 120B (Cerebras)Cerebras · cerebras-gpt-oss-120b | 72 | 557.6 t/s165 ms | $0.000118$0.75/M out | |
Agent grade — 72/100. Leaked a <think> block AND kept the leading "The" in the name ("The DeskPro X1 monitor stand"), so the extraction is slightly off. Verbatim output (139 in / 92 out tokens, 165 ms, $0.000118) <think>
We need to extract product details: name, price, in_stock. Name is "The DeskPro X1 monitor stand". Price is 49.99. in_stock: currently available to ship => true. Return JSON only, no markdown. Ensure keys exact. Provide number, boolean, string.
</think>
{"name":"The DeskPro X1 monitor stand","price":49.99,"in_stock":true} | |||||
Batch 51 · cheap-model-tests-structured-data-extraction evidence contributions. Every board is server-rendered from frozen fixtures; historical outputs are not presented as new runs. Verification date: 2026-09-01.
Deterministic test receipts
These values are read from the exported test records. Prompt hashes identify the immutable test input; raw-output hashes identify each verbatim model response. Version and validation fields are recorded metadata, not a new run or a re-grade.
| Test / model | Prompt SHA-256 | Raw output SHA-256 | Grader | Solver / harness | Validation flags |
|---|---|---|---|---|---|
structured-data-extractiongpt-5.4-nano | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | ebe74112168afc03369b78be17b508792711cb92a84e9c97e48379182dac8206 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractiongemini-3.1-flash-lite | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 3826fa83b2b396c1d710eb5fa27351deab82b5354f768a9743ac558b33a096fe | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractiongrok-4.3 | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 023d73fe1562655765d39c7d7a1306a35666d00426c2d6accfb71cc38ceec5b5 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present |
structured-data-extractionllama-4-scout | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 6755f288b404d7655920b1a485c24cf630a3035286e1fb3f6b6178cd6d6a4761 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractiongpt-oss-120b | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 2e3929c7452e4f92273386be003daae6fb8e13e8ea98598cda2684cdd132afb5 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present |
structured-data-extractiongpt-oss-20b | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 5c48aa19881b50afbbe4df83e7044701da97b6cc8c7715095aa149ae86710a4d | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present |
structured-data-extractionllama-3.3-70b | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397f | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractionllama-3.1-8b | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397f | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractionqwen3-32b | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | cb6eb2d20f562af88a59cbdd63c0b0468efda7ef969d378038ddef1d65697701 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present |
structured-data-extractiondeepseek-v4-flash | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 5b97436361cd389f75d422e6939e7ff9afca4be0a5504c8d9979ab4f59315454 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present |
structured-data-extractiondeepseek-v4-pro | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | d7ceebe74e6c44e5be3735489299cbe9e45cae2554eb1d88198d02bcc3dc9adf | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present |
structured-data-extractionmistral-medium | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397f | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractionmistral-small | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397f | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractionministral-8b | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 6755f288b404d7655920b1a485c24cf630a3035286e1fb3f6b6178cd6d6a4761 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractioncodestral | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397f | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractioncerebras-gpt-oss-120b | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | ba31906fe16dbdf92c8823917e422918f4484f82a3ddb724c37831f9ba6f0937 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present |
structured-data-extractioncerebras-glm-4.7 | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 404c792a7901808747ed616173a4b57c2d04e7587705ffb6dd4866234417d85a | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present |
structured-data-extractionnova-lite | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 6755f288b404d7655920b1a485c24cf630a3035286e1fb3f6b6178cd6d6a4761 | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractionnova-micro | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | 3826fa83b2b396c1d710eb5fa27351deab82b5354f768a9743ac558b33a096fe | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
structured-data-extractionmuse-spark-1.3-contributor | 29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2 | fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397f | claude-agent-grader-v1 | Not applicable | recorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present |
Raw-output JSON and schema validation receipt
Frozen Batch 51 fixture board. Formula / decision rule: strict acceptance = original output parses + schema/type/duplicate/extra-field policy passes Boundary: Outputs are validated as received; repair or cleanup is never silently substituted for validity.
| Frozen fixture / field ID | Identity keys | Deterministic rule | Output / bounded state | Validation |
|---|---|---|---|---|
batch51-cheap-model-tests-structured-data-extraction-m1-r1valid object | suite=cheap; slug=structured-data-extraction; exact model/run; raw hash=joined; source hash=joined; schema=v1 | parse original object and compare to schema v1 | strict acceptance=Unavailable until raw parse receipt joins | UNAVAILABLE — parse receipt |
batch51-cheap-model-tests-structured-data-extraction-m1-r2fenced JSON | suite=cheap; slug=structured-data-extraction; exact model/run; raw hash=joined; extraction rule=code-fence unwrap; schema=v1 | record wrapper removal as extraction, not original validity | original validity=fail or conditional; strict state=Unavailable | CONDITIONAL — wrapper recorded |
batch51-cheap-model-tests-structured-data-extraction-m1-r3prose-prefixed JSON | suite=cheap; slug=structured-data-extraction; exact model/run; raw hash=joined; prefix=present; schema=v1 | prose prefix violates strict answer envelope | parse may be recoverable, strict acceptance=fail | FAIL — prose prefix |
batch51-cheap-model-tests-structured-data-extraction-m1-r4duplicate key | suite=cheap; slug=structured-data-extraction; exact model/run; raw hash=joined; duplicate key=present; schema=v1 | duplicate keys are reported before parser normalization | strict acceptance=fail; preserve key collision | FAIL — duplicate key |
batch51-cheap-model-tests-structured-data-extraction-m1-r5wrong type | suite=cheap; slug=structured-data-extraction; exact model/run; raw hash=joined; field type=wrong; schema=v1 | type validation is field-level and strict | strict acceptance=fail; type error=retained | FAIL — type error |
batch51-cheap-model-tests-structured-data-extraction-m1-r6malformed JSON fixtures | suite=cheap; slug=structured-data-extraction; exact model/run; raw hash=joined; parser=error; schema=v1 | malformed original output cannot be accepted by repair | strict acceptance=fail; original parse=invalid | FAIL — malformed |
Provenance: Batch 51 cheap-model-tests-structured-data-extraction module 1; historical run 2026-06-16; audit verification 2026-09-01. First-party extraction test evidence (JSON). Missing or conflicting joins fail closed.
Field-level semantic diff ledger
Frozen Batch 51 fixture board. Formula / decision rule: field match = explicit source span + typed comparison + allowed normalization; ambiguity never becomes a match Boundary: Every accepted non-null field needs an explicit source span in the frozen source record.
| Frozen fixture / field ID | Identity keys | Deterministic rule | Output / bounded state | Validation |
|---|---|---|---|---|
batch51-cheap-model-tests-structured-data-extraction-m2-r1exact string | suite=cheap; slug=structured-data-extraction; exact model/run; source hash/schema hash=joined; field=string; source span=joined | raw value equals typed source value | semantic match=Yes; normalized value=source value | MATCH — explicit source |
batch51-cheap-model-tests-structured-data-extraction-m2-r2case/whitespace normalization | suite=cheap; slug=structured-data-extraction; exact model/run; field=string; raw differs by case/space; source span=joined | only declared normalization may be applied | semantic match=Conditional; normalization recorded | CONDITIONAL — normalization |
batch51-cheap-model-tests-structured-data-extraction-m2-r3numeric string | suite=cheap; slug=structured-data-extraction; exact model/run; field=numeric; raw="42"; expected=42; source span=joined | numeric coercion is allowed only when schema policy declares it | normalized value=42 if policy joined; otherwise Unavailable | CONDITIONAL — type policy |
batch51-cheap-model-tests-structured-data-extraction-m2-r4missing nullable field | suite=cheap; slug=structured-data-extraction; exact model/run; field=nullable; raw=missing; source says=null; source span=joined | missing and explicit null are distinct unless schema allows omission | semantic match=Conditional; omission rule required | CONDITIONAL — nullable rule |
batch51-cheap-model-tests-structured-data-extraction-m2-r5hallucinated value | suite=cheap; slug=structured-data-extraction; exact model/run; field=non-null; raw value=not in source; source span=none | non-null acceptance requires a source span | semantic match=No; downstream release blocked | FAIL — unsupported value |
batch51-cheap-model-tests-structured-data-extraction-m2-r6conflicting-source fixtures | suite=cheap; slug=structured-data-extraction; exact model/run; field=value; source spans=conflict; schema hash=joined | conflicting sources fail closed until owner resolves precedence | semantic=Unresolved; retain both source spans | UNRESOLVED — source conflict |
Provenance: Batch 51 cheap-model-tests-structured-data-extraction module 2; historical run 2026-06-16; audit verification 2026-09-01. First-party extraction test evidence (JSON). Missing or conflicting joins fail closed.
Downstream blast-radius and release gate
Frozen Batch 51 fixture board. Formula / decision rule: release = strict/semantic validation pass for every decision-critical field + action-specific review policy Boundary: External side effects are blocked when a decision-critical field is missing, invalid, ambiguous, or unsupported.
| Frozen fixture / field ID | Identity keys | Deterministic rule | Output / bounded state | Validation |
|---|---|---|---|---|
batch51-cheap-model-tests-structured-data-extraction-m3-r1analytics-only field | suite=cheap; slug=structured-data-extraction; exact model/run; action=analytics; fields=non-critical; validation=joined | analytics may proceed only under the declared non-critical policy | release=Conditional; retain provenance for downstream users | CONDITIONAL — low blast radius |
batch51-cheap-model-tests-structured-data-extraction-m3-r2billing total | suite=cheap; slug=structured-data-extraction; exact model/run; action=billing; total field=decision-critical; strict validation=Unknown | billing requires strict typed validation and source span | release=blocked until strict validation passes | BLOCKED — financial side effect |
batch51-cheap-model-tests-structured-data-extraction-m3-r3customer routing | suite=cheap; slug=structured-data-extraction; exact model/run; action=routing; route field=critical; semantic=joined | routing requires field-level semantic match and fallback owner | release=Conditional; human review if ambiguity remains | CONDITIONAL — routing gate |
batch51-cheap-model-tests-structured-data-extraction-m3-r4compliance flag | suite=cheap; slug=structured-data-extraction; exact model/run; action=compliance; flag=critical; source precedence=Unknown | unknown precedence blocks compliance decision | release=blocked; human/compliance owner required | BLOCKED — compliance gap |
batch51-cheap-model-tests-structured-data-extraction-m3-r5external write | suite=cheap; slug=structured-data-extraction; exact model/run; action=external write; fields=critical; idempotency=Unknown | external writes require valid fields plus side-effect containment | release=blocked; no automatic write | BLOCKED — external side effect |
batch51-cheap-model-tests-structured-data-extraction-m3-r6ambiguous record fixtures | suite=cheap; slug=structured-data-extraction; exact model/run; action=any; field ambiguity=present; review owner=joined | ambiguity cannot be normalized away | release=blocked pending human review; record retained | BLOCKED — ambiguity |
Provenance: Batch 51 cheap-model-tests-structured-data-extraction module 3; historical run 2026-06-16; audit verification 2026-09-01. First-party extraction test evidence (JSON). Missing or conflicting joins fail closed.
Verdict
For structured data on a budget, GPT-5.4 Nano took the top accuracy score (100/100) in this run. If raw throughput matters most, GLM 4.7 (Cerebras) was fastest at 562.9 tokens/sec, and Amazon Nova Micro was the cheapest to run. Every model here costs under $3 per million output tokens, so the real differentiator is how cleanly each one follows instructions — expand the rows above to judge for yourself.
