← All cheap-model tests
Structured Data · Models under $3/M tokens · Historical run Jun 16, 2026 · Audit verified 2026-09-01

Extracting Structured Data: Which Cheap AI Model Wins?

A structured-output task: read one sentence and return strict JSON with a string name, numeric price, and boolean stock flag — no markdown, no prose. We sent this exact task to all 20 budget models (each under $3 per million output tokens) through the live API and measured real speed, cost, and accuracy. The full output of every model is included below — click any row to read it.

Most accurate
GPT-5.4 Nano
100/100 · OpenAI
Fastest
GLM 4.7 (Cerebras)
562.9 t/s · 723 ms
Cheapest run
Amazon Nova Micro
$0.0000074

The prompt

Every model received exactly this, with default generation settings:

Extract the product details from the following sentence and return ONLY a valid JSON object with exactly these keys: "name" (string), "price" (number), "in_stock" (boolean). Do not include any explanation or markdown. Sentence: "The DeskPro X1 monitor stand normally retails for $49.99 and is currently available to ship."

How accuracy was graded

  • Output is valid, parseable JSON.
  • Extracts the full product name ("DeskPro X1 monitor stand").
  • Correct types: price as a number (49.99), in_stock as a boolean (true).
  • Returns only the JSON — no markdown fences, prose, or leaked reasoning.

Results — sortable by speed, cost & accuracy

Click Accuracy, Speed, or Cost to re-sort. Click any row to read the full model output.
Per-model results: accuracy, speed (tokens/sec), provider cost (USD), and the full prompt-vs-output for each model on this task.
#ModelAccuracy Speed Cost Output
Agent grade — 100/100. Perfect: compact valid JSON, full product name, correct number and boolean types, no markdown or prose.
Verbatim output (78 in / 24 out tokens, 488 ms, $0.0000456)
{"name":"DeskPro X1 monitor stand","price":49.99,"in_stock":true}
Agent grade — 100/100. Valid, compact JSON; full product name; correct number and boolean types; no markdown or prose. Same reasoning tax as the other two tasks — most of the 767 output tokens were hidden chain-of-thought.
Verbatim output (79 in / 767 out tokens, 3704 ms, $0.000161)
{"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true}
Agent grade — 99/100. Valid JSON, full name, correct types, no markdown — just pretty-printed rather than compact.
Verbatim output (75 in / 36 out tokens, 477 ms, $0.0000727)
{
  "name": "DeskPro X1 monitor stand",
  "price": 49.99,
  "in_stock": true
}
Agent grade — 99/100. Valid compact JSON, full name, correct types, no markdown or prose. Fastest run in this test.
Verbatim output (107 in / 26 out tokens, 172 ms, $0.0000837)
{"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true}
Agent grade — 99/100. Valid compact JSON, full name, correct types, no markdown — and the cheapest run here.
Verbatim output (107 in / 26 out tokens, 296 ms, $0.0000074)
{"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true}
Agent grade — 99/100. Valid compact JSON, full name, correct types, no markdown or leakage.
Verbatim output (91 in / 29 out tokens, 808 ms, $0.0000944)
{"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true}
Agent grade — 99/100. Valid compact JSON, full name, correct types — identical quality to Medium for far less.
Verbatim output (91 in / 29 out tokens, 530 ms, $0.000031)
{"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true}
Agent grade — 99/100. Valid compact JSON, full name, correct types, no markdown or prose.
Verbatim output (79 in / 29 out tokens, 470 ms, $0.0000498)
{"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true}
Agent grade — 99/100. Valid JSON, full name, correct types, no markdown or prose — only pretty-printed rather than compact.
Verbatim output (76 in / 34 out tokens, 380 ms, $0.0000074)
{
  "name": "DeskPro X1 monitor stand",
  "price": 49.99,
  "in_stock": true
}
Agent grade — 85/100. Correct values and full name, but wrapped the JSON in a ```json markdown fence after being told "no markdown".
Verbatim output (82 in / 35 out tokens, 350 ms, $0.0000209)
```json
{
  "name": "DeskPro X1 monitor stand",
  "price": 49.99,
  "in_stock": true
}
```
Agent grade — 85/100. Correct values and full name, but wrapped the output in a ```json markdown fence against instructions.
Verbatim output (79 in / 38 out tokens, 599 ms, $0.0000175)
```json
{
  "name": "DeskPro X1 monitor stand",
  "price": 49.99,
  "in_stock": true
}
```
Agent grade — 85/100. Correct values and full name, but wrapped the JSON in a ```json markdown fence after being told "no markdown".
Verbatim output (76 in / 38 out tokens, 340 ms, $0.0000137)
```json
{
  "name": "DeskPro X1 monitor stand",
  "price": 49.99,
  "in_stock": true
}
```
Agent grade — 82/100. Correct JSON with full name and proper types; docked only for the leaked <think> preamble.
Verbatim output (79 in / 123 out tokens, 2033 ms, $0.0000455)
<think>
We need to extract product details: name, price (number), in_stock. Sentence: "The DeskPro X1 monitor stand normally retails for $49.99 and is currently available to ship."

Name: "DeskPro X1 monitor stand" (should be string)
Price: $49.99 -> 49.99 as number
In_stock: "currently available to ship" implies it's in stock, so true.

Return JSON only.
</think>
{"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true}
Agent grade — 82/100. Correct, complete JSON; the leaked <think> block is the only issue.
Verbatim output (79 in / 179 out tokens, 2104 ms, $0.00019)
<think>
We are asked to extract product details from the sentence: "The DeskPro X1 monitor stand normally retails for $49.99 and is currently available to ship."

We need to return ONLY a valid JSON object with keys: "name" (string), "price" (number), "in_stock" (boolean).

From the sentence:
- Name: "DeskPro X1 monitor stand"
- Price: $49.99 -> number 49.99
- In stock: "currently available to ship" implies it is in stock, so true.

So JSON: {"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true}

Make sure no extra text.
</think>
{"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true}
Agent grade — 80/100. Valid JSON with the full name and correct types, but a leaked <think> block violates "return ONLY a valid JSON object".
Verbatim output (81 in / 74 out tokens, 2819 ms, $0.000286)
<think>
The task is: Extract the product details from the sentence and return ONLY a valid JSON object with exactly these keys: "name" (string), "price" (number), "in_stock" (boolean). Do not include any expl...
</think>
{
"name": "DeskPro X1 monitor stand",
"price": 49.99,
"in_stock": true
}
Agent grade — 80/100. Final JSON is correct and complete, but a <think> block leaked in despite the "no explanation" instruction.
Verbatim output (143 in / 81 out tokens, 277 ms, $0.00007)
<think>
We need to output JSON with keys name, price, in_stock. Extract details: name: "DeskPro X1 monitor stand". price: 49.99. in_stock: true (available to ship). Return only JSON, no markdown.
</think>
{"name":"DeskPro X1 monitor stand","price":49.99,"in_stock":true}
Agent grade — 80/100. Correct JSON with the full name, but preceded by a leaked <think> reasoning block.
Verbatim output (143 in / 164 out tokens, 343 ms, $0.0000599)
<think>
The user wants to extract product details: name: "DeskPro X1 monitor stand", price: 49.99, in_stock: true. The sentence: "The DeskPro X1 monitor stand normally retails for $49.99 and is currently available to ship." The name: "DeskPro X1 monitor stand". Price: 49.99. In_stock: currently available to ship => in stock: true. JSON with keys: name, price, in_stock.

Return only JSON object, no explanation, no markdown. So just:

{"name":"DeskPro X1 monitor stand","price":49.99,"in_stock":true}
</think>
{"name":"DeskPro X1 monitor stand","price":49.99,"in_stock":true}
Agent grade — 75/100. Final JSON is correct, but buried under a ~340-token leaked <think> monologue — the opposite of "JSON only".
Verbatim output (82 in / 336 out tokens, 784 ms, $0.000222)
<think>
Okay, let me try to figure this out. The user wants me to extract product details from a sentence and put them into aJSON object with specific keys: name, price, and in_stock. The sentence given is "The DeskPro X1 monitor stand normally retails for $49.99 and is currently available to ship." 

First, I need to identify the product name. The sentence starts with "The DeskPro X1 monitor stand..." so the name is probably "DeskPro X1 monitor stand". I should check if there's any other part of the name, but "DeskPro X1" seems like the model, and "monitor stand" is the product type. So the name key should be that full phrase.

Next, the price is mentioned as "$49.99". The key requires a number, so I need to remove the dollar sign and convert it to a numerical value. That would be 49.99. 

For the in_stock boolean, the sentence says "is currently available to ship". If it's available to ship, that likely means it's in stock. So in_stock should be true. 

I need to make sure there are only those three keys and that the JSON is valid. Let me double-check each part. The name is correct, price is a number without the currency symbol, and in_stock is a boolean. No other keys are needed. Alright, putting that all together into a JSON object with the exact keys.
</think>

{"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true}
Agent grade — 73/100. The final JSON is correct, but the model exposed a ~400-token multi-step <think> process despite "no explanation".
Verbatim output (77 in / 407 out tokens, 723 ms, $0.001292)
<think>
1.  **Analyze the Request:**
    *   Input sentence: "The DeskPro X1 monitor stand normally retails for $49.99 and is currently available to ship."
    *   Task: Extract product details.
    *   Output format: Valid JSON object only.
    *   Required keys: "name" (string), "price" (number), "in_stock" (boolean).
    *   Constraint: No explanation, no markdown code blocks (just the raw JSON).

2.  **Analyze the Sentence to Extract Data:**
    *   *Name:* "The DeskPro X1 monitor stand". I should probably clean this up slightly to just the product name, or keep it as is. "DeskPro X1 monitor stand" seems most appropriate. Let's stick to the core product name.
    *   *Price:* "$49.99". The key requires a number. So, 49.99.
    *   *Stock Status:* "currently available to ship". This implies it is in stock. So, true.

3.  **Construct the JSON Object:**
    *   `name`: "DeskPro X1 monitor stand"
    *   `price`: 49.99
    *   `in_stock`: true

4.  **Verify against Constraints:**
    *   Valid JSON? Yes.
    *   Keys present? "name", "price", "in_stock". Yes.
    *   Types correct? String, Number, Boolean. Yes.
    *   No explanation? Yes.
    *   No markdown? Yes (I need to ensure I don't wrap it in backticks).

5.  **Final Output Generation:**
    {"name": "DeskPro X1 monitor stand", "price": 49.99, "in_stock": true}
</think>
{
  "name": "DeskPro X1 monitor stand",
  "price": 49.99,
  "in_stock": true
}
Agent grade — 72/100. Leaked a <think> block AND kept the leading "The" in the name ("The DeskPro X1 monitor stand"), so the extraction is slightly off.
Verbatim output (139 in / 92 out tokens, 165 ms, $0.000118)
<think>
We need to extract product details: name, price, in_stock. Name is "The DeskPro X1 monitor stand". Price is 49.99. in_stock: currently available to ship => true. Return JSON only, no markdown. Ensure keys exact. Provide number, boolean, string.
</think>
{"name":"The DeskPro X1 monitor stand","price":49.99,"in_stock":true}

Batch 51 · cheap-model-tests-structured-data-extraction evidence contributions. Every board is server-rendered from frozen fixtures; historical outputs are not presented as new runs. Verification date: 2026-09-01.

Deterministic test receipts

These values are read from the exported test records. Prompt hashes identify the immutable test input; raw-output hashes identify each verbatim model response. Version and validation fields are recorded metadata, not a new run or a re-grade.

Test / modelPrompt SHA-256Raw output SHA-256GraderSolver / harnessValidation flags
structured-data-extraction
gpt-5.4-nano
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2ebe74112168afc03369b78be17b508792711cb92a84e9c97e48379182dac8206claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
gemini-3.1-flash-lite
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f23826fa83b2b396c1d710eb5fa27351deab82b5354f768a9743ac558b33a096feclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
grok-4.3
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2023d73fe1562655765d39c7d7a1306a35666d00426c2d6accfb71cc38ceec5b5claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present
structured-data-extraction
llama-4-scout
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f26755f288b404d7655920b1a485c24cf630a3035286e1fb3f6b6178cd6d6a4761claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
gpt-oss-120b
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f22e3929c7452e4f92273386be003daae6fb8e13e8ea98598cda2684cdd132afb5claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present
structured-data-extraction
gpt-oss-20b
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f25c48aa19881b50afbbe4df83e7044701da97b6cc8c7715095aa149ae86710a4dclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present
structured-data-extraction
llama-3.3-70b
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397fclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
llama-3.1-8b
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397fclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
qwen3-32b
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2cb6eb2d20f562af88a59cbdd63c0b0468efda7ef969d378038ddef1d65697701claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present
structured-data-extraction
deepseek-v4-flash
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f25b97436361cd389f75d422e6939e7ff9afca4be0a5504c8d9979ab4f59315454claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present
structured-data-extraction
deepseek-v4-pro
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2d7ceebe74e6c44e5be3735489299cbe9e45cae2554eb1d88198d02bcc3dc9adfclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present
structured-data-extraction
mistral-medium
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397fclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
mistral-small
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397fclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
ministral-8b
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f26755f288b404d7655920b1a485c24cf630a3035286e1fb3f6b6178cd6d6a4761claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
codestral
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397fclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
cerebras-gpt-oss-120b
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2ba31906fe16dbdf92c8823917e422918f4484f82a3ddb724c37831f9ba6f0937claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present
structured-data-extraction
cerebras-glm-4.7
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2404c792a7901808747ed616173a4b57c2d04e7587705ffb6dd4866234417d85aclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=detected, json-object-candidate=present
structured-data-extraction
nova-lite
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f26755f288b404d7655920b1a485c24cf630a3035286e1fb3f6b6178cd6d6a4761claude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
nova-micro
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f23826fa83b2b396c1d710eb5fa27351deab82b5354f768a9743ac558b33a096feclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present
structured-data-extraction
muse-spark-1.3-contributor
29e380a6f065c08b6e549dc6236ab8446addb06036bc710f1c830bce3a6559f2fd7235933ea0af34f4c177a06a5e952cde2f4234c9a527e41f1ae8b615da397fclaude-agent-grader-v1Not applicablerecorded · prompt-hash=sha256, raw-output-hash=sha256, status=completed, accuracy=range-valid, latency=recorded, tokens=recorded, cost=recorded, grader-score=recorded, reasoning-leak=absent, json-object-candidate=present

Raw-output JSON and schema validation receipt

Frozen Batch 51 fixture board. Formula / decision rule: strict acceptance = original output parses + schema/type/duplicate/extra-field policy passes Boundary: Outputs are validated as received; repair or cleanup is never silently substituted for validity.

Frozen fixture / field IDIdentity keysDeterministic ruleOutput / bounded stateValidation
batch51-cheap-model-tests-structured-data-extraction-m1-r1
valid object
suite=cheap; slug=structured-data-extraction; exact model/run; raw hash=joined; source hash=joined; schema=v1parse original object and compare to schema v1strict acceptance=Unavailable until raw parse receipt joinsUNAVAILABLE — parse receipt
batch51-cheap-model-tests-structured-data-extraction-m1-r2
fenced JSON
suite=cheap; slug=structured-data-extraction; exact model/run; raw hash=joined; extraction rule=code-fence unwrap; schema=v1record wrapper removal as extraction, not original validityoriginal validity=fail or conditional; strict state=UnavailableCONDITIONAL — wrapper recorded
batch51-cheap-model-tests-structured-data-extraction-m1-r3
prose-prefixed JSON
suite=cheap; slug=structured-data-extraction; exact model/run; raw hash=joined; prefix=present; schema=v1prose prefix violates strict answer envelopeparse may be recoverable, strict acceptance=failFAIL — prose prefix
batch51-cheap-model-tests-structured-data-extraction-m1-r4
duplicate key
suite=cheap; slug=structured-data-extraction; exact model/run; raw hash=joined; duplicate key=present; schema=v1duplicate keys are reported before parser normalizationstrict acceptance=fail; preserve key collisionFAIL — duplicate key
batch51-cheap-model-tests-structured-data-extraction-m1-r5
wrong type
suite=cheap; slug=structured-data-extraction; exact model/run; raw hash=joined; field type=wrong; schema=v1type validation is field-level and strictstrict acceptance=fail; type error=retainedFAIL — type error
batch51-cheap-model-tests-structured-data-extraction-m1-r6
malformed JSON fixtures
suite=cheap; slug=structured-data-extraction; exact model/run; raw hash=joined; parser=error; schema=v1malformed original output cannot be accepted by repairstrict acceptance=fail; original parse=invalidFAIL — malformed

Provenance: Batch 51 cheap-model-tests-structured-data-extraction module 1; historical run 2026-06-16; audit verification 2026-09-01. First-party extraction test evidence (JSON). Missing or conflicting joins fail closed.

Field-level semantic diff ledger

Frozen Batch 51 fixture board. Formula / decision rule: field match = explicit source span + typed comparison + allowed normalization; ambiguity never becomes a match Boundary: Every accepted non-null field needs an explicit source span in the frozen source record.

Frozen fixture / field IDIdentity keysDeterministic ruleOutput / bounded stateValidation
batch51-cheap-model-tests-structured-data-extraction-m2-r1
exact string
suite=cheap; slug=structured-data-extraction; exact model/run; source hash/schema hash=joined; field=string; source span=joinedraw value equals typed source valuesemantic match=Yes; normalized value=source valueMATCH — explicit source
batch51-cheap-model-tests-structured-data-extraction-m2-r2
case/whitespace normalization
suite=cheap; slug=structured-data-extraction; exact model/run; field=string; raw differs by case/space; source span=joinedonly declared normalization may be appliedsemantic match=Conditional; normalization recordedCONDITIONAL — normalization
batch51-cheap-model-tests-structured-data-extraction-m2-r3
numeric string
suite=cheap; slug=structured-data-extraction; exact model/run; field=numeric; raw="42"; expected=42; source span=joinednumeric coercion is allowed only when schema policy declares itnormalized value=42 if policy joined; otherwise UnavailableCONDITIONAL — type policy
batch51-cheap-model-tests-structured-data-extraction-m2-r4
missing nullable field
suite=cheap; slug=structured-data-extraction; exact model/run; field=nullable; raw=missing; source says=null; source span=joinedmissing and explicit null are distinct unless schema allows omissionsemantic match=Conditional; omission rule requiredCONDITIONAL — nullable rule
batch51-cheap-model-tests-structured-data-extraction-m2-r5
hallucinated value
suite=cheap; slug=structured-data-extraction; exact model/run; field=non-null; raw value=not in source; source span=nonenon-null acceptance requires a source spansemantic match=No; downstream release blockedFAIL — unsupported value
batch51-cheap-model-tests-structured-data-extraction-m2-r6
conflicting-source fixtures
suite=cheap; slug=structured-data-extraction; exact model/run; field=value; source spans=conflict; schema hash=joinedconflicting sources fail closed until owner resolves precedencesemantic=Unresolved; retain both source spansUNRESOLVED — source conflict

Provenance: Batch 51 cheap-model-tests-structured-data-extraction module 2; historical run 2026-06-16; audit verification 2026-09-01. First-party extraction test evidence (JSON). Missing or conflicting joins fail closed.

Downstream blast-radius and release gate

Frozen Batch 51 fixture board. Formula / decision rule: release = strict/semantic validation pass for every decision-critical field + action-specific review policy Boundary: External side effects are blocked when a decision-critical field is missing, invalid, ambiguous, or unsupported.

Frozen fixture / field IDIdentity keysDeterministic ruleOutput / bounded stateValidation
batch51-cheap-model-tests-structured-data-extraction-m3-r1
analytics-only field
suite=cheap; slug=structured-data-extraction; exact model/run; action=analytics; fields=non-critical; validation=joinedanalytics may proceed only under the declared non-critical policyrelease=Conditional; retain provenance for downstream usersCONDITIONAL — low blast radius
batch51-cheap-model-tests-structured-data-extraction-m3-r2
billing total
suite=cheap; slug=structured-data-extraction; exact model/run; action=billing; total field=decision-critical; strict validation=Unknownbilling requires strict typed validation and source spanrelease=blocked until strict validation passesBLOCKED — financial side effect
batch51-cheap-model-tests-structured-data-extraction-m3-r3
customer routing
suite=cheap; slug=structured-data-extraction; exact model/run; action=routing; route field=critical; semantic=joinedrouting requires field-level semantic match and fallback ownerrelease=Conditional; human review if ambiguity remainsCONDITIONAL — routing gate
batch51-cheap-model-tests-structured-data-extraction-m3-r4
compliance flag
suite=cheap; slug=structured-data-extraction; exact model/run; action=compliance; flag=critical; source precedence=Unknownunknown precedence blocks compliance decisionrelease=blocked; human/compliance owner requiredBLOCKED — compliance gap
batch51-cheap-model-tests-structured-data-extraction-m3-r5
external write
suite=cheap; slug=structured-data-extraction; exact model/run; action=external write; fields=critical; idempotency=Unknownexternal writes require valid fields plus side-effect containmentrelease=blocked; no automatic writeBLOCKED — external side effect
batch51-cheap-model-tests-structured-data-extraction-m3-r6
ambiguous record fixtures
suite=cheap; slug=structured-data-extraction; exact model/run; action=any; field ambiguity=present; review owner=joinedambiguity cannot be normalized awayrelease=blocked pending human review; record retainedBLOCKED — ambiguity

Provenance: Batch 51 cheap-model-tests-structured-data-extraction module 3; historical run 2026-06-16; audit verification 2026-09-01. First-party extraction test evidence (JSON). Missing or conflicting joins fail closed.

Run the cheap-model-tests-structured-data-extraction Batch 51 evidence scenario →

Verdict

For structured data on a budget, GPT-5.4 Nano took the top accuracy score (100/100) in this run. If raw throughput matters most, GLM 4.7 (Cerebras) was fastest at 562.9 tokens/sec, and Amazon Nova Micro was the cheapest to run. Every model here costs under $3 per million output tokens, so the real differentiator is how cleanly each one follows instructions — expand the rows above to judge for yourself.

More cheap-model tests

Code Generation
Writing a Code Snippet
View results →
Copywriting
Writing a Short Paragraph
View results →

Test it with your own prompt

Run all these cheap models side by side on the task that actually matters to you.

Open the live playground →