How Much Does Document Extraction Cost per Month?
At production volume (500,000 calls/month), the cheapest effective option is Amazon Nova Micro at $118/month. The most expensive frontier option, GPT-5.4 Pro, runs $113,400/month — Extraction reads a medium-sized document and returns a small, fixed-schema JSON object — the output is small and predictable by design.
How much does document extraction cost per month?
At production volume (500,000 calls/month), the cheapest effective option for document extraction is Amazon Nova Micro at $118 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $113,400 per month for the same workload.
Token shape
| Shape | medium document in, tiny structured JSON out |
| Input / output tokens per call | 6K in / 0K out |
| Cacheable input | 15% |
| Batch-eligible | Yes |
Input is one document per call (schema and instructions are the only stable, cacheable part); output is a compact JSON object matching a fixed schema. This runs at real volume, so batch discounts matter.
What drives this workload's cost?
The main token-volume driver here is input tokens: each call sends 6K input tokens and requests up to 0K output tokens, at 500,000 calls per month in the default volume. 15% of input is modeled as cache-eligible, so repeated prefixes can reduce the input charge. It is also batch-eligible, which can reduce the modeled total when asynchronous processing is acceptable.
Volume
Ranked cost — Production volume (caching + batch applied where available)
| Model | Provider | List monthly | Effective monthly | Verbosity | Rank Δ |
|---|---|---|---|---|---|
| Amazon Nova Microbudget | Amazon | $123 | $59.15 | 0.76× | — |
| GPT-5 Nanobudgetlegacy | OpenAI | $200 | $180 | — | — |
| Amazon Nova Litebudget | Amazon | $210 | $104 | 0.91× | — |
| GPT-OSS 20Bbudget | Groq | $263 | $167 | 2.93× | — |
| Gemini 2.5 Flash Litebudgetlegacy | $350 | $310 | — | ▲1 | |
| Ministral 8Bbudget | Mistral | $469 | $234 | 0.93× | ▲1 |
| Mistral Small 3.1budget | Mistral | $525 | $257 | 0.85× | ▲4 |
| GPT-4o Minibudgetlegacy | OpenAI | $525 | $464 | — | — |
| Grok-3 Minibudgetlegacy | xAI | $525 | $525 | — | — |
| GPT-OSS 120Bbudget | Groq | $525 | $293 | 1.82× | — |
| Muse Spark 1.3 Contributorbudget | Meta | $325 | $624 | 12.95× | ▼6 |
| Llama 4 Maverickbudgetlegacy | Groq | $675 | $337 | — | — |
| GPT-5.4 Nanobudgetlegacy | OpenAI | $756 | $641 | 0.78× | — |
| Gemini 3.1 Flash Litebudgetlegacy | $938 | $812 | 0.87× | — | |
| Codestralbudget | Mistral | $1,013 | $494 | 0.79× | ▲1 |
Show all 68 models
| GPT-5 Minibudgetlegacy | OpenAI | $1,000 | $899 | — | ▼1 |
| Gemini 3.5 Flash Litebudget | $1,213 | $1,091 | — | ▲1 | |
| Gemini 2.5 Flashbudgetlegacy | $1,213 | $1,091 | — | ▲1 | |
| GPT-OSS 120B (Cerebras)budget | Cerebras | $1,144 | $1,268 | 2.32× | ▼2 |
| Mistral Large 3budget | Mistral | $1,688 | $844 | — | ▲1 |
| DeepSeek V4 Flashbudget | DeepSeek | $1,485 | $1,747 | 2.59× | ▼1 |
| GLM-5.1midlegacy | Z.ai | $2,075 | $2,075 | — | — |
| Qwen 3.8 30Bmid | Groq | $2,175 | $1,088 | — | — |
| Qwen 3.6 27Bmidlegacy | Groq | $2,175 | $1,088 | — | — |
| Qwen 3.7 Plusmid | Qwen | $2,650 | $2,650 | — | — |
| Gemini 3.7 Flashmid | $2,719 | $2,415 | — | — | |
| Amazon Nova Promid | Amazon | $2,800 | $1,400 | — | — |
| GPT-5.4 Minimidlegacy | OpenAI | $2,813 | $2,510 | — | — |
| Gemini 3.1 Flashmidlegacy | $2,813 | $2,509 | — | — | |
| Claude Haiku 4.5mid | Anthropic | $3,625 | $3,229 | — | — |
| GPT-5.6 Lunamid | OpenAI | $3,750 | $3,346 | — | — |
| o3-Minimidlegacy | OpenAI | $3,850 | $3,406 | — | — |
| Grok 4.3mid | xAI | $4,063 | $4,191 | 1.41× | — |
| Muse Spark 1.3mid | Meta | $4,281 | $4,281 | — | — |
| GPT-5midlegacy | OpenAI | $5,000 | $4,496 | — | ▲2 |
| Mistral Medium 3mid | Mistral | $5,438 | $2,644 | 0.84× | ▲3 |
| Gemini 3.6 Flashmid | $5,438 | $4,831 | — | ▲1 | |
| DeepSeek V4 Promid | DeepSeek | $4,455 | $5,593 | 3.30× | ▼3 |
| Qwen 3.8 Maxmid | Qwen | $5,600 | $5,600 | — | ▲1 |
| Qwen 3.7 Maxmid | Qwen | $5,600 | $5,600 | — | ▲1 |
| Gemini 3.5 Flashmidlegacy | $5,625 | $5,018 | — | ▲1 | |
| GLM-5.2mid | Z.ai | $4,750 | $6,329 | 3.87× | ▼6 |
| Grok-3midlegacy | xAI | $6,500 | $6,500 | — | — |
| Grok-4.20 Reasoningmid | xAI | $6,750 | $6,750 | — | — |
| Grok-4.20mid | xAI | $6,750 | $6,750 | — | — |
| Grok 4.6mid | xAI | $6,750 | $6,750 | — | — |
| Grok 4.5mid | xAI | $6,750 | $6,750 | — | — |
| Gemini 3.1 Promid | $7,500 | $6,136 | 0.63× | ▲3 | |
| GPT-4.1midlegacy | OpenAI | $7,000 | $6,193 | — | ▼1 |
| Claude Sonnet 5mid | Anthropic | $7,250 | $6,458 | — | — |
| GPT-4omidlegacy | OpenAI | $8,750 | $7,741 | — | ▲1 |
| GLM 4.7 (Cerebras)mid | Cerebras | $7,094 | $9,338 | 7.53× | ▼3 |
| GPT-5.6 Terramid | OpenAI | $9,375 | $8,366 | — | — |
| GPT-5.4midlegacy | OpenAI | $9,375 | $8,366 | — | — |
| Claude Sonnet 4.6mid | Anthropic | $10,875 | $9,687 | — | — |
| Claude Sonnet 4.5midlegacy | Anthropic | $10,875 | $9,687 | — | — |
| Claude Sonnet 4midlegacy | Anthropic | $10,875 | $9,687 | — | — |
| GPT-5.6 Solmid | OpenAI | $14,500 | $12,886 | — | — |
| Claude Opus 4.8mid | Anthropic | $18,125 | $16,020 | 0.96× | — |
| Claude Opus 4.7midlegacy | Anthropic | $18,125 | $16,145 | — | — |
| Claude Opus 4.6midlegacy | Anthropic | $18,125 | $16,145 | — | — |
| Claude Opus 4.5midlegacy | Anthropic | $18,125 | $16,145 | — | — |
| GPT-4 Turbofrontierlegacy | OpenAI | $33,750 | $29,715 | — | — |
| Claude Fable 5frontier | Anthropic | $36,250 | $32,289 | — | — |
| Claude Opus 5frontier | Anthropic | $54,375 | $48,434 | — | — |
| Claude Opus 4.1frontierlegacy | Anthropic | $54,375 | $48,434 | — | — |
| Claude Opus 4frontierlegacy | Anthropic | $54,375 | $48,434 | — | — |
| GPT-5.4 Profrontierlegacy | OpenAI | $112,500 | $101,295 | 1.04× | — |
Batch 39 · server-rendered decision evidence · verified 2026-08-27
Document extraction cost normalization
Every field is tied to a frozen input and a dated provenance record. Unsupported facts fail closed as Unavailable; they are not treated as zero, free, equivalent, current, fastest, cheapest, private, or passing.
Page, image, credit, and token normalization ledger
Formula / rubric: Comparable $/1,000 pages requires same page mix + same pipeline stages + proven conversion from vendor unit to page.
Provenance: Frozen clean-text, table-heavy, scanned, and handwriting mixes; conversions are visible per row. Unsupported fields fail closed as Unavailable.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State / reproducible bill |
|---|---|---|---|---|
clean textbatch39-document-extraction-m1-r1 | 1,000 pages; 650 tokens/page; text token rate $5/M | 650K input tokens → $3.25 model component. | Comparable only when OCR/layout stage is also included on both sides. | NORMALIZED — token units closed. |
table-heavy scanbatch39-document-extraction-m1-r2 | 1,000 pages; 1 image/page; OCR $1.50/1K pages; model image conversion | OCR $1.50; image token conversion Unavailable — provider image-token conversion is not returned | Do not equate one image with one token page. | Unavailable — provider image-token conversion is not returned |
handwriting mixbatch39-document-extraction-m1-r3 | 1,000 pages; 20% handwriting; specialist credit price $0.08/credit | credit-to-page ratio and review scope are Unavailable — not published | No specialist-versus-token crossover without conversion evidence. | Unavailable — credit-to-page ratio and review scope are not published |
Module citations: Google Gemini vision pricing. All AI Ask evidence registry (verified 2026-08-27).
Accepted-record cost tree
Formula / rubric: Accepted record = parse + extraction + validation + retry paths + review exceptions + rejected-document handling; user supplies pass/review rates.
Provenance: Invoice, form, receipt fixtures with one/two retry branches and manual exception sampling. Unsupported fields fail closed as Unavailable.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State / reproducible bill |
|---|---|---|---|---|
invoicebatch39-document-extraction-m2-r1 | parse $0.002/page; model $0.004; validation $0.0004; retry p=.12; review p=.08 | base $0.0064; expected accepted cost needs acceptance denominator. | Use cost per accepted JSON, not cost per uploaded page. | Unavailable — acceptance denominator is user-supplied |
form / one retrybatch39-document-extraction-m2-r2 | 2 pages; 50 fields; first pass $0.009; retry $0.005; review $0.03 | best case $0.009; bounded path $0.044; pass and review rates not invented. | Expected value waits for measured rates. | Unavailable — pass and review rates are user-supplied |
rejected documentbatch39-document-extraction-m2-r3 | corrupt scan; OCR failure; upload retry; human exception queue | rejected handling charge and upload retry debit are Unavailable — not returned | Include rejected work in monthly budget only after charge is observed. | Unavailable — rejected handling charge and upload retry debit are not returned |
Module citations: Anthropic vision documentation. All AI Ask evidence registry (verified 2026-08-27).
Schema-and-layout complexity surface
Formula / rubric: promptTokens = pageTokens + schemaTokens(fields) + instructions; crossover solves tokenAPI = specialistPerPage under identical stages.
Provenance: 10/50/200 fields × 1/5/20 pages; context fit and batch eligibility are evaluated separately. Unsupported fields fail closed as Unavailable.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State / reproducible bill |
|---|---|---|---|---|
10 fields / 1 pagebatch39-document-extraction-m3-r1 | 650 text tokens; schema 180; output 120; batch eligible | context 950 tokens; token API $0.0054/record under frozen rates. | Token API wins if specialist price exceeds $0.0054 with same OCR stage. | CROSSOVER — token estimate closed. |
50 fields / 5 pagesbatch39-document-extraction-m3-r2 | 3,250 text; schema 700; output 520; one image/page | context 4,470; batch eligibility Unavailable — image batch treatment is not returned | No batch discount until image handling is documented. | Unavailable — image batch treatment is not returned |
200 fields / 20 pagesbatch39-document-extraction-m3-r3 | 13K text; schema 2.8K; output 1.8K; 20 images | context 17.6K; specialist crossover Unavailable — per-page service and image conversion units are incompatible | Do not choose a cheaper path from incompatible units. | Unavailable — per-page service and image conversion units are incompatible |
Module citations: OpenAI structured outputs documentation. All AI Ask evidence registry (verified 2026-08-27).
Prompt caching uses the provider's sourced read multiplier and amortises cache writes over the documented TTL; the largest published read discount is 90%. See the full formula and coverage disclosure →
