How Much Does Document Extraction Cost per Month?

At production volume (500,000 calls/month), the cheapest effective option is Amazon Nova Micro at $118/month. The most expensive frontier option, GPT-5.4 Pro, runs $113,400/month — Extraction reads a medium-sized document and returns a small, fixed-schema JSON object — the output is small and predictable by design.

How much does document extraction cost per month?

At production volume (500,000 calls/month), the cheapest effective option for document extraction is Amazon Nova Micro at $118 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $113,400 per month for the same workload.

Verified 2026-08-08
Cheapest ≠ best. This page ranks by effective cost only. See our graded results for this task on Best LLM For →

Token shape

Shapemedium document in, tiny structured JSON out
Input / output tokens per call6K in / 0K out
Cacheable input15%
Batch-eligibleYes

Input is one document per call (schema and instructions are the only stable, cacheable part); output is a compact JSON object matching a fixed schema. This runs at real volume, so batch discounts matter.

What drives this workload's cost?

The main token-volume driver here is input tokens: each call sends 6K input tokens and requests up to 0K output tokens, at 500,000 calls per month in the default volume. 15% of input is modeled as cache-eligible, so repeated prefixes can reduce the input charge. It is also batch-eligible, which can reduce the modeled total when asynchronous processing is acceptable.

Volume

Side project
50,000 calls/mo
$11.83/mo cheapest
Production
500,000 calls/mo
$118/mo cheapest
Scale
5,000,000 calls/mo
$1,183/mo cheapest

Ranked cost — Production volume (caching + batch applied where available)

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$123$59.150.76×
GPT-5 NanobudgetlegacyOpenAI$200$180
Amazon Nova LitebudgetAmazon$210$1040.91×
GPT-OSS 20BbudgetGroq$263$1672.93×
Gemini 2.5 Flash LitebudgetlegacyGoogle$350$3101
Ministral 8BbudgetMistral$469$2340.93×1
Mistral Small 3.1budgetMistral$525$2570.85×4
GPT-4o MinibudgetlegacyOpenAI$525$464
Grok-3 MinibudgetlegacyxAI$525$525
GPT-OSS 120BbudgetGroq$525$2931.82×
Muse Spark 1.3 ContributorbudgetMeta$325$62412.95×6
Llama 4 MaverickbudgetlegacyGroq$675$337
GPT-5.4 NanobudgetlegacyOpenAI$756$6410.78×
Gemini 3.1 Flash LitebudgetlegacyGoogle$938$8120.87×
CodestralbudgetMistral$1,013$4940.79×1
Show all 68 models
GPT-5 MinibudgetlegacyOpenAI$1,000$8991
Gemini 3.5 Flash LitebudgetGoogle$1,213$1,0911
Gemini 2.5 FlashbudgetlegacyGoogle$1,213$1,0911
GPT-OSS 120B (Cerebras)budgetCerebras$1,144$1,2682.32×2
Mistral Large 3budgetMistral$1,688$8441
DeepSeek V4 FlashbudgetDeepSeek$1,485$1,7472.59×1
GLM-5.1midlegacyZ.ai$2,075$2,075
Qwen 3.8 30BmidGroq$2,175$1,088
Qwen 3.6 27BmidlegacyGroq$2,175$1,088
Qwen 3.7 PlusmidQwen$2,650$2,650
Gemini 3.7 FlashmidGoogle$2,719$2,415
Amazon Nova PromidAmazon$2,800$1,400
GPT-5.4 MinimidlegacyOpenAI$2,813$2,510
Gemini 3.1 FlashmidlegacyGoogle$2,813$2,509
Claude Haiku 4.5midAnthropic$3,625$3,229
GPT-5.6 LunamidOpenAI$3,750$3,346
o3-MinimidlegacyOpenAI$3,850$3,406
Grok 4.3midxAI$4,063$4,1911.41×
Muse Spark 1.3midMeta$4,281$4,281
GPT-5midlegacyOpenAI$5,000$4,4962
Mistral Medium 3midMistral$5,438$2,6440.84×3
Gemini 3.6 FlashmidGoogle$5,438$4,8311
DeepSeek V4 PromidDeepSeek$4,455$5,5933.30×3
Qwen 3.8 MaxmidQwen$5,600$5,6001
Qwen 3.7 MaxmidQwen$5,600$5,6001
Gemini 3.5 FlashmidlegacyGoogle$5,625$5,0181
GLM-5.2midZ.ai$4,750$6,3293.87×6
Grok-3midlegacyxAI$6,500$6,500
Grok-4.20 ReasoningmidxAI$6,750$6,750
Grok-4.20midxAI$6,750$6,750
Grok 4.6midxAI$6,750$6,750
Grok 4.5midxAI$6,750$6,750
Gemini 3.1 PromidGoogle$7,500$6,1360.63×3
GPT-4.1midlegacyOpenAI$7,000$6,1931
Claude Sonnet 5midAnthropic$7,250$6,458
GPT-4omidlegacyOpenAI$8,750$7,7411
GLM 4.7 (Cerebras)midCerebras$7,094$9,3387.53×3
GPT-5.6 TerramidOpenAI$9,375$8,366
GPT-5.4midlegacyOpenAI$9,375$8,366
Claude Sonnet 4.6midAnthropic$10,875$9,687
Claude Sonnet 4.5midlegacyAnthropic$10,875$9,687
Claude Sonnet 4midlegacyAnthropic$10,875$9,687
GPT-5.6 SolmidOpenAI$14,500$12,886
Claude Opus 4.8midAnthropic$18,125$16,0200.96×
Claude Opus 4.7midlegacyAnthropic$18,125$16,145
Claude Opus 4.6midlegacyAnthropic$18,125$16,145
Claude Opus 4.5midlegacyAnthropic$18,125$16,145
GPT-4 TurbofrontierlegacyOpenAI$33,750$29,715
Claude Fable 5frontierAnthropic$36,250$32,289
Claude Opus 5frontierAnthropic$54,375$48,434
Claude Opus 4.1frontierlegacyAnthropic$54,375$48,434
Claude Opus 4frontierlegacyAnthropic$54,375$48,434
GPT-5.4 ProfrontierlegacyOpenAI$112,500$101,2951.04×

Batch 39 · server-rendered decision evidence · verified 2026-08-27

Document extraction cost normalization

Every field is tied to a frozen input and a dated provenance record. Unsupported facts fail closed as Unavailable; they are not treated as zero, free, equivalent, current, fastest, cheapest, private, or passing.

Page, image, credit, and token normalization ledger

Formula / rubric: Comparable $/1,000 pages requires same page mix + same pipeline stages + proven conversion from vendor unit to page.

Provenance: Frozen clean-text, table-heavy, scanned, and handwriting mixes; conversions are visible per row. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
clean text
batch39-document-extraction-m1-r1
1,000 pages; 650 tokens/page; text token rate $5/M650K input tokens → $3.25 model component.Comparable only when OCR/layout stage is also included on both sides.NORMALIZED — token units closed.
table-heavy scan
batch39-document-extraction-m1-r2
1,000 pages; 1 image/page; OCR $1.50/1K pages; model image conversionOCR $1.50; image token conversion Unavailable — provider image-token conversion is not returnedDo not equate one image with one token page.Unavailable — provider image-token conversion is not returned
handwriting mix
batch39-document-extraction-m1-r3
1,000 pages; 20% handwriting; specialist credit price $0.08/creditcredit-to-page ratio and review scope are Unavailable — not publishedNo specialist-versus-token crossover without conversion evidence.Unavailable — credit-to-page ratio and review scope are not published

Module citations: Google Gemini vision pricing. All AI Ask evidence registry (verified 2026-08-27).

Accepted-record cost tree

Formula / rubric: Accepted record = parse + extraction + validation + retry paths + review exceptions + rejected-document handling; user supplies pass/review rates.

Provenance: Invoice, form, receipt fixtures with one/two retry branches and manual exception sampling. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
invoice
batch39-document-extraction-m2-r1
parse $0.002/page; model $0.004; validation $0.0004; retry p=.12; review p=.08base $0.0064; expected accepted cost needs acceptance denominator.Use cost per accepted JSON, not cost per uploaded page.Unavailable — acceptance denominator is user-supplied
form / one retry
batch39-document-extraction-m2-r2
2 pages; 50 fields; first pass $0.009; retry $0.005; review $0.03best case $0.009; bounded path $0.044; pass and review rates not invented.Expected value waits for measured rates.Unavailable — pass and review rates are user-supplied
rejected document
batch39-document-extraction-m2-r3
corrupt scan; OCR failure; upload retry; human exception queuerejected handling charge and upload retry debit are Unavailable — not returnedInclude rejected work in monthly budget only after charge is observed.Unavailable — rejected handling charge and upload retry debit are not returned

Module citations: Anthropic vision documentation. All AI Ask evidence registry (verified 2026-08-27).

Schema-and-layout complexity surface

Formula / rubric: promptTokens = pageTokens + schemaTokens(fields) + instructions; crossover solves tokenAPI = specialistPerPage under identical stages.

Provenance: 10/50/200 fields × 1/5/20 pages; context fit and batch eligibility are evaluated separately. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
10 fields / 1 page
batch39-document-extraction-m3-r1
650 text tokens; schema 180; output 120; batch eligiblecontext 950 tokens; token API $0.0054/record under frozen rates.Token API wins if specialist price exceeds $0.0054 with same OCR stage.CROSSOVER — token estimate closed.
50 fields / 5 pages
batch39-document-extraction-m3-r2
3,250 text; schema 700; output 520; one image/pagecontext 4,470; batch eligibility Unavailable — image batch treatment is not returnedNo batch discount until image handling is documented.Unavailable — image batch treatment is not returned
200 fields / 20 pages
batch39-document-extraction-m3-r3
13K text; schema 2.8K; output 1.8K; 20 imagescontext 17.6K; specialist crossover Unavailable — per-page service and image conversion units are incompatibleDo not choose a cheaper path from incompatible units.Unavailable — per-page service and image conversion units are incompatible

Module citations: OpenAI structured outputs documentation. All AI Ask evidence registry (verified 2026-08-27).

Calculate cost per accepted record

Prompt caching uses the provider's sourced read multiplier and amortises cache writes over the documented TTL; the largest published read discount is 90%. See the full formula and coverage disclosure →

Related

Alternatives to Amazon Nova MicroAmazon cost calculator →LLM Chatbot cost →RAG Question Answering cost →

FAQ

Why is only a small share of input cacheable here?
Unlike a chatbot or RAG pipeline, each document is different — only the extraction instructions and schema repeat across calls.
Is this workload a good fit for batch processing?
Often yes — extraction jobs rarely need a synchronous reply, which makes batch APIs (typically 50% off) a straightforward win at this volume.