How Much Does Long-Document Summarization Cost per Month?

At production volume (8,000 calls/month), the cheapest effective option is Amazon Nova Micro at $26.22/month. The most expensive frontier option, GPT-5.4 Pro, runs $23,397/month — Summarization is the most input-heavy shape in this cluster: a large source document in, a medium-length summary out.

How much does long-document summarization cost per month?

At production volume (8,000 calls/month), the cheapest effective option for long-document summarization is Amazon Nova Micro at $26.22 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $23,397 per month for the same workload.

Verified 2026-08-08
Cheapest ≠ best. This page ranks by effective cost only. See our graded results for this task on Best LLM For →

Token shape

Shapevery large document in, medium summary out
Input / output tokens per call90K in / 1K out
Cacheable input10%
Batch-eligibleYes

Input is a full long-form document (report, transcript, contract); output is a structured summary. Each document is usually read once, so little of the input is cache-eligible.

What drives this workload's cost?

The main token-volume driver here is input tokens: each call sends 90K input tokens and requests up to 1K output tokens, at 8,000 calls per month in the default volume. 10% of input is modeled as cache-eligible, so repeated prefixes can reduce the input charge. It is also batch-eligible, which can reduce the modeled total when asynchronous processing is acceptable.

Volume

Side project
800 calls/mo
$2.62/mo cheapest
Production
8,000 calls/mo
$26.22/mo cheapest
Scale
80,000 calls/mo
$262/mo cheapest

Ranked cost — Production volume (caching + batch applied where available)

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$26.54$13.110.76×
GPT-5 NanobudgetlegacyOpenAI$39.84$37.35
Amazon Nova LitebudgetAmazon$45.50$22.650.91×
GPT-OSS 20BbudgetGroq$56.88$31.222.93×
Gemini 2.5 Flash LitebudgetlegacyGoogle$75.84$69.941
Muse Spark 1.3 ContributorbudgetMeta$73.92$96.8612.95×1
Ministral 8BbudgetMistral$109$54.670.93×
Mistral Small 3.1budgetMistral$114$56.450.85×3
GPT-4o MinibudgetlegacyOpenAI$114$1061
Grok-3 MinibudgetlegacyxAI$114$1141
GPT-OSS 120BbudgetGroq$114$59.241.82×1
Llama 4 MaverickbudgetlegacyGroq$150$74.88
GPT-5.4 NanobudgetlegacyOpenAI$156$1430.78×
Gemini 3.1 Flash LitebudgetlegacyGoogle$194$1780.87×
GPT-5 MinibudgetlegacyOpenAI$199$187
Show all 68 models
CodestralbudgetMistral$225$1110.79×
Gemini 3.5 Flash LitebudgetGoogle$240$222
Gemini 2.5 FlashbudgetlegacyGoogle$240$222
GPT-OSS 120B (Cerebras)budgetCerebras$259$2692.32×
DeepSeek V4 FlashbudgetDeepSeek$329$3502.59×
Mistral Large 3budgetMistral$374$187
GLM-5.1midlegacyZ.ai$453$453
Qwen 3.8 30BmidGroq$461$230
Qwen 3.6 27BmidlegacyGroq$461$230
Gemini 3.7 FlashmidGoogle$576$532
GPT-5.4 MinimidlegacyOpenAI$583$546
Gemini 3.1 FlashmidlegacyGoogle$583$539
Qwen 3.7 PlusmidQwen$595$595
Amazon Nova PromidAmazon$607$303
Claude Haiku 4.5midAnthropic$768$786
GPT-5.6 LunamidOpenAI$778$728
o3-MinimidlegacyOpenAI$834$779
Grok 4.3midxAI$924$9341.41×
Muse Spark 1.3midMeta$941$941
GPT-5midlegacyOpenAI$996$9341
DeepSeek V4 PromidDeepSeek$988$1,0763.30×1
Mistral Medium 3midMistral$1,152$5700.84×2
Gemini 3.6 FlashmidGoogle$1,152$1,064
Gemini 3.5 FlashmidlegacyGoogle$1,166$1,0781
GLM-5.2midZ.ai$1,050$1,1713.87×3
Qwen 3.8 MaxmidQwen$1,213$1,213
Qwen 3.7 MaxmidQwen$1,213$1,213
Grok-3midlegacyxAI$1,478$1,478
Grok-4.20 ReasoningmidxAI$1,498$1,498
Grok-4.20midxAI$1,498$1,498
Grok 4.6midxAI$1,498$1,498
Grok 4.5midxAI$1,498$1,498
Gemini 3.1 PromidGoogle$1,555$1,3950.63×2
GPT-4.1midlegacyOpenAI$1,517$1,4171
Claude Sonnet 5midAnthropic$1,536$1,5721
GLM 4.7 (Cerebras)midCerebras$1,646$1,8197.53×
GPT-4omidlegacyOpenAI$1,896$1,771
GPT-5.6 TerramidOpenAI$1,944$1,819
GPT-5.4midlegacyOpenAI$1,944$1,819
Claude Sonnet 4.6midAnthropic$2,304$2,358
Claude Sonnet 4.5midlegacyAnthropic$2,304$2,358
Claude Sonnet 4midlegacyAnthropic$2,304$2,358
GPT-5.6 SolmidOpenAI$3,072$2,872
Claude Opus 4.8midAnthropic$3,840$3,9200.96×
Claude Opus 4.7midlegacyAnthropic$3,840$3,930
Claude Opus 4.6midlegacyAnthropic$3,840$3,930
Claude Opus 4.5midlegacyAnthropic$3,840$3,930
GPT-4 TurbofrontierlegacyOpenAI$7,488$6,989
Claude Fable 5frontierAnthropic$7,680$7,860
Claude Opus 5frontierAnthropic$11,520$11,790
Claude Opus 4.1frontierlegacyAnthropic$11,520$11,790
Claude Opus 4frontierlegacyAnthropic$11,520$11,790
GPT-5.4 ProfrontierlegacyOpenAI$23,328$21,9001.04×

Batch 40 · server-rendered decision evidence · verified 2026-08-27

Long-document summarization cost evidence

Frozen inputs, formulas, provenance, and decision boundaries are visible in the initial HTML. Unsupported evidence fails closed as Unavailable.

Topology-expanded token ledger

Formula / scoring rule: Monthly bill = calls × (input tokens × $/M + output tokens × $/M); chunk overlap and intermediate summaries are counted.

Provenance: Frozen 90K-input / 1.2K-output document, 8K context, $5/M input and $15/M output illustrative rates.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
one-pass
batch40-long-document-summarization-m1-r1
1 document; 90K input; 1,200 output; 8K contextUnavailable — 90K input exceeds the 8K context windowReject before cost math; do not silently truncate.Unavailable — context eligibility fails
map-reduce
batch40-long-document-summarization-m1-r2
12 chunks × 8K + 10% overlap; 12 × 600 intermediate; 1,200 finalInput 105,600; output 8,400; bill = (105,600×5 + 8,400×15)/1M = $1.182.Include overlap and synthesis tokens in every topology.ELIGIBLE — calculated token ledger.
rolling refine
batch40-long-document-summarization-m1-r3
23 serial chunks; 8K chunk; 600 retained summary each passInput 182,400; output 14,400; bill = $1.128; serial depth 23.Latency/concurrency tradeoff remains separate from bill.ELIGIBLE — calculated token ledger.

Module citation: All AI Ask evidence registry (verified 2026-08-27).

Portfolio percentile planner

Formula / scoring rule: Budget = Σ row-level routed bill; mean shortcut = document count × mean-token bill and may hide P99 exclusions.

Provenance: Frozen 1,000-document manifest: P50 9K, P90 42K, P99 180K; scanned share 8%.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
P50 / 500 docs
batch40-long-document-summarization-m2-r1
500 × 9K tokens; map path; 1 retry per 50 docs4.5M source tokens; retry count 10; source bill Unavailable — model rate selected by userDo not convert unavailable rate into zero.Unavailable — model rate selected by user
P90 / 400 docs
batch40-long-document-summarization-m2-r2
400 × 42K; split at 8K; 8% scanned pages2,100 chunks; OCR charge Unavailable — OCR price per page is not suppliedKeep OCR/storage outside token subtotal until priced.Unavailable — OCR price per page is not supplied
P99 / 100 docs
batch40-long-document-summarization-m2-r3
100 × 180K; 100 over-window; split required; retry set 7All 100 route to split topology; successful count Unavailable — post-retry acceptance is not observedNo acceptance rate is inferred from routing.Unavailable — post-retry acceptance is not observed

Module citation: All AI Ask workload registry.

Accepted-summary TCO tree

Formula / scoring rule: Cost/accepted = (API + OCR + checks + retries + review) / accepted summaries; denominator is not processed documents.

Provenance: 8K-document baseline; review is explicitly user-supplied at 6 minutes and $60/hour.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
API + checks
batch40-long-document-summarization-m3-r1
8,000 docs; API subtotal $1,182; factual check $0.002/docAPI + checks = $1,198; 8,000 processed denominator.Accepted denominator still needs measured pass rate.Unavailable — accepted-summary count is not observed
human review
batch40-long-document-summarization-m3-r2
12% review; 6 min/doc; $60/hourExpected review labor = 960 × 0.1 × $60 = $5,760.Labor is an explicit scenario, not an observed rate.CALCULATED — user inputs visible.
OCR placeholder
batch40-long-document-summarization-m3-r3
8% scanned share; 640 docs; page count and OCR tariff missingUnavailable — OCR pages and price are not suppliedDo not publish end-to-end TCO without OCR units.Unavailable — OCR pages and price are not supplied

Module citation: All AI Ask summarization evidence.

Calculate your accepted-summary cost

Prompt caching uses the provider's sourced read multiplier and amortises cache writes over the documented TTL; the largest published read discount is 90%. See the full formula and coverage disclosure →

Related

Alternatives to Amazon Nova MicroAmazon cost calculator →LLM Chatbot cost →RAG Question Answering cost →

FAQ

Why does price scale so directly with input tokens here?
At 90,000 input tokens versus 1,200 output tokens, input is roughly 98% of the token volume on every call — the input rate dominates the bill.
Is chunking cheaper than one large call?
Not usually — chunking adds extra model calls and can lose cross-section context, without changing the total tokens processed.