How Much Does Long-Document Summarization Cost per Month?
At production volume (8,000 calls/month), the cheapest effective option is Amazon Nova Micro at $26.22/month. The most expensive frontier option, GPT-5.4 Pro, runs $23,397/month — Summarization is the most input-heavy shape in this cluster: a large source document in, a medium-length summary out.
How much does long-document summarization cost per month?
At production volume (8,000 calls/month), the cheapest effective option for long-document summarization is Amazon Nova Micro at $26.22 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $23,397 per month for the same workload.
Token shape
| Shape | very large document in, medium summary out |
| Input / output tokens per call | 90K in / 1K out |
| Cacheable input | 10% |
| Batch-eligible | Yes |
Input is a full long-form document (report, transcript, contract); output is a structured summary. Each document is usually read once, so little of the input is cache-eligible.
What drives this workload's cost?
The main token-volume driver here is input tokens: each call sends 90K input tokens and requests up to 1K output tokens, at 8,000 calls per month in the default volume. 10% of input is modeled as cache-eligible, so repeated prefixes can reduce the input charge. It is also batch-eligible, which can reduce the modeled total when asynchronous processing is acceptable.
Volume
Ranked cost — Production volume (caching + batch applied where available)
| Model | Provider | List monthly | Effective monthly | Verbosity | Rank Δ |
|---|---|---|---|---|---|
| Amazon Nova Microbudget | Amazon | $26.54 | $13.11 | 0.76× | — |
| GPT-5 Nanobudgetlegacy | OpenAI | $39.84 | $37.35 | — | — |
| Amazon Nova Litebudget | Amazon | $45.50 | $22.65 | 0.91× | — |
| GPT-OSS 20Bbudget | Groq | $56.88 | $31.22 | 2.93× | — |
| Gemini 2.5 Flash Litebudgetlegacy | $75.84 | $69.94 | — | ▲1 | |
| Muse Spark 1.3 Contributorbudget | Meta | $73.92 | $96.86 | 12.95× | ▼1 |
| Ministral 8Bbudget | Mistral | $109 | $54.67 | 0.93× | — |
| Mistral Small 3.1budget | Mistral | $114 | $56.45 | 0.85× | ▲3 |
| GPT-4o Minibudgetlegacy | OpenAI | $114 | $106 | — | ▼1 |
| Grok-3 Minibudgetlegacy | xAI | $114 | $114 | — | ▼1 |
| GPT-OSS 120Bbudget | Groq | $114 | $59.24 | 1.82× | ▼1 |
| Llama 4 Maverickbudgetlegacy | Groq | $150 | $74.88 | — | — |
| GPT-5.4 Nanobudgetlegacy | OpenAI | $156 | $143 | 0.78× | — |
| Gemini 3.1 Flash Litebudgetlegacy | $194 | $178 | 0.87× | — | |
| GPT-5 Minibudgetlegacy | OpenAI | $199 | $187 | — | — |
Show all 68 models
| Codestralbudget | Mistral | $225 | $111 | 0.79× | — |
| Gemini 3.5 Flash Litebudget | $240 | $222 | — | — | |
| Gemini 2.5 Flashbudgetlegacy | $240 | $222 | — | — | |
| GPT-OSS 120B (Cerebras)budget | Cerebras | $259 | $269 | 2.32× | — |
| DeepSeek V4 Flashbudget | DeepSeek | $329 | $350 | 2.59× | — |
| Mistral Large 3budget | Mistral | $374 | $187 | — | — |
| GLM-5.1midlegacy | Z.ai | $453 | $453 | — | — |
| Qwen 3.8 30Bmid | Groq | $461 | $230 | — | — |
| Qwen 3.6 27Bmidlegacy | Groq | $461 | $230 | — | — |
| Gemini 3.7 Flashmid | $576 | $532 | — | — | |
| GPT-5.4 Minimidlegacy | OpenAI | $583 | $546 | — | — |
| Gemini 3.1 Flashmidlegacy | $583 | $539 | — | — | |
| Qwen 3.7 Plusmid | Qwen | $595 | $595 | — | — |
| Amazon Nova Promid | Amazon | $607 | $303 | — | — |
| Claude Haiku 4.5mid | Anthropic | $768 | $786 | — | — |
| GPT-5.6 Lunamid | OpenAI | $778 | $728 | — | — |
| o3-Minimidlegacy | OpenAI | $834 | $779 | — | — |
| Grok 4.3mid | xAI | $924 | $934 | 1.41× | — |
| Muse Spark 1.3mid | Meta | $941 | $941 | — | — |
| GPT-5midlegacy | OpenAI | $996 | $934 | — | ▲1 |
| DeepSeek V4 Promid | DeepSeek | $988 | $1,076 | 3.30× | ▼1 |
| Mistral Medium 3mid | Mistral | $1,152 | $570 | 0.84× | ▲2 |
| Gemini 3.6 Flashmid | $1,152 | $1,064 | — | — | |
| Gemini 3.5 Flashmidlegacy | $1,166 | $1,078 | — | ▲1 | |
| GLM-5.2mid | Z.ai | $1,050 | $1,171 | 3.87× | ▼3 |
| Qwen 3.8 Maxmid | Qwen | $1,213 | $1,213 | — | — |
| Qwen 3.7 Maxmid | Qwen | $1,213 | $1,213 | — | — |
| Grok-3midlegacy | xAI | $1,478 | $1,478 | — | — |
| Grok-4.20 Reasoningmid | xAI | $1,498 | $1,498 | — | — |
| Grok-4.20mid | xAI | $1,498 | $1,498 | — | — |
| Grok 4.6mid | xAI | $1,498 | $1,498 | — | — |
| Grok 4.5mid | xAI | $1,498 | $1,498 | — | — |
| Gemini 3.1 Promid | $1,555 | $1,395 | 0.63× | ▲2 | |
| GPT-4.1midlegacy | OpenAI | $1,517 | $1,417 | — | ▼1 |
| Claude Sonnet 5mid | Anthropic | $1,536 | $1,572 | — | ▼1 |
| GLM 4.7 (Cerebras)mid | Cerebras | $1,646 | $1,819 | 7.53× | — |
| GPT-4omidlegacy | OpenAI | $1,896 | $1,771 | — | — |
| GPT-5.6 Terramid | OpenAI | $1,944 | $1,819 | — | — |
| GPT-5.4midlegacy | OpenAI | $1,944 | $1,819 | — | — |
| Claude Sonnet 4.6mid | Anthropic | $2,304 | $2,358 | — | — |
| Claude Sonnet 4.5midlegacy | Anthropic | $2,304 | $2,358 | — | — |
| Claude Sonnet 4midlegacy | Anthropic | $2,304 | $2,358 | — | — |
| GPT-5.6 Solmid | OpenAI | $3,072 | $2,872 | — | — |
| Claude Opus 4.8mid | Anthropic | $3,840 | $3,920 | 0.96× | — |
| Claude Opus 4.7midlegacy | Anthropic | $3,840 | $3,930 | — | — |
| Claude Opus 4.6midlegacy | Anthropic | $3,840 | $3,930 | — | — |
| Claude Opus 4.5midlegacy | Anthropic | $3,840 | $3,930 | — | — |
| GPT-4 Turbofrontierlegacy | OpenAI | $7,488 | $6,989 | — | — |
| Claude Fable 5frontier | Anthropic | $7,680 | $7,860 | — | — |
| Claude Opus 5frontier | Anthropic | $11,520 | $11,790 | — | — |
| Claude Opus 4.1frontierlegacy | Anthropic | $11,520 | $11,790 | — | — |
| Claude Opus 4frontierlegacy | Anthropic | $11,520 | $11,790 | — | — |
| GPT-5.4 Profrontierlegacy | OpenAI | $23,328 | $21,900 | 1.04× | — |
Batch 40 · server-rendered decision evidence · verified 2026-08-27
Long-document summarization cost evidence
Frozen inputs, formulas, provenance, and decision boundaries are visible in the initial HTML. Unsupported evidence fails closed as Unavailable.
Topology-expanded token ledger
Formula / scoring rule: Monthly bill = calls × (input tokens × $/M + output tokens × $/M); chunk overlap and intermediate summaries are counted.
Provenance: Frozen 90K-input / 1.2K-output document, 8K context, $5/M input and $15/M output illustrative rates.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State |
|---|---|---|---|---|
one-passbatch40-long-document-summarization-m1-r1 | 1 document; 90K input; 1,200 output; 8K context | Unavailable — 90K input exceeds the 8K context window | Reject before cost math; do not silently truncate. | Unavailable — context eligibility fails |
map-reducebatch40-long-document-summarization-m1-r2 | 12 chunks × 8K + 10% overlap; 12 × 600 intermediate; 1,200 final | Input 105,600; output 8,400; bill = (105,600×5 + 8,400×15)/1M = $1.182. | Include overlap and synthesis tokens in every topology. | ELIGIBLE — calculated token ledger. |
rolling refinebatch40-long-document-summarization-m1-r3 | 23 serial chunks; 8K chunk; 600 retained summary each pass | Input 182,400; output 14,400; bill = $1.128; serial depth 23. | Latency/concurrency tradeoff remains separate from bill. | ELIGIBLE — calculated token ledger. |
Module citation: All AI Ask evidence registry (verified 2026-08-27).
Portfolio percentile planner
Formula / scoring rule: Budget = Σ row-level routed bill; mean shortcut = document count × mean-token bill and may hide P99 exclusions.
Provenance: Frozen 1,000-document manifest: P50 9K, P90 42K, P99 180K; scanned share 8%.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State |
|---|---|---|---|---|
P50 / 500 docsbatch40-long-document-summarization-m2-r1 | 500 × 9K tokens; map path; 1 retry per 50 docs | 4.5M source tokens; retry count 10; source bill Unavailable — model rate selected by user | Do not convert unavailable rate into zero. | Unavailable — model rate selected by user |
P90 / 400 docsbatch40-long-document-summarization-m2-r2 | 400 × 42K; split at 8K; 8% scanned pages | 2,100 chunks; OCR charge Unavailable — OCR price per page is not supplied | Keep OCR/storage outside token subtotal until priced. | Unavailable — OCR price per page is not supplied |
P99 / 100 docsbatch40-long-document-summarization-m2-r3 | 100 × 180K; 100 over-window; split required; retry set 7 | All 100 route to split topology; successful count Unavailable — post-retry acceptance is not observed | No acceptance rate is inferred from routing. | Unavailable — post-retry acceptance is not observed |
Module citation: All AI Ask workload registry.
Accepted-summary TCO tree
Formula / scoring rule: Cost/accepted = (API + OCR + checks + retries + review) / accepted summaries; denominator is not processed documents.
Provenance: 8K-document baseline; review is explicitly user-supplied at 6 minutes and $60/hour.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State |
|---|---|---|---|---|
API + checksbatch40-long-document-summarization-m3-r1 | 8,000 docs; API subtotal $1,182; factual check $0.002/doc | API + checks = $1,198; 8,000 processed denominator. | Accepted denominator still needs measured pass rate. | Unavailable — accepted-summary count is not observed |
human reviewbatch40-long-document-summarization-m3-r2 | 12% review; 6 min/doc; $60/hour | Expected review labor = 960 × 0.1 × $60 = $5,760. | Labor is an explicit scenario, not an observed rate. | CALCULATED — user inputs visible. |
OCR placeholderbatch40-long-document-summarization-m3-r3 | 8% scanned share; 640 docs; page count and OCR tariff missing | Unavailable — OCR pages and price are not supplied | Do not publish end-to-end TCO without OCR units. | Unavailable — OCR pages and price are not supplied |
Module citation: All AI Ask summarization evidence.
Prompt caching uses the provider's sourced read multiplier and amortises cache writes over the documented TTL; the largest published read discount is 90%. See the full formula and coverage disclosure →
