How Much Does LLM Classification at Volume Cost per Month?
At production volume (5,000,000 calls/month), the cheapest effective option is Amazon Nova Micro at $98.14/month. The most expensive frontier option, GPT-5.4 Pro, runs $93,720/month — Classification is the smallest per-call shape in this cluster, but volume is the entire story — a tiny per-call cost still adds up at millions of calls a month.
How much does classification at volume cost per month?
At production volume (5,000,000 calls/month), the cheapest effective option for classification at volume is Amazon Nova Micro at $98.14 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $93,720 per month for the same workload.
Token shape
| Shape | tiny input in, single-label output out |
| Input / output tokens per call | 1K in / 0K out |
| Cacheable input | 40% |
| Batch-eligible | Yes |
Input is one short text plus a stable instruction/label-set prompt; output is a single label or short code. The instruction portion is highly cacheable, and this is the clearest batch-processing candidate in the cluster.
What drives this workload's cost?
The main token-volume driver here is input tokens: each call sends 1K input tokens and requests up to 0K output tokens, at 5,000,000 calls per month in the default volume. 40% of input is modeled as cache-eligible, so repeated prefixes can reduce the input charge. It is also batch-eligible, which can reduce the modeled total when asynchronous processing is acceptable.
Volume
Ranked cost — Production volume (caching + batch applied where available)
| Model | Provider | List monthly | Effective monthly | Verbosity | Rank Δ |
|---|---|---|---|---|---|
| Amazon Nova Microbudget | Amazon | $101 | $49.07 | 0.76× | — |
| GPT-5 Nanobudgetlegacy | OpenAI | $165 | $120 | — | — |
| Amazon Nova Litebudget | Amazon | $174 | $85.92 | 0.91× | — |
| GPT-OSS 20Bbudget | Groq | $218 | $138 | 2.93× | — |
| Gemini 2.5 Flash Litebudgetlegacy | $290 | $200 | — | ▲1 | |
| Ministral 8Bbudget | Mistral | $390 | $194 | 0.93× | ▲1 |
| Mistral Small 3.1budget | Mistral | $435 | $213 | 0.85× | ▲4 |
| GPT-4o Minibudgetlegacy | OpenAI | $435 | $300 | — | — |
| Grok-3 Minibudgetlegacy | xAI | $435 | $435 | — | — |
| GPT-OSS 120Bbudget | Groq | $435 | $242 | 1.82× | — |
| Muse Spark 1.3 Contributorbudget | Meta | $270 | $509 | 12.95× | ▼6 |
| Llama 4 Maverickbudgetlegacy | Groq | $560 | $280 | — | — |
| GPT-5.4 Nanobudgetlegacy | OpenAI | $625 | $418 | 0.78× | — |
| Gemini 3.1 Flash Litebudgetlegacy | $775 | $531 | 0.87× | — | |
| Codestralbudget | Mistral | $840 | $411 | 0.79× | ▲1 |
Show all 68 models
| GPT-5 Minibudgetlegacy | OpenAI | $825 | $600 | — | ▼1 |
| Gemini 3.5 Flash Litebudget | $1,000 | $730 | — | ▲1 | |
| Gemini 2.5 Flashbudgetlegacy | $1,000 | $730 | — | ▲1 | |
| GPT-OSS 120B (Cerebras)budget | Cerebras | $950 | $1,049 | 2.32× | ▼2 |
| Mistral Large 3budget | Mistral | $1,400 | $700 | — | ▲1 |
| DeepSeek V4 Flashbudget | DeepSeek | $1,232 | $1,442 | 2.59× | ▼1 |
| GLM-5.1midlegacy | Z.ai | $1,720 | $1,720 | — | — |
| Qwen 3.8 30Bmid | Groq | $1,800 | $900 | — | — |
| Qwen 3.6 27Bmidlegacy | Groq | $1,800 | $900 | — | — |
| Qwen 3.7 Plusmid | Qwen | $2,200 | $2,200 | — | — |
| Gemini 3.7 Flashmid | $2,250 | $1,575 | — | — | |
| Amazon Nova Promid | Amazon | $2,320 | $1,160 | — | — |
| GPT-5.4 Minimidlegacy | OpenAI | $2,325 | $1,650 | — | — |
| Gemini 3.1 Flashmidlegacy | $2,325 | $1,650 | — | — | |
| Claude Haiku 4.5mid | Anthropic | $3,000 | $2,102 | — | — |
| GPT-5.6 Lunamid | OpenAI | $3,100 | $2,200 | — | — |
| o3-Minimidlegacy | OpenAI | $3,190 | $2,200 | — | — |
| Grok 4.3mid | xAI | $3,375 | $3,478 | 1.41× | — |
| Muse Spark 1.3mid | Meta | $3,550 | $3,550 | — | — |
| GPT-5midlegacy | OpenAI | $4,125 | $3,000 | — | ▲2 |
| Mistral Medium 3mid | Mistral | $4,500 | $2,190 | 0.84× | ▲3 |
| Gemini 3.6 Flashmid | $4,500 | $3,150 | — | ▲1 | |
| DeepSeek V4 Promid | DeepSeek | $3,696 | $4,607 | 3.30× | ▼3 |
| Qwen 3.8 Maxmid | Qwen | $4,640 | $4,640 | — | ▲1 |
| Qwen 3.7 Maxmid | Qwen | $4,640 | $4,640 | — | ▲1 |
| Gemini 3.5 Flashmidlegacy | $4,650 | $3,300 | — | ▲1 | |
| GLM-5.2mid | Z.ai | $3,940 | $5,203 | 3.87× | ▼6 |
| Grok-3midlegacy | xAI | $5,400 | $5,400 | — | — |
| Grok-4.20 Reasoningmid | xAI | $5,600 | $5,600 | — | — |
| Grok-4.20mid | xAI | $5,600 | $5,600 | — | — |
| Grok 4.6mid | xAI | $5,600 | $5,600 | — | — |
| Grok 4.5mid | xAI | $5,600 | $5,600 | — | — |
| Gemini 3.1 Promid | $6,200 | $3,956 | 0.63× | ▲3 | |
| GPT-4.1midlegacy | OpenAI | $5,800 | $4,001 | — | ▼1 |
| Claude Sonnet 5mid | Anthropic | $6,000 | $4,204 | — | — |
| GPT-4omidlegacy | OpenAI | $7,250 | $5,001 | — | ▲1 |
| GLM 4.7 (Cerebras)mid | Cerebras | $5,900 | $7,696 | 7.53× | ▼3 |
| GPT-5.6 Terramid | OpenAI | $7,750 | $5,501 | — | — |
| GPT-5.4midlegacy | OpenAI | $7,750 | $5,501 | — | — |
| Claude Sonnet 4.6mid | Anthropic | $9,000 | $6,306 | — | — |
| Claude Sonnet 4.5midlegacy | Anthropic | $9,000 | $6,306 | — | — |
| Claude Sonnet 4midlegacy | Anthropic | $9,000 | $6,306 | — | — |
| GPT-5.6 Solmid | OpenAI | $12,000 | $8,401 | — | — |
| Claude Opus 4.8mid | Anthropic | $15,000 | $10,410 | 0.96× | — |
| Claude Opus 4.7midlegacy | Anthropic | $15,000 | $10,510 | — | — |
| Claude Opus 4.6midlegacy | Anthropic | $15,000 | $10,510 | — | — |
| Claude Opus 4.5midlegacy | Anthropic | $15,000 | $10,510 | — | — |
| GPT-4 Turbofrontierlegacy | OpenAI | $28,000 | $19,003 | — | — |
| Claude Fable 5frontier | Anthropic | $30,000 | $21,020 | — | — |
| Claude Opus 5frontier | Anthropic | $45,000 | $31,530 | — | — |
| Claude Opus 4.1frontierlegacy | Anthropic | $45,000 | $31,530 | — | — |
| Claude Opus 4frontierlegacy | Anthropic | $45,000 | $31,530 | — | — |
| GPT-5.4 Profrontierlegacy | OpenAI | $93,000 | $66,730 | 1.04× | — |
Batch 40 · server-rendered decision evidence · verified 2026-08-27
Classification-at-volume cost and capacity evidence
Frozen inputs, formulas, provenance, and decision boundaries are visible in the initial HTML. Unsupported evidence fails closed as Unavailable.
Label-taxonomy prompt-growth cube
Formula / scoring rule: Bill/record = (input tokens × input rate + output tokens × output rate)/1M; cached prefix is separated from uncached input.
Provenance: Frozen 500-input / 20-output record; 2/10/50/200 classes and 0/1/5 examples per class.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State |
|---|---|---|---|---|
2 classes / 0 examplesbatch40-classification-at-volume-m1-r1 | prompt 620; output 20; single-label JSON | Per 1M records: 620M×$5 + 20M×$15 = $3,400. | Tokenizer count is fixture-specific; no quality transfer. | CALCULATED — exact fixture. |
50 classes / 5 examplesbatch40-classification-at-volume-m1-r2 | prompt 8,420; output 34; multi-label JSON | Per 1M records: 8,420M×$5 + 34M×$15 = $42,610. | Schema growth must be included before choosing cache. | CALCULATED — exact fixture. |
200 classes / 5 examplesbatch40-classification-at-volume-m1-r3 | prompt 31,200; tokenizer result not stored; cache boundary requested | Unavailable — tokenizer result and cache-read rate are absent | Estimated tokens cannot become an exact bill. | Unavailable — tokenizer result and cache-read rate are absent |
Module citation: OpenAI structured outputs documentation.
Confidence-and-review settlement tree
Formula / scoring rule: Cost/accepted label = total attempted path cost / accepted labels; abstentions and invalid repairs stay in numerator.
Provenance: Five-million-record scenario; confidence bands and routing probabilities are user-supplied, not accuracy observations.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State |
|---|---|---|---|---|
high confidencebatch40-classification-at-volume-m2-r1 | 3.5M records; cheap model $0.003/record; no review | Attempted cost $10,500; accepted labels Unavailable — calibrated acceptance is not observed | Confidence is not accuracy. | Unavailable — calibrated acceptance is not observed |
abstain / retrybatch40-classification-at-volume-m2-r2 | 1M records; retry $0.004; premium escalation $0.02; 8% invalid repair | Path costs can be summed; accepted denominator Unavailable — not supplied | Do not invent pass probabilities. | Unavailable — not supplied |
human adjudicationbatch40-classification-at-volume-m2-r3 | 500K records; 6% review; 3 min at $60/hour | Scenario review labor = 30,000 × 0.05 × $60 = $90,000. | Review cost is a scenario input, not provider billing. | CALCULATED — user-supplied. |
Module citation: All AI Ask extraction evidence registry.
Capacity and deadline planner
Formula / scoring rule: Completion capacity = min(RPM, TPM/input tokens) × active shards × window; missing quota fails closed.
Provenance: 5M and 50M record scenarios; provider quotas and completion guarantees are not inferred.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State |
|---|---|---|---|---|
5M / onlinebatch40-classification-at-volume-m3-r1 | 5M × 500 input tokens; 2,500M tokens; 10 shards; RPM Unavailable — provider-specific quota | Unavailable — online completion deadline cannot be closed | Do not display infinite throughput. | Unavailable — provider-specific quota |
50M / batchbatch40-classification-at-volume-m3-r2 | 50M JSONL rows; 20 shards; submission window 24h | Batch eligibility and completion SLA Unavailable — not established for selected provider | A discount does not prove deadline compliance. | Unavailable — not established for selected provider |
partial failurebatch40-classification-at-volume-m3-r3 | 100K submitted; 2,400 invalid; 700 duplicates; retry subset 1,100 | Unique accepted input = 96,900 before label acceptance. | Deduplicate before cost-per-accepted calculation. | CALCULATED — settlement counts closed. |
Module citation: OpenAI rate limits.
Prompt caching uses the provider's sourced read multiplier and amortises cache writes over the documented TTL; the largest published read discount is 90%. See the full formula and coverage disclosure →
