How Much Does LLM Classification at Volume Cost per Month?

At production volume (5,000,000 calls/month), the cheapest effective option is Amazon Nova Micro at $98.14/month. The most expensive frontier option, GPT-5.4 Pro, runs $93,720/month — Classification is the smallest per-call shape in this cluster, but volume is the entire story — a tiny per-call cost still adds up at millions of calls a month.

How much does classification at volume cost per month?

At production volume (5,000,000 calls/month), the cheapest effective option for classification at volume is Amazon Nova Micro at $98.14 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $93,720 per month for the same workload.

Verified 2026-08-08
Cheapest ≠ best. This page ranks by effective cost only.

Token shape

Shapetiny input in, single-label output out
Input / output tokens per call1K in / 0K out
Cacheable input40%
Batch-eligibleYes

Input is one short text plus a stable instruction/label-set prompt; output is a single label or short code. The instruction portion is highly cacheable, and this is the clearest batch-processing candidate in the cluster.

What drives this workload's cost?

The main token-volume driver here is input tokens: each call sends 1K input tokens and requests up to 0K output tokens, at 5,000,000 calls per month in the default volume. 40% of input is modeled as cache-eligible, so repeated prefixes can reduce the input charge. It is also batch-eligible, which can reduce the modeled total when asynchronous processing is acceptable.

Volume

Side project
500,000 calls/mo
$9.81/mo cheapest
Production
5,000,000 calls/mo
$98.14/mo cheapest
Scale
50,000,000 calls/mo
$981/mo cheapest

Ranked cost — Production volume (caching + batch applied where available)

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$101$49.070.76×
GPT-5 NanobudgetlegacyOpenAI$165$120
Amazon Nova LitebudgetAmazon$174$85.920.91×
GPT-OSS 20BbudgetGroq$218$1382.93×
Gemini 2.5 Flash LitebudgetlegacyGoogle$290$2001
Ministral 8BbudgetMistral$390$1940.93×1
Mistral Small 3.1budgetMistral$435$2130.85×4
GPT-4o MinibudgetlegacyOpenAI$435$300
Grok-3 MinibudgetlegacyxAI$435$435
GPT-OSS 120BbudgetGroq$435$2421.82×
Muse Spark 1.3 ContributorbudgetMeta$270$50912.95×6
Llama 4 MaverickbudgetlegacyGroq$560$280
GPT-5.4 NanobudgetlegacyOpenAI$625$4180.78×
Gemini 3.1 Flash LitebudgetlegacyGoogle$775$5310.87×
CodestralbudgetMistral$840$4110.79×1
Show all 68 models
GPT-5 MinibudgetlegacyOpenAI$825$6001
Gemini 3.5 Flash LitebudgetGoogle$1,000$7301
Gemini 2.5 FlashbudgetlegacyGoogle$1,000$7301
GPT-OSS 120B (Cerebras)budgetCerebras$950$1,0492.32×2
Mistral Large 3budgetMistral$1,400$7001
DeepSeek V4 FlashbudgetDeepSeek$1,232$1,4422.59×1
GLM-5.1midlegacyZ.ai$1,720$1,720
Qwen 3.8 30BmidGroq$1,800$900
Qwen 3.6 27BmidlegacyGroq$1,800$900
Qwen 3.7 PlusmidQwen$2,200$2,200
Gemini 3.7 FlashmidGoogle$2,250$1,575
Amazon Nova PromidAmazon$2,320$1,160
GPT-5.4 MinimidlegacyOpenAI$2,325$1,650
Gemini 3.1 FlashmidlegacyGoogle$2,325$1,650
Claude Haiku 4.5midAnthropic$3,000$2,102
GPT-5.6 LunamidOpenAI$3,100$2,200
o3-MinimidlegacyOpenAI$3,190$2,200
Grok 4.3midxAI$3,375$3,4781.41×
Muse Spark 1.3midMeta$3,550$3,550
GPT-5midlegacyOpenAI$4,125$3,0002
Mistral Medium 3midMistral$4,500$2,1900.84×3
Gemini 3.6 FlashmidGoogle$4,500$3,1501
DeepSeek V4 PromidDeepSeek$3,696$4,6073.30×3
Qwen 3.8 MaxmidQwen$4,640$4,6401
Qwen 3.7 MaxmidQwen$4,640$4,6401
Gemini 3.5 FlashmidlegacyGoogle$4,650$3,3001
GLM-5.2midZ.ai$3,940$5,2033.87×6
Grok-3midlegacyxAI$5,400$5,400
Grok-4.20 ReasoningmidxAI$5,600$5,600
Grok-4.20midxAI$5,600$5,600
Grok 4.6midxAI$5,600$5,600
Grok 4.5midxAI$5,600$5,600
Gemini 3.1 PromidGoogle$6,200$3,9560.63×3
GPT-4.1midlegacyOpenAI$5,800$4,0011
Claude Sonnet 5midAnthropic$6,000$4,204
GPT-4omidlegacyOpenAI$7,250$5,0011
GLM 4.7 (Cerebras)midCerebras$5,900$7,6967.53×3
GPT-5.6 TerramidOpenAI$7,750$5,501
GPT-5.4midlegacyOpenAI$7,750$5,501
Claude Sonnet 4.6midAnthropic$9,000$6,306
Claude Sonnet 4.5midlegacyAnthropic$9,000$6,306
Claude Sonnet 4midlegacyAnthropic$9,000$6,306
GPT-5.6 SolmidOpenAI$12,000$8,401
Claude Opus 4.8midAnthropic$15,000$10,4100.96×
Claude Opus 4.7midlegacyAnthropic$15,000$10,510
Claude Opus 4.6midlegacyAnthropic$15,000$10,510
Claude Opus 4.5midlegacyAnthropic$15,000$10,510
GPT-4 TurbofrontierlegacyOpenAI$28,000$19,003
Claude Fable 5frontierAnthropic$30,000$21,020
Claude Opus 5frontierAnthropic$45,000$31,530
Claude Opus 4.1frontierlegacyAnthropic$45,000$31,530
Claude Opus 4frontierlegacyAnthropic$45,000$31,530
GPT-5.4 ProfrontierlegacyOpenAI$93,000$66,7301.04×

Batch 40 · server-rendered decision evidence · verified 2026-08-27

Classification-at-volume cost and capacity evidence

Frozen inputs, formulas, provenance, and decision boundaries are visible in the initial HTML. Unsupported evidence fails closed as Unavailable.

Label-taxonomy prompt-growth cube

Formula / scoring rule: Bill/record = (input tokens × input rate + output tokens × output rate)/1M; cached prefix is separated from uncached input.

Provenance: Frozen 500-input / 20-output record; 2/10/50/200 classes and 0/1/5 examples per class.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
2 classes / 0 examples
batch40-classification-at-volume-m1-r1
prompt 620; output 20; single-label JSONPer 1M records: 620M×$5 + 20M×$15 = $3,400.Tokenizer count is fixture-specific; no quality transfer.CALCULATED — exact fixture.
50 classes / 5 examples
batch40-classification-at-volume-m1-r2
prompt 8,420; output 34; multi-label JSONPer 1M records: 8,420M×$5 + 34M×$15 = $42,610.Schema growth must be included before choosing cache.CALCULATED — exact fixture.
200 classes / 5 examples
batch40-classification-at-volume-m1-r3
prompt 31,200; tokenizer result not stored; cache boundary requestedUnavailable — tokenizer result and cache-read rate are absentEstimated tokens cannot become an exact bill.Unavailable — tokenizer result and cache-read rate are absent

Module citation: OpenAI structured outputs documentation.

Confidence-and-review settlement tree

Formula / scoring rule: Cost/accepted label = total attempted path cost / accepted labels; abstentions and invalid repairs stay in numerator.

Provenance: Five-million-record scenario; confidence bands and routing probabilities are user-supplied, not accuracy observations.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
high confidence
batch40-classification-at-volume-m2-r1
3.5M records; cheap model $0.003/record; no reviewAttempted cost $10,500; accepted labels Unavailable — calibrated acceptance is not observedConfidence is not accuracy.Unavailable — calibrated acceptance is not observed
abstain / retry
batch40-classification-at-volume-m2-r2
1M records; retry $0.004; premium escalation $0.02; 8% invalid repairPath costs can be summed; accepted denominator Unavailable — not suppliedDo not invent pass probabilities.Unavailable — not supplied
human adjudication
batch40-classification-at-volume-m2-r3
500K records; 6% review; 3 min at $60/hourScenario review labor = 30,000 × 0.05 × $60 = $90,000.Review cost is a scenario input, not provider billing.CALCULATED — user-supplied.

Module citation: All AI Ask extraction evidence registry.

Capacity and deadline planner

Formula / scoring rule: Completion capacity = min(RPM, TPM/input tokens) × active shards × window; missing quota fails closed.

Provenance: 5M and 50M record scenarios; provider quotas and completion guarantees are not inferred.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
5M / online
batch40-classification-at-volume-m3-r1
5M × 500 input tokens; 2,500M tokens; 10 shards; RPM Unavailable — provider-specific quotaUnavailable — online completion deadline cannot be closedDo not display infinite throughput.Unavailable — provider-specific quota
50M / batch
batch40-classification-at-volume-m3-r2
50M JSONL rows; 20 shards; submission window 24hBatch eligibility and completion SLA Unavailable — not established for selected providerA discount does not prove deadline compliance.Unavailable — not established for selected provider
partial failure
batch40-classification-at-volume-m3-r3
100K submitted; 2,400 invalid; 700 duplicates; retry subset 1,100Unique accepted input = 96,900 before label acceptance.Deduplicate before cost-per-accepted calculation.CALCULATED — settlement counts closed.

Module citation: OpenAI rate limits.

Calculate cost per accepted label

Prompt caching uses the provider's sourced read multiplier and amortises cache writes over the documented TTL; the largest published read discount is 90%. See the full formula and coverage disclosure →

Related

Alternatives to Amazon Nova MicroAmazon cost calculator →LLM Chatbot cost →RAG Question Answering cost →

FAQ

Why does a tiny per-call cost matter here?
At tens of millions of calls a month, a fraction-of-a-cent difference per call compounds into a large monthly delta between models.
Should this always run through a batch API?
Whenever the classification does not need a synchronous response, yes — batch discounts apply directly to a workload this uniform.