LLM Cost Calculator — Real Monthly Cost, Not Just Rate

Raw dataset: data.json. Verbosity measured 2026-06-21. Cite this: All AI Ask LLM Verbosity Index and Effective Cost Dataset, retrieved 2026-06-21.

A $/M rate is not your bill. Set your call shape below and see every priced model ranked by effective monthly cost — adjusted for how many output tokens each model actually spends on a job of this size, not its list price alone.

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$26.60$24.250.76×
Amazon Nova LitebudgetAmazon$45.60$44.090.91×
GPT-5 NanobudgetlegacyOpenAI$52.00$52.00
Gemini 2.5 Flash LitebudgetlegacyGoogle$76.00$76.002
Ministral 8BbudgetMistral$82.50$81.770.93×2
GPT-OSS 20BbudgetGroq$57.00$97.532.93×2
Mistral Small 3.1budgetMistral$114$1080.85×4
GPT-4o MinibudgetlegacyOpenAI$114$114
Grok-3 MinibudgetlegacyxAI$114$114
Llama 4 MaverickbudgetlegacyGroq$138$1382
GPT-OSS 120BbudgetGroq$114$1481.82×1
GPT-5.4 NanobudgetlegacyOpenAI$184$1640.78×1
CodestralbudgetMistral$207$1940.79×1
Gemini 3.1 Flash LitebudgetlegacyGoogle$225$2110.87×2
Muse Spark 1.3 ContributorbudgetMeta$62.00$22912.95×10
Show all 68 models
GPT-5 MinibudgetlegacyOpenAI$260$2601
GPT-OSS 120B (Cerebras)budgetCerebras$221$2902.32×2
Gemini 3.5 Flash LitebudgetGoogle$319$3191
Gemini 2.5 FlashbudgetlegacyGoogle$319$3191
Mistral Large 3budgetMistral$345$3451
GLM-5.1midlegacyZ.ai$442$4421
DeepSeek V4 FlashbudgetDeepSeek$304$4512.59×4
Qwen 3.8 30BmidGroq$498$498
Qwen 3.6 27BmidlegacyGroq$498$498
Qwen 3.7 PlusmidQwen$524$524
Amazon Nova PromidAmazon$608$608
Gemini 3.7 FlashmidGoogle$623$623
GPT-5.4 MinimidlegacyOpenAI$675$675
Gemini 3.1 FlashmidlegacyGoogle$675$675
Claude Haiku 4.5midAnthropic$830$8301
o3-MinimidlegacyOpenAI$836$8361
Grok 4.3midxAI$775$8471.41×2
Muse Spark 1.3midMeta$898$898
GPT-5.6 LunamidOpenAI$900$900
Mistral Medium 3midMistral$1,245$1,1610.84×6
Qwen 3.8 MaxmidQwen$1,216$1,2161
Qwen 3.7 MaxmidQwen$1,216$1,2161
Grok-3midlegacyxAI$1,240$1,2401
Gemini 3.6 FlashmidGoogle$1,245$1,2451
GPT-5midlegacyOpenAI$1,300$1,3003
Gemini 3.5 FlashmidlegacyGoogle$1,350$1,3503
Grok-4.20 ReasoningmidxAI$1,380$1,3803
Grok-4.20midxAI$1,380$1,3803
Grok 4.6midxAI$1,380$1,3803
Grok 4.5midxAI$1,380$1,3803
Gemini 3.1 PromidGoogle$1,800$1,4890.63×5
GPT-4.1midlegacyOpenAI$1,520$1,5202
DeepSeek V4 PromidDeepSeek$911$1,5483.30×13
Claude Sonnet 5midAnthropic$1,660$1,6601
GLM-5.2midZ.ai$980$1,8643.87×14
GPT-4omidlegacyOpenAI$1,900$1,9001
GPT-5.6 TerramidOpenAI$2,250$2,2501
GPT-5.4midlegacyOpenAI$2,250$2,2501
Claude Sonnet 4.6midAnthropic$2,490$2,4901
Claude Sonnet 4.5midlegacyAnthropic$2,490$2,4901
Claude Sonnet 4midlegacyAnthropic$2,490$2,4901
GLM 4.7 (Cerebras)midCerebras$1,273$2,5307.53×15
GPT-5.6 SolmidOpenAI$3,320$3,320
Claude Opus 4.8midAnthropic$4,150$4,0800.96×
Claude Opus 4.7midlegacyAnthropic$4,150$4,150
Claude Opus 4.6midlegacyAnthropic$4,150$4,150
Claude Opus 4.5midlegacyAnthropic$4,150$4,150
GPT-4 TurbofrontierlegacyOpenAI$6,900$6,900
Claude Fable 5frontierAnthropic$8,300$8,300
Claude Opus 5frontierAnthropic$12,450$12,450
Claude Opus 4.1frontierlegacyAnthropic$12,450$12,450
Claude Opus 4frontierlegacyAnthropic$12,450$12,450
GPT-5.4 ProfrontierlegacyOpenAI$27,000$27,5041.04×

List price lied to you

These models move the most once verbosity is priced in — a model that talks more costs more, regardless of its list rate.

  • GLM 4.7 (Cerebras) is priced #42 by list rate but #57 once its 7.53× verbosity is billed — $1,273 list vs $2,530 effective.
  • GLM-5.2 is priced #36 by list rate but #50 once its 3.87× verbosity is billed — $980 list vs $1,864 effective.
  • DeepSeek V4 Pro is priced #35 by list rate but #48 once its 3.30× verbosity is billed — $911 list vs $1,548 effective.
  • Muse Spark 1.3 Contributor is priced #5 by list rate but #15 once its 12.95× verbosity is billed — $62.00 list vs $229 effective.
  • Mistral Medium 3 is priced #41 by list rate but #35 once its 0.84× verbosity is billed — $1,245 list vs $1,161 effective.
  • Gemini 3.1 Pro is priced #51 by list rate but #46 once its 0.63× verbosity is billed — $1,800 list vs $1,489 effective.

Batch 39 · server-rendered decision evidence · verified 2026-08-27

Whole-workload cost routing and price replay

Every field is tied to a frozen input and a dated provenance record. Unsupported facts fail closed as Unavailable; they are not treated as zero, free, equivalent, current, fastest, cheapest, private, or passing.

Workload-component router

Formula / rubric: Total scope = base token bill + applicable retrieval + embeddings + storage + tools + media + review + retries + parallelism; route to specialist when component exists.

Provenance: Frozen defaults cover chatbot, RAG, coding agent, extraction, summarization, and tool loops. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
chatbot
batch39-llm-cost-calculator-m1-r1
30K calls; 1,200 input/220 output; no tools; 10% cachebase token estimate $279.00/month.Use global calculator; no non-token component detected.ROUTED — base-token path.
RAG
batch39-llm-cost-calculator-m1-r2
30K calls; embeddings + vector reads + rerank + generationbase token path closes; full RAG path routes to /llm-cost-calculator/rag-question-answering.Do not reuse chatbot estimate as end-to-end RAG cost.ROUTED — specialist calculator.
coding agent
batch39-llm-cost-calculator-m1-r3
5–30 turns; tools; tests; compaction; parallel workerstoken estimate closes; repository fan-out and repair route to /llm-cost-calculator/coding-agent.Generic token output is a lower-bound component, not task cost.ROUTED — specialist calculator.

Module citations: All AI Ask pricing registry. All AI Ask evidence registry (verified 2026-08-27).

Five-axis break-even matrix

Formula / rubric: winner threshold = first axis value where totalCostA ≤ totalCostB, holding other declared inputs fixed; unavailable mechanic blocks threshold.

Provenance: Axes: output expansion, cache hit, batch share, retry rate, and monthly calls; dated units must match. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
output expansion
batch39-llm-cost-calculator-m2-r1
220 → 440 output tokens; input 1,200; 30K callsmodel A crosses model B at 318 output tokens under frozen rates.Threshold is valid only for these model IDs and rates.CALCULATED — axis closed.
cache/batch axes
batch39-llm-cost-calculator-m2-r2
cache 0/50/100%; batch 0/50/100%; same token shapecache-read multiplier for one candidate is Unavailable — not returned in compatible registry unitsNo break-even matrix cell is emitted for that candidate.Unavailable — cache-read multiplier is not returned in compatible registry units
retry/call volume
batch39-llm-cost-calculator-m2-r3
retry 0/5/20%; calls 1K/30K/1M; output 220call-volume scaling is linear; retry debit for candidate C is Unavailable — not returnedRetain a threshold only where retry billing is observed.Unavailable — retry debit is not returned

Module citations: OpenAI API pricing. All AI Ask evidence registry (verified 2026-08-27).

Current versus prior price-change budget ledger

Formula / rubric: Impact = replay(frozen workload, current registry) − replay(frozen workload, prior registry), with catalog/coverage changes separated from rate deltas.

Provenance: Current and immediately prior verified registry snapshots; no live pricing call is made. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
small workload
batch39-llm-cost-calculator-m3-r1
1K calls; 2K input/400 output; current $5/$15; prior $5/$15rate delta $0.00; model catalog unchanged.Attribute only rate movement to budget impact.NO CHANGE — replay closes.
production workload
batch39-llm-cost-calculator-m3-r2
30K calls; 12K input/1.2K output; prior output coverage absentcurrent replay exists; prior verbosity-adjusted output is Unavailable — not covered by prior snapshotDo not label coverage change as price change.Unavailable — prior verbosity-adjusted output is not covered
scale workload
batch39-llm-cost-calculator-m3-r3
1M calls; batch 40%; cache 60%; provider TTL mechanic changedcatalog and cache mechanics changed together; attributable delta is Unavailable — not separablePublish no single price-change percentage.Unavailable — catalog and cache mechanic deltas are not separable

Module citations: All AI Ask model pricing registry. All AI Ask evidence registry (verified 2026-08-27).

Calculate a shareable whole-workload estimate

The formula, published

outputTokensBilled = outputTokens × (verbosityIndex ?? 1)
inputCost  = inputPerM  × inputTokens        / 1e6
outputCost = outputPerM × outputTokensBilled / 1e6
effective  = (inputCost + outputCost) × callsPerMonth

with caching: inputCost × (1 − cacheableInputPct × 0.9)
with batching: (inputCost + outputCost) × (1 − batchDiscountPct/100)

up to 90% cache-read saving is the largest sourced read discount in the provider table; provider-specific estimates use that provider's multiplier and account for cache writes. The free-form estimator assumes 30% of input is cache-eligible; pick a workload preset below for a shape-specific assumption instead.

Coverage: 20 of 68 priced models have a measured verbosity index (2+ graded runs). The rest render with a verbosity of — and are shown at unadjusted list price.

Cost by workload shape

LLM Chatbot
growing conversation history in, short reply out
RAG Question Answering
large retrieved context in, short answer out
Coding Agent
large file context in, large diff out
Document Extraction
medium document in, tiny structured JSON out
Long-Document Summarization
very large document in, medium summary out
Content Generation
tiny brief in, long piece out
Classification at Volume
tiny input in, single-label output out
Agentic Tool Loop
many small round-trips per completed task