LLM Cost Calculator — Real Monthly Cost, Not Just Rate
Raw dataset: data.json. Verbosity measured 2026-06-21. Cite this: All AI Ask LLM Verbosity Index and Effective Cost Dataset, retrieved 2026-06-21.
A $/M rate is not your bill. Set your call shape below and see every priced model ranked by effective monthly cost — adjusted for how many output tokens each model actually spends on a job of this size, not its list price alone.
| Model | Provider | List monthly | Effective monthly | Verbosity | Rank Δ |
|---|---|---|---|---|---|
| Amazon Nova Microbudget | Amazon | $26.60 | $24.25 | 0.76× | — |
| Amazon Nova Litebudget | Amazon | $45.60 | $44.09 | 0.91× | — |
| GPT-5 Nanobudgetlegacy | OpenAI | $52.00 | $52.00 | — | — |
| Gemini 2.5 Flash Litebudgetlegacy | $76.00 | $76.00 | — | ▲2 | |
| Ministral 8Bbudget | Mistral | $82.50 | $81.77 | 0.93× | ▲2 |
| GPT-OSS 20Bbudget | Groq | $57.00 | $97.53 | 2.93× | ▼2 |
| Mistral Small 3.1budget | Mistral | $114 | $108 | 0.85× | ▲4 |
| GPT-4o Minibudgetlegacy | OpenAI | $114 | $114 | — | — |
| Grok-3 Minibudgetlegacy | xAI | $114 | $114 | — | — |
| Llama 4 Maverickbudgetlegacy | Groq | $138 | $138 | — | ▲2 |
| GPT-OSS 120Bbudget | Groq | $114 | $148 | 1.82× | ▼1 |
| GPT-5.4 Nanobudgetlegacy | OpenAI | $184 | $164 | 0.78× | ▲1 |
| Codestralbudget | Mistral | $207 | $194 | 0.79× | ▲1 |
| Gemini 3.1 Flash Litebudgetlegacy | $225 | $211 | 0.87× | ▲2 | |
| Muse Spark 1.3 Contributorbudget | Meta | $62.00 | $229 | 12.95× | ▼10 |
Show all 68 models
| GPT-5 Minibudgetlegacy | OpenAI | $260 | $260 | — | ▲1 |
| GPT-OSS 120B (Cerebras)budget | Cerebras | $221 | $290 | 2.32× | ▼2 |
| Gemini 3.5 Flash Litebudget | $319 | $319 | — | ▲1 | |
| Gemini 2.5 Flashbudgetlegacy | $319 | $319 | — | ▲1 | |
| Mistral Large 3budget | Mistral | $345 | $345 | — | ▲1 |
| GLM-5.1midlegacy | Z.ai | $442 | $442 | — | ▲1 |
| DeepSeek V4 Flashbudget | DeepSeek | $304 | $451 | 2.59× | ▼4 |
| Qwen 3.8 30Bmid | Groq | $498 | $498 | — | — |
| Qwen 3.6 27Bmidlegacy | Groq | $498 | $498 | — | — |
| Qwen 3.7 Plusmid | Qwen | $524 | $524 | — | — |
| Amazon Nova Promid | Amazon | $608 | $608 | — | — |
| Gemini 3.7 Flashmid | $623 | $623 | — | — | |
| GPT-5.4 Minimidlegacy | OpenAI | $675 | $675 | — | — |
| Gemini 3.1 Flashmidlegacy | $675 | $675 | — | — | |
| Claude Haiku 4.5mid | Anthropic | $830 | $830 | — | ▲1 |
| o3-Minimidlegacy | OpenAI | $836 | $836 | — | ▲1 |
| Grok 4.3mid | xAI | $775 | $847 | 1.41× | ▼2 |
| Muse Spark 1.3mid | Meta | $898 | $898 | — | — |
| GPT-5.6 Lunamid | OpenAI | $900 | $900 | — | — |
| Mistral Medium 3mid | Mistral | $1,245 | $1,161 | 0.84× | ▲6 |
| Qwen 3.8 Maxmid | Qwen | $1,216 | $1,216 | — | ▲1 |
| Qwen 3.7 Maxmid | Qwen | $1,216 | $1,216 | — | ▲1 |
| Grok-3midlegacy | xAI | $1,240 | $1,240 | — | ▲1 |
| Gemini 3.6 Flashmid | $1,245 | $1,245 | — | ▲1 | |
| GPT-5midlegacy | OpenAI | $1,300 | $1,300 | — | ▲3 |
| Gemini 3.5 Flashmidlegacy | $1,350 | $1,350 | — | ▲3 | |
| Grok-4.20 Reasoningmid | xAI | $1,380 | $1,380 | — | ▲3 |
| Grok-4.20mid | xAI | $1,380 | $1,380 | — | ▲3 |
| Grok 4.6mid | xAI | $1,380 | $1,380 | — | ▲3 |
| Grok 4.5mid | xAI | $1,380 | $1,380 | — | ▲3 |
| Gemini 3.1 Promid | $1,800 | $1,489 | 0.63× | ▲5 | |
| GPT-4.1midlegacy | OpenAI | $1,520 | $1,520 | — | ▲2 |
| DeepSeek V4 Promid | DeepSeek | $911 | $1,548 | 3.30× | ▼13 |
| Claude Sonnet 5mid | Anthropic | $1,660 | $1,660 | — | ▲1 |
| GLM-5.2mid | Z.ai | $980 | $1,864 | 3.87× | ▼14 |
| GPT-4omidlegacy | OpenAI | $1,900 | $1,900 | — | ▲1 |
| GPT-5.6 Terramid | OpenAI | $2,250 | $2,250 | — | ▲1 |
| GPT-5.4midlegacy | OpenAI | $2,250 | $2,250 | — | ▲1 |
| Claude Sonnet 4.6mid | Anthropic | $2,490 | $2,490 | — | ▲1 |
| Claude Sonnet 4.5midlegacy | Anthropic | $2,490 | $2,490 | — | ▲1 |
| Claude Sonnet 4midlegacy | Anthropic | $2,490 | $2,490 | — | ▲1 |
| GLM 4.7 (Cerebras)mid | Cerebras | $1,273 | $2,530 | 7.53× | ▼15 |
| GPT-5.6 Solmid | OpenAI | $3,320 | $3,320 | — | — |
| Claude Opus 4.8mid | Anthropic | $4,150 | $4,080 | 0.96× | — |
| Claude Opus 4.7midlegacy | Anthropic | $4,150 | $4,150 | — | — |
| Claude Opus 4.6midlegacy | Anthropic | $4,150 | $4,150 | — | — |
| Claude Opus 4.5midlegacy | Anthropic | $4,150 | $4,150 | — | — |
| GPT-4 Turbofrontierlegacy | OpenAI | $6,900 | $6,900 | — | — |
| Claude Fable 5frontier | Anthropic | $8,300 | $8,300 | — | — |
| Claude Opus 5frontier | Anthropic | $12,450 | $12,450 | — | — |
| Claude Opus 4.1frontierlegacy | Anthropic | $12,450 | $12,450 | — | — |
| Claude Opus 4frontierlegacy | Anthropic | $12,450 | $12,450 | — | — |
| GPT-5.4 Profrontierlegacy | OpenAI | $27,000 | $27,504 | 1.04× | — |
List price lied to you
These models move the most once verbosity is priced in — a model that talks more costs more, regardless of its list rate.
- GLM 4.7 (Cerebras) is priced #42 by list rate but #57 once its 7.53× verbosity is billed — $1,273 list vs $2,530 effective.
- GLM-5.2 is priced #36 by list rate but #50 once its 3.87× verbosity is billed — $980 list vs $1,864 effective.
- DeepSeek V4 Pro is priced #35 by list rate but #48 once its 3.30× verbosity is billed — $911 list vs $1,548 effective.
- Muse Spark 1.3 Contributor is priced #5 by list rate but #15 once its 12.95× verbosity is billed — $62.00 list vs $229 effective.
- Mistral Medium 3 is priced #41 by list rate but #35 once its 0.84× verbosity is billed — $1,245 list vs $1,161 effective.
- Gemini 3.1 Pro is priced #51 by list rate but #46 once its 0.63× verbosity is billed — $1,800 list vs $1,489 effective.
Batch 39 · server-rendered decision evidence · verified 2026-08-27
Whole-workload cost routing and price replay
Every field is tied to a frozen input and a dated provenance record. Unsupported facts fail closed as Unavailable; they are not treated as zero, free, equivalent, current, fastest, cheapest, private, or passing.
Workload-component router
Formula / rubric: Total scope = base token bill + applicable retrieval + embeddings + storage + tools + media + review + retries + parallelism; route to specialist when component exists.
Provenance: Frozen defaults cover chatbot, RAG, coding agent, extraction, summarization, and tool loops. Unsupported fields fail closed as Unavailable.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State / reproducible bill |
|---|---|---|---|---|
chatbotbatch39-llm-cost-calculator-m1-r1 | 30K calls; 1,200 input/220 output; no tools; 10% cache | base token estimate $279.00/month. | Use global calculator; no non-token component detected. | ROUTED — base-token path. |
RAGbatch39-llm-cost-calculator-m1-r2 | 30K calls; embeddings + vector reads + rerank + generation | base token path closes; full RAG path routes to /llm-cost-calculator/rag-question-answering. | Do not reuse chatbot estimate as end-to-end RAG cost. | ROUTED — specialist calculator. |
coding agentbatch39-llm-cost-calculator-m1-r3 | 5–30 turns; tools; tests; compaction; parallel workers | token estimate closes; repository fan-out and repair route to /llm-cost-calculator/coding-agent. | Generic token output is a lower-bound component, not task cost. | ROUTED — specialist calculator. |
Module citations: All AI Ask pricing registry. All AI Ask evidence registry (verified 2026-08-27).
Five-axis break-even matrix
Formula / rubric: winner threshold = first axis value where totalCostA ≤ totalCostB, holding other declared inputs fixed; unavailable mechanic blocks threshold.
Provenance: Axes: output expansion, cache hit, batch share, retry rate, and monthly calls; dated units must match. Unsupported fields fail closed as Unavailable.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State / reproducible bill |
|---|---|---|---|---|
output expansionbatch39-llm-cost-calculator-m2-r1 | 220 → 440 output tokens; input 1,200; 30K calls | model A crosses model B at 318 output tokens under frozen rates. | Threshold is valid only for these model IDs and rates. | CALCULATED — axis closed. |
cache/batch axesbatch39-llm-cost-calculator-m2-r2 | cache 0/50/100%; batch 0/50/100%; same token shape | cache-read multiplier for one candidate is Unavailable — not returned in compatible registry units | No break-even matrix cell is emitted for that candidate. | Unavailable — cache-read multiplier is not returned in compatible registry units |
retry/call volumebatch39-llm-cost-calculator-m2-r3 | retry 0/5/20%; calls 1K/30K/1M; output 220 | call-volume scaling is linear; retry debit for candidate C is Unavailable — not returned | Retain a threshold only where retry billing is observed. | Unavailable — retry debit is not returned |
Module citations: OpenAI API pricing. All AI Ask evidence registry (verified 2026-08-27).
Current versus prior price-change budget ledger
Formula / rubric: Impact = replay(frozen workload, current registry) − replay(frozen workload, prior registry), with catalog/coverage changes separated from rate deltas.
Provenance: Current and immediately prior verified registry snapshots; no live pricing call is made. Unsupported fields fail closed as Unavailable.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State / reproducible bill |
|---|---|---|---|---|
small workloadbatch39-llm-cost-calculator-m3-r1 | 1K calls; 2K input/400 output; current $5/$15; prior $5/$15 | rate delta $0.00; model catalog unchanged. | Attribute only rate movement to budget impact. | NO CHANGE — replay closes. |
production workloadbatch39-llm-cost-calculator-m3-r2 | 30K calls; 12K input/1.2K output; prior output coverage absent | current replay exists; prior verbosity-adjusted output is Unavailable — not covered by prior snapshot | Do not label coverage change as price change. | Unavailable — prior verbosity-adjusted output is not covered |
scale workloadbatch39-llm-cost-calculator-m3-r3 | 1M calls; batch 40%; cache 60%; provider TTL mechanic changed | catalog and cache mechanics changed together; attributable delta is Unavailable — not separable | Publish no single price-change percentage. | Unavailable — catalog and cache mechanic deltas are not separable |
Module citations: All AI Ask model pricing registry. All AI Ask evidence registry (verified 2026-08-27).
The formula, published
outputTokensBilled = outputTokens × (verbosityIndex ?? 1) inputCost = inputPerM × inputTokens / 1e6 outputCost = outputPerM × outputTokensBilled / 1e6 effective = (inputCost + outputCost) × callsPerMonth with caching: inputCost × (1 − cacheableInputPct × 0.9) with batching: (inputCost + outputCost) × (1 − batchDiscountPct/100)
up to 90% cache-read saving is the largest sourced read discount in the provider table; provider-specific estimates use that provider's multiplier and account for cache writes. The free-form estimator assumes 30% of input is cache-eligible; pick a workload preset below for a shape-specific assumption instead.
Coverage: 20 of 68 priced models have a measured verbosity index (2+ graded runs). The rest render with a verbosity of — and are shown at unadjusted list price.
