How Much Does an LLM Chatbot Cost per Month?
At production volume (200,000 calls/month), the cheapest effective option is Amazon Nova Micro at $24.25/month. The most expensive frontier option, GPT-5.4 Pro, runs $27,504/month — A support or product chatbot re-sends a growing conversation history on every turn, so the input side of the bill grows even though each individual reply stays short.
How much does llm chatbot cost per month?
At production volume (200,000 calls/month), the cheapest effective option for llm chatbot is Amazon Nova Micro at $24.25 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $27,504 per month for the same workload.
Token shape
| Shape | growing conversation history in, short reply out |
| Input / output tokens per call | 2K in / 0K out |
| Cacheable input | 35% |
| Batch-eligible | No |
Input is a system prompt plus roughly 8-10 turns of history by the middle of a conversation; output is a typical single-paragraph reply. Only the system prompt is stable enough to cache.
What drives this workload's cost?
The main token-volume driver here is input tokens: each call sends 2K input tokens and requests up to 0K output tokens, at 200,000 calls per month in the default volume. 35% of input is modeled as cache-eligible, so repeated prefixes can reduce the input charge. It is not marked batch-eligible because the profile assumes a synchronous response.
Volume
Ranked cost — Production volume
| Model | Provider | List monthly | Effective monthly | Verbosity | Rank Δ |
|---|---|---|---|---|---|
| Amazon Nova Microbudget | Amazon | $26.60 | $24.25 | 0.76× | — |
| Amazon Nova Litebudget | Amazon | $45.60 | $44.09 | 0.91× | — |
| GPT-5 Nanobudgetlegacy | OpenAI | $52.00 | $44.51 | — | — |
| Gemini 2.5 Flash Litebudgetlegacy | $76.00 | $60.93 | — | ▲2 | |
| Ministral 8Bbudget | Mistral | $82.50 | $81.77 | 0.93× | ▲2 |
| GPT-OSS 20Bbudget | Groq | $57.00 | $97.53 | 2.93× | ▼2 |
| Mistral Small 3.1budget | Mistral | $114 | $108 | 0.85× | ▲4 |
| GPT-4o Minibudgetlegacy | OpenAI | $114 | $91.53 | — | — |
| Grok-3 Minibudgetlegacy | xAI | $114 | $114 | — | — |
| Llama 4 Maverickbudgetlegacy | Groq | $138 | $138 | — | ▲2 |
| GPT-OSS 120Bbudget | Groq | $114 | $148 | 1.82× | ▼1 |
| GPT-5.4 Nanobudgetlegacy | OpenAI | $184 | $134 | 0.78× | ▲1 |
| Codestralbudget | Mistral | $207 | $194 | 0.79× | ▲1 |
| Gemini 3.1 Flash Litebudgetlegacy | $225 | $174 | 0.87× | ▲2 | |
| Muse Spark 1.3 Contributorbudget | Meta | $62.00 | $229 | 12.95× | ▼10 |
Show all 68 models
| GPT-5 Minibudgetlegacy | OpenAI | $260 | $223 | — | ▲1 |
| GPT-OSS 120B (Cerebras)budget | Cerebras | $221 | $290 | 2.32× | ▼2 |
| Gemini 3.5 Flash Litebudget | $319 | $274 | — | ▲1 | |
| Gemini 2.5 Flashbudgetlegacy | $319 | $274 | — | ▲1 | |
| Mistral Large 3budget | Mistral | $345 | $345 | — | ▲1 |
| GLM-5.1midlegacy | Z.ai | $442 | $442 | — | ▲1 |
| DeepSeek V4 Flashbudget | DeepSeek | $304 | $451 | 2.59× | ▼4 |
| Qwen 3.8 30Bmid | Groq | $498 | $498 | — | — |
| Qwen 3.6 27Bmidlegacy | Groq | $498 | $498 | — | — |
| Qwen 3.7 Plusmid | Qwen | $524 | $524 | — | — |
| Amazon Nova Promid | Amazon | $608 | $608 | — | — |
| Gemini 3.7 Flashmid | $623 | $510 | — | — | |
| GPT-5.4 Minimidlegacy | OpenAI | $675 | $563 | — | — |
| Gemini 3.1 Flashmidlegacy | $675 | $562 | — | — | |
| Claude Haiku 4.5mid | Anthropic | $830 | $687 | — | ▲1 |
| o3-Minimidlegacy | OpenAI | $836 | $671 | — | ▲1 |
| Grok 4.3mid | xAI | $775 | $847 | 1.41× | ▼2 |
| Muse Spark 1.3mid | Meta | $898 | $898 | — | — |
| GPT-5.6 Lunamid | OpenAI | $900 | $750 | — | — |
| Mistral Medium 3mid | Mistral | $1,245 | $1,161 | 0.84× | ▲6 |
| Qwen 3.8 Maxmid | Qwen | $1,216 | $1,216 | — | ▲1 |
| Qwen 3.7 Maxmid | Qwen | $1,216 | $1,216 | — | ▲1 |
| Grok-3midlegacy | xAI | $1,240 | $1,240 | — | ▲1 |
| Gemini 3.6 Flashmid | $1,245 | $1,019 | — | ▲1 | |
| GPT-5midlegacy | OpenAI | $1,300 | $1,113 | — | ▲3 |
| Gemini 3.5 Flashmidlegacy | $1,350 | $1,124 | — | ▲3 | |
| Grok-4.20 Reasoningmid | xAI | $1,380 | $1,380 | — | ▲3 |
| Grok-4.20mid | xAI | $1,380 | $1,380 | — | ▲3 |
| Grok 4.6mid | xAI | $1,380 | $1,380 | — | ▲3 |
| Grok 4.5mid | xAI | $1,380 | $1,380 | — | ▲3 |
| Gemini 3.1 Promid | $1,800 | $1,188 | 0.63× | ▲5 | |
| GPT-4.1midlegacy | OpenAI | $1,520 | $1,220 | — | ▲2 |
| DeepSeek V4 Promid | DeepSeek | $911 | $1,548 | 3.30× | ▼13 |
| Claude Sonnet 5mid | Anthropic | $1,660 | $1,374 | — | ▲1 |
| GLM-5.2mid | Z.ai | $980 | $1,864 | 3.87× | ▼14 |
| GPT-4omidlegacy | OpenAI | $1,900 | $1,525 | — | ▲1 |
| GPT-5.6 Terramid | OpenAI | $2,250 | $1,875 | — | ▲1 |
| GPT-5.4midlegacy | OpenAI | $2,250 | $1,875 | — | ▲1 |
| Claude Sonnet 4.6mid | Anthropic | $2,490 | $2,061 | — | ▲1 |
| Claude Sonnet 4.5midlegacy | Anthropic | $2,490 | $2,061 | — | ▲1 |
| Claude Sonnet 4midlegacy | Anthropic | $2,490 | $2,061 | — | ▲1 |
| GLM 4.7 (Cerebras)mid | Cerebras | $1,273 | $2,530 | 7.53× | ▼15 |
| GPT-5.6 Solmid | OpenAI | $3,320 | $2,721 | — | — |
| Claude Opus 4.8mid | Anthropic | $4,150 | $3,366 | 0.96× | — |
| Claude Opus 4.7midlegacy | Anthropic | $4,150 | $3,436 | — | — |
| Claude Opus 4.6midlegacy | Anthropic | $4,150 | $3,436 | — | — |
| Claude Opus 4.5midlegacy | Anthropic | $4,150 | $3,436 | — | — |
| GPT-4 Turbofrontierlegacy | OpenAI | $6,900 | $5,402 | — | — |
| Claude Fable 5frontier | Anthropic | $8,300 | $6,871 | — | — |
| Claude Opus 5frontier | Anthropic | $12,450 | $10,307 | — | — |
| Claude Opus 4.1frontierlegacy | Anthropic | $12,450 | $10,307 | — | — |
| Claude Opus 4frontierlegacy | Anthropic | $12,450 | $10,307 | — | — |
| GPT-5.4 Profrontierlegacy | OpenAI | $27,000 | $23,010 | 1.04× | — |
Batch 40 · server-rendered decision evidence · verified 2026-08-27
Chatbot conversation cost evidence
Frozen inputs, formulas, provenance, and decision boundaries are visible in the initial HTML. Unsupported evidence fails closed as Unavailable.
Conversation-cohort history ledger
Formula / scoring rule: Conversation bill = Σ(turn input history × input rate + turn output × output rate); history policy changes the denominator.
Provenance: Frozen 30K monthly conversations, 8 turns, 1,200 initial input, 220 output/turn; unresolved chats are retained.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State |
|---|---|---|---|---|
full historybatch40-chatbot-m1-r1 | 8 turns; 1,200 initial; 220 output/turn; history retained | Approx input 9,680; output 1,760; bill = (9,680×5 + 1,760×15)/1M = $0.0748/conversation. | Do not multiply the first turn by eight. | CALCULATED — growing history. |
summary after turn 4batch40-chatbot-m1-r2 | turn 4 summary 600 tokens; turns 5–8 use summary + recent 2 turns | Estimated input 6,480; output 1,760; bill = $0.0588/conversation. | Summary quality and summary-call cost must be added when measured. | CALCULATED — summary path. |
unresolved cohortbatch40-chatbot-m1-r3 | 30K conversations; 12% unresolved; 2 extra turns; resolution log absent | Extra turns are counted; resolved denominator Unavailable — conversation outcome log is absent | Do not report cost per resolved conversation as cost per chat. | Unavailable — conversation outcome log is absent |
Module citation: All AI Ask evidence registry (verified 2026-08-27).
Turn-policy crossover surface
Formula / scoring rule: Policy cost = model turns + summary/compaction calls + retries; crossover is first policy with lower cost at the same resolution floor.
Provenance: Matched 4/8/12-turn scenarios; compaction is modeled as a separate call, not free state.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State |
|---|---|---|---|---|
4-turn capbatch40-chatbot-m2-r1 | 4 turns; 1,200 input; 220 output; stop reason=cap | Bill subtotal = $0.0408; unresolved rate Unavailable — observed resolution is absent | A lower cap cannot claim equivalent service. | Unavailable — observed resolution is absent |
8-turn + summarybatch40-chatbot-m2-r2 | 8 turns; summary call at turn 4; 600 summary tokens | Model bill calculable; summary acceptance Unavailable — not graded in fixture | Summary overhead remains in numerator. | Unavailable — not graded in fixture |
12-turn overflowbatch40-chatbot-m2-r3 | 12 turns; 3 retries; duplicate user messages 2% | Duplicate and retry tokens counted; exact bill Unavailable — retry usage export is absent | Do not hide overflow in average turn count. | Unavailable — retry usage export is absent |
Module citation: All AI Ask chatbot workload registry.
Resolved-conversation TCO tree
Formula / scoring rule: TCO/resolved = (model + moderation + retrieval + human escalation) / resolved conversations; unsupported components remain Unavailable.
Provenance: 30K conversation baseline; escalation minutes/rate are user-supplied; chatbot quality is out of scope.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State |
|---|---|---|---|---|
model subtotalbatch40-chatbot-m3-r1 | 30K × $0.0748 full-history conversation | Model subtotal = 30,000 × $0.0748 = $2,244. | This is token cost, not resolved-conversation TCO. | CALCULATED — token component. |
retrieval / moderationbatch40-chatbot-m3-r2 | retrieval calls 20%; moderation endpoint and vector fee | Additional fees Unavailable — provider/tool tariffs are not supplied | Do not treat ancillary calls as included for free. | Unavailable — provider/tool tariffs are not supplied |
human escalationbatch40-chatbot-m3-r3 | 6% escalated; 4 minutes; $60/hour; resolution denominator absent | Scenario labor = 1,800 × 4/60 × $60 = $7,200. | Publish per-resolved only after outcome counts close. | Unavailable — resolution denominator absent |
Module citation: All AI Ask chatbot calculator.
Prompt caching uses the provider's sourced read multiplier and amortises cache writes over the documented TTL; the largest published read discount is 90%. See the full formula and coverage disclosure →
