How Much Does an LLM Chatbot Cost per Month?

At production volume (200,000 calls/month), the cheapest effective option is Amazon Nova Micro at $24.25/month. The most expensive frontier option, GPT-5.4 Pro, runs $27,504/month — A support or product chatbot re-sends a growing conversation history on every turn, so the input side of the bill grows even though each individual reply stays short.

How much does llm chatbot cost per month?

At production volume (200,000 calls/month), the cheapest effective option for llm chatbot is Amazon Nova Micro at $24.25 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $27,504 per month for the same workload.

Verified 2026-08-08
Cheapest ≠ best. This page ranks by effective cost only. See our graded results for this task on Best LLM For →

Token shape

Shapegrowing conversation history in, short reply out
Input / output tokens per call2K in / 0K out
Cacheable input35%
Batch-eligibleNo

Input is a system prompt plus roughly 8-10 turns of history by the middle of a conversation; output is a typical single-paragraph reply. Only the system prompt is stable enough to cache.

What drives this workload's cost?

The main token-volume driver here is input tokens: each call sends 2K input tokens and requests up to 0K output tokens, at 200,000 calls per month in the default volume. 35% of input is modeled as cache-eligible, so repeated prefixes can reduce the input charge. It is not marked batch-eligible because the profile assumes a synchronous response.

Volume

Side project
20,000 calls/mo
$2.42/mo cheapest
Production
200,000 calls/mo
$24.25/mo cheapest
Scale
2,000,000 calls/mo
$242/mo cheapest

Ranked cost — Production volume

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$26.60$24.250.76×
Amazon Nova LitebudgetAmazon$45.60$44.090.91×
GPT-5 NanobudgetlegacyOpenAI$52.00$44.51
Gemini 2.5 Flash LitebudgetlegacyGoogle$76.00$60.932
Ministral 8BbudgetMistral$82.50$81.770.93×2
GPT-OSS 20BbudgetGroq$57.00$97.532.93×2
Mistral Small 3.1budgetMistral$114$1080.85×4
GPT-4o MinibudgetlegacyOpenAI$114$91.53
Grok-3 MinibudgetlegacyxAI$114$114
Llama 4 MaverickbudgetlegacyGroq$138$1382
GPT-OSS 120BbudgetGroq$114$1481.82×1
GPT-5.4 NanobudgetlegacyOpenAI$184$1340.78×1
CodestralbudgetMistral$207$1940.79×1
Gemini 3.1 Flash LitebudgetlegacyGoogle$225$1740.87×2
Muse Spark 1.3 ContributorbudgetMeta$62.00$22912.95×10
Show all 68 models
GPT-5 MinibudgetlegacyOpenAI$260$2231
GPT-OSS 120B (Cerebras)budgetCerebras$221$2902.32×2
Gemini 3.5 Flash LitebudgetGoogle$319$2741
Gemini 2.5 FlashbudgetlegacyGoogle$319$2741
Mistral Large 3budgetMistral$345$3451
GLM-5.1midlegacyZ.ai$442$4421
DeepSeek V4 FlashbudgetDeepSeek$304$4512.59×4
Qwen 3.8 30BmidGroq$498$498
Qwen 3.6 27BmidlegacyGroq$498$498
Qwen 3.7 PlusmidQwen$524$524
Amazon Nova PromidAmazon$608$608
Gemini 3.7 FlashmidGoogle$623$510
GPT-5.4 MinimidlegacyOpenAI$675$563
Gemini 3.1 FlashmidlegacyGoogle$675$562
Claude Haiku 4.5midAnthropic$830$6871
o3-MinimidlegacyOpenAI$836$6711
Grok 4.3midxAI$775$8471.41×2
Muse Spark 1.3midMeta$898$898
GPT-5.6 LunamidOpenAI$900$750
Mistral Medium 3midMistral$1,245$1,1610.84×6
Qwen 3.8 MaxmidQwen$1,216$1,2161
Qwen 3.7 MaxmidQwen$1,216$1,2161
Grok-3midlegacyxAI$1,240$1,2401
Gemini 3.6 FlashmidGoogle$1,245$1,0191
GPT-5midlegacyOpenAI$1,300$1,1133
Gemini 3.5 FlashmidlegacyGoogle$1,350$1,1243
Grok-4.20 ReasoningmidxAI$1,380$1,3803
Grok-4.20midxAI$1,380$1,3803
Grok 4.6midxAI$1,380$1,3803
Grok 4.5midxAI$1,380$1,3803
Gemini 3.1 PromidGoogle$1,800$1,1880.63×5
GPT-4.1midlegacyOpenAI$1,520$1,2202
DeepSeek V4 PromidDeepSeek$911$1,5483.30×13
Claude Sonnet 5midAnthropic$1,660$1,3741
GLM-5.2midZ.ai$980$1,8643.87×14
GPT-4omidlegacyOpenAI$1,900$1,5251
GPT-5.6 TerramidOpenAI$2,250$1,8751
GPT-5.4midlegacyOpenAI$2,250$1,8751
Claude Sonnet 4.6midAnthropic$2,490$2,0611
Claude Sonnet 4.5midlegacyAnthropic$2,490$2,0611
Claude Sonnet 4midlegacyAnthropic$2,490$2,0611
GLM 4.7 (Cerebras)midCerebras$1,273$2,5307.53×15
GPT-5.6 SolmidOpenAI$3,320$2,721
Claude Opus 4.8midAnthropic$4,150$3,3660.96×
Claude Opus 4.7midlegacyAnthropic$4,150$3,436
Claude Opus 4.6midlegacyAnthropic$4,150$3,436
Claude Opus 4.5midlegacyAnthropic$4,150$3,436
GPT-4 TurbofrontierlegacyOpenAI$6,900$5,402
Claude Fable 5frontierAnthropic$8,300$6,871
Claude Opus 5frontierAnthropic$12,450$10,307
Claude Opus 4.1frontierlegacyAnthropic$12,450$10,307
Claude Opus 4frontierlegacyAnthropic$12,450$10,307
GPT-5.4 ProfrontierlegacyOpenAI$27,000$23,0101.04×

Batch 40 · server-rendered decision evidence · verified 2026-08-27

Chatbot conversation cost evidence

Frozen inputs, formulas, provenance, and decision boundaries are visible in the initial HTML. Unsupported evidence fails closed as Unavailable.

Conversation-cohort history ledger

Formula / scoring rule: Conversation bill = Σ(turn input history × input rate + turn output × output rate); history policy changes the denominator.

Provenance: Frozen 30K monthly conversations, 8 turns, 1,200 initial input, 220 output/turn; unresolved chats are retained.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
full history
batch40-chatbot-m1-r1
8 turns; 1,200 initial; 220 output/turn; history retainedApprox input 9,680; output 1,760; bill = (9,680×5 + 1,760×15)/1M = $0.0748/conversation.Do not multiply the first turn by eight.CALCULATED — growing history.
summary after turn 4
batch40-chatbot-m1-r2
turn 4 summary 600 tokens; turns 5–8 use summary + recent 2 turnsEstimated input 6,480; output 1,760; bill = $0.0588/conversation.Summary quality and summary-call cost must be added when measured.CALCULATED — summary path.
unresolved cohort
batch40-chatbot-m1-r3
30K conversations; 12% unresolved; 2 extra turns; resolution log absentExtra turns are counted; resolved denominator Unavailable — conversation outcome log is absentDo not report cost per resolved conversation as cost per chat.Unavailable — conversation outcome log is absent

Module citation: All AI Ask evidence registry (verified 2026-08-27).

Turn-policy crossover surface

Formula / scoring rule: Policy cost = model turns + summary/compaction calls + retries; crossover is first policy with lower cost at the same resolution floor.

Provenance: Matched 4/8/12-turn scenarios; compaction is modeled as a separate call, not free state.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
4-turn cap
batch40-chatbot-m2-r1
4 turns; 1,200 input; 220 output; stop reason=capBill subtotal = $0.0408; unresolved rate Unavailable — observed resolution is absentA lower cap cannot claim equivalent service.Unavailable — observed resolution is absent
8-turn + summary
batch40-chatbot-m2-r2
8 turns; summary call at turn 4; 600 summary tokensModel bill calculable; summary acceptance Unavailable — not graded in fixtureSummary overhead remains in numerator.Unavailable — not graded in fixture
12-turn overflow
batch40-chatbot-m2-r3
12 turns; 3 retries; duplicate user messages 2%Duplicate and retry tokens counted; exact bill Unavailable — retry usage export is absentDo not hide overflow in average turn count.Unavailable — retry usage export is absent

Module citation: All AI Ask chatbot workload registry.

Resolved-conversation TCO tree

Formula / scoring rule: TCO/resolved = (model + moderation + retrieval + human escalation) / resolved conversations; unsupported components remain Unavailable.

Provenance: 30K conversation baseline; escalation minutes/rate are user-supplied; chatbot quality is out of scope.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState
model subtotal
batch40-chatbot-m3-r1
30K × $0.0748 full-history conversationModel subtotal = 30,000 × $0.0748 = $2,244.This is token cost, not resolved-conversation TCO.CALCULATED — token component.
retrieval / moderation
batch40-chatbot-m3-r2
retrieval calls 20%; moderation endpoint and vector feeAdditional fees Unavailable — provider/tool tariffs are not suppliedDo not treat ancillary calls as included for free.Unavailable — provider/tool tariffs are not supplied
human escalation
batch40-chatbot-m3-r3
6% escalated; 4 minutes; $60/hour; resolution denominator absentScenario labor = 1,800 × 4/60 × $60 = $7,200.Publish per-resolved only after outcome counts close.Unavailable — resolution denominator absent

Module citation: All AI Ask chatbot calculator.

Calculate cost per resolved conversation

Prompt caching uses the provider's sourced read multiplier and amortises cache writes over the documented TTL; the largest published read discount is 90%. See the full formula and coverage disclosure →

Related

Alternatives to Amazon Nova MicroAmazon cost calculator →RAG Question Answering cost →Coding Agent cost →

FAQ

Why does conversation history dominate the cost?
Most chat APIs are stateless — the full history is re-sent on every turn, so a long conversation costs more per reply even though the model only generates one short answer at a time.
Does prompt caching help chatbot costs?
Only for the stable part of the prompt (system instructions). The conversation history itself changes every turn, so it is rarely cache-eligible.