← All tasks

Best LLM for Chatbots & Support in 2026

For chatbots & support, Muse Spark 1.3 Contributor is our pick: $0.13/M tokens on a Support chat turn workload, 1.0M context.

Support chat is a latency-sensitive, extremely high-volume workload — users notice a slow reply far more than a slightly less polished one, and the per-call cost compounds fast at real support volumes.

Verdict: We have not run a controlled quality test for this task — ranking here is a requirements match on measured speed and price, not a conversation-quality comparison.

Quick answer: What is the best LLM for chatbots and customer support?

Muse Spark 1.3 Contributor, from Meta, is the best fit for chatbots & support at $0.13 per million task tokens on a Support chat turn workload, with a 1.0M-token context window. No cheaper value pick beats it for this task.

Verified 2026-08-08
Best overall
Muse Spark 1.3 Contributor
Meta · $0.13/M
Fit 87/100 — the top requirements match for this task.
Best value
Muse Spark 1.3 Contributor
Meta · $0.13/M
The strongest fit among budget and mid-tier priced models.
Fastest
GPT-OSS 120B (Cerebras)
Cerebras · $0.48/M
2450 tokens/sec measured.
Longest context
Gemini 3.1 Pro
Google · $5.33/M
2M token context window.

What evidence supports the Chatbots & Support recommendation?

We have not run a controlled test for chatbots & support. This ranking is a requirements match on price, measured throughput, and context window — not a quality comparison. Models that fit this task's requirements are ranked; which one performs best on your prompt is a question you should answer by running it. Run all three side by side →

Reproducible Chatbots & Support evidence and decision rubric

Test / runPrompt and verificationHard rule
Code Snippetexact prompt + 20 recorded runsCorrect iterative algorithm and code-only output
Hard Algorithmexact prompt + 4 recorded runs5,000-case harness; O(log n) partition and correct edge cases

Run dates: budget suite 2026-06-16T20:31:30.728Z; premium suite 2026-06-21T00:00:00.000Z. Results are not a claim about every repository or prompt.

Two-test rubric, failure analysis, and task-shaped ranking

ModelAccuracyLatencyOutput tokensRun costFailure / qualification note
GPT-5.4 Pro100/100206116 ms548$0.103Passes all 5,000+ randomized cases and every edge case; genuine O(log(min(m,n))) partition search, returns a float, code-only, and explicitly raises on two empty lists. Correct — but the slowest run by far (over three minutes), and now that real usage is reported, comfortably the most expensive.
Claude Opus 4.8100/1003790 ms382$0.011Passes all 5,000+ randomized cases and every edge case; genuine O(log(min(m,n))) partition, float return, code-only. The fastest correct solution in this task.
GLM 5.2 (Max)100/10047713 ms2148$0.010Passes all 5,000+ randomized cases and every edge case with a genuine O(log) partition and float return. The <think> block ahead of the code is GLM's reasoning channel surfaced by our gateway, not reasoning dumped into the answer — GLM's actual content is the clean code block — so it scores level with the other correct solutions, as the cheapest of them.
Gemini 3.1 Pro99/10017413 ms320$0.004Passes all tests with a clean, minimal O(log) partition and float return, code-only. Docked one point only because two empty lists yield NaN rather than an explicit guard (not required by the prompt).

Availability caveat: short code tests do not establish repository-scale debugging, multi-file tool use, or agent reliability. The fastest acceptable verdict must therefore clear the correctness rule before speed is considered.

Task-shaped cost ranking (20,000 tasks/month)

RankModelEffective monthlyMeasured verbosity
1Amazon Nova Micro$18.260.76×
2Amazon Nova Lite$32.740.91×
3GPT-5 Nano$36.00Unavailable; neutral fallback
4Gemini 2.5 Flash Lite$56.00Unavailable; neutral fallback
5GPT-OSS 20B$65.162.93×
6Ministral 8B$65.580.93×

Verified 2026-08-08. full prompt/run evidence

Try these models for Chatbots & Support

Volume, requirement-gate, and pricing cross-check for Chatbots & Support

Monthly spend ladder at Support chat turn shape

Calls / monthOverall pick monthlyBudget pick monthlyOverall − budget delta
250,000$20.00$20.00$0.0000
500,000$40.00$40.00$0.0000
1,000,000$80.00$80.00$0.0000
2,500,000$200.00$200.00$0.0000

Monthly cost = task price/M × (400 input + 200 output tokens) × calls ÷ 1,000,000, at 0.5×, 1×, 2×, 5× the published 500,000-call/month baseline.

Requirement-gate margin for the picked models

ModelRequirementMeasured valueMargin / result
No hard requirement filter for this taskrequires: {}All qualified candidates pass by default

Requirement gates are hard filters, not down-ranking: a model failing any row here is excluded from Chatbots & Support candidates entirely, regardless of price or speed.

Pricing-page cross-link for each pick

ModelTask-weighted $/MMonthly at published volumeMeasured throughput
Muse Spark 1.3 Contributor$0.13$40.00Unavailable
GPT-OSS 120B (Cerebras)$0.48$145.002450 tok/s
Gemini 3.1 Pro$5.33$1600.0055 tok/s

Evidence coverage: 0 of 49 candidates have a graded run. No graded accuracy evidence exists for this task; the ranking above is a requirements-and-price match, not a quality claim.

Test the Chatbots & Support picks side by side →

Verified 2026-08-08. "Unavailable" means no compatible dated evidence was found for that field; it is never treated as zero. Dated source · Full ranking and rubric

Batch 10 chatbot economics

1. Conversation-history cost curve

TurnsRetained input assumptionTotal cost / turnCurrent-turn spendHistory spend
1400 current + 0 history$0.0001$0.0001$0.0000
5400 current + 1,600 history$0.0002$0.0001$0.0002
10400 current + 3,600 history$0.0004$0.0001$0.0004
20400 current + 7,600 history$0.0008$0.0001$0.0008

Formula: cost = exact model input + output token rates from MODEL_PRICING. Current-turn spend includes current input and turn output tokens; history spend is input tokens for prior turns. Total cost = Current-turn spend + History spend.

2. Support capacity envelope

Declared arrival rateCompletion lengthTurn latency (TTFT + output ÷ tok/s)Required concurrent seatsMonthly workloadWindow headroom
10 req/min50 tokensUnavailableUnavailable432K turns/mo1,046,126 tokens
10 req/min200 tokensUnavailableUnavailable432K turns/mo1,045,976 tokens
10 req/min500 tokensUnavailableUnavailable432K turns/mo1,045,676 tokens
50 req/min50 tokensUnavailableUnavailable2160K turns/mo1,046,126 tokens
50 req/min200 tokensUnavailableUnavailable2160K turns/mo1,045,976 tokens
50 req/min500 tokensUnavailableUnavailable2160K turns/mo1,045,676 tokens
200 req/min50 tokensUnavailableUnavailable8640K turns/mo1,046,126 tokens
200 req/min200 tokensUnavailableUnavailable8640K turns/mo1,045,976 tokens
200 req/min500 tokensUnavailableUnavailable8640K turns/mo1,045,676 tokens

Required concurrent seats = ceil((arrival rate ÷ 60) × turn latency). Completion times are calculated from dated speed records for 50, 200, and 500 output tokens. Window headroom = context window − (2,000 system + 400 input + output tokens).

3. Routine-to-escalation router

Escalation sharePremium modelBlended monthly costBlended latencyMix
0%Claude Opus 4.8$40.00Unavailable0% premium + 100% routine
5%Claude Opus 4.8$213.00Unavailable5% premium + 95% routine
10%Claude Opus 4.8$386.00Unavailable10% premium + 90% routine
20%Claude Opus 4.8$732.00Unavailable20% premium + 80% routine

Blended cost and latency are weighted across routine Muse Spark 1.3 Contributor and the named premium candidate. Premium quality remains Unavailable.

Stale-data behavior: dated registry values are snapshots. If a source is older than the page verification date, or a provider changes its policy/pricing/model, re-verify before production use; unknown values remain Unavailable.

Verified 2026-08-08. Luna is the data owner. “Unavailable” means no compatible dated evidence was found; it is never zero, an estimate, or a guessed policy. Stale-data behavior: dated registry values are snapshots. If a source is older than the page verification date, or a provider changes its policy/pricing/model, re-verify before production use; unknown values remain Unavailable. Dated registry source · Run this scenario yourself →

Which models rank highest for Chatbots & Support?

"Fit" is a requirements match, not a quality benchmark — it combines price, measured speed, context window, and (where we have run it) graded accuracy on this task. Formula below.

#ModelProviderFitEvidenceTask price/MTokens/secContextScored on
1Muse Spark 1.3 ContributorMeta87$0.131.0Mprice, context
2Gemini 2.5 Flash LiteLegacyGoogle82$0.201Mprice, context
3GPT-OSS 120B (Cerebras)Cerebras76$0.482450131Kprice, context, speed
4GPT-4o MiniLegacyOpenAI63$0.30128Kprice, context
5Gemini 2.5 FlashLegacyGoogle60$1.031Mprice, context
6GPT-OSS 20BGroq60$0.151120131Kprice, context, speed
7GLM 4.7 (Cerebras)Cerebras57$2.421980200Kprice, context, speed
8Gemini 3.7 FlashGoogle53$1.751.0Mprice, context

What will Chatbots & Support cost?

At 500,000 support chat turn calls/month:

ModelTask price/MEst. monthly cost
Muse Spark 1.3 Contributor$0.13$40.00
Gemini 2.5 Flash Lite$0.20$60.00
GPT-OSS 120B (Cerebras)$0.48$145.00

How is the best LLM for Chatbots & Support ranked?

Weights: evidence 0%, price 45%, speed 45%, context 10%.

Requirements: none — every current model is eligible. 49 models eligible.

Price and context sub-scores are min-max normalised (log-scaled) within this task's eligible set only. Speed uses measured tokens/sec only — estimated rows are excluded. A model missing a measurement is never scored as zero: its weight is redistributed across the components we do have, and "Scored on" in the table above shows exactly which ones.

Prices verified 2026-08-08.

What related resources help with Chatbots & Support?

Meta provider hubMuse Spark 1.3 Contributor pricingBest LLM for CodingBest LLM for Math & ReasoningBest LLM for Structured Data Extraction

What are common questions about the best LLM for Chatbots & Support?

Why not use a flagship model for every support ticket?

At 500,000 turns a month the price gap between a flagship and a fast budget model is enormous, and most support turns don't need frontier-level reasoning.

How fast does a chatbot model need to be?

Aim for sub-second time-to-first-token where possible — measured tokens/sec on this page is a proxy for how quickly a reply starts streaming.

Should I escalate hard questions to a bigger model?

Yes — a common pattern is routing routine turns to a fast, cheap model and escalating low-confidence or complex turns to a stronger one.

Run this exact prompt against the top 3

Don't take a ranking's word for it — try Muse Spark 1.3 Contributor and its closest alternatives on your own prompt.

Try It Free