How Much Does RAG Question Answering Cost per Month?

At production volume (60,000 calls/month), the cheapest effective option is Amazon Nova Micro at $27.75/month. The most expensive frontier option, GPT-5.4 Pro, runs $26,093/month — RAG pays mostly for the retrieved context you stuff into the prompt — the answer itself is usually short relative to the chunks that justify it.

How much does rag question answering cost per month?

At production volume (60,000 calls/month), the cheapest effective option for rag question answering is Amazon Nova Micro at $27.75 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $26,093 per month for the same workload.

Verified 2026-08-08
Cheapest ≠ best. This page ranks by effective cost only. See our graded results for this task on Best LLM For →

Token shape

Shapelarge retrieved context in, short answer out
Input / output tokens per call12K in / 0K out
Cacheable input50%
Batch-eligibleNo

Input is a system prompt plus several retrieved chunks (roughly 12K tokens); a stable system prompt and frequently-reused chunks make about half the input cacheable in a well-tuned pipeline.

What drives this workload's cost?

The main token-volume driver here is input tokens: each call sends 12K input tokens and requests up to 0K output tokens, at 60,000 calls per month in the default volume. 50% of input is modeled as cache-eligible, so repeated prefixes can reduce the input charge. It is not marked batch-eligible because the profile assumes a synchronous response.

Volume

Side project
5,000 calls/mo
$2.31/mo cheapest
Production
60,000 calls/mo
$27.75/mo cheapest
Scale
600,000 calls/mo
$278/mo cheapest

Ranked cost — Production volume

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Amazon Nova MicrobudgetAmazon$28.56$27.750.76×
GPT-5 NanobudgetlegacyOpenAI$45.60$29.90
Amazon Nova LitebudgetAmazon$48.96$48.440.91×
GPT-OSS 20BbudgetGroq$61.20$75.102.93×
Gemini 2.5 Flash LitebudgetlegacyGoogle$81.60$49.591
Ministral 8BbudgetMistral$112$1110.93×1
Mistral Small 3.1budgetMistral$122$1200.85×4
GPT-4o MinibudgetlegacyOpenAI$122$75.29
Grok-3 MinibudgetlegacyxAI$122$122
Muse Spark 1.3 ContributorbudgetMeta$76.80$13412.95×5
GPT-OSS 120BbudgetGroq$122$1341.82×1
Llama 4 MaverickbudgetlegacyGroq$158$158
GPT-5.4 NanobudgetlegacyOpenAI$174$1050.78×
Gemini 3.1 Flash LitebudgetlegacyGoogle$216$1310.87×
GPT-5 MinibudgetlegacyOpenAI$228$149
Show all 68 models
CodestralbudgetMistral$238$2330.79×
Gemini 3.5 Flash LitebudgetGoogle$276$1801
Gemini 2.5 FlashbudgetlegacyGoogle$276$1801
GPT-OSS 120B (Cerebras)budgetCerebras$270$2942.32×2
Mistral Large 3budgetMistral$396$3961
DeepSeek V4 FlashbudgetDeepSeek$348$3992.59×1
GLM-5.1midlegacyZ.ai$485$485
Qwen 3.8 30BmidGroq$504$504
Qwen 3.6 27BmidlegacyGroq$504$504
Qwen 3.7 PlusmidQwen$624$624
Gemini 3.7 FlashmidGoogle$630$390
GPT-5.4 MinimidlegacyOpenAI$648$412
Gemini 3.1 FlashmidlegacyGoogle$648$408
Amazon Nova PromidAmazon$653$653
Claude Haiku 4.5midAnthropic$840$576
GPT-5.6 LunamidOpenAI$864$550
o3-MinimidlegacyOpenAI$898$552
Grok 4.3midxAI$960$9851.41×
Muse Spark 1.3midMeta$1,002$1,002
GPT-5midlegacyOpenAI$1,140$7472
Mistral Medium 3midMistral$1,260$1,2310.84×3
Gemini 3.6 FlashmidGoogle$1,260$7801
DeepSeek V4 PromidDeepSeek$1,045$1,2643.30×3
Gemini 3.5 FlashmidlegacyGoogle$1,296$8161
Qwen 3.8 MaxmidQwen$1,306$1,3061
Qwen 3.7 MaxmidQwen$1,306$1,3061
GLM-5.2midZ.ai$1,114$1,4173.87×6
Grok-3midlegacyxAI$1,536$1,536
Grok-4.20 ReasoningmidxAI$1,584$1,584
Grok-4.20midxAI$1,584$1,584
Grok 4.6midxAI$1,584$1,584
Grok 4.5midxAI$1,584$1,584
Gemini 3.1 PromidGoogle$1,728$9810.63×3
GPT-4.1midlegacyOpenAI$1,632$1,0041
Claude Sonnet 5midAnthropic$1,680$1,1511
GPT-4omidlegacyOpenAI$2,040$1,2551
GLM 4.7 (Cerebras)midCerebras$1,686$2,1177.53×2
GPT-5.6 TerramidOpenAI$2,160$1,375
GPT-5.4midlegacyOpenAI$2,160$1,375
Claude Sonnet 4.6midAnthropic$2,520$1,727
Claude Sonnet 4.5midlegacyAnthropic$2,520$1,727
Claude Sonnet 4midlegacyAnthropic$2,520$1,727
GPT-5.6 SolmidOpenAI$3,360$2,104
Claude Opus 4.8midAnthropic$4,200$2,8540.96×
Claude Opus 4.7midlegacyAnthropic$4,200$2,878
Claude Opus 4.6midlegacyAnthropic$4,200$2,878
Claude Opus 4.5midlegacyAnthropic$4,200$2,878
GPT-4 TurbofrontierlegacyOpenAI$7,920$4,779
Claude Fable 5frontierAnthropic$8,400$5,756
Claude Opus 5frontierAnthropic$12,600$8,634
Claude Opus 4.1frontierlegacyAnthropic$12,600$8,634
Claude Opus 4frontierlegacyAnthropic$12,600$8,634
GPT-5.4 ProfrontierlegacyOpenAI$25,920$16,6711.04×

Batch 39 · server-rendered decision evidence · verified 2026-08-27

End-to-end RAG cost waterfall

Every field is tied to a frozen input and a dated provenance record. Unsupported facts fail closed as Unavailable; they are not treated as zero, free, equivalent, current, fastest, cheapest, private, or passing.

End-to-end RAG cost waterfall

Formula / rubric: Monthly = parse + initial/incremental embeddings + storage + vector reads + rerank + prompt/cache + output + evaluation + retries.

Provenance: Frozen corpus 1.2M tokens, 30K queries/month, top-k 5; incompatible page/vector/token units stay separate. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
initial ingestion
batch39-rag-question-answering-m1-r1
1.2M tokens; 2,400 chunks; 768 dimensions; parse $0.40; embedding $0.13/Mparse $0.40 + embedding $0.16 = $0.56 one-time.Do not blend one-time ingest with recurring query cost.CALCULATED — one-time.
monthly query path
batch39-rag-question-answering-m1-r2
30K queries; 5 vectors/read; rerank 30K; 900 input/180 output each; 20% cachegeneration and rerank close; vector-read unit is Unavailable — provider/vector-store compatible unit is not sourcedRecurring total remains unavailable until vector unit is compatible.Unavailable — provider/vector-store compatible unit is not sourced
evaluation/retry
batch39-rag-question-answering-m1-r3
10% sampled; 3% retries; grounded acceptance denominator requiredevaluation call count = 3,000; accepted-answer cost Unavailable — grounded acceptance denominator is not returnedNever call cost per grounded answer equal to cost per API response.Unavailable — grounded acceptance denominator is not returned

Module citations: Google embedding and generative AI pricing. All AI Ask evidence registry (verified 2026-08-27).

Chunking and retrieval sensitivity cube

Formula / rubric: chunks = corpusTokens/(chunkTokens×(1−overlap)); retrievedPrompt = topK×chunkTokens + questionTokens; monthly cost uses observed units.

Provenance: Nine frozen combinations: 256/512/1,024 tokens × 0/10/20% overlap × top-k 3/5/10. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
256 tokens / 0% / k3
batch39-rag-question-answering-m2-r1
1.2M corpus; 4,688 vectors; 768 dimensions; 30K queriesretrieved prompt ≈ 792 tokens/query; vector count is reproducible.Quality metric still required: answer support recall at k3.Unavailable — support-recall run is not present
512 tokens / 10% / k5
batch39-rag-question-answering-m2-r2
1.2M corpus; 2,604 vectors; 768 dimensions; 30K queriesretrieved prompt ≈ 2,660 tokens/query; monthly generation bill can be calculated.Prefer only if support recall meets declared floor.Unavailable — support-recall run is not present
1,024 tokens / 20% / k10
batch39-rag-question-answering-m2-r3
1.2M corpus; 1,465 vectors; 768 dimensions; 30K queriesretrieved prompt ≈ 10,280 tokens/query; context fit is eligible.Reject if context or support-recall constraint fails.Unavailable — support-recall run is not present

Module citations: Pinecone pricing documentation. All AI Ask evidence registry (verified 2026-08-27).

Reindex and tenant attribution ledger

Formula / rubric: Tenant monthly = shared storage/index allocation + tenant churn embeddings + tenant vector reads + tenant generation; isolation/egress remain explicit.

Provenance: 1/10/100 tenant allocations at 1/10/30% monthly corpus churn; shared overhead allocated by indexed tokens. Unsupported fields fail closed as Unavailable.

Frozen fixture / runVisible inputsField-level resultDecision boundaryState / reproducible bill
1 tenant / 1% churn
batch39-rag-question-answering-m3-r1
1.2M tokens; 12K incremental tokens; shared storage 1 indexincremental embedding 12K tokens; shared allocation 100%; egress Unavailable — egress unit is not sourcedPublish only token and storage components that share units.Unavailable — egress unit is not sourced
10 tenants / 10% churn
batch39-rag-question-answering-m3-r2
120K incremental tokens; 10% shared index allocation; 30K queries split evenlytenant generation allocations close; minimum monthly charge is Unavailable — provider minimum is not sourcedNo per-tenant total until minimums and isolation are known.Unavailable — provider minimum is not sourced
100 tenants / 30% churn
batch39-rag-question-answering-m3-r3
360K incremental tokens; 100 namespaces; per-tenant egress and isolation requestedchurn embedding count closes; namespace isolation and egress are not observed.Keep shared and tenant-specific spend separate.Unavailable — namespace isolation and egress are not observed

Module citations: AWS OpenSearch pricing. All AI Ask evidence registry (verified 2026-08-27).

Model your grounded-answer workload

Prompt caching uses the provider's sourced read multiplier and amortises cache writes over the documented TTL; the largest published read discount is 90%. See the full formula and coverage disclosure →

Related

Alternatives to Amazon Nova MicroAmazon cost calculator →LLM Chatbot cost →Coding Agent cost →

FAQ

Why is RAG input-heavy?
The point of retrieval is to give the model enough grounding text to answer accurately — that context is usually many times larger than the answer it produces.
Does a bigger context window make RAG cheaper?
No — window size is a capacity limit, not a price discount. You still pay per input token regardless of how large the window is.