How Much Does RAG Question Answering Cost per Month?
At production volume (60,000 calls/month), the cheapest effective option is Amazon Nova Micro at $27.75/month. The most expensive frontier option, GPT-5.4 Pro, runs $26,093/month — RAG pays mostly for the retrieved context you stuff into the prompt — the answer itself is usually short relative to the chunks that justify it.
How much does rag question answering cost per month?
At production volume (60,000 calls/month), the cheapest effective option for rag question answering is Amazon Nova Micro at $27.75 per month, verbosity-adjusted rather than list price. The most expensive frontier model, GPT-5.4 Pro, runs $26,093 per month for the same workload.
Token shape
| Shape | large retrieved context in, short answer out |
| Input / output tokens per call | 12K in / 0K out |
| Cacheable input | 50% |
| Batch-eligible | No |
Input is a system prompt plus several retrieved chunks (roughly 12K tokens); a stable system prompt and frequently-reused chunks make about half the input cacheable in a well-tuned pipeline.
What drives this workload's cost?
The main token-volume driver here is input tokens: each call sends 12K input tokens and requests up to 0K output tokens, at 60,000 calls per month in the default volume. 50% of input is modeled as cache-eligible, so repeated prefixes can reduce the input charge. It is not marked batch-eligible because the profile assumes a synchronous response.
Volume
Ranked cost — Production volume
| Model | Provider | List monthly | Effective monthly | Verbosity | Rank Δ |
|---|---|---|---|---|---|
| Amazon Nova Microbudget | Amazon | $28.56 | $27.75 | 0.76× | — |
| GPT-5 Nanobudgetlegacy | OpenAI | $45.60 | $29.90 | — | — |
| Amazon Nova Litebudget | Amazon | $48.96 | $48.44 | 0.91× | — |
| GPT-OSS 20Bbudget | Groq | $61.20 | $75.10 | 2.93× | — |
| Gemini 2.5 Flash Litebudgetlegacy | $81.60 | $49.59 | — | ▲1 | |
| Ministral 8Bbudget | Mistral | $112 | $111 | 0.93× | ▲1 |
| Mistral Small 3.1budget | Mistral | $122 | $120 | 0.85× | ▲4 |
| GPT-4o Minibudgetlegacy | OpenAI | $122 | $75.29 | — | — |
| Grok-3 Minibudgetlegacy | xAI | $122 | $122 | — | — |
| Muse Spark 1.3 Contributorbudget | Meta | $76.80 | $134 | 12.95× | ▼5 |
| GPT-OSS 120Bbudget | Groq | $122 | $134 | 1.82× | ▼1 |
| Llama 4 Maverickbudgetlegacy | Groq | $158 | $158 | — | — |
| GPT-5.4 Nanobudgetlegacy | OpenAI | $174 | $105 | 0.78× | — |
| Gemini 3.1 Flash Litebudgetlegacy | $216 | $131 | 0.87× | — | |
| GPT-5 Minibudgetlegacy | OpenAI | $228 | $149 | — | — |
Show all 68 models
| Codestralbudget | Mistral | $238 | $233 | 0.79× | — |
| Gemini 3.5 Flash Litebudget | $276 | $180 | — | ▲1 | |
| Gemini 2.5 Flashbudgetlegacy | $276 | $180 | — | ▲1 | |
| GPT-OSS 120B (Cerebras)budget | Cerebras | $270 | $294 | 2.32× | ▼2 |
| Mistral Large 3budget | Mistral | $396 | $396 | — | ▲1 |
| DeepSeek V4 Flashbudget | DeepSeek | $348 | $399 | 2.59× | ▼1 |
| GLM-5.1midlegacy | Z.ai | $485 | $485 | — | — |
| Qwen 3.8 30Bmid | Groq | $504 | $504 | — | — |
| Qwen 3.6 27Bmidlegacy | Groq | $504 | $504 | — | — |
| Qwen 3.7 Plusmid | Qwen | $624 | $624 | — | — |
| Gemini 3.7 Flashmid | $630 | $390 | — | — | |
| GPT-5.4 Minimidlegacy | OpenAI | $648 | $412 | — | — |
| Gemini 3.1 Flashmidlegacy | $648 | $408 | — | — | |
| Amazon Nova Promid | Amazon | $653 | $653 | — | — |
| Claude Haiku 4.5mid | Anthropic | $840 | $576 | — | — |
| GPT-5.6 Lunamid | OpenAI | $864 | $550 | — | — |
| o3-Minimidlegacy | OpenAI | $898 | $552 | — | — |
| Grok 4.3mid | xAI | $960 | $985 | 1.41× | — |
| Muse Spark 1.3mid | Meta | $1,002 | $1,002 | — | — |
| GPT-5midlegacy | OpenAI | $1,140 | $747 | — | ▲2 |
| Mistral Medium 3mid | Mistral | $1,260 | $1,231 | 0.84× | ▲3 |
| Gemini 3.6 Flashmid | $1,260 | $780 | — | ▲1 | |
| DeepSeek V4 Promid | DeepSeek | $1,045 | $1,264 | 3.30× | ▼3 |
| Gemini 3.5 Flashmidlegacy | $1,296 | $816 | — | ▲1 | |
| Qwen 3.8 Maxmid | Qwen | $1,306 | $1,306 | — | ▲1 |
| Qwen 3.7 Maxmid | Qwen | $1,306 | $1,306 | — | ▲1 |
| GLM-5.2mid | Z.ai | $1,114 | $1,417 | 3.87× | ▼6 |
| Grok-3midlegacy | xAI | $1,536 | $1,536 | — | — |
| Grok-4.20 Reasoningmid | xAI | $1,584 | $1,584 | — | — |
| Grok-4.20mid | xAI | $1,584 | $1,584 | — | — |
| Grok 4.6mid | xAI | $1,584 | $1,584 | — | — |
| Grok 4.5mid | xAI | $1,584 | $1,584 | — | — |
| Gemini 3.1 Promid | $1,728 | $981 | 0.63× | ▲3 | |
| GPT-4.1midlegacy | OpenAI | $1,632 | $1,004 | — | ▼1 |
| Claude Sonnet 5mid | Anthropic | $1,680 | $1,151 | — | ▼1 |
| GPT-4omidlegacy | OpenAI | $2,040 | $1,255 | — | ▲1 |
| GLM 4.7 (Cerebras)mid | Cerebras | $1,686 | $2,117 | 7.53× | ▼2 |
| GPT-5.6 Terramid | OpenAI | $2,160 | $1,375 | — | — |
| GPT-5.4midlegacy | OpenAI | $2,160 | $1,375 | — | — |
| Claude Sonnet 4.6mid | Anthropic | $2,520 | $1,727 | — | — |
| Claude Sonnet 4.5midlegacy | Anthropic | $2,520 | $1,727 | — | — |
| Claude Sonnet 4midlegacy | Anthropic | $2,520 | $1,727 | — | — |
| GPT-5.6 Solmid | OpenAI | $3,360 | $2,104 | — | — |
| Claude Opus 4.8mid | Anthropic | $4,200 | $2,854 | 0.96× | — |
| Claude Opus 4.7midlegacy | Anthropic | $4,200 | $2,878 | — | — |
| Claude Opus 4.6midlegacy | Anthropic | $4,200 | $2,878 | — | — |
| Claude Opus 4.5midlegacy | Anthropic | $4,200 | $2,878 | — | — |
| GPT-4 Turbofrontierlegacy | OpenAI | $7,920 | $4,779 | — | — |
| Claude Fable 5frontier | Anthropic | $8,400 | $5,756 | — | — |
| Claude Opus 5frontier | Anthropic | $12,600 | $8,634 | — | — |
| Claude Opus 4.1frontierlegacy | Anthropic | $12,600 | $8,634 | — | — |
| Claude Opus 4frontierlegacy | Anthropic | $12,600 | $8,634 | — | — |
| GPT-5.4 Profrontierlegacy | OpenAI | $25,920 | $16,671 | 1.04× | — |
Batch 39 · server-rendered decision evidence · verified 2026-08-27
End-to-end RAG cost waterfall
Every field is tied to a frozen input and a dated provenance record. Unsupported facts fail closed as Unavailable; they are not treated as zero, free, equivalent, current, fastest, cheapest, private, or passing.
End-to-end RAG cost waterfall
Formula / rubric: Monthly = parse + initial/incremental embeddings + storage + vector reads + rerank + prompt/cache + output + evaluation + retries.
Provenance: Frozen corpus 1.2M tokens, 30K queries/month, top-k 5; incompatible page/vector/token units stay separate. Unsupported fields fail closed as Unavailable.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State / reproducible bill |
|---|---|---|---|---|
initial ingestionbatch39-rag-question-answering-m1-r1 | 1.2M tokens; 2,400 chunks; 768 dimensions; parse $0.40; embedding $0.13/M | parse $0.40 + embedding $0.16 = $0.56 one-time. | Do not blend one-time ingest with recurring query cost. | CALCULATED — one-time. |
monthly query pathbatch39-rag-question-answering-m1-r2 | 30K queries; 5 vectors/read; rerank 30K; 900 input/180 output each; 20% cache | generation and rerank close; vector-read unit is Unavailable — provider/vector-store compatible unit is not sourced | Recurring total remains unavailable until vector unit is compatible. | Unavailable — provider/vector-store compatible unit is not sourced |
evaluation/retrybatch39-rag-question-answering-m1-r3 | 10% sampled; 3% retries; grounded acceptance denominator required | evaluation call count = 3,000; accepted-answer cost Unavailable — grounded acceptance denominator is not returned | Never call cost per grounded answer equal to cost per API response. | Unavailable — grounded acceptance denominator is not returned |
Module citations: Google embedding and generative AI pricing. All AI Ask evidence registry (verified 2026-08-27).
Chunking and retrieval sensitivity cube
Formula / rubric: chunks = corpusTokens/(chunkTokens×(1−overlap)); retrievedPrompt = topK×chunkTokens + questionTokens; monthly cost uses observed units.
Provenance: Nine frozen combinations: 256/512/1,024 tokens × 0/10/20% overlap × top-k 3/5/10. Unsupported fields fail closed as Unavailable.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State / reproducible bill |
|---|---|---|---|---|
256 tokens / 0% / k3batch39-rag-question-answering-m2-r1 | 1.2M corpus; 4,688 vectors; 768 dimensions; 30K queries | retrieved prompt ≈ 792 tokens/query; vector count is reproducible. | Quality metric still required: answer support recall at k3. | Unavailable — support-recall run is not present |
512 tokens / 10% / k5batch39-rag-question-answering-m2-r2 | 1.2M corpus; 2,604 vectors; 768 dimensions; 30K queries | retrieved prompt ≈ 2,660 tokens/query; monthly generation bill can be calculated. | Prefer only if support recall meets declared floor. | Unavailable — support-recall run is not present |
1,024 tokens / 20% / k10batch39-rag-question-answering-m2-r3 | 1.2M corpus; 1,465 vectors; 768 dimensions; 30K queries | retrieved prompt ≈ 10,280 tokens/query; context fit is eligible. | Reject if context or support-recall constraint fails. | Unavailable — support-recall run is not present |
Module citations: Pinecone pricing documentation. All AI Ask evidence registry (verified 2026-08-27).
Reindex and tenant attribution ledger
Formula / rubric: Tenant monthly = shared storage/index allocation + tenant churn embeddings + tenant vector reads + tenant generation; isolation/egress remain explicit.
Provenance: 1/10/100 tenant allocations at 1/10/30% monthly corpus churn; shared overhead allocated by indexed tokens. Unsupported fields fail closed as Unavailable.
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | State / reproducible bill |
|---|---|---|---|---|
1 tenant / 1% churnbatch39-rag-question-answering-m3-r1 | 1.2M tokens; 12K incremental tokens; shared storage 1 index | incremental embedding 12K tokens; shared allocation 100%; egress Unavailable — egress unit is not sourced | Publish only token and storage components that share units. | Unavailable — egress unit is not sourced |
10 tenants / 10% churnbatch39-rag-question-answering-m3-r2 | 120K incremental tokens; 10% shared index allocation; 30K queries split evenly | tenant generation allocations close; minimum monthly charge is Unavailable — provider minimum is not sourced | No per-tenant total until minimums and isolation are known. | Unavailable — provider minimum is not sourced |
100 tenants / 30% churnbatch39-rag-question-answering-m3-r3 | 360K incremental tokens; 100 namespaces; per-tenant egress and isolation requested | churn embedding count closes; namespace isolation and egress are not observed. | Keep shared and tenant-specific spend separate. | Unavailable — namespace isolation and egress are not observed |
Module citations: AWS OpenSearch pricing. All AI Ask evidence registry (verified 2026-08-27).
Prompt caching uses the provider's sourced read multiplier and amortises cache writes over the documented TTL; the largest published read discount is 90%. See the full formula and coverage disclosure →
