← All tasks

Best LLM for Long Documents & RAG in 2026

For long documents & rag, Muse Spark 1.3 Contributor is our pick: $0.10/M tokens on a Long-document Q&A workload, 1.0M context.

Feeding a whole document or codebase into a single call needs headroom well beyond the text itself — retrieval overhead, system prompts, and chat history all eat into the window. We require at least 200K tokens of context and rank primarily on window size.

Verdict: We have not run a controlled quality test for this task — every current model handles a well-formed long-context prompt competently, so context window and price are what actually separate the candidates.

Quick answer: What is the best LLM for long documents and RAG?

Muse Spark 1.3 Contributor, from Meta, is the best fit for long documents & rag at $0.10 per million task tokens on a Long-document Q&A workload, with a 1.0M-token context window. No cheaper value pick beats it for this task.

Verified 2026-08-08
Best overall
Muse Spark 1.3 Contributor
Meta · $0.10/M
Fit 78/100 — the top requirements match for this task.
Best value
Muse Spark 1.3 Contributor
Meta · $0.10/M
The strongest fit among budget and mid-tier priced models.
Fastest
GLM 4.7 (Cerebras)
Cerebras · $2.25/M
1980 tokens/sec measured.
Longest context
Gemini 3.1 Pro
Google · $2.07/M
2M token context window.

What evidence supports the Long Documents & RAG recommendation?

We have not run a controlled test for long documents & rag. This ranking is a requirements match on price, measured throughput, and context window — not a quality comparison. Models that fit this task's requirements are ranked; which one performs best on your prompt is a question you should answer by running it. Run all three side by side →

Reproducible Long Documents & RAG evidence and decision rubric

Test / runPrompt and verificationHard rule
Code Snippetexact prompt + 20 recorded runsCorrect iterative algorithm and code-only output
Hard Algorithmexact prompt + 4 recorded runs5,000-case harness; O(log n) partition and correct edge cases

Run dates: budget suite 2026-06-16T20:31:30.728Z; premium suite 2026-06-21T00:00:00.000Z. Results are not a claim about every repository or prompt.

Two-test rubric, failure analysis, and task-shaped ranking

ModelAccuracyLatencyOutput tokensRun costFailure / qualification note
GPT-5.4 Pro100/100206116 ms548$0.103Passes all 5,000+ randomized cases and every edge case; genuine O(log(min(m,n))) partition search, returns a float, code-only, and explicitly raises on two empty lists. Correct — but the slowest run by far (over three minutes), and now that real usage is reported, comfortably the most expensive.
Claude Opus 4.8100/1003790 ms382$0.011Passes all 5,000+ randomized cases and every edge case; genuine O(log(min(m,n))) partition, float return, code-only. The fastest correct solution in this task.
GLM 5.2 (Max)100/10047713 ms2148$0.010Passes all 5,000+ randomized cases and every edge case with a genuine O(log) partition and float return. The <think> block ahead of the code is GLM's reasoning channel surfaced by our gateway, not reasoning dumped into the answer — GLM's actual content is the clean code block — so it scores level with the other correct solutions, as the cheapest of them.
Gemini 3.1 Pro99/10017413 ms320$0.004Passes all tests with a clean, minimal O(log) partition and float return, code-only. Docked one point only because two empty lists yield NaN rather than an explicit guard (not required by the prompt).

Availability caveat: short code tests do not establish repository-scale debugging, multi-file tool use, or agent reliability. The fastest acceptable verdict must therefore clear the correctness rule before speed is considered.

Task-shaped cost ranking (20,000 tasks/month)

RankModelEffective monthlyMeasured verbosity
1Amazon Nova Micro$18.260.76×
2Amazon Nova Lite$32.740.91×
3GPT-5 Nano$36.00Unavailable; neutral fallback
4Gemini 2.5 Flash Lite$56.00Unavailable; neutral fallback
5GPT-OSS 20B$65.162.93×
6Ministral 8B$65.580.93×

Verified 2026-08-08. full prompt/run evidence

Try these models for Long Documents & RAG

Volume, requirement-gate, and pricing cross-check for Long Documents & RAG

Monthly spend ladder at Long-document Q&A shape

Calls / monthOverall pick monthlyBudget pick monthlyOverall − budget delta
1,000$15.20$15.20$0.0000
2,000$30.40$30.40$0.0000
4,000$60.80$60.80$0.0000
10,000$152.00$152.00$0.0000

Monthly cost = task price/M × (150,000 input + 1,000 output tokens) × calls ÷ 1,000,000, at 0.5×, 1×, 2×, 5× the published 2,000-call/month baseline.

Requirement-gate margin for the picked models

ModelRequirementMeasured valueMargin / result
Muse Spark 1.3 Contributor200K required1.0M measured+849K headroom
Muse Spark 1.3 Contributor200K required1.0M measured+849K headroom
Gemini 3.1 Pro200K required2M measured+1.8M headroom

Requirement gates are hard filters, not down-ranking: a model failing any row here is excluded from Long Documents & RAG candidates entirely, regardless of price or speed.

Pricing-page cross-link for each pick

ModelTask-weighted $/MMonthly at published volumeMeasured throughput
Muse Spark 1.3 Contributor$0.10$30.40Unavailable
GLM 4.7 (Cerebras)$2.25$680.501980 tok/s
Gemini 3.1 Pro$2.07$624.0055 tok/s

Evidence coverage: 0 of 40 candidates have a graded run. No graded accuracy evidence exists for this task; the ranking above is a requirements-and-price match, not a quality claim.

Test the Long Documents & RAG picks side by side →

Verified 2026-08-08. "Unavailable" means no compatible dated evidence was found for that field; it is never treated as zero. Dated source · Full ranking and rubric

Batch 10 long-document architecture envelope

1. Usable-window occupancy

ModelOccupancy %Raw window fractionFixed overhead (2K sys + 4K ret + 2K out)1.4x headroom floorUsable document budgetWinner fit
Muse Spark 1.3 Contributor25%262,1442,000 + 4,000 + 2,000 = 8,0001.4×181,531Fits workload
Muse Spark 1.3 Contributor50%524,2882,000 + 4,000 + 2,000 = 8,0001.4×368,777Fits workload
Muse Spark 1.3 Contributor75%786,4322,000 + 4,000 + 2,000 = 8,0001.4×556,022Fits workload
Muse Spark 1.3 Contributor90%943,718.42,000 + 4,000 + 2,000 = 8,0001.4×668,370Fits workload
Gemini 2.5 Flash Lite25%250,0002,000 + 4,000 + 2,000 = 8,0001.4×172,857Fits workload
Gemini 2.5 Flash Lite50%500,0002,000 + 4,000 + 2,000 = 8,0001.4×351,428Fits workload
Gemini 2.5 Flash Lite75%750,0002,000 + 4,000 + 2,000 = 8,0001.4×530,000Fits workload
Gemini 2.5 Flash Lite90%900,0002,000 + 4,000 + 2,000 = 8,0001.4×637,142Fits workload
Gemini 2.5 Flash25%250,0002,000 + 4,000 + 2,000 = 8,0001.4×172,857Fits workload
Gemini 2.5 Flash50%500,0002,000 + 4,000 + 2,000 = 8,0001.4×351,428Fits workload
Gemini 2.5 Flash75%750,0002,000 + 4,000 + 2,000 = 8,0001.4×530,000Fits workload
Gemini 2.5 Flash90%900,0002,000 + 4,000 + 2,000 = 8,0001.4×637,142Fits workload
Gemini 3.1 Pro25%500,0002,000 + 4,000 + 2,000 = 8,0001.4×351,428Fits workload
Gemini 3.1 Pro50%1,000,0002,000 + 4,000 + 2,000 = 8,0001.4×708,571Fits workload
Gemini 3.1 Pro75%1,500,0002,000 + 4,000 + 2,000 = 8,0001.4×1,065,714Fits workload
Gemini 3.1 Pro90%1,800,0002,000 + 4,000 + 2,000 = 8,0001.4×1,280,000Fits workload
Gemini 3.7 Flash25%262,1442,000 + 4,000 + 2,000 = 8,0001.4×181,531Fits workload
Gemini 3.7 Flash50%524,2882,000 + 4,000 + 2,000 = 8,0001.4×368,777Fits workload
Gemini 3.7 Flash75%786,4322,000 + 4,000 + 2,000 = 8,0001.4×556,022Fits workload
Gemini 3.7 Flash90%943,718.42,000 + 4,000 + 2,000 = 8,0001.4×668,370Fits workload

Usable document budget = floor((raw window fraction − 2,000 system − 4,000 retrieval − 2,000 output) ÷ 1.4). Context capacity is not recall quality.

2. Chunk-overlap inflation

Chunk sizeOverlap %Net effective tokens/chunkChunks for 100K docTokens resent (overlap)Monthly bill (1K docs)
8K0%8,000130 tokens$12.08
8K10%7,2001410,400 tokens$13.28
8K20%6,4001624,000 tokens$14.96
32K0%32,00040 tokens$10.64
32K10%28,80049,600 tokens$11.60
32K20%25,600419,200 tokens$12.56
64K0%64,00020 tokens$10.32
64K10%57,60026,400 tokens$10.96
64K20%51,200212,800 tokens$11.60

Net effective tokens/chunk = chunk × (1 − overlap). Chunks for 100K doc = ceil(100,000 ÷ net effective tokens). Tokens resent = (chunks − 1) × overlap tokens. Monthly bill is for 1,000 documents using exact MODEL_PRICING input rates on (100,000 + tokens resent) input tokens plus output rates on (chunks × 800) output tokens.

3. Single-pass vs hierarchical map-reduce

CallsTopologyReference costCritical pathQuality
1single pass$0.02UnavailableRecall/fidelity: Unavailable
21 map (concurrent) + 1 reduce$0.02UnavailableRecall/fidelity: Unavailable
43 map (concurrent) + 1 reduce$0.02UnavailableRecall/fidelity: Unavailable

Critical-path latency assumes concurrent map worker execution plus serial final reduce call. Map-reduce reference cost accounts for both map inputs/outputs and reduce input (concatenated map outputs) + final output. Stale-data behavior: dated registry values are snapshots. If a source is older than the page verification date, or a provider changes its policy/pricing/model, re-verify before production use; unknown values remain Unavailable.

Verified 2026-08-08. Luna is the data owner. “Unavailable” means no compatible dated evidence was found; it is never zero, an estimate, or a guessed policy. Stale-data behavior: dated registry values are snapshots. If a source is older than the page verification date, or a provider changes its policy/pricing/model, re-verify before production use; unknown values remain Unavailable. Dated registry source · Run this scenario yourself →

Batch 59 · server-rendered evidence boards · verified 2026-09-07

Intent answer: Analyzing long documents (100K to 2M tokens) requires high retrieval fidelity across the entire context window, resistance to middle-context degradation, and cost-effective input pricing. Gemini and Claude lead this capability. Verified 2026-09-07.

Demand evidence: Qualitative demand: long-context document analysis benchmarks reviewed 2026-09-07; exact US monthly volume is unavailable.

Scope boundary: Compare top AI models for long-context document analysis, needle-in-a-haystack retrieval, legal/financial PDF comprehension, and long-input pricing. Exact joins required; unresolved joins render Unavailable.

Needle-in-a-haystack retrieval & context fidelity frontier

Deterministic formula / rule: retrieval_rate = correct_needle_answers / total_needle_tests across 0%, 25%, 50%, 75%, 100% depth.

Boundary: Owns long-context retrieval benchmark evaluation.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch59-task-longdocs-m1-r1
needle at 50K tokens (start)
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$needle at 50K tokens (start); model ID; context capacity; 0-25% depth score; 25-75% middle-depth score; 75-100% end-depth score; lost-in-the-middle vulnerability grade; verified=2026-09-07Unavailable — frozen task-longdocs fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-task-longdocs-m1-r2
needle at 128K tokens (middle)
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$needle at 128K tokens (middle); model ID; context capacity; 0-25% depth score; 25-75% middle-depth score; 75-100% end-depth score; lost-in-the-middle vulnerability grade; verified=2026-09-07Unavailable — frozen task-longdocs fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-task-longdocs-m1-r3
needle at 500K tokens (middle-deep)
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$needle at 500K tokens (middle-deep); model ID; context capacity; 0-25% depth score; 25-75% middle-depth score; 75-100% end-depth score; lost-in-the-middle vulnerability grade; verified=2026-09-07Unavailable — frozen task-longdocs fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-task-longdocs-m1-r4
needle at 1M tokens (end)
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$needle at 1M tokens (end); model ID; context capacity; 0-25% depth score; 25-75% middle-depth score; 75-100% end-depth score; lost-in-the-middle vulnerability grade; verified=2026-09-07Unavailable — frozen task-longdocs fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-task-longdocs-m1-r5
multi-needle cross-reference test
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$multi-needle cross-reference test; model ID; context capacity; 0-25% depth score; 25-75% middle-depth score; 75-100% end-depth score; lost-in-the-middle vulnerability grade; verified=2026-09-07Unavailable — frozen task-longdocs fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-task-longdocs-m1-r6
unsupported context length
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$unsupported context length; model ID; context capacity; 0-25% depth score; 25-75% middle-depth score; 75-100% end-depth score; lost-in-the-middle vulnerability grade; verified=2026-09-07Unavailable — exact task-longdocs evidence join is not closed for "unsupported context length"FAIL CLOSED — manual, probe, or source evidence required

First-party citation: All AI Ask model specifications. Verified 2026-09-07; missing or conflicting joins fail closed.

Large document batch analysis unit economics

Deterministic formula / rule: doc_cost = (page_count * tokens_per_page * in_rate) + (synthesis_tokens * out_rate); cache/batch discount.

Boundary: Owns large document processing cost calculations.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch59-task-longdocs-m2-r1
50-page legal brief (25K tokens)
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$50-page legal brief (25K tokens); document archetype; total input tokens; standard processing cost; cached/batch processing cost; lowest-cost model; recommendation tier; verified=2026-09-07Unavailable — frozen task-longdocs fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-task-longdocs-m2-r2
200-page SEC 10-K report (100K tokens)
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$200-page SEC 10-K report (100K tokens); document archetype; total input tokens; standard processing cost; cached/batch processing cost; lowest-cost model; recommendation tier; verified=2026-09-07Unavailable — frozen task-longdocs fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-task-longdocs-m2-r3
500-page technical manual (250K tokens)
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$500-page technical manual (250K tokens); document archetype; total input tokens; standard processing cost; cached/batch processing cost; lowest-cost model; recommendation tier; verified=2026-09-07Unavailable — frozen task-longdocs fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-task-longdocs-m2-r4
multi-file code repository (750K tokens)
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$multi-file code repository (750K tokens); document archetype; total input tokens; standard processing cost; cached/batch processing cost; lowest-cost model; recommendation tier; verified=2026-09-07Unavailable — frozen task-longdocs fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-task-longdocs-m2-r5
multi-book library archive (1.5M tokens)
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$multi-book library archive (1.5M tokens); document archetype; total input tokens; standard processing cost; cached/batch processing cost; lowest-cost model; recommendation tier; verified=2026-09-07Unavailable — frozen task-longdocs fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-task-longdocs-m2-r6
unsupported document volume
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$unsupported document volume; document archetype; total input tokens; standard processing cost; cached/batch processing cost; lowest-cost model; recommendation tier; verified=2026-09-07Unavailable — exact task-longdocs evidence join is not closed for "unsupported document volume"FAIL CLOSED — manual, probe, or source evidence required

First-party citation: All AI Ask pricing registry and formulas. Verified 2026-09-07; missing or conflicting joins fail closed.

Whole-document vs RAG chunking decision framework

Deterministic formula / rule: architecture = (doc_tokens <= context_threshold && query_frequency == high) ? full_context_cache : vector_rag.

Boundary: Owns architectural decision between full-context LLM and vector RAG.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch59-task-longdocs-m3-r1
frequently queried single 100K doc
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$frequently queried single 100K doc; corpus size; monthly query volume; full-context monthly cost; vector RAG monthly cost; accuracy trade-off; architectural recommendation; verified=2026-09-07Unavailable — frozen task-longdocs fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-task-longdocs-m3-r2
infrequently queried 500K doc collection
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$infrequently queried 500K doc collection; corpus size; monthly query volume; full-context monthly cost; vector RAG monthly cost; accuracy trade-off; architectural recommendation; verified=2026-09-07Unavailable — frozen task-longdocs fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-task-longdocs-m3-r3
massive 100M token multi-tenant library
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$massive 100M token multi-tenant library; corpus size; monthly query volume; full-context monthly cost; vector RAG monthly cost; accuracy trade-off; architectural recommendation; verified=2026-09-07Unavailable — frozen task-longdocs fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-task-longdocs-m3-r4
complex cross-document synthesis query
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$complex cross-document synthesis query; corpus size; monthly query volume; full-context monthly cost; vector RAG monthly cost; accuracy trade-off; architectural recommendation; verified=2026-09-07Unavailable — frozen task-longdocs fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-task-longdocs-m3-r5
high-precision verbatim quote extraction
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$high-precision verbatim quote extraction; corpus size; monthly query volume; full-context monthly cost; vector RAG monthly cost; accuracy trade-off; architectural recommendation; verified=2026-09-07Unavailable — frozen task-longdocs fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-task-longdocs-m3-r6
unsupported corpus architecture
route=$/best-llm-for/long-documents; owner=$task-longdocs; scenario=$unsupported corpus architecture; corpus size; monthly query volume; full-context monthly cost; vector RAG monthly cost; accuracy trade-off; architectural recommendation; verified=2026-09-07Unavailable — exact task-longdocs evidence join is not closed for "unsupported corpus architecture"FAIL CLOSED — manual, probe, or source evidence required

First-party citation: All AI Ask measured speed dataset. Verified 2026-09-07; missing or conflicting joins fail closed.

Method and limitations: this board exposes deterministic rules, first-party citations, and dated evidence identities. It does not invent volume, coverage, entitlement, retention, residency, quota, capacity, feature support, quality, price, reliability, or legal conclusions. Run the task-longdocs evidence flow →

Which models rank highest for Long Documents & RAG?

"Fit" is a requirements match, not a quality benchmark — it combines price, measured speed, context window, and (where we have run it) graded accuracy on this task. Formula below.

#ModelProviderFitEvidenceTask price/MTokens/secContextScored on
1Muse Spark 1.3 ContributorMeta78$0.101.0Mprice, context
2Gemini 2.5 Flash LiteLegacyGoogle76$0.101Mprice, context
3Gemini 2.5 FlashLegacyGoogle70$0.311Mprice, context
4Gemini 3.1 ProGoogle69$2.07552Mprice, context, speed
5Gemini 3.7 FlashGoogle67$0.771.0Mprice, context
6Muse Spark 1.3Meta64$1.271.0Mprice, context
7GLM-5.2Z.ai62$1.421Mprice, context
8Gemini 3.5 Flash LiteGoogle60$0.311621Mprice, context, speed

What will Long Documents & RAG cost?

At 2,000 long-document q&a calls/month:

ModelTask price/MEst. monthly cost
Muse Spark 1.3 Contributor$0.10$30.40
Gemini 2.5 Flash Lite$0.10$30.80
Gemini 2.5 Flash$0.31$95.00

How is the best LLM for Long Documents & RAG ranked?

Weights: evidence 0%, price 25%, speed 15%, context 60%.

Requirements: ≥200K context. 40 models eligible.

Price and context sub-scores are min-max normalised (log-scaled) within this task's eligible set only. Speed uses measured tokens/sec only — estimated rows are excluded. A model missing a measurement is never scored as zero: its weight is redistributed across the components we do have, and "Scored on" in the table above shows exactly which ones.

Prices verified 2026-08-08.

What related resources help with Long Documents & RAG?

Meta provider hubMuse Spark 1.3 Contributor pricingBest LLM for CodingBest LLM for Math & ReasoningBest LLM for Chatbots & Support

What are common questions about the best LLM for Long Documents & RAG?

How much context headroom do I actually need?

Budget for your document plus retrieval overhead, system prompt, and conversation history — a 150K-token document comfortably needs a 200K+ window, not exactly 150K.

Does a bigger context window mean better recall inside it?

Not necessarily — window size is a hard capacity limit, not a quality guarantee. Very long prompts can still see recall degrade in the middle of the context ("lost in the middle").

Is RAG still worth it if the context window is huge?

Often yes — retrieval keeps cost and latency down by sending only relevant passages instead of the whole corpus, even when the model could technically fit everything.

Run this exact prompt against the top 3

Don't take a ranking's word for it — try Muse Spark 1.3 Contributor and its closest alternatives on your own prompt.

Try It Free