← All tasks

Best LLM for Image Understanding in 2026

For image understanding, Gemini 3.7 Flash is our pick: $1.38/M tokens on a Image analysis call workload, 1.0M context.

Vision tasks — reading a screenshot, describing a photo, parsing a chart — need a model that accepts image input at all, which is a hard requirement, not a nice-to-have. Among vision-capable models we rank on context window and price.

Verdict: We have not run a controlled quality test for this task — ranking here is a requirements match among vision-capable models on context window and price, not a visual-accuracy comparison.

Quick answer: What is the best LLM for image understanding?

Gemini 3.7 Flash, from Google, is the best fit for image understanding at $1.38 per million task tokens on a Image analysis call workload, with a 1.0M-token context window. No cheaper value pick beats it for this task.

Verified 2026-08-08
Best overall
Gemini 3.7 Flash
Google · $1.38/M
Fit 66/100 — the top requirements match for this task.
Best value
Gemini 3.7 Flash
Google · $1.38/M
The strongest fit among budget and mid-tier priced models.
Fastest
Qwen 3.8 30B
Groq · $1.11/M
690 tokens/sec measured.
Longest context
Gemini 3.1 Pro
Google · $4.11/M
2M token context window.

Can't use Gemini 3.7 Flash? See Gemini 3.7 Flash alternatives.

What evidence supports the Image Understanding recommendation?

We have not run a controlled test for image understanding. This ranking is a requirements match on price, measured throughput, and context window — not a quality comparison. Models that fit this task's requirements are ranked; which one performs best on your prompt is a question you should answer by running it. Run all three side by side →

Reproducible Image Understanding evidence and decision rubric

Test / runPrompt and verificationHard rule
Code Snippetexact prompt + 20 recorded runsCorrect iterative algorithm and code-only output
Hard Algorithmexact prompt + 4 recorded runs5,000-case harness; O(log n) partition and correct edge cases

Run dates: budget suite 2026-06-16T20:31:30.728Z; premium suite 2026-06-21T00:00:00.000Z. Results are not a claim about every repository or prompt.

Two-test rubric, failure analysis, and task-shaped ranking

ModelAccuracyLatencyOutput tokensRun costFailure / qualification note
GPT-5.4 Pro100/100206116 ms548$0.103Passes all 5,000+ randomized cases and every edge case; genuine O(log(min(m,n))) partition search, returns a float, code-only, and explicitly raises on two empty lists. Correct — but the slowest run by far (over three minutes), and now that real usage is reported, comfortably the most expensive.
Claude Opus 4.8100/1003790 ms382$0.011Passes all 5,000+ randomized cases and every edge case; genuine O(log(min(m,n))) partition, float return, code-only. The fastest correct solution in this task.
GLM 5.2 (Max)100/10047713 ms2148$0.010Passes all 5,000+ randomized cases and every edge case with a genuine O(log) partition and float return. The <think> block ahead of the code is GLM's reasoning channel surfaced by our gateway, not reasoning dumped into the answer — GLM's actual content is the clean code block — so it scores level with the other correct solutions, as the cheapest of them.
Gemini 3.1 Pro99/10017413 ms320$0.004Passes all tests with a clean, minimal O(log) partition and float return, code-only. Docked one point only because two empty lists yield NaN rather than an explicit guard (not required by the prompt).

Availability caveat: short code tests do not establish repository-scale debugging, multi-file tool use, or agent reliability. The fastest acceptable verdict must therefore clear the correctness rule before speed is considered.

Task-shaped cost ranking (20,000 tasks/month)

RankModelEffective monthlyMeasured verbosity
1Amazon Nova Micro$18.260.76×
2Amazon Nova Lite$32.740.91×
3GPT-5 Nano$36.00Unavailable; neutral fallback
4Gemini 2.5 Flash Lite$56.00Unavailable; neutral fallback
5GPT-OSS 20B$65.162.93×
6Ministral 8B$65.580.93×

Verified 2026-08-08. full prompt/run evidence

Try these models for Image Understanding

Volume, requirement-gate, and pricing cross-check for Image Understanding

Monthly spend ladder at Image analysis call shape

Calls / monthOverall pick monthlyBudget pick monthlyOverall − budget delta
10,000$26.25$26.25$0.0000
20,000$52.50$52.50$0.0000
40,000$105.00$105.00$0.0000
100,000$262.50$262.50$0.0000

Monthly cost = task price/M × (1,500 input + 400 output tokens) × calls ÷ 1,000,000, at 0.5×, 1×, 2×, 5× the published 20,000-call/month baseline.

Requirement-gate margin for the picked models

ModelRequirementMeasured valueMargin / result
Gemini 3.7 Flashvision requiredtext, vision, audioGate passed
Gemini 3.7 Flashvision requiredtext, vision, audioGate passed
Gemini 3.1 Provision requiredtext, vision, audioGate passed

Requirement gates are hard filters, not down-ranking: a model failing any row here is excluded from Image Understanding candidates entirely, regardless of price or speed.

Pricing-page cross-link for each pick

ModelTask-weighted $/MMonthly at published volumeMeasured throughput
Gemini 3.7 Flash$1.38$52.50Unavailable
Qwen 3.8 30B$1.11$42.00690 tok/s
Gemini 3.1 Pro$4.11$156.0055 tok/s

Evidence coverage: 0 of 37 candidates have a graded run. No graded accuracy evidence exists for this task; the ranking above is a requirements-and-price match, not a quality claim.

Test the Image Understanding picks side by side →

Verified 2026-08-08. "Unavailable" means no compatible dated evidence was found for that field; it is never treated as zero. Dated source · Full ranking and rubric

Batch 10 image-understanding evidence boundary

1. Vision billing and gateway matrix

CandidateModality gateGateway/providerText-token task priceImage billing
Gemini 2.5 Flash LiteVision eligibleGoogle$0.16Image units / resolution tier: Unavailable
Gemini 2.5 FlashVision eligibleGoogle$0.76Image units / resolution tier: Unavailable
Gemini 3.7 FlashVision eligibleGoogle$1.38Image units / resolution tier: Unavailable
Gemini 3.5 Flash LiteVision eligibleGoogle$0.76Image units / resolution tier: Unavailable
Gemini 3.1 ProVision eligibleGoogle$4.11Image units / resolution tier: Unavailable
Grok-3Vision eligiblexAI$2.42Image units / resolution tier: Unavailable

Image units, resolution tiers, and per-image prices are Unavailable unless sourced; missing image prices never become zero.

2. Multi-image context planner

ImagesGemini 2.5 Flash LiteGemini 2.5 FlashGemini 3.7 Flash
1 imagesGemini 2.5 Flash Lite: 997KGemini 2.5 Flash: 997KGemini 3.7 Flash: 1.0M
5 imagesGemini 2.5 Flash Lite: 993KGemini 2.5 Flash: 993KGemini 3.7 Flash: 1.0M
10 imagesGemini 2.5 Flash Lite: 988KGemini 2.5 Flash: 988KGemini 3.7 Flash: 1.0M
25 imagesGemini 2.5 Flash Lite: 973KGemini 2.5 Flash: 973KGemini 3.7 Flash: 1.0M

Remaining context = sourced window − (images × declared 1,000 tokens/image) − 1,500 text tokens − 400 output tokens. Per-image tokens are an estimate, not a provider fact.

3. Task-transfer test matrix

TaskCoverageRequired runVerdict
OCRUnavailableNeed a dated task run with exact images and promptNo quality verdict
ChartsUnavailableNeed a dated task run with exact images and promptNo quality verdict
ScreenshotsUnavailableNeed a dated task run with exact images and promptNo quality verdict
PhotographsUnavailableNeed a dated task run with exact images and promptNo quality verdict
Multi-image comparisonUnavailableNeed a dated task run with exact images and promptNo quality verdict

A single unrelated test cannot establish OCR, chart, screenshot, photograph, or multi-image accuracy.

Verified 2026-08-08. Luna is the data owner. “Unavailable” means no compatible dated evidence was found; it is never zero, an estimate, or a guessed policy. Stale-data behavior: dated registry values are snapshots. If a source is older than the page verification date, or a provider changes its policy/pricing/model, re-verify before production use; unknown values remain Unavailable. Dated registry source · Run this scenario yourself →

Which models rank highest for Image Understanding?

"Fit" is a requirements match, not a quality benchmark — it combines price, measured speed, context window, and (where we have run it) graded accuracy on this task. Formula below.

#ModelProviderFitEvidenceTask price/MTokens/secContextScored on
1Gemini 2.5 Flash LiteLegacyGoogle82$0.161Mprice, context
2Gemini 2.5 FlashLegacyGoogle70$0.761Mprice, context
3Gemini 3.7 FlashGoogle66$1.381.0Mprice, context
4Gemini 3.5 Flash LiteGoogle65$0.761621Mprice, context, speed
5Gemini 3.1 ProGoogle64$4.11552Mprice, context, speed
6Grok-3LegacyxAI61$2.421Mprice, context
7Grok 4.3xAI59$1.51981Mprice, context, speed
8GPT-5.6 LunaOpenAI57$2.051261Mprice, context, speed

What will Image Understanding cost?

At 20,000 image analysis call calls/month:

ModelTask price/MEst. monthly cost
Gemini 2.5 Flash Lite$0.16$6.20
Gemini 2.5 Flash$0.76$29.00
Gemini 3.7 Flash$1.38$52.50

How is the best LLM for Image Understanding ranked?

Weights: evidence 0%, price 40%, speed 10%, context 50%.

Requirements: vision input. 37 models eligible.

Price and context sub-scores are min-max normalised (log-scaled) within this task's eligible set only. Speed uses measured tokens/sec only — estimated rows are excluded. A model missing a measurement is never scored as zero: its weight is redistributed across the components we do have, and "Scored on" in the table above shows exactly which ones.

Prices verified 2026-08-08.

What related resources help with Image Understanding?

Google provider hubGemini 3.7 Flash pricingBest LLM for CodingBest LLM for Math & ReasoningBest LLM for Chatbots & Support

What are common questions about the best LLM for Image Understanding?

Does context window matter for a single image?

Mostly for multi-image or image-plus-long-text prompts — a single image and short prompt fits comfortably in any vision-capable model's window here.

What about audio or video understanding?

A few models on this page also accept audio input (noted on their pricing page) — this ranking filters on vision support specifically, since that is the more universal requirement.

Is a bigger model always more accurate on images?

Not reliably — vision accuracy depends on training, not just parameter count. We have not graded this directly; verify against your own images before committing at volume.

Run this exact prompt against the top 3

Don't take a ranking's word for it — try Gemini 3.7 Flash and its closest alternatives on your own prompt.

Try It Free