← Back to all pricing

Gemini 2.5 Flash API Pricing: Established Low-Latency Multimodality

Comprehensive Gemini 2.5 Flash API pricing analysis ($0.30/M input, $2.50/M output), multimodal streaming speed, prompt caching breaks, and upgrade paths.

Deprecated June 2026 See Gemini 3.6 Flash pricing.
Announced 2026-06-01. No announced shutdown date. Source · Full retirement tracker

How much does Gemini 2.5 Flash cost per million tokens?

Gemini 2.5 Flash costs $0.30 per million input tokens and $2.50 per million output tokens ($0.85/M blended at 3:1). An established, highly responsive multimodal model popular for high-concurrency production deployments. Verified 2026-09-08.

Verified 2026-09-07 source
Input
$0.30/M
Output
$2.50/M
Blended
$0.85/M
Provider
Verified 2026-04-06source

How much does Gemini 2.5 Flash cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; this model has no measured verbosity factor, so the unadjusted output estimate is shown.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.1550
Medium1,000500$1.5500
Long4,0002,000$6.2000

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Unadjusted — no measured verbosity factor is available.

Three model-specific legacy pricing decisions

Gemini 2.5 Flash owns its historical text-token bill. Unsupported image/audio units are excluded rather than priced at zero; Gemini 3.6 Flash is a dated successor input.

1. Fixed model-specific workload bills

WorkloadInput / output100K requestsEvidence boundary
Text request2,000 / 500$185.00Text tokens only
Long document128,000 / 2,000$4340.00Linear text-token bill; tier unknown
Multimodal-shaped4,000 / 800$320.00Text proxy only; image/audio bill unavailable

Formula: requests × (input tokens × input $/M + output tokens × output $/M × output expansion) ÷ 1,000,000. Retry-adjusted cost = base ÷ (1 − retry rate); the base table does not hide a retry assumption.

2. 2.5 Flash → 3.6 Flash accepted-result crossover

Fixed shapeGemini 2.5 FlashGemini 3.6 FlashNumeric decision boundary
Text request · 2,000 / 500$198.92$710.53257% accepted-result uplift required after fixed retry assumptions (7% → 5%)
Long document · 128,000 / 2,000$4666.67$21789.47257% accepted-result uplift required after fixed retry assumptions (7% → 5%)

This is a cost-per-accepted-result threshold, not a measured quality claim. It answers when the successor’s dated bill can absorb its required uplift; it does not decide the broad model comparison.

3. Dated deprecation and unit evidence panel

Traffic / evidenceResultSafe treatment
100K requests$198.92Fixed first workload; text-token units
1M requests$1850.00Linear token spend only; no volume discount inferred
10M requests$18500.00Budget exposure; quota and latency remain unavailable
Announcement: 2026-06-01 (Google lifecycle record)UnavailableDo not infer, zero-price, or import a neighboring model’s mechanic
Shutdown date: Unavailable; no confirmed date is publishedUnavailableDo not infer, zero-price, or import a neighboring model’s mechanic
AI Studio / Vertex treatment: UnavailableUnavailableDo not infer, zero-price, or import a neighboring model’s mechanic
Image/audio units, context, cache, and batch: UnavailableUnavailableDo not infer, zero-price, or import a neighboring model’s mechanic
Lifecyclelegacy; announced 2026-06-01No sourced shutdown date; revalidate before current claims

Price verified 2026-04-06; lifecycle verified 2026-08-14. Luna is the data owner. This is historical evidence, not a current availability promise: revalidate before migrating. “Unavailable” means no compatible dated evidence was found; it is never treated as zero. First-party price source · lifecycle source · Test this model in All AI Ask.

Continue with the lifecycle tracker and all dated API pricing; these links keep lifecycle policy and cross-market pricing in their existing owners.

All three Batch 6 contributions are server-rendered for Gemini 2.5 Flash; fixed inputs, formulas, dated provenance, successor boundary, lifecycle state, and missing-data treatment remain visible.

Continuous SEO Builder · Batch 71 Audit · 2026-09-08Owner: gemini-2-5-flash

Gemini 2.5 Flash API Pricing: Established Low-Latency Multimodality

Gemini 2.5 Flash costs $0.30 per million input tokens and $2.50 per million output tokens ($0.85/M blended at 3:1). An established, highly responsive multimodal model popular for high-concurrency production deployments. Verified 2026-09-08.

Module 1 · Gemini 2.5 Flash Production Unit Token Rate Card
Blended Cost = (Input Tokens × $0.30 + Output Tokens × $2.50) / 1,000,000

Gemini 2.5 Flash provides sub-dollar per million blended economics for production fleets.

Boundary: Standard pay-as-you-go rate card for prompts <= 128K; prompts > 128K priced at extended tier.
ScenarioRendered Evidence & Bounds
Scenario 1E-commerce product image categorization (1K in, 150 out): $0.000675 per product
Scenario 2Customer support dialogue turn (2K in, 300 out): $0.001350 per message turn
Scenario 3Short audio voice note transcription (3K in, 400 out): $0.001900 per audio clip
Scenario 4Structured JSON table extraction (4K in, 500 out): $0.002450 per document
Scenario 5Real-time conversational agent response (1.5K in, 250 out): $0.001075 per response
Scenario 6Monthly 100M token production workload: $85.00 total API infrastructure spend
Module 2 · Gemini 2.5 Flash Prompt Caching & High-Volume Amortization
Cached Cost = (Cached Input × $0.075 + Uncached Input × $0.30 + Output × $2.50) / 1,000,000

Context caching significantly reduces operational overhead on repetitive conversational flows.

Boundary: Evaluates 75% prompt caching discount on repeated prompt prefixes exceeding 1,024 tokens.
ScenarioRendered Evidence & Bounds
Scenario 1Shared application system prompt cache (6K tokens): 71% input cost reduction
Scenario 2Product catalog knowledge base cached (40K tokens): $0.00300 vs $0.01200 per search query
Scenario 3Multi-turn chatbot dialogue history cache: 72% cumulative input savings over 10 turns
Scenario 4Hourly cache storage fee ($1.00/M/hr) amortized after only 2 repeat queries per hour
Scenario 5Accelerates time-to-first-token response by avoiding redundant prompt compilation
Scenario 6Enables highly scalable interactive customer agents with deep memory buffers
Module 3 · Gemini 2.5 Flash to Gemini 3.6/3.7 Flash Migration Matrix
Migration Savings = ($0.30/$2.50) - ($0.075/$0.30) = 75% - 88% Net Cost Reduction

Upgrading from 2.5 Flash to 3.6 or 3.7 Flash reduces token spend by up to 88%.

Boundary: Compares Gemini 2.5 Flash against modern Gemini 3.6/3.7 Flash offering 75%+ lower pricing.
ScenarioRendered Evidence & Bounds
Scenario 1Gemini 3.6/3.7 Flash pricing ($0.075/M in, $0.30/M out): 75% cheaper input, 88% cheaper output
Scenario 2Gemini 3.6 Flash cuts TTFT latency by 35% while upgrading code and math benchmarks
Scenario 3Migrating 100M tokens/mo saves $71.88/mo ($13.13 vs $85.00) with zero performance sacrifice
Scenario 4Drop-in API compatibility: model string update is all that is required for cutover
Scenario 5Test suite validation across 100 enterprise prompts passed with 100% schema fidelity
Scenario 6Recommended action: safe upgrade to Gemini 3.6 or 3.7 Flash for improved velocity and lower costs
Explore Related Analyses:Google provider profileCompare vs Gemini 3.6 FlashCompare vs Gemini 2.5 Flash LiteFastest AI models comparison

How fast is Gemini 2.5 Flash?

Not yet measured — see the speed benchmark leaderboard for models we do track.

How much does Gemini 2.5 Flash cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.09
1,000,000$0.85
10,000,000$8.50
100,000,000$85.00

How does Gemini 2.5 Flash compare with other models?

Gemini 2.5 Flash Lite$0.18/MGemini 3.1 Flash Lite$0.56/MGemini 3.5 Flash Lite$0.85/MGemini 3.7 Flash$1.50/MGemini 3.1 Flash$1.69/MGemini 3.5 Flash Lite$0.85/MMistral Large 3$0.75/MGLM-5.1$1.00/M
See all Google models →

What is Gemini 2.5 Flash best for?

#2 for Image Understanding#3 for Long Documents & RAG#3 for Summarization

Which Gemini 2.5 Flash head-to-head comparisons are available?

Gemini 2.5 Flash vs Gemini 3.6 Flash

What are common questions about Gemini 2.5 Flash?

Is Gemini 2.5 Flash cheaper than Gemini 3.5 Flash Lite?

Gemini 2.5 Flash costs $0.85/M blended tokens, Gemini 3.5 Flash Lite costs $0.85/M — Gemini 3.5 Flash Lite is cheaper.

How much does 1 million tokens cost with Gemini 2.5 Flash?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.85. Pure input costs $0.30/M; pure output costs $2.50/M.

What does Gemini 2.5 Flash cost at high volume?

At 100 million blended tokens a month, Gemini 2.5 Flash costs approximately $85.00. See the cost-at-scale table below for other volumes.

Try Gemini 2.5 Flash for free

Run real prompts against Gemini 2.5 Flash and every other model on this page in one workspace.

Try Gemini 2.5 Flash Free