Gemini 2.5 Flash API Pricing: Established Low-Latency Multimodality
Comprehensive Gemini 2.5 Flash API pricing analysis ($0.30/M input, $2.50/M output), multimodal streaming speed, prompt caching breaks, and upgrade paths.
How much does Gemini 2.5 Flash cost per million tokens?
Gemini 2.5 Flash costs $0.30 per million input tokens and $2.50 per million output tokens ($0.85/M blended at 3:1). An established, highly responsive multimodal model popular for high-concurrency production deployments. Verified 2026-09-08.
How much does Gemini 2.5 Flash cost per 1,000 requests?
Computed from generated token pricing. Each row assumes the listed input and output tokens per request; this model has no measured verbosity factor, so the unadjusted output estimate is shown.
| Request shape | Input tokens | Output tokens | Cost / 1,000 requests |
|---|---|---|---|
| Short | 100 | 50 | $0.1550 |
| Medium | 1,000 | 500 | $1.5500 |
| Long | 4,000 | 2,000 | $6.2000 |
Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Unadjusted — no measured verbosity factor is available.
Three model-specific legacy pricing decisions
Gemini 2.5 Flash owns its historical text-token bill. Unsupported image/audio units are excluded rather than priced at zero; Gemini 3.6 Flash is a dated successor input.
1. Fixed model-specific workload bills
| Workload | Input / output | 100K requests | Evidence boundary |
|---|---|---|---|
| Text request | 2,000 / 500 | $185.00 | Text tokens only |
| Long document | 128,000 / 2,000 | $4340.00 | Linear text-token bill; tier unknown |
| Multimodal-shaped | 4,000 / 800 | $320.00 | Text proxy only; image/audio bill unavailable |
Formula: requests × (input tokens × input $/M + output tokens × output $/M × output expansion) ÷ 1,000,000. Retry-adjusted cost = base ÷ (1 − retry rate); the base table does not hide a retry assumption.
2. 2.5 Flash → 3.6 Flash accepted-result crossover
| Fixed shape | Gemini 2.5 Flash | Gemini 3.6 Flash | Numeric decision boundary |
|---|---|---|---|
| Text request · 2,000 / 500 | $198.92 | $710.53 | 257% accepted-result uplift required after fixed retry assumptions (7% → 5%) |
| Long document · 128,000 / 2,000 | $4666.67 | $21789.47 | 257% accepted-result uplift required after fixed retry assumptions (7% → 5%) |
This is a cost-per-accepted-result threshold, not a measured quality claim. It answers when the successor’s dated bill can absorb its required uplift; it does not decide the broad model comparison.
3. Dated deprecation and unit evidence panel
| Traffic / evidence | Result | Safe treatment |
|---|---|---|
| 100K requests | $198.92 | Fixed first workload; text-token units |
| 1M requests | $1850.00 | Linear token spend only; no volume discount inferred |
| 10M requests | $18500.00 | Budget exposure; quota and latency remain unavailable |
| Announcement: 2026-06-01 (Google lifecycle record) | Unavailable | Do not infer, zero-price, or import a neighboring model’s mechanic |
| Shutdown date: Unavailable; no confirmed date is published | Unavailable | Do not infer, zero-price, or import a neighboring model’s mechanic |
| AI Studio / Vertex treatment: Unavailable | Unavailable | Do not infer, zero-price, or import a neighboring model’s mechanic |
| Image/audio units, context, cache, and batch: Unavailable | Unavailable | Do not infer, zero-price, or import a neighboring model’s mechanic |
| Lifecycle | legacy; announced 2026-06-01 | No sourced shutdown date; revalidate before current claims |
Price verified 2026-04-06; lifecycle verified 2026-08-14. Luna is the data owner. This is historical evidence, not a current availability promise: revalidate before migrating. “Unavailable” means no compatible dated evidence was found; it is never treated as zero. First-party price source · lifecycle source · Test this model in All AI Ask.
Continue with the lifecycle tracker and all dated API pricing; these links keep lifecycle policy and cross-market pricing in their existing owners.
All three Batch 6 contributions are server-rendered for Gemini 2.5 Flash; fixed inputs, formulas, dated provenance, successor boundary, lifecycle state, and missing-data treatment remain visible.
gemini-2-5-flashGemini 2.5 Flash API Pricing: Established Low-Latency Multimodality
Gemini 2.5 Flash costs $0.30 per million input tokens and $2.50 per million output tokens ($0.85/M blended at 3:1). An established, highly responsive multimodal model popular for high-concurrency production deployments. Verified 2026-09-08.
Gemini 2.5 Flash provides sub-dollar per million blended economics for production fleets.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | E-commerce product image categorization (1K in, 150 out): $0.000675 per product |
| Scenario 2 | Customer support dialogue turn (2K in, 300 out): $0.001350 per message turn |
| Scenario 3 | Short audio voice note transcription (3K in, 400 out): $0.001900 per audio clip |
| Scenario 4 | Structured JSON table extraction (4K in, 500 out): $0.002450 per document |
| Scenario 5 | Real-time conversational agent response (1.5K in, 250 out): $0.001075 per response |
| Scenario 6 | Monthly 100M token production workload: $85.00 total API infrastructure spend |
Context caching significantly reduces operational overhead on repetitive conversational flows.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Shared application system prompt cache (6K tokens): 71% input cost reduction |
| Scenario 2 | Product catalog knowledge base cached (40K tokens): $0.00300 vs $0.01200 per search query |
| Scenario 3 | Multi-turn chatbot dialogue history cache: 72% cumulative input savings over 10 turns |
| Scenario 4 | Hourly cache storage fee ($1.00/M/hr) amortized after only 2 repeat queries per hour |
| Scenario 5 | Accelerates time-to-first-token response by avoiding redundant prompt compilation |
| Scenario 6 | Enables highly scalable interactive customer agents with deep memory buffers |
Upgrading from 2.5 Flash to 3.6 or 3.7 Flash reduces token spend by up to 88%.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Gemini 3.6/3.7 Flash pricing ($0.075/M in, $0.30/M out): 75% cheaper input, 88% cheaper output |
| Scenario 2 | Gemini 3.6 Flash cuts TTFT latency by 35% while upgrading code and math benchmarks |
| Scenario 3 | Migrating 100M tokens/mo saves $71.88/mo ($13.13 vs $85.00) with zero performance sacrifice |
| Scenario 4 | Drop-in API compatibility: model string update is all that is required for cutover |
| Scenario 5 | Test suite validation across 100 enterprise prompts passed with 100% schema fidelity |
| Scenario 6 | Recommended action: safe upgrade to Gemini 3.6 or 3.7 Flash for improved velocity and lower costs |
How fast is Gemini 2.5 Flash?
How much does Gemini 2.5 Flash cost at scale?
| Tokens / month | Est. cost (blended 3:1) |
|---|---|
| 100,000 | $0.09 |
| 1,000,000 | $0.85 |
| 10,000,000 | $8.50 |
| 100,000,000 | $85.00 |
How does Gemini 2.5 Flash compare with other models?
What is Gemini 2.5 Flash best for?
What should you explore next for Gemini 2.5 Flash?
Which Gemini 2.5 Flash head-to-head comparisons are available?
What are common questions about Gemini 2.5 Flash?
Is Gemini 2.5 Flash cheaper than Gemini 3.5 Flash Lite?
Gemini 2.5 Flash costs $0.85/M blended tokens, Gemini 3.5 Flash Lite costs $0.85/M — Gemini 3.5 Flash Lite is cheaper.
How much does 1 million tokens cost with Gemini 2.5 Flash?
At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.85. Pure input costs $0.30/M; pure output costs $2.50/M.
What does Gemini 2.5 Flash cost at high volume?
At 100 million blended tokens a month, Gemini 2.5 Flash costs approximately $85.00. See the cost-at-scale table below for other volumes.
