← Back to all pricing

Gemini 2.5 Flash Lite API Pricing: Extreme Low-Cost High-Speed Inference

Comprehensive Gemini 2.5 Flash Lite API pricing analysis ($0.10/M input, $0.40/M output), sub-80ms first-token latency, classification efficiency, and upgrade comparisons.

Deprecated June 2026 See Gemini 3.5 Flash Lite pricing.
Announced 2026-06-01. No announced shutdown date. Source · Full retirement tracker

How much does Gemini 2.5 Flash Lite cost per million tokens?

Gemini 2.5 Flash Lite costs $0.10 per million input tokens and $0.40 per million output tokens ($0.175/M blended at 3:1). Designed for ultra-high-volume micro-tasks, classification, and real-time voice latency. Verified 2026-09-08.

Verified 2026-09-07 source
Input
$0.10/M
Output
$0.40/M
Blended
$0.18/M
Provider
Verified 2026-04-06source

How much does Gemini 2.5 Flash Lite cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; this model has no measured verbosity factor, so the unadjusted output estimate is shown.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.0300
Medium1,000500$0.3000
Long4,0002,000$1.2000

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Unadjusted — no measured verbosity factor is available.

Migration delta, task-ranking cross-check, and lifecycle ledger for Gemini 2.5 Flash Lite

Successor rate delta

RateGemini 2.5 Flash LiteGemini 3.5 Flash LiteDelta ($ / %)
Input $/M$0.10$0.30$0.20 (200.0%)
Output $/M$0.40$2.50$2.10 (525.0%)

Delta = successor rate − Gemini 2.5 Flash Lite rate, both taken from the current dated pricing registry. Cache, batch, and context-tier pricing are not substituted across models; each figure is model-specific or marked Unavailable.

Where Gemini 2.5 Flash Lite still ranks across every task page

TaskRank / statusTask-weighted $/MPick status
CodingRank 3 of 49$0.16Dated — excluded from picks
Structured Data ExtractionRank 2 of 49$0.16Dated — excluded from picks
Writing & ContentRank 2 of 49$0.26Dated — excluded from picks
Math & ReasoningNot a qualifying candidateExcluded by a hard requirement filter
Agents & Tool UseNot a qualifying candidateExcluded by a hard requirement filter
Long Documents & RAGRank 2 of 40$0.10Dated — excluded from picks
SummarizationRank 2 of 49$0.10Dated — excluded from picks
Chatbots & SupportRank 2 of 49$0.20Dated — excluded from picks
TranslationRank 2 of 49$0.25Dated — excluded from picks
Image UnderstandingRank 1 of 37$0.16Dated — excluded from picks

Rank is computed live from the same {price, speed, context, evidence} formula published on each /best-llm-for page; a dated model can still rank by price/speed/context but is excluded from the "overall pick" by policy.

Lifecycle and availability ledger

FieldRecorded valueNote
Lifecycle statuslegacyVerified 2026-08-14
Deprecation announced2026-06-01No inference beyond the dated record
Shutdown dateUnavailableNull/unavailable is not a promise of indefinite availability
SuccessorGemini 3.5 Flash Litegemini-3-5-flash-lite priced in this registry
Context window1M tokensVerified 2026-08-14
Compare Gemini 2.5 Flash Lite against its successor →

Verified 2026-04-06. "Unavailable" means no compatible dated evidence was found for that field; it is never treated as zero. Dated source · Lifecycle source

Continuous SEO Builder · Batch 71 Audit · 2026-09-08Owner: gemini-2-5-flash-lite

Gemini 2.5 Flash Lite API Pricing: Extreme Low-Cost High-Speed Inference

Gemini 2.5 Flash Lite costs $0.10 per million input tokens and $0.40 per million output tokens ($0.175/M blended at 3:1). Designed for ultra-high-volume micro-tasks, classification, and real-time voice latency. Verified 2026-09-08.

Module 1 · Gemini 2.5 Flash Lite Micro-Budget Unit Token Economics
Blended Cost = (Input Tokens × $0.10 + Output Tokens × $0.40) / 1,000,000

Gemini 2.5 Flash Lite delivers industry-leading cost efficiency at $0.175/M blended tokens.

Boundary: Standard pay-as-you-go rate card for prompts <= 128K; prompts > 128K priced at extended tier.
ScenarioRendered Evidence & Bounds
Scenario 1Live voice assistant dialogue turn (400 in, 80 out): $0.000072 per turn
Scenario 2Real-time text autocomplete suggestion (150 in, 20 out): $0.000023 per keystroke completion
Scenario 3Support ticket intent routing (600 in, 40 out): $0.000076 per ticket
Scenario 4Document sentiment scoring (1K in, 50 out): $0.000120 per document
Scenario 5Automated data field normalization (500 in, 60 out): $0.000074 per record
Scenario 6Monthly 100M token classification fleet: $17.50 total API infrastructure spend
Module 2 · Gemini 2.5 Flash Lite Low-Latency Voice & Streaming Optimization
Fleet Efficiency = Throughput (tokens/sec) / Blended Price ($/M)

Microscopic latency and rock-bottom token pricing make it the ideal engine for voice bots.

Boundary: Evaluates sub-80ms time-to-first-token performance on interactive streaming applications.
ScenarioRendered Evidence & Bounds
Scenario 1Sub-80ms TTFT enables natural, human-like voice conversation turn-taking
Scenario 2Streaming throughput exceeding 150 tps prevents UI buffering on client devices
Scenario 3Zero-latency prompt processing handles rapid burst traffic without queue buildup
Scenario 4High-volume webhook ingestion: processes 10,000 webhooks for less than $0.01
Scenario 5Micro-memory footprint allows high parallel connection limits on cloud gateways
Scenario 6Optimal choice for automated phone bots, live chat routing, and realtime moderation
Module 3 · Gemini 2.5 Flash Lite to 3.5 Flash Lite Generational Upgrade
Generational Comparison = ($0.10/$0.40) vs ($0.0375/$0.15) = 62.5% Cost Reduction

Upgrading to Gemini 3.5 Flash Lite reduces token spend by 62.5% with higher precision.

Boundary: Compares Gemini 2.5 Flash Lite against next-gen 3.5 Flash Lite offering a 62.5% price decrease.
ScenarioRendered Evidence & Bounds
Scenario 1Gemini 3.5 Flash Lite pricing ($0.0375/M in, $0.15/M out): 62.5% lower cost across all tiers
Scenario 2Gemini 3.5 Flash Lite improves multilingual accuracy and complex instruction following
Scenario 3Migrating 100M tokens/mo saves $10.94/mo ($6.56 vs $17.50) while boosting accuracy
Scenario 4Drop-in SDK compatibility: zero code modifications needed beyond updating model ID
Scenario 5Benchmark accuracy: 3.5 Flash Lite shows 5% higher intent classification precision
Scenario 6Recommended action: safe immediate migration to Gemini 3.5 Flash Lite for lowest total spend
Explore Related Analyses:Google provider profileCompare vs Gemini 3.5 Flash LiteCompare vs Gemini 2.5 FlashCheapest AI API comparison

How fast is Gemini 2.5 Flash Lite?

Not yet measured — see the speed benchmark leaderboard for models we do track.

How much does Gemini 2.5 Flash Lite cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.02
1,000,000$0.18
10,000,000$1.75
100,000,000$17.50

How does Gemini 2.5 Flash Lite compare with other models?

Gemini 3.1 Flash Lite$0.56/MGemini 3.5 Flash Lite$0.85/MGemini 2.5 Flash$0.85/MGemini 3.7 Flash$1.50/MGemini 3.1 Flash$1.69/MMinistral 8B$0.15/MGPT-5 Nano$0.14/MGPT-OSS 20B$0.13/M
See all Google models →

What is Gemini 2.5 Flash Lite best for?

#1 for Image Understanding#2 for Structured Data Extraction#2 for Writing & Content

Which Gemini 2.5 Flash Lite head-to-head comparisons are available?

Gemini 2.5 Flash Lite vs Gemini 3.5 Flash Lite

What are common questions about Gemini 2.5 Flash Lite?

Is Gemini 2.5 Flash Lite cheaper than Ministral 8B?

Gemini 2.5 Flash Lite costs $0.18/M blended tokens, Ministral 8B costs $0.15/M — Ministral 8B is cheaper.

How much does 1 million tokens cost with Gemini 2.5 Flash Lite?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.18. Pure input costs $0.10/M; pure output costs $0.40/M.

What does Gemini 2.5 Flash Lite cost at high volume?

At 100 million blended tokens a month, Gemini 2.5 Flash Lite costs approximately $17.50. See the cost-at-scale table below for other volumes.

Try Gemini 2.5 Flash Lite for free

Run real prompts against Gemini 2.5 Flash Lite and every other model on this page in one workspace.

Try Gemini 2.5 Flash Lite Free