← Back to all pricing

Gemini 3.1 Flash Lite API Pricing: Ultra-Low Latency Inference

Comprehensive Gemini 3.1 Flash Lite API pricing analysis ($0.25/M input, $1.50/M output), sub-100ms first-token latency, classification efficiency, and upgrade comparisons.

Legacy — superseded by Gemini 3.5 Flash Lite See Gemini 3.5 Flash Lite pricing.
No announced shutdown date. Source · Full retirement tracker

How much does Gemini 3.1 Flash Lite cost per million tokens?

Gemini 3.1 Flash Lite costs $0.25 per million input tokens and $1.50 per million output tokens ($0.5625/M blended at 3:1). Engineered for sub-100ms time-to-first-token responses and high-volume classification workloads. Verified 2026-09-08.

Verified 2026-09-07 source
Input
$0.25/M
Output
$1.50/M
Blended
$0.56/M
Provider
Verified 2026-04-06source

How much does Gemini 3.1 Flash Lite cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 0.87× verbosity factor.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.0903
Medium1,000500$0.9025
Long4,0002,000$3.6100

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.

Volume ladder, output-cost sensitivity, and migration ledger for Gemini 3.1 Flash Lite

1. Fixed-shape monthly cost ladder

WorkloadInput / output tokens100K requests1M requests10M requests
Short chat500 / 150$35.00$350.00$3500.00
Code review4,000 / 800$220.00$2200.00$22000.00
Document summary16,000 / 2,000$700.00$7000.00$70000.00

Cost = requests × (input tokens × $0.25/M + output tokens × $1.50/M) ÷ 1,000,000, at three fixed text-token workload shapes. No volume discount, cache, or batch rate is inferred — Gemini 3.1 Flash Lite's registry entry does not price those tiers.

2. Output-token cost-sensitivity band

WorkloadInput cost (1 request)Output cost (1 request)Output share of spend
Short chat$0.0001$0.000264.3%
Code review$0.0010$0.001254.5%
Document summary$0.0040$0.003042.9%

Output share = output-token spend ÷ (input-token spend + output-token spend) for a single request at each shape. Across these three shapes, Gemini 3.1 Flash Lite's output share spans 42.9% to 64.3% — a 21.4%-point swing driven entirely by workload shape, not by any change in the $0.25/$1.50 per-1M rates.

3. Successor rate delta and lifecycle ledger

RateGemini 3.1 Flash LiteGemini 3.5 Flash LiteDelta ($ / %)
Input $/M$0.25$0.30$0.05 (20.0%)
Output $/M$1.50$2.50$1.00 (66.7%)

Delta = successor rate − Gemini 3.1 Flash Lite rate, both taken from the current dated pricing registry. Cache, batch, and context-tier pricing are not substituted across models; each figure is model-specific or marked Unavailable.

FieldRecorded valueNote
Lifecycle statuslegacyVerified 2026-08-14
Deprecation announcedUnavailableNo inference beyond the dated record
Shutdown dateUnavailableNull/unavailable is not a promise of indefinite availability
SuccessorGemini 3.5 Flash Litegemini-3-5-flash-lite priced in this registry
Context windowUnavailableUnavailable — no model-specific spec sourced
Compare Gemini 3.1 Flash Lite against its successor →

Price verified 2026-04-06; lifecycle verified 2026-08-14. Luna is the data owner. "Unavailable" means no compatible dated evidence was found for that field; it is never treated as zero. Dated price source · lifecycle source · Test Gemini 3.1 Flash Lite in All AI Ask.

Continuous SEO Builder · Batch 71 Audit · 2026-09-08Owner: gemini-3-1-flash-lite

Gemini 3.1 Flash Lite API Pricing: Ultra-Low Latency Inference

Gemini 3.1 Flash Lite costs $0.25 per million input tokens and $1.50 per million output tokens ($0.5625/M blended at 3:1). Engineered for sub-100ms time-to-first-token responses and high-volume classification workloads. Verified 2026-09-08.

Module 1 · Gemini 3.1 Flash Lite Micro-Latency Token Economics
Blended Cost = (Input Tokens × $0.25 + Output Tokens × $1.50) / 1,000,000

Gemini 3.1 Flash Lite combines sub-dollar per million economics with sub-100ms TTFT latency.

Boundary: Standard rate card for prompts <= 128K; prompts > 128K billed at extended tier.
ScenarioRendered Evidence & Bounds
Scenario 1Voice agent conversational turn (500 in, 100 out): $0.000275 per exchange
Scenario 2Real-time search autocomplete suggestion (200 in, 30 out): $0.000095 per query
Scenario 3Customer support ticket routing (800 in, 50 out): $0.000275 per ticket
Scenario 4Document sentiment classification (1.5K in, 100 out): $0.000525 per review
Scenario 5Form input validation and autocorrection (400 in, 80 out): $0.000220 per validation
Scenario 6Monthly 100M token high-speed triage fleet: $56.25 total infrastructure spend
Module 2 · Gemini 3.1 Flash Lite High-Throughput Batch Processing
Batch Spend = On-Demand Spend × 0.50 (Asynchronous Batch API)

Batch API queuing cuts already-low token costs in half for asynchronous background pipelines.

Boundary: Evaluates cost reduction when running asynchronous classification pipelines through Batch API.
ScenarioRendered Evidence & Bounds
Scenario 110M token overnight customer survey tagging: $2.81 batch cost vs $5.63 standard
Scenario 250M token product catalog category normalization: $14.06 batch vs $28.13 standard
Scenario 3100M token log analysis and anomaly categorization: $28.13 batch vs $56.25 standard
Scenario 4Batch queues execute within 24-hour turnaround SLA with 100% throughput predictability
Scenario 5Ideal for non-urgent background data structuring, tagging, and entity extraction
Scenario 6Reduces effective blended token price to $0.281/M on large asynchronous datasets
Module 3 · Gemini 3.1 Flash Lite to 3.5 Flash Lite Generational Upgrade
Generational Savings = ($0.25/$1.50) - ($0.0375/$0.15) = 85% - 90% Net Cost Reduction

Upgrading to Gemini 3.5 Flash Lite slashes token costs by 85% while boosting inference velocity.

Boundary: Compares Gemini 3.1 Flash Lite against next-generation 3.5 Flash Lite offering an 85% price drop.
ScenarioRendered Evidence & Bounds
Scenario 1Gemini 3.5 Flash Lite pricing ($0.0375/M in, $0.15/M out): 85% cheaper input, 90% cheaper output
Scenario 2Gemini 3.5 Flash Lite delivers even lower first-token latency (sub-80ms TTFT)
Scenario 3Migrating 100M tokens/mo saves $49.69/mo ($6.56 vs $56.25) while upgrading responsiveness
Scenario 4Drop-in SDK compatibility: model ID swap requires zero code refactoring
Scenario 5Audited benchmark accuracy: 3.5 Flash Lite matches or exceeds 3.1 on all classification tasks
Scenario 6Recommended action: immediate migration to Gemini 3.5 Flash Lite for maximum cost efficiency
Explore Related Analyses:Google provider profileCompare vs Gemini 3.5 Flash LiteCompare vs Gemini 2.5 Flash LiteCheapest AI API comparison

How fast is Gemini 3.1 Flash Lite?

Not yet measured — see the speed benchmark leaderboard for models we do track.

How much does Gemini 3.1 Flash Lite cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.06
1,000,000$0.56
10,000,000$5.63
100,000,000$56.25

How does Gemini 3.1 Flash Lite compare with other models?

Gemini 2.5 Flash Lite$0.18/MGemini 3.5 Flash Lite$0.85/MGemini 2.5 Flash$0.85/MGemini 3.7 Flash$1.50/MGemini 3.1 Flash$1.69/MDeepSeek V4 Flash$0.66/MGPT-5.4 Nano$0.46/MCodestral$0.45/M
See all Google models →

What are common questions about Gemini 3.1 Flash Lite?

Is Gemini 3.1 Flash Lite cheaper than DeepSeek V4 Flash?

Gemini 3.1 Flash Lite costs $0.56/M blended tokens, DeepSeek V4 Flash costs $0.66/M — Gemini 3.1 Flash Lite is cheaper.

How much does 1 million tokens cost with Gemini 3.1 Flash Lite?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.56. Pure input costs $0.25/M; pure output costs $1.50/M.

What does Gemini 3.1 Flash Lite cost at high volume?

At 100 million blended tokens a month, Gemini 3.1 Flash Lite costs approximately $56.25. See the cost-at-scale table below for other volumes.

Try Gemini 3.1 Flash Lite for free

Run real prompts against Gemini 3.1 Flash Lite and every other model on this page in one workspace.

Try Gemini 3.1 Flash Lite Free