Gemini 3.1 Flash Lite API Pricing: Ultra-Low Latency Inference
Comprehensive Gemini 3.1 Flash Lite API pricing analysis ($0.25/M input, $1.50/M output), sub-100ms first-token latency, classification efficiency, and upgrade comparisons.
How much does Gemini 3.1 Flash Lite cost per million tokens?
Gemini 3.1 Flash Lite costs $0.25 per million input tokens and $1.50 per million output tokens ($0.5625/M blended at 3:1). Engineered for sub-100ms time-to-first-token responses and high-volume classification workloads. Verified 2026-09-08.
How much does Gemini 3.1 Flash Lite cost per 1,000 requests?
Computed from generated token pricing. Each row assumes the listed input and output tokens per request; output is adjusted by this model's measured 0.87× verbosity factor.
| Request shape | Input tokens | Output tokens | Cost / 1,000 requests |
|---|---|---|---|
| Short | 100 | 50 | $0.0903 |
| Medium | 1,000 | 500 | $0.9025 |
| Long | 4,000 | 2,000 | $3.6100 |
Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Verbosity run: 2026-06-16T20:31:30.728Z.
Volume ladder, output-cost sensitivity, and migration ledger for Gemini 3.1 Flash Lite
1. Fixed-shape monthly cost ladder
| Workload | Input / output tokens | 100K requests | 1M requests | 10M requests |
|---|---|---|---|---|
| Short chat | 500 / 150 | $35.00 | $350.00 | $3500.00 |
| Code review | 4,000 / 800 | $220.00 | $2200.00 | $22000.00 |
| Document summary | 16,000 / 2,000 | $700.00 | $7000.00 | $70000.00 |
Cost = requests × (input tokens × $0.25/M + output tokens × $1.50/M) ÷ 1,000,000, at three fixed text-token workload shapes. No volume discount, cache, or batch rate is inferred — Gemini 3.1 Flash Lite's registry entry does not price those tiers.
2. Output-token cost-sensitivity band
| Workload | Input cost (1 request) | Output cost (1 request) | Output share of spend |
|---|---|---|---|
| Short chat | $0.0001 | $0.0002 | 64.3% |
| Code review | $0.0010 | $0.0012 | 54.5% |
| Document summary | $0.0040 | $0.0030 | 42.9% |
Output share = output-token spend ÷ (input-token spend + output-token spend) for a single request at each shape. Across these three shapes, Gemini 3.1 Flash Lite's output share spans 42.9% to 64.3% — a 21.4%-point swing driven entirely by workload shape, not by any change in the $0.25/$1.50 per-1M rates.
3. Successor rate delta and lifecycle ledger
| Rate | Gemini 3.1 Flash Lite | Gemini 3.5 Flash Lite | Delta ($ / %) |
|---|---|---|---|
| Input $/M | $0.25 | $0.30 | $0.05 (20.0%) |
| Output $/M | $1.50 | $2.50 | $1.00 (66.7%) |
Delta = successor rate − Gemini 3.1 Flash Lite rate, both taken from the current dated pricing registry. Cache, batch, and context-tier pricing are not substituted across models; each figure is model-specific or marked Unavailable.
| Field | Recorded value | Note |
|---|---|---|
| Lifecycle status | legacy | Verified 2026-08-14 |
| Deprecation announced | Unavailable | No inference beyond the dated record |
| Shutdown date | Unavailable | Null/unavailable is not a promise of indefinite availability |
| Successor | Gemini 3.5 Flash Lite | gemini-3-5-flash-lite priced in this registry |
| Context window | Unavailable | Unavailable — no model-specific spec sourced |
Price verified 2026-04-06; lifecycle verified 2026-08-14. Luna is the data owner. "Unavailable" means no compatible dated evidence was found for that field; it is never treated as zero. Dated price source · lifecycle source · Test Gemini 3.1 Flash Lite in All AI Ask.
gemini-3-1-flash-liteGemini 3.1 Flash Lite API Pricing: Ultra-Low Latency Inference
Gemini 3.1 Flash Lite costs $0.25 per million input tokens and $1.50 per million output tokens ($0.5625/M blended at 3:1). Engineered for sub-100ms time-to-first-token responses and high-volume classification workloads. Verified 2026-09-08.
Gemini 3.1 Flash Lite combines sub-dollar per million economics with sub-100ms TTFT latency.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Voice agent conversational turn (500 in, 100 out): $0.000275 per exchange |
| Scenario 2 | Real-time search autocomplete suggestion (200 in, 30 out): $0.000095 per query |
| Scenario 3 | Customer support ticket routing (800 in, 50 out): $0.000275 per ticket |
| Scenario 4 | Document sentiment classification (1.5K in, 100 out): $0.000525 per review |
| Scenario 5 | Form input validation and autocorrection (400 in, 80 out): $0.000220 per validation |
| Scenario 6 | Monthly 100M token high-speed triage fleet: $56.25 total infrastructure spend |
Batch API queuing cuts already-low token costs in half for asynchronous background pipelines.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | 10M token overnight customer survey tagging: $2.81 batch cost vs $5.63 standard |
| Scenario 2 | 50M token product catalog category normalization: $14.06 batch vs $28.13 standard |
| Scenario 3 | 100M token log analysis and anomaly categorization: $28.13 batch vs $56.25 standard |
| Scenario 4 | Batch queues execute within 24-hour turnaround SLA with 100% throughput predictability |
| Scenario 5 | Ideal for non-urgent background data structuring, tagging, and entity extraction |
| Scenario 6 | Reduces effective blended token price to $0.281/M on large asynchronous datasets |
Upgrading to Gemini 3.5 Flash Lite slashes token costs by 85% while boosting inference velocity.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Gemini 3.5 Flash Lite pricing ($0.0375/M in, $0.15/M out): 85% cheaper input, 90% cheaper output |
| Scenario 2 | Gemini 3.5 Flash Lite delivers even lower first-token latency (sub-80ms TTFT) |
| Scenario 3 | Migrating 100M tokens/mo saves $49.69/mo ($6.56 vs $56.25) while upgrading responsiveness |
| Scenario 4 | Drop-in SDK compatibility: model ID swap requires zero code refactoring |
| Scenario 5 | Audited benchmark accuracy: 3.5 Flash Lite matches or exceeds 3.1 on all classification tasks |
| Scenario 6 | Recommended action: immediate migration to Gemini 3.5 Flash Lite for maximum cost efficiency |
How fast is Gemini 3.1 Flash Lite?
How much does Gemini 3.1 Flash Lite cost at scale?
| Tokens / month | Est. cost (blended 3:1) |
|---|---|
| 100,000 | $0.06 |
| 1,000,000 | $0.56 |
| 10,000,000 | $5.63 |
| 100,000,000 | $56.25 |
How does Gemini 3.1 Flash Lite compare with other models?
What should you explore next for Gemini 3.1 Flash Lite?
What are common questions about Gemini 3.1 Flash Lite?
Is Gemini 3.1 Flash Lite cheaper than DeepSeek V4 Flash?
Gemini 3.1 Flash Lite costs $0.56/M blended tokens, DeepSeek V4 Flash costs $0.66/M — Gemini 3.1 Flash Lite is cheaper.
How much does 1 million tokens cost with Gemini 3.1 Flash Lite?
At a 3:1 input:output ratio, 1 million blended tokens costs approximately $0.56. Pure input costs $0.25/M; pure output costs $1.50/M.
What does Gemini 3.1 Flash Lite cost at high volume?
At 100 million blended tokens a month, Gemini 3.1 Flash Lite costs approximately $56.25. See the cost-at-scale table below for other volumes.
