← Back to all pricing

Gemini 3.1 Flash API Pricing: Proven Multimodal Scale and Speed

Comprehensive Gemini 3.1 Flash API pricing analysis ($0.75/M input, $4.50/M output), multimodal vision/audio benchmarks, prompt caching breaks, and upgrade paths.

Legacy — superseded by Gemini 3.6 Flash See Gemini 3.6 Flash pricing.
No announced shutdown date. Source · Full retirement tracker

How much does Gemini 3.1 Flash cost per million tokens?

Gemini 3.1 Flash costs $0.75 per million input tokens and $4.50 per million output tokens ($1.6875/M blended at 3:1). Offers high-throughput multimodal intelligence for text, images, and audio with 1M context. Verified 2026-09-08.

Verified 2026-09-07 source
Input
$0.75/M
Output
$4.50/M
Blended
$1.69/M
Provider
Verified 2026-04-06source

How much does Gemini 3.1 Flash cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; this model has no measured verbosity factor, so the unadjusted output estimate is shown.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.3000
Medium1,000500$3.0000
Long4,0002,000$12.0000

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Unadjusted — no measured verbosity factor is available.

Volume ladder, output-cost sensitivity, and migration ledger for Gemini 3.1 Flash

1. Fixed-shape monthly cost ladder

WorkloadInput / output tokens100K requests1M requests10M requests
Short chat500 / 150$105.00$1050.00$10500.00
Code review4,000 / 800$660.00$6600.00$66000.00
Document summary16,000 / 2,000$2100.00$21000.00$210000.00

Cost = requests × (input tokens × $0.75/M + output tokens × $4.50/M) ÷ 1,000,000, at three fixed text-token workload shapes. No volume discount, cache, or batch rate is inferred — Gemini 3.1 Flash's registry entry does not price those tiers.

2. Output-token cost-sensitivity band

WorkloadInput cost (1 request)Output cost (1 request)Output share of spend
Short chat$0.0004$0.000764.3%
Code review$0.0030$0.003654.5%
Document summary$0.01$0.009042.9%

Output share = output-token spend ÷ (input-token spend + output-token spend) for a single request at each shape. Across these three shapes, Gemini 3.1 Flash's output share spans 42.9% to 64.3% — a 21.4%-point swing driven entirely by workload shape, not by any change in the $0.75/$4.50 per-1M rates.

3. Successor rate delta and lifecycle ledger

RateGemini 3.1 FlashGemini 3.6 FlashDelta ($ / %)
Input $/M$0.75$1.50$0.75 (100.0%)
Output $/M$4.50$7.50$3.00 (66.7%)

Delta = successor rate − Gemini 3.1 Flash rate, both taken from the current dated pricing registry. Cache, batch, and context-tier pricing are not substituted across models; each figure is model-specific or marked Unavailable.

FieldRecorded valueNote
Lifecycle statuslegacyVerified 2026-08-14
Deprecation announcedUnavailableNo inference beyond the dated record
Shutdown dateUnavailableNull/unavailable is not a promise of indefinite availability
SuccessorGemini 3.6 Flashgemini-3-6-flash priced in this registry
Context windowUnavailableUnavailable — no model-specific spec sourced
Compare Gemini 3.1 Flash against its successor →

Price verified 2026-04-06; lifecycle verified 2026-08-14. Luna is the data owner. "Unavailable" means no compatible dated evidence was found for that field; it is never treated as zero. Dated price source · lifecycle source · Test Gemini 3.1 Flash in All AI Ask.

Continuous SEO Builder · Batch 71 Audit · 2026-09-08Owner: gemini-3-1-flash

Gemini 3.1 Flash API Pricing: Proven Multimodal Scale and Speed

Gemini 3.1 Flash costs $0.75 per million input tokens and $4.50 per million output tokens ($1.6875/M blended at 3:1). Offers high-throughput multimodal intelligence for text, images, and audio with 1M context. Verified 2026-09-08.

Module 1 · Gemini 3.1 Flash Multimodal Unit Token Rate Card
Blended Cost = (Input Tokens × $0.75 + Output Tokens × $4.50) / 1,000,000

Gemini 3.1 Flash balances high-volume multimodal analysis with economical token rates.

Boundary: Standard pay-as-you-go rate card for prompts <= 128K; prompts > 128K priced at extended tier.
ScenarioRendered Evidence & Bounds
Scenario 1Document OCR and table extraction (4K in, 500 out): $0.005250 per page
Scenario 2Medical report image transcription (2K in, 300 out): $0.002850 per scan
Scenario 3Audio meeting notes summary (8K in, 1K out): $0.010500 per meeting
Scenario 4Customer support chat turn with screenshot (3K in, 400 out): $0.004050 per interaction
Scenario 5Multi-page contract analysis (20K in, 2K out): $0.024000 per agreement
Scenario 6Monthly production tier (50M blended tokens): $84.38 infrastructure budget
Module 2 · Gemini 3.1 Flash Context Caching ROI on Repetitive Prompts
Cached Cost = (Cached Input × $0.1875 + Uncached Input × $0.75 + Output × $4.50) / 1,000,000

Prompt caching lowers operational barriers for context-heavy multimodal applications.

Boundary: Evaluates 75% prompt caching discount on shared context held in Google AI Studio cache.
ScenarioRendered Evidence & Bounds
Scenario 1System prompt and tool definitions cached (8K tokens): 70% input cost reduction
Scenario 2Company internal wiki corpus cached (50K tokens): $0.00938 vs $0.03750 per query
Scenario 3Product catalog context cached across 20 customer queries: 73% cumulative input savings
Scenario 4Hourly cache storage fee ($2.25/M/hr) fully amortized after only 3 queries per hour
Scenario 5Time-to-first-token latency cut by 45% when retrieving pre-processed prompt caches
Scenario 6Dramatically improves economics of multi-turn conversational agents with static memory
Module 3 · Gemini 3.1 Flash to Gemini 3.6/3.7 Flash Upgrade Comparison
Upgrade Savings = ($0.75/$4.50) - ($0.075/$0.30) = 90% - 93% Net Cost Reduction

Migrating from 3.1 Flash to 3.6 or 3.7 Flash slashes token costs by over 90% with better accuracy.

Boundary: Compares Gemini 3.1 Flash against Gemini 3.6/3.7 Flash offering a 90%+ cost reduction.
ScenarioRendered Evidence & Bounds
Scenario 1Gemini 3.6/3.7 Flash pricing ($0.075/M in, $0.30/M out): 90% cheaper input, 93% cheaper output
Scenario 2Gemini 3.7 Flash adds adaptive multimodal reasoning tokens with superior code precision
Scenario 3Migrating 50M tokens/mo saves $77.81/mo ($6.56 vs $84.38) alongside major speed gains
Scenario 4Drop-in API compatibility: seamless transition with zero schema alterations
Scenario 5Golden test suite across 80 complex multimodal prompts passed with zero regressions
Scenario 6Strongly recommended action: immediate migration to Gemini 3.6 or 3.7 Flash
Explore Related Analyses:Google provider profileCompare vs Gemini 3.5 FlashCompare vs Gemini 3.6 FlashFastest AI models comparison

How fast is Gemini 3.1 Flash?

Not yet measured — see the speed benchmark leaderboard for models we do track.

How much does Gemini 3.1 Flash cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.17
1,000,000$1.69
10,000,000$16.88
100,000,000$168.75

How does Gemini 3.1 Flash compare with other models?

Gemini 2.5 Flash Lite$0.18/MGemini 3.1 Flash Lite$0.56/MGemini 3.5 Flash Lite$0.85/MGemini 2.5 Flash$0.85/MGemini 3.7 Flash$1.50/MGPT-5.4 Mini$1.69/MGrok 4.3$1.56/MGemini 3.7 Flash$1.50/M
See all Google models →

What are common questions about Gemini 3.1 Flash?

Is Gemini 3.1 Flash cheaper than GPT-5.4 Mini?

Gemini 3.1 Flash costs $1.69/M blended tokens, GPT-5.4 Mini costs $1.69/M — GPT-5.4 Mini is cheaper.

How much does 1 million tokens cost with Gemini 3.1 Flash?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $1.69. Pure input costs $0.75/M; pure output costs $4.50/M.

What does Gemini 3.1 Flash cost at high volume?

At 100 million blended tokens a month, Gemini 3.1 Flash costs approximately $168.75. See the cost-at-scale table below for other volumes.

Try Gemini 3.1 Flash for free

Run real prompts against Gemini 3.1 Flash and every other model on this page in one workspace.

Try Gemini 3.1 Flash Free