← Back to all pricing

Gemini 3.5 Flash API Pricing: Fast Multimodal Intelligence at Scale

Comprehensive Gemini 3.5 Flash API pricing analysis ($1.50/M input, $9.00/M output), 1M context caching, multimodal audio/video processing, and Google AI Studio SLAs.

Legacy — superseded by Gemini 3.6 Flash See Gemini 3.6 Flash pricing.
No announced shutdown date. Source · Full retirement tracker

How much does Gemini 3.5 Flash cost per million tokens?

Gemini 3.5 Flash costs $1.50 per million input tokens and $9.00 per million output tokens ($3.375/M blended at 3:1). Delivers low-latency multimodal processing across text, audio, images, and video with up to 1M context. Verified 2026-09-08.

Verified 2026-09-07 source
Input
$1.50/M
Output
$9.00/M
Blended
$3.38/M
Provider
Verified 2026-08-14source

How much does Gemini 3.5 Flash cost per 1,000 requests?

Computed from generated token pricing. Each row assumes the listed input and output tokens per request; this model has no measured verbosity factor, so the unadjusted output estimate is shown.

Request shapeInput tokensOutput tokensCost / 1,000 requests
Short10050$0.6000
Medium1,000500$6.0000
Long4,0002,000$24.0000

Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Unadjusted — no measured verbosity factor is available.

Volume ladder, output-cost sensitivity, and migration ledger for Gemini 3.5 Flash

1. Fixed-shape monthly cost ladder

WorkloadInput / output tokens100K requests1M requests10M requests
Short chat500 / 150$210.00$2100.00$21000.00
Code review4,000 / 800$1320.00$13200.00$132000.00
Document summary16,000 / 2,000$4200.00$42000.00$420000.00

Cost = requests × (input tokens × $1.50/M + output tokens × $9.00/M) ÷ 1,000,000, at three fixed text-token workload shapes. No volume discount, cache, or batch rate is inferred — Gemini 3.5 Flash's registry entry does not price those tiers.

2. Output-token cost-sensitivity band

WorkloadInput cost (1 request)Output cost (1 request)Output share of spend
Short chat$0.0008$0.001464.3%
Code review$0.0060$0.007254.5%
Document summary$0.02$0.0242.9%

Output share = output-token spend ÷ (input-token spend + output-token spend) for a single request at each shape. Across these three shapes, Gemini 3.5 Flash's output share spans 42.9% to 64.3% — a 21.4%-point swing driven entirely by workload shape, not by any change in the $1.50/$9.00 per-1M rates.

3. Successor rate delta and lifecycle ledger

RateGemini 3.5 FlashGemini 3.6 FlashDelta ($ / %)
Input $/M$1.50$1.50$0.0000 (0.0%)
Output $/M$9.00$7.50$-1.5000 (-16.7%)

Delta = successor rate − Gemini 3.5 Flash rate, both taken from the current dated pricing registry. Cache, batch, and context-tier pricing are not substituted across models; each figure is model-specific or marked Unavailable.

FieldRecorded valueNote
Lifecycle statuslegacyVerified 2026-08-14
Deprecation announcedUnavailableNo inference beyond the dated record
Shutdown dateUnavailableNull/unavailable is not a promise of indefinite availability
SuccessorGemini 3.6 Flashgemini-3-6-flash priced in this registry
Context windowUnavailableUnavailable — no model-specific spec sourced
Compare Gemini 3.5 Flash against its successor →

Price verified 2026-08-14; lifecycle verified 2026-08-14. Luna is the data owner. "Unavailable" means no compatible dated evidence was found for that field; it is never treated as zero. Dated price source · lifecycle source · Test Gemini 3.5 Flash in All AI Ask.

Continuous SEO Builder · Batch 71 Audit · 2026-09-08Owner: gemini-3-5-flash

Gemini 3.5 Flash API Pricing: Fast Multimodal Intelligence at Scale

Gemini 3.5 Flash costs $1.50 per million input tokens and $9.00 per million output tokens ($3.375/M blended at 3:1). Delivers low-latency multimodal processing across text, audio, images, and video with up to 1M context. Verified 2026-09-08.

Module 1 · Gemini 3.5 Flash Multimodal Unit Token Economics
Blended Cost = (Input Tokens × $1.50 + Output Tokens × $9.00) / 1,000,000

Gemini 3.5 Flash provides native multimodal token handling across text, audio, and visual inputs.

Boundary: Standard pay-as-you-go rate card for prompts <= 128K; prompts > 128K priced at extended tier.
ScenarioRendered Evidence & Bounds
Scenario 1High-resolution image analysis (1,500 tokens in, 200 out): $0.004050 per image
Scenario 21-minute audio recording transcription & summary (4K in, 500 out): $0.010500 per minute
Scenario 35-minute video keyframe extraction (20K in, 1K out): $0.039000 per video clip
Scenario 4PDF technical manual analysis (30K in, 2K out): $0.063000 per document
Scenario 5Customer video KYC verification (15K in, 500 out): $0.027000 per verification
Scenario 6Monthly multimodal processing tier (50M blended tokens): $168.75 infrastructure budget
Module 2 · Gemini 3.5 Flash Context Caching Amortization Thresholds
Cached Cost = (Cached Input × $0.375 + Uncached Input × $1.50 + Output × $9.00) / 1,000,000

Google context caching slashes repetitive input processing costs on large media collections.

Boundary: Evaluates 75% prompt caching discount on context stored in Google AI Studio / Vertex AI cache.
ScenarioRendered Evidence & Bounds
Scenario 1Cached video tutorial corpus (100K cached tokens, 2K query): 73% input cost reduction
Scenario 2Large product catalog cache (200K cached tokens): $0.07500 vs $0.30000 per search query
Scenario 3Multi-turn enterprise knowledge base Q&A: 70% net input savings across query bursts
Scenario 4Hourly storage fee ($4.50/M/hr) fully amortized after only 4 queries per hour
Scenario 5Latency reduction: cached queries bypass prompt processing stage, cutting TTFT by 50%
Scenario 6Enables interactive conversational exploration of heavy video and document archives
Module 3 · Gemini 3.5 Flash to Gemini 3.6/3.7 Flash Upgrade Comparison
Upgrade Savings = Generational Price Drop ($0.075/$0.30 vs $1.50/$9.00 = 95% Cost Reduction)

Migrating to Gemini 3.6 or 3.7 Flash unlocks next-generation speed and a 95% token cost reduction.

Boundary: Compares Gemini 3.5 Flash against modern Gemini 3.6/3.7 Flash offering a 95% price reduction.
ScenarioRendered Evidence & Bounds
Scenario 1Gemini 3.6/3.7 Flash pricing ($0.075/M in, $0.30/M out): 95% cheaper than 3.5 Flash
Scenario 2Gemini 3.7 Flash introduces native multimodal thinking tokens with superior reasoning
Scenario 3Migrating 50M tokens/mo saves $162.19/mo ($6.56 vs $168.75) with major latency gains
Scenario 4API compatibility: seamless migration on Google AI Studio and Vertex AI SDKs
Scenario 5Zero breaking changes in multimodal response formats or schema parameters
Scenario 6Strongly recommended action: immediate migration to Gemini 3.6 or 3.7 Flash
Explore Related Analyses:Google provider profileCompare vs Gemini 3.6 FlashCompare vs Gemini 3.5 Flash LiteFastest AI models comparison

How fast is Gemini 3.5 Flash?

Not yet measured — see the speed benchmark leaderboard for models we do track.

How much does Gemini 3.5 Flash cost at scale?

Tokens / monthEst. cost (blended 3:1)
100,000$0.34
1,000,000$3.38
10,000,000$33.75
100,000,000$337.50

How does Gemini 3.5 Flash compare with other models?

Gemini 2.5 Flash Lite$0.18/MGemini 3.1 Flash Lite$0.56/MGemini 3.5 Flash Lite$0.85/MGemini 2.5 Flash$0.85/MGemini 3.7 Flash$1.50/MGPT-5$3.44/MGPT-4.1$3.50/MGrok-4.20 Reasoning$3.00/M
See all Google models →

What are common questions about Gemini 3.5 Flash?

Is Gemini 3.5 Flash cheaper than GPT-5?

Gemini 3.5 Flash costs $3.38/M blended tokens, GPT-5 costs $3.44/M — Gemini 3.5 Flash is cheaper.

How much does 1 million tokens cost with Gemini 3.5 Flash?

At a 3:1 input:output ratio, 1 million blended tokens costs approximately $3.38. Pure input costs $1.50/M; pure output costs $9.00/M.

What does Gemini 3.5 Flash cost at high volume?

At 100 million blended tokens a month, Gemini 3.5 Flash costs approximately $337.50. See the cost-at-scale table below for other volumes.

Try Gemini 3.5 Flash for free

Run real prompts against Gemini 3.5 Flash and every other model on this page in one workspace.

Try Gemini 3.5 Flash Free