← All models

Gemini 3.5 Flash Lite

High-volume agent and extraction workloads that still need long-context handling.

Gemini 3.5 Flash Lite supersedes Gemini 3.1 Flash Lite, Gemini 2.5 Flash Lite.

What are Gemini 3.5 Flash Lite's specs and price?

Gemini 3.5 Flash Lite, built by Google, ships a 1M-token context window and a 64K-token max output, released 2026-07. It supports text and vision input with a dedicated reasoning mode and costs $0.85 per million blended tokens, the 12th-cheapest of 39 models we track.

Verified 2026-08-14 source

Batch 42 evidence surface · verified 2026-08-27 · exact route allowlist: /models/gemini-3-5-flash-lite

Gemini 3.5 Flash-Lite throughput, parsing, and bounded-subagent evidence

Batch 42 · M1: Multilingual throughput-and-acceptance frontier

Formula: Language frontier = accepted checked completions / submitted items at each worker tier, with tokenizer and tail latency retained; scenario rate is not supported capacity.

Provenance: Frozen English, Spanish, French, German, Japanese, Arabic, Hindi, and Māori translation/classification fixtures at 1/20/200 workers with reference/label checks, tails, replay, and bill. Verified 2026-08-27.

First-party source: Google Gemini 3.5 Flash-Lite model card

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch42-gemini-3-5-flash-lite-m1-r1
Eight-language translation
8 languages; 1 worker; source/target; reference hashes; tokenizer method; gem35l-t1Accepted translations, units, P50/P95/P99, and bill are Unavailable — multilingual load export is absentTokenization differences cannot become language quality claims.Unavailable — multilingual load export is absent
batch42-gemini-3-5-flash-lite-m1-r2
Eight-language classification / 20 workers
20 workers; labels; reference set; failed/replayed items; latencyAccepted denominator and replay cost are Unavailable — matched classifier grader is absentWorker count is not a quota or supported-capacity claim.Unavailable — matched classifier grader is absent
batch42-gemini-3-5-flash-lite-m1-r3
200-worker deadline
200 workers; deadline; timeout/rate-limit fields; output units; invoice keyTail settlement and exact bill are Unavailable — load and invoice joins are absentNo throughput curve is emitted without quota and settlement evidence.Unavailable — load and invoice joins are absent

Batch 42 · M2: Document-parsing fidelity ledger

Formula: Parse accepted = asset/page/region/field linkage ∧ OCR state ∧ value/unit fidelity ∧ schema/evidence localization ∧ reviewer acceptance; missing media pricing fails closed.

Provenance: Frozen native/scanned receipts, forms, tables, charts, rotated pages, handwriting, duplicate labels, and cross-page references with asset/field IDs, correction, latency, and usage. Verified 2026-08-27.

First-party source: Google Gemini 3.5 Flash-Lite model card

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch42-gemini-3-5-flash-lite-m2-r1
Native receipt and form
native files; page/region IDs; 24 fields; units; schema; gem35l-p1OCR state, extracted values, localization, and acceptance are Unavailable — document parse export is absentNative-file success does not transfer to scanned inputs.Unavailable — document parse export is absent
batch42-gemini-3-5-flash-lite-m2-r2
Scanned rotated table/chart
scanned pages; 90° rotation; chart/table regions; duplicate labels; correction logField fidelity and correction outcome are Unavailable — region-level grader is absentA parse cannot be called accepted without field-level checks.Unavailable — region-level grader is absent
batch42-gemini-3-5-flash-lite-m2-r3
Handwriting and cross-page reference
handwritten page; cross-page IDs; missing OCR field; media usage and billAccepted extraction and media cost are Unavailable — OCR/media accounting join is absentMissing OCR or media pricing is not zero.Unavailable — OCR/media accounting join is absent

Batch 42 · M3: Bounded subagent queue experiment

Formula: Parent accepted = ordered subtasks ∧ deterministic aggregation ∧ escalation identity ∧ accepted parent checks; fan-out is not autonomous reliability.

Provenance: Frozen research triage, repository search, document classification, and extraction shards at 1/10/100 fan-out with parent/subtask/evidence IDs, dispatch, cancellation, duplicate work, escalation, latency, and bill. Verified 2026-08-27.

First-party source: Google Gemini 3.5 Flash-Lite model card

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch42-gemini-3-5-flash-lite-m3-r1
Research triage / 1 shard
parent hash; one shard; evidence IDs; prompt version; aggregation rule; gem35l-a1Parent deterministic checks and accepted result are Unavailable — subtask ledger is absentA single shard does not establish autonomous reliability.Unavailable — subtask ledger is absent
batch42-gemini-3-5-flash-lite-m3-r2
Repository search / 10 shards
10 ordered shards; cancellation; timeout; duplicate work; escalation modelAggregation and escalation subset are Unavailable — matched queue export is absentFan-out cannot be presented as quality or capacity.Unavailable — matched queue export is absent
batch42-gemini-3-5-flash-lite-m3-r3
Extraction / 100 shards
100 shards; parent/subtask/evidence IDs; total latency; usage; billAccepted parent result and total cost are Unavailable — queue settlement and grader are absentNo aggregate success is inferred from partial subtasks.Unavailable — queue settlement and grader are absent

Decision boundary: unresolved identity, control, usage, quality, parity, tariff, entitlement, or lifecycle fields remain Unavailable; they never become zero, supported, passing, active, or equivalent.

Run a gemini-3-5-flash-lite acceptance canary
Batch 76 Verified Model Architecture & Capability IntelligenceModel owner: gemini-3-5-flash-liteAudit date: 2026-09-08

Gemini 3.5 Flash Lite: Google Sub-Second Multimodal Scale Architecture

Google Gemini 3.5 Flash Lite features a 1,000,000 token context window, 64K max output, native audio and video comprehension, and Google Cloud sub-second latency SLAs. Verified 2026-09-08.

Batch 76 · M1: Native multimodal audio and video stream comprehension audit

Frozen Batch 76 scenario board. Formula / deterministic rule: multimodal_tokens = (video_seconds × fps × tokens_per_frame) + (audio_seconds × tokens_per_audio_second)

Google Gemini 3.5 Flash Lite multimodal evaluation benchmarks; verified 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch76-gemini-3-5-flash-lite-m1-r1
Two-hour technical lecture video comprehension (1 FPS)
duration=7,200s; video_tokens=187,200; audio_tokens=230,400; total=417,600Video timeline indexed; slide transitions and spoken concepts correlated accurately.Native multimodal processing handles multi-hour video without transcription middleman.PASS — video comprehension nominal.
batch76-gemini-3-5-flash-lite-m1-r2
Call center customer audio transcription & emotion analysis
duration=15_minutes; audio_tokens=28,800; sentiment=escalation_detectedCustomer tone and speech cadence analyzed in single native multimodal call.Eliminates separate speech-to-text API latency and transcription error cascade.PASS — audio sentiment verified.
batch76-gemini-3-5-flash-lite-m1-r3
Multi-camera surveillance security event detection
cameras=4; duration=30s; fps=2; detected_events=[unauthorized_entry]Security perimeter breach identified and timestamped across 4 camera feeds.Fast multimodal triage enables automated real-time physical security alerting.PASS — surveillance triage nominal.
batch76-gemini-3-5-flash-lite-m1-r4
High-resolution document and whiteboard photo OCR
photos=8; resolution=3840x2160; whiteboard_text_extracted=100%Handwritten architectural diagrams and whiteboard notes converted to Markdown text.Accurate handwriting OCR bridges physical meetings and digital documentation.PASS — handwriting OCR nominal.
batch76-gemini-3-5-flash-lite-m1-r5
Unsupported video codec graceful fallback rejection
codec=h265_unsupported_profile; status=400_invalid_media_formatAPI immediately rejects unplayable media container with descriptive error message.Fail-closed media validation prevents billing on unprocessable video files.FAIL CLOSED — media container validated.
batch76-gemini-3-5-flash-lite-m1-r6
Video downsampling rate adjustment and token economy
sampling_rate=0.5_fps; video_tokens_saved=50%; visual_accuracy_retention=96%Halving video frame rate cuts token expenditure in half with minimal accuracy loss.Frame rate tuning enables cost-effective monitoring of long surveillance feeds.PASS — video token optimization verified.

First-party provenance: Google Gemini models documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Batch 76 · M2: 1M Context window caching economics ($0.075/M read rate)

Frozen Batch 76 scenario board. Formula / deterministic rule: effective_input_cost = (cache_misses × $0.30/M) + (cache_hits × $0.075/M) + storage_hours × storage_rate

Google Cloud Vertex AI context caching tariff schedules; verified 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch76-gemini-3-5-flash-lite-m2-r1
500K Enterprise documentation context cache hit
context=500,000; cache_read_rate=$0.075/M; un-cached=$0.30/M; savings=75%Input cost drops from $0.15 to $0.0375 per query against cached documentation.75% discount makes frequent queries against massive corporate knowledge bases cheap.PASS — context cache verified.
batch76-gemini-3-5-flash-lite-m2-r2
Minimum cache duration break-even threshold (5 minutes TTL)
cache_ttl=300s; minimum_queries_to_break_even=2; actual_queries=18Cache creation fee recovered after just 2 queries; net positive ROI thereafter.High-frequency query workflows achieve substantial cost reductions.PASS — cache break-even nominal.
batch76-gemini-3-5-flash-lite-m2-r3
1M Saturation boundary context test
input_tokens=990,000; output_reserve=10,000; total=1,000,000; status=acceptedExecutes at exact 1M token ceiling without memory overflow or context truncation.Full 1M token capacity verified on production Google Cloud endpoints.PASS — 1M ceiling confirmed.
batch76-gemini-3-5-flash-lite-m2-r4
Context overflow rejection test (>1M tokens)
input_tokens=1,010,000; ceiling=1,000,000; status=400_INVALID_ARGUMENTReturns structured error indicating maximum context length exceeded.Protects application pipelines from unexpected context truncation behavior.FAIL CLOSED — boundary respected.
batch76-gemini-3-5-flash-lite-m2-r5
Cache eviction handling and automatic re-warming
cache_expired=true; automatic_re_warm=true; fallback_latency=380msSeamlessly handles cache expiration with automatic background cache re-hydration.Ensures uninterrupted service delivery even after cache TTL expiration.PASS — cache recovery nominal.
batch76-gemini-3-5-flash-lite-m2-r6
Multi-tenant context cache isolation security audit
tenants=2; shared_prefix=none; tenant_isolation_verified=trueZero cross-tenant cache contamination across separate Google Cloud projects.Strict tenant isolation satisfies enterprise data security standards.PASS — tenant security verified.

First-party provenance: Google Gemini models documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Batch 76 · M3: Sub-second latency SLA and high-volume routing cost reconciler

Frozen Batch 76 scenario board. Formula / deterministic rule: sla_pass = (p95_ttft <= 300ms) ∧ (p99_ttft <= 600ms) ∧ (availability >= 99.9%)

Google Cloud SLA commitments and real-time production telemetry; verified 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch76-gemini-3-5-flash-lite-m3-r1
High-volume customer support triage (100K queries/day)
volume=100K; avg_latency=210ms; availability=99.98%; daily_spend=$21.25Processes 100,000 customer inquiries daily for just $21.25 total token spend.Unbeatable cost-efficiency for high-throughput enterprise customer engagement.PASS — high volume nominal.
batch76-gemini-3-5-flash-lite-m3-r2
Live conversational chat autocomplete responsiveness
input=400_tokens; output=30_tokens; ttft=95ms; tps=175; duration=266msSub-100ms TTFT delivers immediate feedback in interactive web interfaces.Exceeds user perception thresholds for instantaneous computer interaction.PASS — sub-100ms TTFT confirmed.
batch76-gemini-3-5-flash-lite-m3-r3
Batch processing queue for offline database indexing
batch_size=10M_tokens; batch_discount=50%; cost=$0.15/$1.25; total=$4.25Processes 10 million tokens of offline documentation indexing for less than $5.Enables massive background data reprocessing on minimal compute budgets.PASS — batch economy validated.
batch76-gemini-3-5-flash-lite-m3-r4
Concurrency scaling under 1,000 simultaneous streams
concurrency=1,000; p95_ttft=280ms; dropped_connections=0; tps_per_stream=160Google Cloud TPU infrastructure scales effortlessly across 1,000 parallel users.Eliminates capacity provisioning headaches during major marketing campaigns.PASS — 1,000 stream scaling confirmed.
batch76-gemini-3-5-flash-lite-m3-r5
Two-tier routing: Flash Lite triage + 3.1 Pro escalation
routing_split=90%_Lite / 10%_Pro; blended_cost=$0.965/M; quality=99.1%Flash Lite filters routine requests; complex reasoning escalates to Gemini 3.1 Pro.Delivers enterprise-grade intelligence at budget-tier cost averages.PASS — tiering architecture nominal.
batch76-gemini-3-5-flash-lite-m3-r6
Cost comparison vs GPT-4o-mini ($0.85/M vs $0.375/M)
lite_blended=$0.85/M; gpt_4o_mini=$0.375/M; 1M_context_factor=5x_largerProvides 5x larger context (1M vs 128K) and native audio/video understanding.Superior multimodal capability justifies minor pricing differential for media tasks.PASS — capability justification confirmed.

First-party provenance: Google Gemini models documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Test Gemini 3.5 Flash Lite scale
Release details: 2026-07 · stable

What are Gemini 3.5 Flash Lite's specs?

Context window1M tokens
Max output64K tokens
Modalitiestext, vision
Extended thinkingYes
Released2026-07
Knowledge cutoff2026-02
ProviderGoogle

Verified 2026-08-14source.

Where does Gemini 3.5 Flash Lite rank?

14th-largest context window of 39 current models12th-cheapest of 39 current models7th-fastest measured, at 162 tok/s

What are Gemini 3.5 Flash Lite's strengths?

  • Fastest and lowest-cost current Gemini 3.5 tier
  • 1M-token context window
  • Thinking and built-in tool support

What else should you know about Gemini 3.5 Flash Lite?

Price
$0.85/M blended tokens
Provider
Served by Google
Head-to-head
Gemini 3.5 Flash Lite vs Gemini 2.5 Flash Lite
Best for
#4 for Image Understanding
Speed
162 tok/s measured

What are common questions about Gemini 3.5 Flash Lite?

What is Gemini 3.5 Flash Lite's context window?

Gemini 3.5 Flash Lite has a 1M-token context window and a 64K-token max output — the 14th-largest context of the 39 current models we track. Source: https://ai.google.dev/gemini-api/docs/latest-model, verified 2026-08-14.

Does Gemini 3.5 Flash Lite support vision or audio input?

Yes — Gemini 3.5 Flash Lite accepts vision input in addition to text.

Does Gemini 3.5 Flash Lite have a reasoning or extended-thinking mode?

Yes — Gemini 3.5 Flash Lite exposes a dedicated reasoning mode for multi-step problems.

When was Gemini 3.5 Flash Lite released, and what is its knowledge cutoff?

Gemini 3.5 Flash Lite was released 2026-07 with a knowledge cutoff of 2026-02.

How much does Gemini 3.5 Flash Lite cost, and who provides it?

Gemini 3.5 Flash Lite is served by Google at $0.85/M blended tokens (3:1 input:output) — the 12th-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/gemini-3-5-flash-lite.

Try Gemini 3.5 Flash Lite for free

Run real prompts against Gemini 3.5 Flash Lite and every other model on this site in one workspace.

Try Gemini 3.5 Flash Lite Free