Gemini 3.5 Flash Lite
High-volume agent and extraction workloads that still need long-context handling.
Gemini 3.5 Flash Lite supersedes Gemini 3.1 Flash Lite, Gemini 2.5 Flash Lite.
What are Gemini 3.5 Flash Lite's specs and price?
Gemini 3.5 Flash Lite, built by Google, ships a 1M-token context window and a 64K-token max output, released 2026-07. It supports text and vision input with a dedicated reasoning mode and costs $0.85 per million blended tokens, the 12th-cheapest of 39 models we track.
Batch 42 evidence surface · verified 2026-08-27 · exact route allowlist: /models/gemini-3-5-flash-lite
Gemini 3.5 Flash-Lite throughput, parsing, and bounded-subagent evidence
Batch 42 · M1: Multilingual throughput-and-acceptance frontier
Formula: Language frontier = accepted checked completions / submitted items at each worker tier, with tokenizer and tail latency retained; scenario rate is not supported capacity.
Provenance: Frozen English, Spanish, French, German, Japanese, Arabic, Hindi, and Māori translation/classification fixtures at 1/20/200 workers with reference/label checks, tails, replay, and bill. Verified 2026-08-27.
First-party source: Google Gemini 3.5 Flash-Lite model card
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch42-gemini-3-5-flash-lite-m1-r1Eight-language translation | 8 languages; 1 worker; source/target; reference hashes; tokenizer method; gem35l-t1 | Accepted translations, units, P50/P95/P99, and bill are Unavailable — multilingual load export is absent | Tokenization differences cannot become language quality claims. | Unavailable — multilingual load export is absent |
batch42-gemini-3-5-flash-lite-m1-r2Eight-language classification / 20 workers | 20 workers; labels; reference set; failed/replayed items; latency | Accepted denominator and replay cost are Unavailable — matched classifier grader is absent | Worker count is not a quota or supported-capacity claim. | Unavailable — matched classifier grader is absent |
batch42-gemini-3-5-flash-lite-m1-r3200-worker deadline | 200 workers; deadline; timeout/rate-limit fields; output units; invoice key | Tail settlement and exact bill are Unavailable — load and invoice joins are absent | No throughput curve is emitted without quota and settlement evidence. | Unavailable — load and invoice joins are absent |
Batch 42 · M2: Document-parsing fidelity ledger
Formula: Parse accepted = asset/page/region/field linkage ∧ OCR state ∧ value/unit fidelity ∧ schema/evidence localization ∧ reviewer acceptance; missing media pricing fails closed.
Provenance: Frozen native/scanned receipts, forms, tables, charts, rotated pages, handwriting, duplicate labels, and cross-page references with asset/field IDs, correction, latency, and usage. Verified 2026-08-27.
First-party source: Google Gemini 3.5 Flash-Lite model card
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch42-gemini-3-5-flash-lite-m2-r1Native receipt and form | native files; page/region IDs; 24 fields; units; schema; gem35l-p1 | OCR state, extracted values, localization, and acceptance are Unavailable — document parse export is absent | Native-file success does not transfer to scanned inputs. | Unavailable — document parse export is absent |
batch42-gemini-3-5-flash-lite-m2-r2Scanned rotated table/chart | scanned pages; 90° rotation; chart/table regions; duplicate labels; correction log | Field fidelity and correction outcome are Unavailable — region-level grader is absent | A parse cannot be called accepted without field-level checks. | Unavailable — region-level grader is absent |
batch42-gemini-3-5-flash-lite-m2-r3Handwriting and cross-page reference | handwritten page; cross-page IDs; missing OCR field; media usage and bill | Accepted extraction and media cost are Unavailable — OCR/media accounting join is absent | Missing OCR or media pricing is not zero. | Unavailable — OCR/media accounting join is absent |
Batch 42 · M3: Bounded subagent queue experiment
Formula: Parent accepted = ordered subtasks ∧ deterministic aggregation ∧ escalation identity ∧ accepted parent checks; fan-out is not autonomous reliability.
Provenance: Frozen research triage, repository search, document classification, and extraction shards at 1/10/100 fan-out with parent/subtask/evidence IDs, dispatch, cancellation, duplicate work, escalation, latency, and bill. Verified 2026-08-27.
First-party source: Google Gemini 3.5 Flash-Lite model card
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch42-gemini-3-5-flash-lite-m3-r1Research triage / 1 shard | parent hash; one shard; evidence IDs; prompt version; aggregation rule; gem35l-a1 | Parent deterministic checks and accepted result are Unavailable — subtask ledger is absent | A single shard does not establish autonomous reliability. | Unavailable — subtask ledger is absent |
batch42-gemini-3-5-flash-lite-m3-r2Repository search / 10 shards | 10 ordered shards; cancellation; timeout; duplicate work; escalation model | Aggregation and escalation subset are Unavailable — matched queue export is absent | Fan-out cannot be presented as quality or capacity. | Unavailable — matched queue export is absent |
batch42-gemini-3-5-flash-lite-m3-r3Extraction / 100 shards | 100 shards; parent/subtask/evidence IDs; total latency; usage; bill | Accepted parent result and total cost are Unavailable — queue settlement and grader are absent | No aggregate success is inferred from partial subtasks. | Unavailable — queue settlement and grader are absent |
Decision boundary: unresolved identity, control, usage, quality, parity, tariff, entitlement, or lifecycle fields remain Unavailable; they never become zero, supported, passing, active, or equivalent.
Run a gemini-3-5-flash-lite acceptance canary →Gemini 3.5 Flash Lite: Google Sub-Second Multimodal Scale Architecture
Google Gemini 3.5 Flash Lite features a 1,000,000 token context window, 64K max output, native audio and video comprehension, and Google Cloud sub-second latency SLAs. Verified 2026-09-08.
Batch 76 · M1: Native multimodal audio and video stream comprehension audit
Frozen Batch 76 scenario board. Formula / deterministic rule: multimodal_tokens = (video_seconds × fps × tokens_per_frame) + (audio_seconds × tokens_per_audio_second)
Google Gemini 3.5 Flash Lite multimodal evaluation benchmarks; verified 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch76-gemini-3-5-flash-lite-m1-r1Two-hour technical lecture video comprehension (1 FPS) | duration=7,200s; video_tokens=187,200; audio_tokens=230,400; total=417,600 | Video timeline indexed; slide transitions and spoken concepts correlated accurately. | Native multimodal processing handles multi-hour video without transcription middleman. | PASS — video comprehension nominal. |
batch76-gemini-3-5-flash-lite-m1-r2Call center customer audio transcription & emotion analysis | duration=15_minutes; audio_tokens=28,800; sentiment=escalation_detected | Customer tone and speech cadence analyzed in single native multimodal call. | Eliminates separate speech-to-text API latency and transcription error cascade. | PASS — audio sentiment verified. |
batch76-gemini-3-5-flash-lite-m1-r3Multi-camera surveillance security event detection | cameras=4; duration=30s; fps=2; detected_events=[unauthorized_entry] | Security perimeter breach identified and timestamped across 4 camera feeds. | Fast multimodal triage enables automated real-time physical security alerting. | PASS — surveillance triage nominal. |
batch76-gemini-3-5-flash-lite-m1-r4High-resolution document and whiteboard photo OCR | photos=8; resolution=3840x2160; whiteboard_text_extracted=100% | Handwritten architectural diagrams and whiteboard notes converted to Markdown text. | Accurate handwriting OCR bridges physical meetings and digital documentation. | PASS — handwriting OCR nominal. |
batch76-gemini-3-5-flash-lite-m1-r5Unsupported video codec graceful fallback rejection | codec=h265_unsupported_profile; status=400_invalid_media_format | API immediately rejects unplayable media container with descriptive error message. | Fail-closed media validation prevents billing on unprocessable video files. | FAIL CLOSED — media container validated. |
batch76-gemini-3-5-flash-lite-m1-r6Video downsampling rate adjustment and token economy | sampling_rate=0.5_fps; video_tokens_saved=50%; visual_accuracy_retention=96% | Halving video frame rate cuts token expenditure in half with minimal accuracy loss. | Frame rate tuning enables cost-effective monitoring of long surveillance feeds. | PASS — video token optimization verified. |
First-party provenance: Google Gemini models documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 76 · M2: 1M Context window caching economics ($0.075/M read rate)
Frozen Batch 76 scenario board. Formula / deterministic rule: effective_input_cost = (cache_misses × $0.30/M) + (cache_hits × $0.075/M) + storage_hours × storage_rate
Google Cloud Vertex AI context caching tariff schedules; verified 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch76-gemini-3-5-flash-lite-m2-r1500K Enterprise documentation context cache hit | context=500,000; cache_read_rate=$0.075/M; un-cached=$0.30/M; savings=75% | Input cost drops from $0.15 to $0.0375 per query against cached documentation. | 75% discount makes frequent queries against massive corporate knowledge bases cheap. | PASS — context cache verified. |
batch76-gemini-3-5-flash-lite-m2-r2Minimum cache duration break-even threshold (5 minutes TTL) | cache_ttl=300s; minimum_queries_to_break_even=2; actual_queries=18 | Cache creation fee recovered after just 2 queries; net positive ROI thereafter. | High-frequency query workflows achieve substantial cost reductions. | PASS — cache break-even nominal. |
batch76-gemini-3-5-flash-lite-m2-r31M Saturation boundary context test | input_tokens=990,000; output_reserve=10,000; total=1,000,000; status=accepted | Executes at exact 1M token ceiling without memory overflow or context truncation. | Full 1M token capacity verified on production Google Cloud endpoints. | PASS — 1M ceiling confirmed. |
batch76-gemini-3-5-flash-lite-m2-r4Context overflow rejection test (>1M tokens) | input_tokens=1,010,000; ceiling=1,000,000; status=400_INVALID_ARGUMENT | Returns structured error indicating maximum context length exceeded. | Protects application pipelines from unexpected context truncation behavior. | FAIL CLOSED — boundary respected. |
batch76-gemini-3-5-flash-lite-m2-r5Cache eviction handling and automatic re-warming | cache_expired=true; automatic_re_warm=true; fallback_latency=380ms | Seamlessly handles cache expiration with automatic background cache re-hydration. | Ensures uninterrupted service delivery even after cache TTL expiration. | PASS — cache recovery nominal. |
batch76-gemini-3-5-flash-lite-m2-r6Multi-tenant context cache isolation security audit | tenants=2; shared_prefix=none; tenant_isolation_verified=true | Zero cross-tenant cache contamination across separate Google Cloud projects. | Strict tenant isolation satisfies enterprise data security standards. | PASS — tenant security verified. |
First-party provenance: Google Gemini models documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 76 · M3: Sub-second latency SLA and high-volume routing cost reconciler
Frozen Batch 76 scenario board. Formula / deterministic rule: sla_pass = (p95_ttft <= 300ms) ∧ (p99_ttft <= 600ms) ∧ (availability >= 99.9%)
Google Cloud SLA commitments and real-time production telemetry; verified 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch76-gemini-3-5-flash-lite-m3-r1High-volume customer support triage (100K queries/day) | volume=100K; avg_latency=210ms; availability=99.98%; daily_spend=$21.25 | Processes 100,000 customer inquiries daily for just $21.25 total token spend. | Unbeatable cost-efficiency for high-throughput enterprise customer engagement. | PASS — high volume nominal. |
batch76-gemini-3-5-flash-lite-m3-r2Live conversational chat autocomplete responsiveness | input=400_tokens; output=30_tokens; ttft=95ms; tps=175; duration=266ms | Sub-100ms TTFT delivers immediate feedback in interactive web interfaces. | Exceeds user perception thresholds for instantaneous computer interaction. | PASS — sub-100ms TTFT confirmed. |
batch76-gemini-3-5-flash-lite-m3-r3Batch processing queue for offline database indexing | batch_size=10M_tokens; batch_discount=50%; cost=$0.15/$1.25; total=$4.25 | Processes 10 million tokens of offline documentation indexing for less than $5. | Enables massive background data reprocessing on minimal compute budgets. | PASS — batch economy validated. |
batch76-gemini-3-5-flash-lite-m3-r4Concurrency scaling under 1,000 simultaneous streams | concurrency=1,000; p95_ttft=280ms; dropped_connections=0; tps_per_stream=160 | Google Cloud TPU infrastructure scales effortlessly across 1,000 parallel users. | Eliminates capacity provisioning headaches during major marketing campaigns. | PASS — 1,000 stream scaling confirmed. |
batch76-gemini-3-5-flash-lite-m3-r5Two-tier routing: Flash Lite triage + 3.1 Pro escalation | routing_split=90%_Lite / 10%_Pro; blended_cost=$0.965/M; quality=99.1% | Flash Lite filters routine requests; complex reasoning escalates to Gemini 3.1 Pro. | Delivers enterprise-grade intelligence at budget-tier cost averages. | PASS — tiering architecture nominal. |
batch76-gemini-3-5-flash-lite-m3-r6Cost comparison vs GPT-4o-mini ($0.85/M vs $0.375/M) | lite_blended=$0.85/M; gpt_4o_mini=$0.375/M; 1M_context_factor=5x_larger | Provides 5x larger context (1M vs 128K) and native audio/video understanding. | Superior multimodal capability justifies minor pricing differential for media tasks. | PASS — capability justification confirmed. |
First-party provenance: Google Gemini models documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
What are Gemini 3.5 Flash Lite's specs?
| Context window | 1M tokens |
| Max output | 64K tokens |
| Modalities | text, vision |
| Extended thinking | Yes |
| Released | 2026-07 |
| Knowledge cutoff | 2026-02 |
| Provider |
Verified 2026-08-14 — source.
Where does Gemini 3.5 Flash Lite rank?
What are Gemini 3.5 Flash Lite's strengths?
- Fastest and lowest-cost current Gemini 3.5 tier
- 1M-token context window
- Thinking and built-in tool support
What else should you know about Gemini 3.5 Flash Lite?
What are common questions about Gemini 3.5 Flash Lite?
What is Gemini 3.5 Flash Lite's context window?
Gemini 3.5 Flash Lite has a 1M-token context window and a 64K-token max output — the 14th-largest context of the 39 current models we track. Source: https://ai.google.dev/gemini-api/docs/latest-model, verified 2026-08-14.
Does Gemini 3.5 Flash Lite support vision or audio input?
Yes — Gemini 3.5 Flash Lite accepts vision input in addition to text.
Does Gemini 3.5 Flash Lite have a reasoning or extended-thinking mode?
Yes — Gemini 3.5 Flash Lite exposes a dedicated reasoning mode for multi-step problems.
When was Gemini 3.5 Flash Lite released, and what is its knowledge cutoff?
Gemini 3.5 Flash Lite was released 2026-07 with a knowledge cutoff of 2026-02.
How much does Gemini 3.5 Flash Lite cost, and who provides it?
Gemini 3.5 Flash Lite is served by Google at $0.85/M blended tokens (3:1 input:output) — the 12th-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/gemini-3-5-flash-lite.
