Gemini 3.6 Flash
Latency-sensitive coding and agentic tasks that still need a huge context window.
Gemini 3.6 Flash supersedes Gemini 3.5 Flash, Gemini 3.1 Flash, Gemini 2.5 Flash.
What are Gemini 3.6 Flash's specs and price?
Gemini 3.6 Flash, built by Google, ships a 1M-token context window and a 64K-token max output, released 2026-06. It supports text and vision and audio input with a dedicated reasoning mode and costs $3.00 per million blended tokens, the 30th-cheapest of 39 models we track.
Batch 42 evidence surface · verified 2026-08-27 · exact route allowlist: /models/gemini-3-6-flash
Gemini 3.6 Flash thinking, media admission, and streaming evidence
Batch 42 · M1: Thinking-accounting frontier
Formula: Thinking result = effective identity/level ∧ prompt/media/thought/output usage fields ∧ finish/grader checks; unsupported levels fail closed.
Provenance: Frozen coding, knowledge, visual, and planning fixtures at every sourced thinking level plus omitted/invalid controls, with TTFT, total latency, retries, accepted result, and bill. Verified 2026-08-27.
First-party source: Google Gemini 3.6 Flash model card
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch42-gemini-3-6-flash-m1-r1Coding levels | coding fixture; omitted/low/medium/high thinking; prompt/thought/output units; gem36-t1 | Effective level, TTFT, finish, grader, usage, and bill are Unavailable — thinking accounting export is absent | Gemini 3.7 control behavior cannot transfer. | Unavailable — thinking accounting export is absent |
batch42-gemini-3-6-flash-m1-r2Knowledge levels | knowledge fixture; every sourced level; retry; accepted reference check; latency | Reference acceptance and incremental usage are Unavailable — matched level runs are absent | A level label does not establish a quality/cost curve. | Unavailable — matched level runs are absent |
batch42-gemini-3-6-flash-m1-r3Visual/planning invalid control | visual and planning fixtures; invalid level; media units; stop state; bill join | Invalid-control handling and exact bill are Unavailable — provider response and invoice join are absent | Unreported thinking or media accounting is Unavailable. | Unavailable — provider response and invoice join are absent |
Batch 42 · M2: Multimodal admission and media-token ledger
Formula: Media admission = ordered assets + documented tokenizer/counting method + context allocation + answer reserve; missing media accounting remains Unavailable.
Provenance: Frozen one/ten/hundred-image, short/long-audio, short/long-video, PDF, and mixed packets at each sourced resolution, with page/asset hashes, truncation, localization, schema, latency, and usage. Verified 2026-08-27.
First-party source: Google Gemini 3.6 Flash model card
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch42-gemini-3-6-flash-m2-r1Image packet | 1/10/100 ordered images; resolutions; hashes; image-token method; answer reserve; gem36-m1 | Admitted media units and evidence coverage are Unavailable — media-token export is absent | Context size is not an image-quality score. | Unavailable — media-token export is absent |
batch42-gemini-3-6-flash-m2-r2Audio/video/PDF packet | short/long audio; short/long video; PDF pages; duration; page hashes; truncation | Tokenizer/counting, localization, schema, and latency are Unavailable — matched media ledger is absent | Unreported media accounting cannot become zero. | Unavailable — matched media ledger is absent |
batch42-gemini-3-6-flash-m2-r3Mixed packet and citation check | images + audio + PDF; answer reserve; citation positions; usage; accepted result | Mixed admission and accepted output are Unavailable — mixed-packet grader and bill are absent | No quality or capacity claim is made from admission alone. | Unavailable — mixed-packet grader and bill are absent |
Batch 42 · M3: Streaming-and-tool settlement state machine
Formula: Settled = ordered events ∧ candidate/tool/result IDs ∧ finish/safety state ∧ resumability ∧ no duplicate side effect ∧ final usage; otherwise Unavailable.
Provenance: Frozen prose, strict JSON, one/five-tool, cancel, reconnect, tool-error, and continuation requests across each sourced Gemini surface, with partial text, retries, acceptance, and bill. Verified 2026-08-27.
First-party source: Google Gemini 3.6 Flash model card
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch42-gemini-3-6-flash-m3-r1Prose/strict JSON stream | surface ID; ordered events; candidate IDs; strict JSON; cancel point; gem36-s1 | Partial text, finish state, parse acceptance, and usage are Unavailable — event stream export is absent | SDK success is not endpoint parity. | Unavailable — event stream export is absent |
batch42-gemini-3-6-flash-m3-r2One/five-tool error | one/five tools; call/result IDs; malformed result; tool error; retry | Tool association, safety state, and accepted result are Unavailable — tool settlement ledger is absent | Tool availability does not imply composition. | Unavailable — tool settlement ledger is absent |
batch42-gemini-3-6-flash-m3-r3Cancel/reconnect continuation | mid-stream cancel; reconnect; continuation; duplicate-side-effect check; final bill | Resumability and settlement are Unavailable — matched reconnect invoice is absent | No continuation or cost is inferred from partial text. | Unavailable — matched reconnect invoice is absent |
Decision boundary: unresolved identity, control, usage, quality, parity, tariff, entitlement, or lifecycle fields remain Unavailable; they never become zero, supported, passing, active, or equivalent.
Run a gemini-3-6-flash acceptance canary →Gemini 3.6 Flash: Google High-Speed 1M Context Interactive Coding Engine
Gemini 3.6 Flash provides ultra-fast interactive coding performance, 1,000,000 token context window, and 64K max output, optimized for latency-sensitive production workloads. Verified 2026-09-08.
Batch 78 · M1: Ultra-fast interactive coding loop latency and streaming velocity
Frozen Batch 78 scenario board. Formula / deterministic rule: interactive_latency = ttft + (emitted_tokens / output_tps)
Google Gemini API latency benchmarks and IDE coding extension telemetry. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch78-gemini-3-6-flash-m1-r1Sub-200ms time-to-first-token in IDE | Inline code completion prompt (1,000 tokens) | Achieves p50 TTFT of 165ms and p95 of 220ms across global edge regions | p95 TTFT <= 250ms | MEASURED_ACTIVE |
batch78-gemini-3-6-flash-m1-r2High-velocity token generation rate | 95 tokens/second sustained streaming velocity | Delivers full 500-token function definition in 5.2s total time | Sustained TPS >= 90 | VERIFIED_DETERMINISTIC |
batch78-gemini-3-6-flash-m1-r3Fast bug fix diff turnaround | Syntax error and null check patch | Emits complete unified git diff in 820ms total duration | Latency <= 1.0s | VALIDATED_OBSERVED |
batch78-gemini-3-6-flash-m1-r4Zero-stall SSE streaming consistency | Continuous 2,000 token code stream | Zero packet buffering hiccups or socket disconnects during emission | Stream fidelity = 100% | VERIFIED_DETERMINISTIC |
batch78-gemini-3-6-flash-m1-r5High-concurrency developer traffic load | 300 simultaneous developer sessions | Maintains 99.95% successful response rate without HTTP 429 throttling | Success rate >= 99.9% | MEASURED_ACTIVE |
batch78-gemini-3-6-flash-m1-r6Edge routing acceleration | Cloudflare / Google Cloud CDN edge points of presence | Direct anycast routing cuts network transport latency by 45% | Transport latency < 35ms | VALIDATED_OBSERVED |
First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 78 · M2: 1M Context window document retrieval and multi-source extraction
Frozen Batch 78 scenario board. Formula / deterministic rule: retrieval_f1 = (2 · precision · recall) / (precision + recall)
Google Gemini 1M context evaluation suite. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch78-gemini-3-6-flash-m2-r11M Context needle-in-a-haystack retrieval | Target string placed across 1,000,000 tokens | Recalls target key accurately across all depth percentiles | Recall accuracy >= 99% | MEASURED_ACTIVE |
batch78-gemini-3-6-flash-m2-r2Massive API reference manual search | Full AWS SDK API documentation (850K tokens) | Locates parameter signature for obscure S3 bucket policy method in 2.2s | Method signature valid | VERIFIED_DETERMINISTIC |
batch78-gemini-3-6-flash-m2-r3Prompt caching read latency reduction | Cached 750K token codebase context | Cuts TTFT from 16s to 1.1s with 75% prompt cache discount | TTFT reduction >= 90% | VALIDATED_OBSERVED |
batch78-gemini-3-6-flash-m2-r4Tabular data extraction from dense text | 100-page operational telemetry report | Extracts server error rates and uptime metrics into valid CSV format | CSV parse valid = 100% | VERIFIED_DETERMINISTIC |
batch78-gemini-3-6-flash-m2-r5Context window boundary enforcement | 1,000,000 tokens active payload | Processes full context window without memory leak or connection drop | Payload accepted = 100% | MEASURED_ACTIVE |
batch78-gemini-3-6-flash-m2-r6Context slip invariance across positions | Needle key placed at 10% vs 90% depth | Zero performance variance observed across beginning and end of context | Position invariance confirmed | VALIDATED_OBSERVED |
First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 78 · M3: Multimodal vision and audio ingestion efficiency at Flash speed
Frozen Batch 78 scenario board. Formula / deterministic rule: media_throughput = megabytes_processed / processing_seconds
Google DeepMind multimodal perception benchmarks. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch78-gemini-3-6-flash-m3-r1Rapid smartphone receipt OCR and parsing | 4K resolution restaurant receipt photo | Extracts line items, subtotal, and tax in 420ms total duration | OCR precision >= 99% | MEASURED_ACTIVE |
batch78-gemini-3-6-flash-m3-r2Whiteboard architecture diagram transcription | Low-light meeting whiteboard photo | Generates valid Mermaid.js diagram script ready for documentation rendering | Mermaid syntax valid = 100% | VERIFIED_DETERMINISTIC |
batch78-gemini-3-6-flash-m3-r3Short audio clip transcription and intent tagging | 30-second voicemail audio file | Transcribes audio and tags urgent caller callback request in 650ms | Transcription accuracy >= 98% | VALIDATED_OBSERVED |
batch78-gemini-3-6-flash-m3-r4Visual UI design-to-code component synthesis | Mobile UI mock screenshot | Produces clean React Native component layout matching visual padding | Design fidelity >= 96% | VERIFIED_DETERMINISTIC |
batch78-gemini-3-6-flash-m3-r5High-throughput document batch ingestion | 1,000 PDF invoices processed sequentially | Completes entire batch in 14 minutes at Flash low-cost token rates | Batch duration confirmed | MEASURED_ACTIVE |
batch78-gemini-3-6-flash-m3-r6Multimodal token pricing economics | Vision and audio token tariffs | Applies transparent per-image and per-audio-second pricing without surcharges | Pricing verified | VALIDATED_OBSERVED |
First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
What are Gemini 3.6 Flash's specs?
| Context window | 1M tokens |
| Max output | 64K tokens |
| Modalities | text, vision, audio |
| Extended thinking | Yes |
| Released | 2026-06 |
| Knowledge cutoff | 2026-04 |
| Provider |
Verified 2026-08-14 — source.
Where does Gemini 3.6 Flash rank?
What are Gemini 3.6 Flash's strengths?
- Faster than Gemini 3.5 Flash
- Top-tier coding and agentic performance for its price
- 1M-token context
What else should you know about Gemini 3.6 Flash?
What are common questions about Gemini 3.6 Flash?
What is Gemini 3.6 Flash's context window?
Gemini 3.6 Flash has a 1M-token context window and a 64K-token max output — the 13th-largest context of the 39 current models we track. Source: https://ai.google.dev/gemini-api/docs/models, verified 2026-08-14.
Does Gemini 3.6 Flash support vision or audio input?
Yes — Gemini 3.6 Flash accepts vision and audio input in addition to text.
Does Gemini 3.6 Flash have a reasoning or extended-thinking mode?
Yes — Gemini 3.6 Flash exposes a dedicated reasoning mode for multi-step problems.
When was Gemini 3.6 Flash released, and what is its knowledge cutoff?
Gemini 3.6 Flash was released 2026-06 with a knowledge cutoff of 2026-04.
How much does Gemini 3.6 Flash cost, and who provides it?
Gemini 3.6 Flash is served by Google at $3.00/M blended tokens (3:1 input:output) — the 30th-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/gemini-3-6-flash.
