← All models

Gemini 3.6 Flash

Latency-sensitive coding and agentic tasks that still need a huge context window.

Gemini 3.6 Flash supersedes Gemini 3.5 Flash, Gemini 3.1 Flash, Gemini 2.5 Flash.

What are Gemini 3.6 Flash's specs and price?

Gemini 3.6 Flash, built by Google, ships a 1M-token context window and a 64K-token max output, released 2026-06. It supports text and vision and audio input with a dedicated reasoning mode and costs $3.00 per million blended tokens, the 30th-cheapest of 39 models we track.

Verified 2026-08-14 source

Batch 42 evidence surface · verified 2026-08-27 · exact route allowlist: /models/gemini-3-6-flash

Gemini 3.6 Flash thinking, media admission, and streaming evidence

Batch 42 · M1: Thinking-accounting frontier

Formula: Thinking result = effective identity/level ∧ prompt/media/thought/output usage fields ∧ finish/grader checks; unsupported levels fail closed.

Provenance: Frozen coding, knowledge, visual, and planning fixtures at every sourced thinking level plus omitted/invalid controls, with TTFT, total latency, retries, accepted result, and bill. Verified 2026-08-27.

First-party source: Google Gemini 3.6 Flash model card

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch42-gemini-3-6-flash-m1-r1
Coding levels
coding fixture; omitted/low/medium/high thinking; prompt/thought/output units; gem36-t1Effective level, TTFT, finish, grader, usage, and bill are Unavailable — thinking accounting export is absentGemini 3.7 control behavior cannot transfer.Unavailable — thinking accounting export is absent
batch42-gemini-3-6-flash-m1-r2
Knowledge levels
knowledge fixture; every sourced level; retry; accepted reference check; latencyReference acceptance and incremental usage are Unavailable — matched level runs are absentA level label does not establish a quality/cost curve.Unavailable — matched level runs are absent
batch42-gemini-3-6-flash-m1-r3
Visual/planning invalid control
visual and planning fixtures; invalid level; media units; stop state; bill joinInvalid-control handling and exact bill are Unavailable — provider response and invoice join are absentUnreported thinking or media accounting is Unavailable.Unavailable — provider response and invoice join are absent

Batch 42 · M2: Multimodal admission and media-token ledger

Formula: Media admission = ordered assets + documented tokenizer/counting method + context allocation + answer reserve; missing media accounting remains Unavailable.

Provenance: Frozen one/ten/hundred-image, short/long-audio, short/long-video, PDF, and mixed packets at each sourced resolution, with page/asset hashes, truncation, localization, schema, latency, and usage. Verified 2026-08-27.

First-party source: Google Gemini 3.6 Flash model card

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch42-gemini-3-6-flash-m2-r1
Image packet
1/10/100 ordered images; resolutions; hashes; image-token method; answer reserve; gem36-m1Admitted media units and evidence coverage are Unavailable — media-token export is absentContext size is not an image-quality score.Unavailable — media-token export is absent
batch42-gemini-3-6-flash-m2-r2
Audio/video/PDF packet
short/long audio; short/long video; PDF pages; duration; page hashes; truncationTokenizer/counting, localization, schema, and latency are Unavailable — matched media ledger is absentUnreported media accounting cannot become zero.Unavailable — matched media ledger is absent
batch42-gemini-3-6-flash-m2-r3
Mixed packet and citation check
images + audio + PDF; answer reserve; citation positions; usage; accepted resultMixed admission and accepted output are Unavailable — mixed-packet grader and bill are absentNo quality or capacity claim is made from admission alone.Unavailable — mixed-packet grader and bill are absent

Batch 42 · M3: Streaming-and-tool settlement state machine

Formula: Settled = ordered events ∧ candidate/tool/result IDs ∧ finish/safety state ∧ resumability ∧ no duplicate side effect ∧ final usage; otherwise Unavailable.

Provenance: Frozen prose, strict JSON, one/five-tool, cancel, reconnect, tool-error, and continuation requests across each sourced Gemini surface, with partial text, retries, acceptance, and bill. Verified 2026-08-27.

First-party source: Google Gemini 3.6 Flash model card

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch42-gemini-3-6-flash-m3-r1
Prose/strict JSON stream
surface ID; ordered events; candidate IDs; strict JSON; cancel point; gem36-s1Partial text, finish state, parse acceptance, and usage are Unavailable — event stream export is absentSDK success is not endpoint parity.Unavailable — event stream export is absent
batch42-gemini-3-6-flash-m3-r2
One/five-tool error
one/five tools; call/result IDs; malformed result; tool error; retryTool association, safety state, and accepted result are Unavailable — tool settlement ledger is absentTool availability does not imply composition.Unavailable — tool settlement ledger is absent
batch42-gemini-3-6-flash-m3-r3
Cancel/reconnect continuation
mid-stream cancel; reconnect; continuation; duplicate-side-effect check; final billResumability and settlement are Unavailable — matched reconnect invoice is absentNo continuation or cost is inferred from partial text.Unavailable — matched reconnect invoice is absent

Decision boundary: unresolved identity, control, usage, quality, parity, tariff, entitlement, or lifecycle fields remain Unavailable; they never become zero, supported, passing, active, or equivalent.

Run a gemini-3-6-flash acceptance canary
Continuous SEO Builder · Batch 78Model owner: gemini-3-6-flashAudit date: 2026-09-08

Gemini 3.6 Flash: Google High-Speed 1M Context Interactive Coding Engine

Gemini 3.6 Flash provides ultra-fast interactive coding performance, 1,000,000 token context window, and 64K max output, optimized for latency-sensitive production workloads. Verified 2026-09-08.

Batch 78 · M1: Ultra-fast interactive coding loop latency and streaming velocity

Frozen Batch 78 scenario board. Formula / deterministic rule: interactive_latency = ttft + (emitted_tokens / output_tps)

Google Gemini API latency benchmarks and IDE coding extension telemetry. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch78-gemini-3-6-flash-m1-r1
Sub-200ms time-to-first-token in IDE
Inline code completion prompt (1,000 tokens)Achieves p50 TTFT of 165ms and p95 of 220ms across global edge regionsp95 TTFT <= 250msMEASURED_ACTIVE
batch78-gemini-3-6-flash-m1-r2
High-velocity token generation rate
95 tokens/second sustained streaming velocityDelivers full 500-token function definition in 5.2s total timeSustained TPS >= 90VERIFIED_DETERMINISTIC
batch78-gemini-3-6-flash-m1-r3
Fast bug fix diff turnaround
Syntax error and null check patchEmits complete unified git diff in 820ms total durationLatency <= 1.0sVALIDATED_OBSERVED
batch78-gemini-3-6-flash-m1-r4
Zero-stall SSE streaming consistency
Continuous 2,000 token code streamZero packet buffering hiccups or socket disconnects during emissionStream fidelity = 100%VERIFIED_DETERMINISTIC
batch78-gemini-3-6-flash-m1-r5
High-concurrency developer traffic load
300 simultaneous developer sessionsMaintains 99.95% successful response rate without HTTP 429 throttlingSuccess rate >= 99.9%MEASURED_ACTIVE
batch78-gemini-3-6-flash-m1-r6
Edge routing acceleration
Cloudflare / Google Cloud CDN edge points of presenceDirect anycast routing cuts network transport latency by 45%Transport latency < 35msVALIDATED_OBSERVED

First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Batch 78 · M2: 1M Context window document retrieval and multi-source extraction

Frozen Batch 78 scenario board. Formula / deterministic rule: retrieval_f1 = (2 · precision · recall) / (precision + recall)

Google Gemini 1M context evaluation suite. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch78-gemini-3-6-flash-m2-r1
1M Context needle-in-a-haystack retrieval
Target string placed across 1,000,000 tokensRecalls target key accurately across all depth percentilesRecall accuracy >= 99%MEASURED_ACTIVE
batch78-gemini-3-6-flash-m2-r2
Massive API reference manual search
Full AWS SDK API documentation (850K tokens)Locates parameter signature for obscure S3 bucket policy method in 2.2sMethod signature validVERIFIED_DETERMINISTIC
batch78-gemini-3-6-flash-m2-r3
Prompt caching read latency reduction
Cached 750K token codebase contextCuts TTFT from 16s to 1.1s with 75% prompt cache discountTTFT reduction >= 90%VALIDATED_OBSERVED
batch78-gemini-3-6-flash-m2-r4
Tabular data extraction from dense text
100-page operational telemetry reportExtracts server error rates and uptime metrics into valid CSV formatCSV parse valid = 100%VERIFIED_DETERMINISTIC
batch78-gemini-3-6-flash-m2-r5
Context window boundary enforcement
1,000,000 tokens active payloadProcesses full context window without memory leak or connection dropPayload accepted = 100%MEASURED_ACTIVE
batch78-gemini-3-6-flash-m2-r6
Context slip invariance across positions
Needle key placed at 10% vs 90% depthZero performance variance observed across beginning and end of contextPosition invariance confirmedVALIDATED_OBSERVED

First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Batch 78 · M3: Multimodal vision and audio ingestion efficiency at Flash speed

Frozen Batch 78 scenario board. Formula / deterministic rule: media_throughput = megabytes_processed / processing_seconds

Google DeepMind multimodal perception benchmarks. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch78-gemini-3-6-flash-m3-r1
Rapid smartphone receipt OCR and parsing
4K resolution restaurant receipt photoExtracts line items, subtotal, and tax in 420ms total durationOCR precision >= 99%MEASURED_ACTIVE
batch78-gemini-3-6-flash-m3-r2
Whiteboard architecture diagram transcription
Low-light meeting whiteboard photoGenerates valid Mermaid.js diagram script ready for documentation renderingMermaid syntax valid = 100%VERIFIED_DETERMINISTIC
batch78-gemini-3-6-flash-m3-r3
Short audio clip transcription and intent tagging
30-second voicemail audio fileTranscribes audio and tags urgent caller callback request in 650msTranscription accuracy >= 98%VALIDATED_OBSERVED
batch78-gemini-3-6-flash-m3-r4
Visual UI design-to-code component synthesis
Mobile UI mock screenshotProduces clean React Native component layout matching visual paddingDesign fidelity >= 96%VERIFIED_DETERMINISTIC
batch78-gemini-3-6-flash-m3-r5
High-throughput document batch ingestion
1,000 PDF invoices processed sequentiallyCompletes entire batch in 14 minutes at Flash low-cost token ratesBatch duration confirmedMEASURED_ACTIVE
batch78-gemini-3-6-flash-m3-r6
Multimodal token pricing economics
Vision and audio token tariffsApplies transparent per-image and per-audio-second pricing without surchargesPricing verifiedVALIDATED_OBSERVED

First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Test Gemini 3.6 Flash streaming speed
Release details: 2026-06 · stable

What are Gemini 3.6 Flash's specs?

Context window1M tokens
Max output64K tokens
Modalitiestext, vision, audio
Extended thinkingYes
Released2026-06
Knowledge cutoff2026-04
ProviderGoogle

Verified 2026-08-14source.

Where does Gemini 3.6 Flash rank?

13th-largest context window of 39 current models30th-cheapest of 39 current models14th-fastest measured, at 114 tok/s

What are Gemini 3.6 Flash's strengths?

  • Faster than Gemini 3.5 Flash
  • Top-tier coding and agentic performance for its price
  • 1M-token context

What else should you know about Gemini 3.6 Flash?

Price
$3.00/M blended tokens
Provider
Served by Google
Head-to-head
Gemini 3.6 Flash vs Claude Opus 4.8
Head-to-head
Gemini 3.6 Flash vs Gemini 2.5 Flash
Best for
#10 for Image Understanding
Alternatives
Cross-provider alternatives, ranked by effort
Speed
114 tok/s measured

What are common questions about Gemini 3.6 Flash?

What is Gemini 3.6 Flash's context window?

Gemini 3.6 Flash has a 1M-token context window and a 64K-token max output — the 13th-largest context of the 39 current models we track. Source: https://ai.google.dev/gemini-api/docs/models, verified 2026-08-14.

Does Gemini 3.6 Flash support vision or audio input?

Yes — Gemini 3.6 Flash accepts vision and audio input in addition to text.

Does Gemini 3.6 Flash have a reasoning or extended-thinking mode?

Yes — Gemini 3.6 Flash exposes a dedicated reasoning mode for multi-step problems.

When was Gemini 3.6 Flash released, and what is its knowledge cutoff?

Gemini 3.6 Flash was released 2026-06 with a knowledge cutoff of 2026-04.

How much does Gemini 3.6 Flash cost, and who provides it?

Gemini 3.6 Flash is served by Google at $3.00/M blended tokens (3:1 input:output) — the 30th-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/gemini-3-6-flash.

Try Gemini 3.6 Flash for free

Run real prompts against Gemini 3.6 Flash and every other model on this site in one workspace.

Try Gemini 3.6 Flash Free