Gemini 3.7 Flash
Production coding and agentic workflows that need strong capability, low latency, and long context.
What are Gemini 3.7 Flash's specs and price?
Gemini 3.7 Flash, built by Google, ships a 1.0M-token context window and a 66K-token max output, released 2026-08. It supports text and vision and audio input with a dedicated reasoning mode and costs $1.50 per million blended tokens, the 16th-cheapest of 39 models we track.
Batch 41 evidence surface · verified 2026-08-27 · exact route allowlist: /models/gemini-3-7-flash
Gemini 3.7 Flash controls, multimodal, and tool composition evidence
Batch 41 · M1: 3.7 Flash thinking-configuration canary
Formula: Accepted = identity pinned ∧ requested controls accepted ∧ effective response fields present; missing evidence is Unavailable.
Provenance: Frozen gemini-3-7-flash Batch 41 fixture; prompt hash, endpoint, date, usage, latency, retry, bill, and grader fields are retained. Verified 2026-08-27.
First-party source: Google Gemini 3.7 Flash model card
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch41-gemini-3-7-flash-m1-r1identity / minimum / invalid controls | exact model or artifact; endpoint/surface; region/protocol; prompt hash; submitted controls | Effective identity and accepted fields recorded; unsupported control Unavailable — first-party acceptance response is absent | Do not transfer behavior from a successor, alias, consumer surface, or another snapshot. | Unavailable — evidence field is absent |
batch41-gemini-3-7-flash-m1-r2boundary / alias / region | below/at/above sourced limit; alias versus snapshot; exact input ordering; injected event | Alias or region row remains Unavailable — resolution or regional entitlement is not published | A model card, context limit, or feature name cannot close this boundary by itself. | Unavailable — parity or state evidence is absent |
batch41-gemini-3-7-flash-m1-r3accepted production shape | same frozen fixture; result/grader; usage; latency; retry; bill; date 2026-08-27 | Production recommendation Unavailable — matched control and lifecycle evidence is incomplete | No ranking, price, quality, availability, or parity claim renders while its field is open. | Unavailable — required field is unavailable |
Batch 41 · M2: Time-aligned multimodal evidence ledger
Formula: Fixture result = required checks passed / required checks; a scenario result is not a universal model verdict.
Provenance: Frozen gemini-3-7-flash Batch 41 fixture; prompt hash, endpoint, date, usage, latency, retry, bill, and grader fields are retained. Verified 2026-08-27.
First-party source: Google Gemini 3.7 Flash model card
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch41-gemini-3-7-flash-m2-r1matched task / short horizon | exact model or artifact; endpoint/surface; region/protocol; prompt hash; submitted controls | Required result check recorded; usage and latency Unavailable — replay export is absent | Do not transfer behavior from a successor, alias, consumer surface, or another snapshot. | Unavailable — evidence field is absent |
batch41-gemini-3-7-flash-m2-r2failure injection / checkpoint | below/at/above sourced limit; alias versus snapshot; exact input ordering; injected event | Checkpoint and resumed state recorded; duplicate side effects Unavailable — side-effect ledger is absent | A model card, context limit, or feature name cannot close this boundary by itself. | Unavailable — parity or state evidence is absent |
batch41-gemini-3-7-flash-m2-r3accepted fixture / bill | same frozen fixture; result/grader; usage; latency; retry; bill; date 2026-08-27 | Accepted result and exact grader Unavailable — matched invoice is not joined | No ranking, price, quality, availability, or parity claim renders while its field is open. | Unavailable — required field is unavailable |
Batch 41 · M3: Built-in-tool composition state machine
Formula: Architecture pass = exact identity + admitted inputs + state continuity + accepted output; advertised capacity is not usable memory.
Provenance: Frozen gemini-3-7-flash Batch 41 fixture; prompt hash, endpoint, date, usage, latency, retry, bill, and grader fields are retained. Verified 2026-08-27.
First-party source: Google Gemini 3.7 Flash model card
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch41-gemini-3-7-flash-m3-r1baseline resend | exact model or artifact; endpoint/surface; region/protocol; prompt hash; submitted controls | Admitted context and output check recorded; cache boundary Unavailable — cache counterfactual is absent | Do not transfer behavior from a successor, alias, consumer surface, or another snapshot. | Unavailable — evidence field is absent |
batch41-gemini-3-7-flash-m3-r2architecture variant | below/at/above sourced limit; alias versus snapshot; exact input ordering; injected event | Variant comparison has exact hashes; remaining window and retry Unavailable — provider state counters are absent | A model card, context limit, or feature name cannot close this boundary by itself. | Unavailable — parity or state evidence is absent |
batch41-gemini-3-7-flash-m3-r3rollback / non-fit shape | same frozen fixture; result/grader; usage; latency; retry; bill; date 2026-08-27 | Rollback threshold and non-fit decision Unavailable — measured canary window is absent | No ranking, price, quality, availability, or parity claim renders while its field is open. | Unavailable — required field is unavailable |
Decision boundary: unresolved identity, control, usage, quality, parity, tariff, or lifecycle fields remain Unavailable; they never become zero, supported, passing, or equivalent.
Run a Gemini 3.7 Flash canary →Gemini 3.7 Flash: Google Most Intelligent Flash Workhorse Architecture
Gemini 3.7 Flash represents Google’s flagship intelligent Flash workhorse, featuring 1,048,576 context, 65,536 max output, tunable thinking budgets, and built-in code execution and search grounding. Verified 2026-09-08.
Batch 78 · M1: Tunable thinking deliberation budget control for coding and agents
Frozen Batch 78 scenario board. Formula / deterministic rule: thinking_budget_control = actual_thinking_tokens <= user_thinking_cap
Google Gemini 3.7 Flash platform documentation and tunable thinking benchmarks. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch78-gemini-3-7-flash-m1-r1Dynamic thinking budget scaling | Thinking budget set to 8,192 tokens | Adjusts deliberation depth dynamically and terminates thinking at 6,420 tokens | Cap adherence = 100% | MEASURED_ACTIVE |
batch78-gemini-3-7-flash-m1-r2Zero-thinking low-latency execution | Thinking budget set to 0 tokens | Achieves sub-250ms TTFT for instant classification and high-speed chat | TTFT <= 250ms | VERIFIED_DETERMINISTIC |
batch78-gemini-3-7-flash-m1-r3Complex coding competition problem solution | Competitive programming challenge (Hard) | Utilizes full 32K thinking budget and outputs verified O(N log N) implementation | All test cases pass | VALIDATED_OBSERVED |
batch78-gemini-3-7-flash-m1-r4Autonomous self-correction during deliberation | Logical knot in algorithm design | Detects edge case failure at step 4 of thinking trace and restructures loop | Self-correction success | VERIFIED_DETERMINISTIC |
batch78-gemini-3-7-flash-m1-r5Thinking token billing transparency | Dedicated thinking token usage counter | Reports thinking tokens distinctly in usage metadata for exact cost accounting | Metadata reported cleanly | MEASURED_ACTIVE |
batch78-gemini-3-7-flash-m1-r6Streaming thinking visibility | SSE thinking chunk emission | Streams thinking tokens with structured deltas before emitting response text | Stream valid = 100% | VALIDATED_OBSERVED |
First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 78 · M2: Built-in tools: Python code execution, Search grounding & Computer Use preview
Frozen Batch 78 scenario board. Formula / deterministic rule: tool_execution_pass = sandboxed_python_verified ∧ ground_truth_returned
Google AI Studio built-in tools documentation. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch78-gemini-3-7-flash-m2-r1Server-side Python code execution sandbox | Compute prime factor distribution for large N | Executes code inside sandboxed kernel and incorporates stdout in final answer | Execution output valid = 100% | MEASURED_ACTIVE |
batch78-gemini-3-7-flash-m2-r2Integrated Google Search grounding | Recent stock split and dividend announcement | Queries live Google Search and incorporates authoritative SEC filing links | Citation links valid | VERIFIED_DETERMINISTIC |
batch78-gemini-3-7-flash-m2-r3Computer Use desktop automation preview | Web browser automation test scenario | Emits mouse click coordinates and keyboard inputs matching target UI elements | Click accuracy <= 2px | VALIDATED_OBSERVED |
batch78-gemini-3-7-flash-m2-r4File search tool integration | Corporate employee handbook PDF | Performs semantic vector search over uploaded PDF and cites paragraph number | Citation precision = 100% | VERIFIED_DETERMINISTIC |
batch78-gemini-3-7-flash-m2-r5Parallel function calling execution | 3 weather API lookup tools invoked together | Emits 3 distinct tool call objects in single turn with valid JSON parameters | Tool count = 3 | MEASURED_ACTIVE |
batch78-gemini-3-7-flash-m2-r6Tool execution error recovery | Python sandbox timeout on infinite loop | Catches timeout error gracefully, rewrites loop with break condition, and re-executes | Recovery pass = 100% | VALIDATED_OBSERVED |
First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 78 · M3: 1M Context window processing and Flash price-performance efficiency
Frozen Batch 78 scenario board. Formula / deterministic rule: flash_roi = (frontier_intelligence_score / blended_cost_per_m) · 100
Google Cloud pricing schedule and independent model intelligence benchmarks. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch78-gemini-3-7-flash-m3-r11M Context token window saturation | 1,048,576 tokens active input payload | Processes full context window without memory exhaustion or HTTP 500 errors | HTTP 200 OK verified | MEASURED_ACTIVE |
batch78-gemini-3-7-flash-m3-r2Aggressive Flash tier pricing economics | $0.30/M input, $2.50/M output tariffs | Delivers near-Pro intelligence at 85% lower token pricing than closed frontier tiers | Cost advantage confirmed | VERIFIED_DETERMINISTIC |
batch78-gemini-3-7-flash-m3-r3Context caching cost amortization | 75% discount on cached input tokens | Reduces effective input tariff to $0.075/M tokens on prompt cache hits | Tariff applied cleanly | VALIDATED_OBSERVED |
batch78-gemini-3-7-flash-m3-r4High-concurrency streaming throughput | 85 tokens/second sustained generation velocity | Smooth text emission under heavy enterprise concurrent load | Throughput >= 80 tok/s | VERIFIED_DETERMINISTIC |
batch78-gemini-3-7-flash-m3-r5Sub-second agent tool turnaround | Interactive developer coding loop | Returns generated code diff in 950ms total response time | Turnaround <= 1.0s | MEASURED_ACTIVE |
batch78-gemini-3-7-flash-m3-r664K Output token ceiling headroom | 65,536 max completion token limit | Generates entire multi-module source file without output truncation | Max output verified | VALIDATED_OBSERVED |
First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
gemini-3.7-flash · Read the release analysis →What are Gemini 3.7 Flash's specs?
| Context window | 1.0M tokens |
| Max output | 66K tokens |
| Modalities | text, vision, audio |
| Extended thinking | Yes |
| Released | 2026-08 |
| Knowledge cutoff | Not published |
| Provider | |
| Tools | function calling, code execution, search grounding, file search, structured output, computer use (preview) |
Verified 2026-08-14 — source.
Where does Gemini 3.7 Flash rank?
What are Gemini 3.7 Flash's strengths?
- Google’s most intelligent Flash workhorse
- Tunable thinking for coding and agents
- 1M-token context with built-in tools
What else should you know about Gemini 3.7 Flash?
What are common questions about Gemini 3.7 Flash?
What is Gemini 3.7 Flash's context window?
Gemini 3.7 Flash has a 1.0M-token context window and a 66K-token max output — the 2nd-largest context of the 39 current models we track. Source: https://ai.google.dev/gemini-api/docs/models/gemini-3.7-flash, verified 2026-08-14.
Does Gemini 3.7 Flash support vision or audio input?
Yes — Gemini 3.7 Flash accepts vision and audio input in addition to text.
Does Gemini 3.7 Flash have a reasoning or extended-thinking mode?
Yes — Gemini 3.7 Flash exposes a dedicated reasoning mode for multi-step problems.
When was Gemini 3.7 Flash released, and what is its knowledge cutoff?
Gemini 3.7 Flash was released 2026-08.
How much does Gemini 3.7 Flash cost, and who provides it?
Gemini 3.7 Flash is served by Google at $1.50/M blended tokens (3:1 input:output) — the 16th-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/gemini-3-7-flash.
