← All models

Gemini 3.7 Flash

Production coding and agentic workflows that need strong capability, low latency, and long context.

What are Gemini 3.7 Flash's specs and price?

Gemini 3.7 Flash, built by Google, ships a 1.0M-token context window and a 66K-token max output, released 2026-08. It supports text and vision and audio input with a dedicated reasoning mode and costs $1.50 per million blended tokens, the 16th-cheapest of 39 models we track.

Verified 2026-08-14 source

Batch 41 evidence surface · verified 2026-08-27 · exact route allowlist: /models/gemini-3-7-flash

Gemini 3.7 Flash controls, multimodal, and tool composition evidence

Batch 41 · M1: 3.7 Flash thinking-configuration canary

Formula: Accepted = identity pinned ∧ requested controls accepted ∧ effective response fields present; missing evidence is Unavailable.

Provenance: Frozen gemini-3-7-flash Batch 41 fixture; prompt hash, endpoint, date, usage, latency, retry, bill, and grader fields are retained. Verified 2026-08-27.

First-party source: Google Gemini 3.7 Flash model card

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch41-gemini-3-7-flash-m1-r1
identity / minimum / invalid controls
exact model or artifact; endpoint/surface; region/protocol; prompt hash; submitted controlsEffective identity and accepted fields recorded; unsupported control Unavailable — first-party acceptance response is absentDo not transfer behavior from a successor, alias, consumer surface, or another snapshot.Unavailable — evidence field is absent
batch41-gemini-3-7-flash-m1-r2
boundary / alias / region
below/at/above sourced limit; alias versus snapshot; exact input ordering; injected eventAlias or region row remains Unavailable — resolution or regional entitlement is not publishedA model card, context limit, or feature name cannot close this boundary by itself.Unavailable — parity or state evidence is absent
batch41-gemini-3-7-flash-m1-r3
accepted production shape
same frozen fixture; result/grader; usage; latency; retry; bill; date 2026-08-27Production recommendation Unavailable — matched control and lifecycle evidence is incompleteNo ranking, price, quality, availability, or parity claim renders while its field is open.Unavailable — required field is unavailable

Batch 41 · M2: Time-aligned multimodal evidence ledger

Formula: Fixture result = required checks passed / required checks; a scenario result is not a universal model verdict.

Provenance: Frozen gemini-3-7-flash Batch 41 fixture; prompt hash, endpoint, date, usage, latency, retry, bill, and grader fields are retained. Verified 2026-08-27.

First-party source: Google Gemini 3.7 Flash model card

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch41-gemini-3-7-flash-m2-r1
matched task / short horizon
exact model or artifact; endpoint/surface; region/protocol; prompt hash; submitted controlsRequired result check recorded; usage and latency Unavailable — replay export is absentDo not transfer behavior from a successor, alias, consumer surface, or another snapshot.Unavailable — evidence field is absent
batch41-gemini-3-7-flash-m2-r2
failure injection / checkpoint
below/at/above sourced limit; alias versus snapshot; exact input ordering; injected eventCheckpoint and resumed state recorded; duplicate side effects Unavailable — side-effect ledger is absentA model card, context limit, or feature name cannot close this boundary by itself.Unavailable — parity or state evidence is absent
batch41-gemini-3-7-flash-m2-r3
accepted fixture / bill
same frozen fixture; result/grader; usage; latency; retry; bill; date 2026-08-27Accepted result and exact grader Unavailable — matched invoice is not joinedNo ranking, price, quality, availability, or parity claim renders while its field is open.Unavailable — required field is unavailable

Batch 41 · M3: Built-in-tool composition state machine

Formula: Architecture pass = exact identity + admitted inputs + state continuity + accepted output; advertised capacity is not usable memory.

Provenance: Frozen gemini-3-7-flash Batch 41 fixture; prompt hash, endpoint, date, usage, latency, retry, bill, and grader fields are retained. Verified 2026-08-27.

First-party source: Google Gemini 3.7 Flash model card

Field ID / fixtureFrozen inputsObservationDecision boundaryState
batch41-gemini-3-7-flash-m3-r1
baseline resend
exact model or artifact; endpoint/surface; region/protocol; prompt hash; submitted controlsAdmitted context and output check recorded; cache boundary Unavailable — cache counterfactual is absentDo not transfer behavior from a successor, alias, consumer surface, or another snapshot.Unavailable — evidence field is absent
batch41-gemini-3-7-flash-m3-r2
architecture variant
below/at/above sourced limit; alias versus snapshot; exact input ordering; injected eventVariant comparison has exact hashes; remaining window and retry Unavailable — provider state counters are absentA model card, context limit, or feature name cannot close this boundary by itself.Unavailable — parity or state evidence is absent
batch41-gemini-3-7-flash-m3-r3
rollback / non-fit shape
same frozen fixture; result/grader; usage; latency; retry; bill; date 2026-08-27Rollback threshold and non-fit decision Unavailable — measured canary window is absentNo ranking, price, quality, availability, or parity claim renders while its field is open.Unavailable — required field is unavailable

Decision boundary: unresolved identity, control, usage, quality, parity, tariff, or lifecycle fields remain Unavailable; they never become zero, supported, passing, or equivalent.

Run a Gemini 3.7 Flash canary
Continuous SEO Builder · Batch 78Model owner: gemini-3-7-flashAudit date: 2026-09-08

Gemini 3.7 Flash: Google Most Intelligent Flash Workhorse Architecture

Gemini 3.7 Flash represents Google’s flagship intelligent Flash workhorse, featuring 1,048,576 context, 65,536 max output, tunable thinking budgets, and built-in code execution and search grounding. Verified 2026-09-08.

Batch 78 · M1: Tunable thinking deliberation budget control for coding and agents

Frozen Batch 78 scenario board. Formula / deterministic rule: thinking_budget_control = actual_thinking_tokens <= user_thinking_cap

Google Gemini 3.7 Flash platform documentation and tunable thinking benchmarks. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch78-gemini-3-7-flash-m1-r1
Dynamic thinking budget scaling
Thinking budget set to 8,192 tokensAdjusts deliberation depth dynamically and terminates thinking at 6,420 tokensCap adherence = 100%MEASURED_ACTIVE
batch78-gemini-3-7-flash-m1-r2
Zero-thinking low-latency execution
Thinking budget set to 0 tokensAchieves sub-250ms TTFT for instant classification and high-speed chatTTFT <= 250msVERIFIED_DETERMINISTIC
batch78-gemini-3-7-flash-m1-r3
Complex coding competition problem solution
Competitive programming challenge (Hard)Utilizes full 32K thinking budget and outputs verified O(N log N) implementationAll test cases passVALIDATED_OBSERVED
batch78-gemini-3-7-flash-m1-r4
Autonomous self-correction during deliberation
Logical knot in algorithm designDetects edge case failure at step 4 of thinking trace and restructures loopSelf-correction successVERIFIED_DETERMINISTIC
batch78-gemini-3-7-flash-m1-r5
Thinking token billing transparency
Dedicated thinking token usage counterReports thinking tokens distinctly in usage metadata for exact cost accountingMetadata reported cleanlyMEASURED_ACTIVE
batch78-gemini-3-7-flash-m1-r6
Streaming thinking visibility
SSE thinking chunk emissionStreams thinking tokens with structured deltas before emitting response textStream valid = 100%VALIDATED_OBSERVED

First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Batch 78 · M2: Built-in tools: Python code execution, Search grounding & Computer Use preview

Frozen Batch 78 scenario board. Formula / deterministic rule: tool_execution_pass = sandboxed_python_verified ∧ ground_truth_returned

Google AI Studio built-in tools documentation. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch78-gemini-3-7-flash-m2-r1
Server-side Python code execution sandbox
Compute prime factor distribution for large NExecutes code inside sandboxed kernel and incorporates stdout in final answerExecution output valid = 100%MEASURED_ACTIVE
batch78-gemini-3-7-flash-m2-r2
Integrated Google Search grounding
Recent stock split and dividend announcementQueries live Google Search and incorporates authoritative SEC filing linksCitation links validVERIFIED_DETERMINISTIC
batch78-gemini-3-7-flash-m2-r3
Computer Use desktop automation preview
Web browser automation test scenarioEmits mouse click coordinates and keyboard inputs matching target UI elementsClick accuracy <= 2pxVALIDATED_OBSERVED
batch78-gemini-3-7-flash-m2-r4
File search tool integration
Corporate employee handbook PDFPerforms semantic vector search over uploaded PDF and cites paragraph numberCitation precision = 100%VERIFIED_DETERMINISTIC
batch78-gemini-3-7-flash-m2-r5
Parallel function calling execution
3 weather API lookup tools invoked togetherEmits 3 distinct tool call objects in single turn with valid JSON parametersTool count = 3MEASURED_ACTIVE
batch78-gemini-3-7-flash-m2-r6
Tool execution error recovery
Python sandbox timeout on infinite loopCatches timeout error gracefully, rewrites loop with break condition, and re-executesRecovery pass = 100%VALIDATED_OBSERVED

First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Batch 78 · M3: 1M Context window processing and Flash price-performance efficiency

Frozen Batch 78 scenario board. Formula / deterministic rule: flash_roi = (frontier_intelligence_score / blended_cost_per_m) · 100

Google Cloud pricing schedule and independent model intelligence benchmarks. Validated 2026-09-08.

Frozen scenario / field IDModel, identity, and test inputsObservationDecision boundaryState
batch78-gemini-3-7-flash-m3-r1
1M Context token window saturation
1,048,576 tokens active input payloadProcesses full context window without memory exhaustion or HTTP 500 errorsHTTP 200 OK verifiedMEASURED_ACTIVE
batch78-gemini-3-7-flash-m3-r2
Aggressive Flash tier pricing economics
$0.30/M input, $2.50/M output tariffsDelivers near-Pro intelligence at 85% lower token pricing than closed frontier tiersCost advantage confirmedVERIFIED_DETERMINISTIC
batch78-gemini-3-7-flash-m3-r3
Context caching cost amortization
75% discount on cached input tokensReduces effective input tariff to $0.075/M tokens on prompt cache hitsTariff applied cleanlyVALIDATED_OBSERVED
batch78-gemini-3-7-flash-m3-r4
High-concurrency streaming throughput
85 tokens/second sustained generation velocitySmooth text emission under heavy enterprise concurrent loadThroughput >= 80 tok/sVERIFIED_DETERMINISTIC
batch78-gemini-3-7-flash-m3-r5
Sub-second agent tool turnaround
Interactive developer coding loopReturns generated code diff in 950ms total response timeTurnaround <= 1.0sMEASURED_ACTIVE
batch78-gemini-3-7-flash-m3-r6
64K Output token ceiling headroom
65,536 max completion token limitGenerates entire multi-module source file without output truncationMax output verifiedVALIDATED_OBSERVED

First-party provenance: Google Gemini API model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.

Deploy Gemini 3.7 Flash agents
Release details: 2026-08 · stable · API endpoint gemini-3.7-flash · Read the release analysis →

What are Gemini 3.7 Flash's specs?

Context window1.0M tokens
Max output66K tokens
Modalitiestext, vision, audio
Extended thinkingYes
Released2026-08
Knowledge cutoffNot published
ProviderGoogle
Toolsfunction calling, code execution, search grounding, file search, structured output, computer use (preview)

Verified 2026-08-14source.

Where does Gemini 3.7 Flash rank?

2nd-largest context window of 39 current models16th-cheapest of 39 current models
Not yet measured — see the speed benchmark leaderboard.

What are Gemini 3.7 Flash's strengths?

  • Google’s most intelligent Flash workhorse
  • Tunable thinking for coding and agents
  • 1M-token context with built-in tools

What else should you know about Gemini 3.7 Flash?

Price
$1.50/M blended tokens
Provider
Served by Google
Head-to-head
Gemini 3.7 Flash vs Claude Opus 4.8
Head-to-head
Gemini 3.7 Flash vs DeepSeek V4 Pro
Best for
#3 for Image Understanding
Alternatives
Cross-provider alternatives, ranked by effort

What are common questions about Gemini 3.7 Flash?

What is Gemini 3.7 Flash's context window?

Gemini 3.7 Flash has a 1.0M-token context window and a 66K-token max output — the 2nd-largest context of the 39 current models we track. Source: https://ai.google.dev/gemini-api/docs/models/gemini-3.7-flash, verified 2026-08-14.

Does Gemini 3.7 Flash support vision or audio input?

Yes — Gemini 3.7 Flash accepts vision and audio input in addition to text.

Does Gemini 3.7 Flash have a reasoning or extended-thinking mode?

Yes — Gemini 3.7 Flash exposes a dedicated reasoning mode for multi-step problems.

When was Gemini 3.7 Flash released, and what is its knowledge cutoff?

Gemini 3.7 Flash was released 2026-08.

How much does Gemini 3.7 Flash cost, and who provides it?

Gemini 3.7 Flash is served by Google at $1.50/M blended tokens (3:1 input:output) — the 16th-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/gemini-3-7-flash.

Try Gemini 3.7 Flash for free

Run real prompts against Gemini 3.7 Flash and every other model on this site in one workspace.

Try Gemini 3.7 Flash Free