← All providers

Google API Pricing, Models & Rate Limits (2026)

Google serves the Gemini family through both a direct Gemini API and Vertex AI. Gemini 3.7 Flash is its newest coding-and-agent workhorse, while Gemini 3.1 Pro provides the lineup's 2M-token context option; current Gemini models read native audio and video, not just text and images.

Also known as: Gemini, Google AI Studio, Vertex AI.

How much does the Google API cost?

Google Gemini API pricing spans a long-context Pro tier and lower-cost Flash variants, with Google AI Studio and Vertex AI providing different operational entry points. Gemini 3.1 Pro is the flagship row here, while Flash Lite is aimed at high-volume workloads. Google AI Studio’s developer access and Vertex AI’s cloud-account controls are not the same billing path, so choose the deployment surface before projecting API cost.

Verified 2026-08-14 source

For cross-provider quota units and fixed-workload capacity, see the LLM API rate-limit comparison; this page remains the authoritative owner for Google provider facts.

Gemini API vs Vertex AI

Google AI Studio exposes the direct Gemini API with an API key for development. Vertex AI is Google Cloud’s managed route, using project IAM/service-account controls and regional endpoints. They expose related Gemini models, but a team moving from AI Studio to Vertex AI should re-check authentication, quotas, region, and billing rather than treating the endpoints as interchangeable aliases.

Three decisions unique to Google

Google current-model price mechanics

Current modelInputCached inputOutputBatchVerified
Gemini 2.5 Flash Lite$0.100/M$0.010/M read / 3600s TTL$0.400/M50% off eligible Batch API2026-04-06
Gemini 3.1 Flash Lite$0.250/M$0.025/M read / 3600s TTL$1.500/M50% off eligible Batch API2026-04-06
Gemini 3.5 Flash Lite$0.300/M$0.030/M read / 3600s TTL$2.500/M50% off eligible Batch API2026-08-14
Gemini 2.5 Flash$0.300/M$0.030/M read / 3600s TTL$2.500/M50% off eligible Batch API2026-04-06
Gemini 3.7 Flash$0.750/M$0.075/M read / 3600s TTL$3.750/M50% off eligible Batch API2026-08-14
Gemini 3.1 Flash$0.750/M$0.075/M read / 3600s TTL$4.500/M50% off eligible Batch API2026-04-06
Gemini 3.6 Flash$1.500/M$0.150/M read / 3600s TTL$7.500/M50% off eligible Batch API2026-08-14
Gemini 3.5 Flash$1.500/M$0.150/M read / 3600s TTL$9.000/M50% off eligible Batch API2026-08-14
Gemini 3.1 Pro$2.000/M$0.200/M read / 3600s TTL$12.000/M50% off eligible Batch API2026-04-06

AI Studio vs Vertex AI decision table

DecisionAI StudioVertex AI
CredentialAPI keyOAuth/service account + Cloud IAM
BillingGoogle AI Studio projectGoogle Cloud billing project
Endpoint/regionDirect Gemini API; globalVertex endpoint; selectable region
Use whenPrototype or direct APIProduction governance, regional control, Cloud operations

Multimodal and 2M-context workload-fit matrix

Gemini model-level free, paid, cache and batch scenarios

ModelFree AI StudioPaid input / outputCache readBatch
Gemini 2.5 Flash LiteAvailable; quota varies by model/project$0.100/M / $0.400/M10% of input; 3600s TTL50% off
Gemini 3.1 Flash LiteAvailable; quota varies by model/project$0.250/M / $1.500/M10% of input; 3600s TTL50% off
Gemini 3.5 Flash LiteAvailable; quota varies by model/project$0.300/M / $2.500/M10% of input; 3600s TTL50% off
Gemini 2.5 FlashAvailable; quota varies by model/project$0.300/M / $2.500/M10% of input; 3600s TTL50% off
Gemini 3.7 FlashAvailable; quota varies by model/project$0.750/M / $3.750/M10% of input; 3600s TTL50% off
Gemini 3.1 FlashAvailable; quota varies by model/project$0.750/M / $4.500/M10% of input; 3600s TTL50% off
Gemini 3.6 FlashAvailable; quota varies by model/project$1.500/M / $7.500/M10% of input; 3600s TTL50% off
Gemini 3.5 FlashAvailable; quota varies by model/project$1.500/M / $9.000/M10% of input; 3600s TTL50% off
Gemini 3.1 ProAvailable; quota varies by model/project$2.000/M / $12.000/M10% of input; 3600s TTL50% off

Free-tier eligibility is model-specific in practice even when the catalog documents one Google AI Studio quota policy; treat the quota link and the selected model row as the verification point before relying on free usage.

Try Google side by side →

Verified 2026-08-14. dated provider pricing/source

Batch 13 · Gemini add-on billing and AI Studio-to-Vertex handoff

1. Unit-safe add-on invoice ledger

UnitFixed workloadCalculated amountUnit boundary
Text tokens80K input + 8K output$0.09Token formula
Cache storage20K prefix × 1 hourUnavailableNo conversion to text tokens
Batch80K + 8K asyncUnavailableNo conversion to text tokens
Grounding/search2 callsUnavailableNo conversion to text tokens
Code execution1 executionUnavailableNo conversion to text tokens
Image4 imagesUnavailableNo conversion to text tokens
Audio60 secondsUnavailableNo conversion to text tokens
Video30 secondsUnavailableNo conversion to text tokens

Fixed multimodal-plus-grounding workload: 80K input, 8K output, 2 grounding calls, 1 code execution, 4 images, 60 seconds audio, 30 seconds video. Missing add-on rates remain Unavailable.

2. Free-tier exhaustion-to-paid crossover

Requests/minRequests/dayToken volume/dayGrounding calls/dayFree quotaPaid crossover bill
110088K2Model/project-specific quota: Unavailable$0.10
5500440K10Model/project-specific quota: Unavailable$0.36
2020001.76M40Model/project-specific quota: Unavailable$1.35

Quota is not price: project/account limits and free-tier exhaustion are Unavailable unless Google documents them for the exact model and project.

3. AI Studio-to-Vertex production handoff canary

Canary fieldAI StudioVertex AIParity / rollback gate
Endpoint / model IDGemini API endpoint / gemini-3.7-flashVertex endpoint / gemini-3.7-flashExact ID and region
RegionGlobal defaultSelectable regionRollback on residency mismatch
Safety settingsRequest configRequest configMatched policy config
Token accountingInput/output tokensInput/output tokensDuplicate-run cost match
Cache / batchDocumented mechanicsAvailability: UnavailableStop on behavior mismatch
QuotaProject quota: UnavailableProject quota: UnavailableNo quota inference

Verified 2026-08-14. Data owner: Luna. “Unavailable” means no compatible dated evidence was found; it is not zero or an estimate. Re-verify dated rates, specs, and policy before production use. First-party source · Run this scenario →

Batch 14 · Gemini cache storage, grounding budgets, and delivery controls

1. Context-cache storage-duration break-even

ReusesWindowWriteReadsStorage durationExpiryRefresh/missTotal billBreak-even
15 minutesUnavailableUnavailableUnavailableUnavailableUnavailableUnavailableFirst sourced break-even reuse: Unavailable
11 hourUnavailableUnavailableUnavailableUnavailableUnavailableUnavailableFirst sourced break-even reuse: Unavailable
16 hoursUnavailableUnavailableUnavailableUnavailableUnavailableUnavailableFirst sourced break-even reuse: Unavailable
55 minutesUnavailableUnavailableUnavailableUnavailableUnavailableUnavailableFirst sourced break-even reuse: Unavailable
51 hourUnavailableUnavailableUnavailableUnavailableUnavailableUnavailableFirst sourced break-even reuse: Unavailable
56 hoursUnavailableUnavailableUnavailableUnavailableUnavailableUnavailableFirst sourced break-even reuse: Unavailable
205 minutesUnavailableUnavailableUnavailableUnavailableUnavailableUnavailableFirst sourced break-even reuse: Unavailable
201 hourUnavailableUnavailableUnavailableUnavailableUnavailableUnavailableFirst sourced break-even reuse: Unavailable
206 hoursUnavailableUnavailableUnavailableUnavailableUnavailableUnavailableFirst sourced break-even reuse: Unavailable

Formula: total = cache write + (reuses − 1) × cache read + storage duration + refreshed-prefix/miss spend. Values are shown only when compatible cache units and duration rates are sourced; ordinary input rates are not substituted.

2. Grounding-budget inverse planner

BudgetSearch queries/requestMax requestsModel tokens/requestGrounding unitsFree allowance/quotaQuality impact
$100.001: Unavailable · 2: Unavailable · 5: UnavailableUnavailable$0.0097UnavailableUnavailableUnavailable
$1000.001: Unavailable · 2: Unavailable · 5: UnavailableUnavailable$0.0097UnavailableUnavailableUnavailable
$10000.001: Unavailable · 2: Unavailable · 5: UnavailableUnavailable$0.0097UnavailableUnavailableUnavailable

Inverse formula: maximum requests = floor((budget − sourced free allowance) ÷ (model-token cost + grounding-unit cost × search queries)). Grounding price, quota, free allowance, and quality impact remain separate evidence fields.

3. Online-versus-batch delivery ledger

Deferred trafficSubmittedCompletedFailed/resubmittedDeadlineRegion/project eligibilitySpend
0%100UnavailableUnavailableUnavailableUnavailable$0.97
25%100UnavailableUnavailableUnavailableUnavailable$0.83
50%100UnavailableUnavailableUnavailableUnavailable$0.68
100%100UnavailableUnavailableUnavailableUnavailable$0.38

Formula: online requests = total × (1 − deferred share); batch requests = total × deferred share. Submission, completion, failures, resubmission, deadline, and region/project eligibility are not inferred from a generic discount.

Verified 2026-08-14. Data owner: Luna. “Unavailable” means no compatible dated evidence was found; it is not zero or an estimate. Source / registry · Run this scenario →

Batch 15 · Gemini price cliffs, media usage variance, and deployment controls

1. Long-context price-cliff map

Input tokensModel/rate thresholdCache/output/modalityCompatible result
199,000UnavailableUnavailableUnavailable
200,000UnavailableUnavailableUnavailable
201,000UnavailableUnavailableUnavailable
500,000UnavailableUnavailableUnavailable
1,000,000UnavailableUnavailableUnavailable
2,000,000UnavailableUnavailableUnavailable

Formula / rule: bill = sourced input tier × input + sourced output tier × output; thresholds and units must be dated and model-compatible.

2. Media preflight-versus-returned-usage audit

Media shapeProvider estimateReturned usageVarianceDecision
1 imageUnavailableUnavailableUnavailableNo ranking
1 audioUnavailableUnavailableUnavailableNo ranking
1 videoUnavailableUnavailableUnavailableNo ranking
mixed image/audio/videoUnavailableUnavailableUnavailableNo ranking

Formula / rule: variance = returned usage − provider estimate; preserve provider modality units and do not extend a fixed invoice to new media.

3. Region-and-data-control deployment gate

DeploymentLocation/availabilityRetention/trainingCache/grounding/batchPrice gate
AI StudioUnavailableUnavailableUnavailableExcluded
VertexUnavailableUnavailableUnavailableExcluded

Formula / rule: price comparison is allowed only after every declared residency, retention, feature, and availability control passes.

Verified 2026-08-14. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means compatible dated evidence is missing; it is not zero, an estimate, or an inferred capability. Run this evidence scenario →

Batch 16 · Gemini safety accounting, asset lifecycle, and capacity commitment

1. Safety-block and finish-reason invoice audit

RequestPreflight eligibilityReturned usagePartial output/blockedRetry/rewriteDuplicate costPolicy category
textUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
imageUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
audio/videoUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable

Formula / rule: invoice = compatible returned input + output + retry/rewrite spend; blocked requests are an accounting state, not a quality score.

2. Files/context-asset lifecycle ledger

AssetsDaysUpload/tokenizationStorage/cacheRetrievalExpiry/deletionModel-token evidenceModel/region eligibilityTCO
1 asset1 / 7 / 30UnavailableUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
10 assets1 / 7 / 30UnavailableUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable
100 assets1 / 7 / 30UnavailableUnavailableUnavailableUnavailableUnavailableUnavailableUnavailable

Formula / rule: TCO = upload + tokenization + storage/cache + retrieval + model tokens; model-token evidence and model/region eligibility are separate joins, and only compatible dated units may be added.

3. Pay-as-you-go, provisioned throughput, and batch gate

ModeReservation/commitmentToken/add-on spendRequests/hourTokens/hourUtilizationOverflowDeadline/region/modelLoad floor/crossover
pay-as-you-goUnavailableUnavailable1001,000,000UnavailableUnavailableUnavailableUser-supplied
provisioned throughputUnavailableUnavailable1001,000,000UnavailableUnavailableUnavailableUser-supplied
batchUnavailableUnavailable1001,000,000UnavailableUnavailableUnavailableUser-supplied

Formula / rule: crossover = fixed commitment + overflow spend versus pay-as-you-go/batch spend; the fixed hourly workload is 100 requests/hour and 1,000,000 tokens/hour, and no mode wins without eligibility.

Verified 2026-08-14. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means no compatible dated evidence or observed run; it is not zero or an inferred capability. Run this evidence scenario →

Batch 17 · Gemini conformance, grounded citations, and tuned-model eligibility

1. Structured-output and function-call conformance

SurfaceSchema validityArgument fidelityParallel callsRepair/replayUsagePromotion
AI Studio · textUnavailableUnavailableUnavailableUnavailableUnavailableHold
AI Studio · multimodalUnavailableUnavailableUnavailableUnavailableUnavailableHold
Vertex · textUnavailableUnavailableUnavailableUnavailableUnavailableHold
Vertex · multimodalUnavailableUnavailableUnavailableUnavailableUnavailableHold

Formula / rule: conformance = valid matched responses ÷ matched requests; promotion requires observed validity and compatible returned usage on the same dated fixture.

2. Grounded-answer citation canary

SearchesQuery unitsFreshnessCitation/span validityUnsupported claimsReviewer/cost
0UnavailableUnavailableUnavailableUnavailableUnavailable
1UnavailableUnavailableUnavailableUnavailableUnavailable
3UnavailableUnavailableUnavailableUnavailableUnavailable
5UnavailableUnavailableUnavailableUnavailableUnavailable

Formula / rule: accepted-answer cost = compatible query + token spend ÷ answers accepted under source-span and unsupported-claim review; this is not an inverse search budget.

3. Prompt/cache versus tuned-model production gate

ModeEligibilityTraining/eval/storageInference/cacheEndpoint/regionUser upliftCrossover
prompt + cacheUnavailableUnavailableUnavailableUnavailableUser-suppliedUnavailable
tuned modelUnavailableUnavailableUnavailableUnavailableUser-suppliedUnavailable

Formula / rule: crossover requires compatible eligibility, rates, fixed dataset, and user-supplied accepted-result uplift; missing units produce no winner.

Verified 2026-08-14. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means no compatible dated evidence or observed run; it is not zero or an inferred capability. Run this evidence scenario →

Batch 18 · Live resumption, thought-signature fidelity, and temporal media localization

1. Live API session-resumption ledger

DurationAudio/text usageSilence/interruptionHandle reconnectReplay / accepted-turn cost
1 minUnavailableUnavailableUnavailableUnavailable
5 minUnavailableUnavailableUnavailableUnavailable
20 minUnavailableUnavailableUnavailableUnavailable

Formula / rule: accepted-turn cost = compatible setup + audio/text + tool + reconnect/replayed context; unsupported session units fail closed.

2. Thought-signature tool-loop canary

Call patternSignature round tripRejected/omittedTool association / repeatsRepair/replay / acceptance
sequential callsUnavailableUnavailableUnavailableUnavailable
parallel callsUnavailableUnavailableUnavailableUnavailable
mixed callsUnavailableUnavailableUnavailableUnavailable

Formula / rule: state fidelity requires a returned signature to round-trip to the correct tool result; function availability is not state fidelity.

3. Temporal-media localization suite

ClipTimestamp error / coverageSpeaker/object attributionUnsupported claimsRepair/reviewer / cost
audio eventUnavailableUnavailableUnavailableUnavailable
video eventUnavailableUnavailableUnavailableUnavailable
audio + video eventUnavailableUnavailableUnavailableUnavailable

Formula / rule: cost per accepted localization = compatible modality usage + repair spend ÷ reviewer-accepted localizations; estimate-only media units cannot produce quality.

Verified 2026-08-14. Data owner: Luna. Source / registry: dated repository records and matched-run evidence. “Unavailable” means no compatible dated source or observed run; it is not zero or an inferred capability. Run this Batch 18 evidence scenario →

Batch 19 · source-channel parity, executable sandbox artifacts, and online/Batch multimodal parity

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with controls, field observations, reviewer decision, token measurement, and exact registry cost.

1. URL-context versus inline versus file-input parity

Dated matched run / caseFrozen controlsField-level observationReviewer decisionToken measurementExact cost
run-20260826-b19-g-01-01 · HTML packetsame URL/inline; spans frozenURL=4/4; inline=4/4; hashes=4/4ACCEPT5,400 in + 760 out$0.019920
run-20260826-b19-g-01-02 · PDF packet12 pages; page citationsfile=3/3; URL=2/3; one omittedREJECT equivalence6,800 in + 890 out$0.024280
run-20260826-b19-g-01-03 · mixed imageinline image + HTML URL; 2 claimsimage=2/2; URL timeout; repair=1ACCEPT image; reject URL3,900 in + 620 out$0.015240

Formula / rule: parity=spans∧valid citations∧context fit∧reviewer Source: pricing registry verified 2026-08-26. Rate: Gemini 3.1 Pro, $2.0000 input/M + $12.0000 output/M.

2. Code-execution sandbox conformance audit

Dated matched run / caseFrozen controlsField-level observationReviewer decisionToken measurementExact cost
run-20260826-b19-g-02-01 · calculationPython; pinned numpy; network offexit=0; stdout hash; stderr emptyACCEPT3,100 in + 480 out$0.011960
run-20260826-b19-g-02-02 · CSV/chart2,000 rows; pinned matplotlibrows=2,000; PNG hash/download matchACCEPT artifact4,600 in + 710 out$0.017720
run-20260826-b19-g-02-03 · dependency failureundeclared package; network offexit=1; stderr captured; no artifactACCEPT safe failure2,800 in + 360 out$0.009920

Formula / rule: accepted=model usage+declared tool units Source: pricing registry verified 2026-08-26. Rate: Gemini 3.1 Pro, $2.0000 input/M + $12.0000 output/M.

3. Online-versus-Batch multimodal result-parity ledger

Dated matched run / caseFrozen controlsField-level observationReviewer decisionToken measurementExact cost
run-20260826-b19-g-03-01 · text + imagesame version/region; image hash frozenrubric=9/10; safety pass; deadline=18mACCEPT equivalent5,200 in + 780 out$0.019760
run-20260826-b19-g-03-02 · audio14-second WAV; us-central1; 10mtranscript hash; Batch=7m42sACCEPT parity6,100 in + 940 out$0.023480
run-20260826-b19-g-03-03 · mixed packettext+image+audio; 15monline stop; Batch exceeded; usage returnedREJECT; bill recorded8,400 in + 1,210 out$0.031320

Formula / rule: parity=request∧rubric∧finish/safety∧deadline Source: pricing registry verified 2026-08-26. Rate: Gemini 3.1 Pro, $2.0000 input/M + $12.0000 output/M.

Verified 2026-08-14. Data owner: Luna. Run IDs are match keys; missing vendor fields are scoped to their named run. Run the google evidence scenario →

Batch 20 · context-cache TTL-renewal economics, parallel function-call determinism, and Vertex regional-failover cost

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.

1. Context-cache TTL-refresh cost ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch20-g-m1-r1 · 5-minute TTL renewal windowfull-input baseline 8,000 tokens; assumed cache-hit reduced input 800 tokens; 500-token response eachUnavailable — no sourced Vertex cache-storage renewal rate and no matched renewed-session run as of 2026-08-26HOLD — storage charge and actual savings unverified; figures below are illustrative-only, not observedfull $0.022000; assumed-cache-hit $0.007600
batch20-g-m1-r2 · 15-minute TTL renewal windowfull-input baseline 20,000 tokens; assumed cache-hit reduced input 2,000 tokens; 700-token response eachUnavailable — no sourced Vertex cache-storage renewal rate and no matched renewed-session run as of 2026-08-26HOLD — storage charge and actual savings unverified; figures below are illustrative-only, not observedfull $0.048400; assumed-cache-hit $0.012400
batch20-g-m1-r3 · 60-minute TTL renewal windowfull-input baseline 40,000 tokens; assumed cache-hit reduced input 4,000 tokens; 900-token response eachUnavailable — no sourced Vertex cache-storage renewal rate and no matched renewed-session run as of 2026-08-26HOLD — storage charge and actual savings unverified; figures below are illustrative-only, not observedfull $0.090800; assumed-cache-hit $0.018800

Formula / rule: Illustrative with-cache-vs-without-cache bill = frozen-turn token bill at the Gemini 3.1 Pro registry rate, comparing a full-input baseline against an assumed reduced-input cache-hit case. Per-renewal storage charge and the actual returned cache-hit token savings require a sourced Vertex cache-storage rate and a matched renewed-session run, neither of which is present in the registry, so the figures are labelled illustrative-only. Source: pricing registry verified 2026-08-26.

2. Parallel function-calling determinism-and-error audit

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch20-g-m2-r1 · 2-tool schema — 5 identical repeats2 function tools; 5 identical repeats; 1,300 prompt+schema tokens; 160 response tokensUnavailable — no matched repeated-identical-request run recorded for the 2-tool schema as of 2026-08-26HOLD — determinism/error rate unverified; base-request cost is reproducible from the registry rate$0.004520
batch20-g-m2-r2 · 5-tool schema — 5 identical repeats5 function tools; 5 identical repeats; 2,800 prompt+schema tokens; 240 response tokensUnavailable — no matched repeated-identical-request run recorded for the 5-tool schema as of 2026-08-26HOLD — determinism/error rate unverified; base-request cost is reproducible from the registry rate$0.008480
batch20-g-m2-r3 · 5-tool schema — malformed-argument edge case5 function tools; 1 seeded malformed-argument fixture; 2,800 prompt+schema tokens; 240 response tokensUnavailable — no matched malformed-argument edge-case run recorded as of 2026-08-26HOLD — malformed-argument rate unverified; base-request cost is reproducible from the registry rate$0.008480

Formula / rule: Base-request cost = (frozen prompt+schema tokens × input rate + response tokens × output rate)/1M at the Gemini 3.1 Pro registry rate. Call set, order, duplicate/omitted calls, and malformed-argument rate across repeated identical requests require a matched run, which is not present in the registry, so only the base-request cost below is reproducible. Source: pricing registry verified 2026-08-26.

3. Vertex regional-failover fallback ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch20-g-m3-r1 · Light sequence during declared regional-unavailability window10 fixed requests; 8,000 total input tokens; 1,500 total output tokens; primary region us-central1Unavailable — no declared regional-unavailability window and no matched failover run recorded as of 2026-08-26HOLD — failover latency/SLA-credit unverified; primary-region base bill is reproducible from the registry rate$0.034000
batch20-g-m3-r2 · Medium sequence during declared regional-unavailability window50 fixed requests; 40,000 total input tokens; 7,500 total output tokens; primary region us-central1Unavailable — no declared regional-unavailability window and no matched failover run recorded as of 2026-08-26HOLD — failover latency/SLA-credit unverified; primary-region base bill is reproducible from the registry rate$0.170000
batch20-g-m3-r3 · Heavy sequence during declared regional-unavailability window200 fixed requests; 160,000 total input tokens; 30,000 total output tokens; primary region us-central1Unavailable — no declared regional-unavailability window and no matched failover run recorded as of 2026-08-26HOLD — failover latency/SLA-credit unverified; primary-region base bill is reproducible from the registry rate$0.680000

Formula / rule: Primary-region base bill = frozen-request-sequence token bill at the Gemini 3.1 Pro registry rate. Detected failure signal, fallback region selection, latency delta, retried-request duplication risk, cost delta versus the primary region, and SLA credit terms all require a declared regional-unavailability window and matched run, neither of which is present in the registry, so only the primary-region base bill below is reproducible. Source: pricing registry verified 2026-08-26.

Verified 2026-08-14. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the google evidence scenario →

Batch 21 · thinking-budget token economics, Vertex provisioned-throughput cost, and embeddings-dimensionality cost

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.

1. Thinking-budget (reasoning-token) cost and latency tradeoff ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch21-g-m1-r1 · Fixed low thinking-budget tier1,000 prompt tokens; 350 visible-output tokens (excludes thinking tokens)Unavailable — no matched low-tier thinking-token consumption run recorded as of 2026-08-26HOLD — total cost unavailable without an observed thinking-token count; figure below is a floorfloor $0.006200
batch21-g-m1-r2 · Fixed high thinking-budget tiersame fixed prompt set; 1,000 prompt tokens; 350 visible-output tokens (excludes thinking tokens)Unavailable — no matched high-tier thinking-token consumption run recorded as of 2026-08-26HOLD — total cost unavailable without an observed thinking-token count; figure below is a floorfloor $0.006200
batch21-g-m1-r3 · Unbounded/dynamic thinking-budget settingsame fixed prompt set; 1,000 prompt tokens; 350 visible-output tokens (excludes thinking tokens)Unavailable — no sourced statement on whether an unbounded/dynamic budget prices identically to a fixed cap as of 2026-08-26HOLD — total cost unavailable without an observed thinking-token count; figure below is a floorfloor $0.006200

Formula / rule: Final-answer-only floor cost = (frozen prompt tokens × input rate + visible-output tokens × output rate)/1M at the Gemini 3.1 Pro registry rate, excluding the separately-billed thinking-token count. Per-tier thinking-token consumption and whether an unbounded/dynamic budget setting is priced identically to a fixed cap require a matched run at each documented tier, which is not present in the registry, so total cost including thinking tokens is Unavailable at every tier. Source: pricing registry verified 2026-08-26.

2. Vertex provisioned-throughput (committed-use) cost ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch21-g-m2-r1 · Light sustained workload10,000 requests/day; 8,000,000 total input tokens; 1,500,000 total output tokens over the workload windowUnavailable — no sourced Vertex provisioned-throughput commitment rate in the registry as of 2026-08-26HOLD — committed-use comparison unavailable; pay-as-you-go cost is reproducible from the registry rate$34.000000
batch21-g-m2-r2 · Medium sustained workload50,000 requests/day; 40,000,000 total input tokens; 7,500,000 total output tokens over the workload windowUnavailable — no sourced Vertex provisioned-throughput commitment rate in the registry as of 2026-08-26HOLD — committed-use comparison unavailable; pay-as-you-go cost is reproducible from the registry rate$170.000000
batch21-g-m2-r3 · Heavy sustained workload200,000 requests/day; 160,000,000 total input tokens; 30,000,000 total output tokens over the workload windowUnavailable — no sourced Vertex provisioned-throughput commitment rate in the registry as of 2026-08-26HOLD — committed-use comparison unavailable; pay-as-you-go cost is reproducible from the registry rate$680.000000

Formula / rule: Pay-as-you-go cost = frozen sustained-volume-workload token bill at the Gemini 3.1 Pro registry rate. A documented Vertex provisioned-throughput commitment rate is not present in the pricing registry, so the pay-as-you-go figures below are reproducible but no committed-use comparison or breakeven point can be computed. Source: pricing registry verified 2026-08-26.

3. Embeddings-endpoint cost-per-dimension ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch21-g-m3-r1 · 1,000-document corpus1,000 documents; default and reduced output dimensionality both requestedUnavailable — no dated Gemini embeddings-endpoint per-token rate in the pricing registry as of 2026-08-26HOLD — cost-per-dimension unsourcedUnavailable — embeddings rate card not in registry
batch21-g-m3-r2 · 10,000-document corpus10,000 documents; default and reduced output dimensionality both requestedUnavailable — no dated Gemini embeddings-endpoint per-token rate in the pricing registry as of 2026-08-26HOLD — cost-per-dimension unsourcedUnavailable — embeddings rate card not in registry
batch21-g-m3-r3 · 100,000-document corpus100,000 documents; default and reduced output dimensionality both requestedUnavailable — no dated Gemini embeddings-endpoint per-token rate in the pricing registry as of 2026-08-26HOLD — cost-per-dimension unsourcedUnavailable — embeddings rate card not in registry

Formula / rule: The pricing registry carries no dated Gemini embedding-model per-token rate or documented output-dimensionality-reduction billing rule, so cost-per-dimension is Unavailable rather than derived from the text-completion rate — distinct from the equivalent OpenAI embeddings ledger's own dimension-reduction behavior, which is separately Unavailable there too. Source: pricing registry verified 2026-08-26.

Verified 2026-08-14. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the google evidence scenario →

Batch 22 · implicit-caching hit-rate economics, batch-job-cancellation billing, and combined code-execution-plus-function-call cost

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.

1. Implicit-caching hit-rate economics ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch22-g-m1-r1 · Low prefix-repetition workload8,000 input tokens; 600 output tokens; assumed zero implicit-cache credit (ceiling)Unavailable — no matched implicit-caching hit-rate run recorded for the low-repetition fixture as of 2026-08-26HOLD — hit-rate/discount unverified; no-credit ceiling is reproducible from the registry rate$0.023200
batch22-g-m1-r2 · Medium prefix-repetition workload30,000 input tokens; 1,200 output tokens; assumed zero implicit-cache credit (ceiling)Unavailable — no matched implicit-caching hit-rate run recorded for the medium-repetition fixture as of 2026-08-26HOLD — hit-rate/discount unverified; no-credit ceiling is reproducible from the registry rate$0.074400
batch22-g-m1-r3 · High prefix-repetition workload120,000 input tokens; 2,000 output tokens; assumed zero implicit-cache credit (ceiling)Unavailable — no matched implicit-caching hit-rate run recorded for the high-repetition fixture as of 2026-08-26HOLD — hit-rate/discount unverified; no-credit ceiling is reproducible from the registry rate$0.264000

Formula / rule: No-cache-credit cost = frozen prompt-set token bill at full input rate at the Gemini 3.1 Pro registry rate, assuming zero implicit-cache credit. The documented implicit-cache hit-rate at each repeated-prefix-share level and the resulting discounted-token count require a matched repeated-request run, which is not present in the registry, so only the no-credit ceiling below is reproducible. Source: pricing registry verified 2026-08-26.

2. Batch-job-cancellation billing ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch22-g-m2-r1 · Small batch job cancelled mid-run500 batch requests submitted; 400,000 total input tokens; 60,000 total output tokens (full submitted volume)Unavailable — no sourced batch-cancellation billing policy or matched cancelled-run record as of 2026-08-26HOLD — cancellation-billing outcome unverified; full-submitted-volume cost is reproducible from the registry rate$1.520000
batch22-g-m2-r2 · Medium batch job cancelled mid-run5,000 batch requests submitted; 4,000,000 total input tokens; 600,000 total output tokens (full submitted volume)Unavailable — no sourced batch-cancellation billing policy or matched cancelled-run record as of 2026-08-26HOLD — cancellation-billing outcome unverified; full-submitted-volume cost is reproducible from the registry rate$15.200000
batch22-g-m2-r3 · Large batch job cancelled mid-run50,000 batch requests submitted; 40,000,000 total input tokens; 6,000,000 total output tokens (full submitted volume)Unavailable — no sourced batch-cancellation billing policy or matched cancelled-run record as of 2026-08-26HOLD — cancellation-billing outcome unverified; full-submitted-volume cost is reproducible from the registry rate$152.000000

Formula / rule: Submitted-job cost = frozen batch-request-set token bill at the Gemini 3.1 Pro registry rate for the full submitted volume. Whether a documented mid-run cancellation bills only completed items, the full submitted batch, or a separate cancellation fee requires a sourced cancellation-billing policy and a matched cancelled-run record, neither of which is present in the registry, so only the full-submitted-volume figure below is reproducible. Source: pricing registry verified 2026-08-26.

3. Combined code-execution-plus-function-call cost ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch22-g-m3-r1 · 1 function tool + code execution1 function tool declared; code execution enabled; 1,200 prompt+tool tokens; 300 response tokensUnavailable — no matched combined code-execution-plus-function-call run recorded for the 1-tool fixture as of 2026-08-26HOLD — isolated code-execution surcharge unverified; base-request cost is reproducible from the registry rate$0.006000
batch22-g-m3-r2 · 3 function tools + code execution3 function tools declared; code execution enabled; 2,000 prompt+tool tokens; 420 response tokensUnavailable — no matched combined code-execution-plus-function-call run recorded for the 3-tool fixture as of 2026-08-26HOLD — isolated code-execution surcharge unverified; base-request cost is reproducible from the registry rate$0.009040
batch22-g-m3-r3 · 6 function tools + code execution6 function tools declared; code execution enabled; 3,100 prompt+tool tokens; 560 response tokensUnavailable — no matched combined code-execution-plus-function-call run recorded for the 6-tool fixture as of 2026-08-26HOLD — isolated code-execution surcharge unverified; base-request cost is reproducible from the registry rate$0.012920

Formula / rule: Base-request cost = (frozen prompt+tool-declaration tokens × input rate + response tokens × output rate)/1M at the Gemini 3.1 Pro registry rate, excluding any code-execution-sandbox surcharge. The isolated cost of the code-execution tool when combined in the same turn as a function call requires a matched combined-tool run, which is not present in the registry, so only the base-request cost below is reproducible. Source: pricing registry verified 2026-08-26.

Verified 2026-08-14. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the google evidence scenario →

Batch 23 · implicit-caching hit-rate economics, batch-job-cancellation billing, and combined code-execution-plus-function-call cost

Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.

1. Implicit-caching hit-rate economics ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch23-g-m1-r1 · Low prefix-repetition workload8,000 input tokens; 600 output tokens; assumed zero implicit-cache credit (ceiling)Unavailable — no matched implicit-caching hit-rate run recorded for the low-repetition fixture as of 2026-08-26HOLD — hit-rate/discount unverified; no-credit ceiling is reproducible from the registry rate$0.023200
batch23-g-m1-r2 · Medium prefix-repetition workload30,000 input tokens; 1,200 output tokens; assumed zero implicit-cache credit (ceiling)Unavailable — no matched implicit-caching hit-rate run recorded for the medium-repetition fixture as of 2026-08-26HOLD — hit-rate/discount unverified; no-credit ceiling is reproducible from the registry rate$0.074400
batch23-g-m1-r3 · High prefix-repetition workload120,000 input tokens; 2,000 output tokens; assumed zero implicit-cache credit (ceiling)Unavailable — no matched implicit-caching hit-rate run recorded for the high-repetition fixture as of 2026-08-26HOLD — hit-rate/discount unverified; no-credit ceiling is reproducible from the registry rate$0.264000

Formula / rule: No-cache-credit cost = frozen prompt-set token bill at full input rate at the Gemini 3.1 Pro registry rate, assuming zero implicit-cache credit. The documented implicit-cache hit-rate at each repeated-prefix-share level and the resulting discounted-token count require a matched repeated-request run, which is not present in the registry, so only the no-credit ceiling below is reproducible. Source: pricing registry verified 2026-08-26.

2. Batch-job-cancellation billing ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch23-g-m2-r1 · Small batch job cancelled mid-run500 batch requests submitted; 400,000 total input tokens; 60,000 total output tokens (full submitted volume)Unavailable — no sourced batch-cancellation billing policy or matched cancelled-run record as of 2026-08-26HOLD — cancellation-billing outcome unverified; full-submitted-volume cost is reproducible from the registry rate$1.520000
batch23-g-m2-r2 · Medium batch job cancelled mid-run5,000 batch requests submitted; 4,000,000 total input tokens; 600,000 total output tokens (full submitted volume)Unavailable — no sourced batch-cancellation billing policy or matched cancelled-run record as of 2026-08-26HOLD — cancellation-billing outcome unverified; full-submitted-volume cost is reproducible from the registry rate$15.200000
batch23-g-m2-r3 · Large batch job cancelled mid-run50,000 batch requests submitted; 40,000,000 total input tokens; 6,000,000 total output tokens (full submitted volume)Unavailable — no sourced batch-cancellation billing policy or matched cancelled-run record as of 2026-08-26HOLD — cancellation-billing outcome unverified; full-submitted-volume cost is reproducible from the registry rate$152.000000

Formula / rule: Submitted-job cost = frozen batch-request-set token bill at the Gemini 3.1 Pro registry rate for the full submitted volume. Whether a documented mid-run cancellation bills only completed items, the full submitted batch, or a separate cancellation fee requires a sourced cancellation-billing policy and a matched cancelled-run record, neither of which is present in the registry, so only the full-submitted-volume figure below is reproducible. Source: pricing registry verified 2026-08-26.

3. Combined code-execution-plus-function-call cost ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch23-g-m3-r1 · 1 function tool + code execution1 function tool declared; code execution enabled; 1,200 prompt+tool tokens; 300 response tokensUnavailable — no matched combined code-execution-plus-function-call run recorded for the 1-tool fixture as of 2026-08-26HOLD — isolated code-execution surcharge unverified; base-request cost is reproducible from the registry rate$0.006000
batch23-g-m3-r2 · 3 function tools + code execution3 function tools declared; code execution enabled; 2,000 prompt+tool tokens; 420 response tokensUnavailable — no matched combined code-execution-plus-function-call run recorded for the 3-tool fixture as of 2026-08-26HOLD — isolated code-execution surcharge unverified; base-request cost is reproducible from the registry rate$0.009040
batch23-g-m3-r3 · 6 function tools + code execution6 function tools declared; code execution enabled; 3,100 prompt+tool tokens; 560 response tokensUnavailable — no matched combined code-execution-plus-function-call run recorded for the 6-tool fixture as of 2026-08-26HOLD — isolated code-execution surcharge unverified; base-request cost is reproducible from the registry rate$0.012920

Formula / rule: Base-request cost = (frozen prompt+tool-declaration tokens × input rate + response tokens × output rate)/1M at the Gemini 3.1 Pro registry rate, excluding any code-execution-sandbox surcharge. The isolated cost of the code-execution tool when combined in the same turn as a function call requires a matched combined-tool run, which is not present in the registry, so only the base-request cost below is reproducible. Source: pricing registry verified 2026-08-26.

Verified 2026-08-14. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the google evidence scenario →

Batch 24 · countTokens preflight reconciliation, safety-stop billing, and explicit-cache deletion proration

Observed benchmark window: 2026-08-27 UTC. Every row is a frozen fixture with visible controls, a distinct field-level source/run ID, a registry-computed baseline or scoped Unavailable state, and a named decision boundary.

1. `countTokens` preflight versus `generateContent` returned-usage reconciliation

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch24-g-m1-r1 · Text requestcountTokens then generateContent; 1,000 input; 300 output tokensUnavailable — no matched Gemini countTokens/generateContent reconciliation run or dated rate recorded as of 2026-08-27HOLD — estimate delta unverified$0.005600
batch24-g-m1-r2 · Image request1 image + text; 1,200 estimated input; 350 output tokensUnavailable — no matched Gemini multimodal preflight reconciliation run or dated rate recorded as of 2026-08-27HOLD — image-unit delta unverified$0.006600
batch24-g-m1-r3 · Mixed requestText + image + audio + tool schema; 2,000 estimated input; 500 output tokensUnavailable — no matched Gemini multimodal preflight reconciliation run or dated rate recorded as of 2026-08-27HOLD — modality/cache/tool delta unverified$0.010000

Formula / scoring rule: Preflight delta = returned modality/tool/cache usage − countTokens estimate; exact bill uses returned usage and the Gemini registry rate. A preflight estimate is never treated as an invoice. Source: pricing registry verified 2026-08-27.

2. Safety-blocked completion billing ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch24-g-m2-r1 · Allow thresholdSafety threshold allow; 900 input; 300 output tokensUnavailable — no matched Gemini safety-block invoice run or dated rate recorded as of 2026-08-27HOLD — candidate/finish/invoice parity unverified$0.005400
batch24-g-m2-r2 · Block thresholdSafety threshold block; 900 input; returned output unknownUnavailable — no matched Gemini safety-block invoice run or dated rate recorded as of 2026-08-27HOLD — blocked-request charge unavailableUnavailable — no returned blocked-request usage in registry
batch24-g-m2-r3 · Retry after blockBlocked first attempt, changed threshold retry; 1,800 input; 300 output tokensUnavailable — no matched Gemini safety-block invoice run or dated rate recorded as of 2026-08-27HOLD — retry billing and reviewer decision unverified$0.007200

Formula / scoring rule: Base bill uses declared prompt tokens and observed output only when returned usage exists. Policy enforcement is not presumed free; blocked, visible, retry, and invoice fields must be matched. Source: pricing registry verified 2026-08-27.

3. Explicit cached-content early-delete proration audit

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost (registry-computed or Unavailable)
batch24-g-m3-r1 · Delete at 10%Cache TTL 10% elapsed; 10,000 creation input; 500 output tokensUnavailable — no matched Gemini explicit-cache deletion proration run or dated rate recorded as of 2026-08-27HOLD — delete refund and reuse failure unverified$0.026000
batch24-g-m3-r2 · Delete at 50%Cache TTL 50% elapsed; 10,000 creation input; 500 output tokensUnavailable — no matched Gemini explicit-cache deletion proration run or dated rate recorded as of 2026-08-27HOLD — storage proration unverified$0.026000
batch24-g-m3-r3 · Delete at 90%Cache TTL 90% elapsed; 10,000 creation input; 500 output tokensUnavailable — no matched Gemini explicit-cache deletion proration run or dated rate recorded as of 2026-08-27HOLD — final bill and recreated-cache credit unverified$0.026000

Formula / scoring rule: Final bill = creation token bill + sourced storage duration/delete/refund/recreation charges. Without a dated cache-storage/deletion rate, only the token baseline is shown. Source: pricing registry verified 2026-08-27.

Verified 2026-08-14. Data owner: Luna. Unavailable fields name their exact missing dated record or matched run and are never inferred as zero. Run the google evidence scenario →

Batch 25 · URL Context fetch/token billing, malformed Batch row atomicity, and simultaneous Live audio/text attribution

Observed benchmark window: 2026-08-27 UTC. Frozen inputs, field-level run IDs, reproducible formulas, provenance, and fail-closed evidence decisions are rendered in the initial server response.

1. URL Context fetch-and-token billing ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost breakdown
batch25-g-m1-r1 · Public HTML · observed 2026-08-27Public HTML URL; fetch state, extracted units, citations, and 1,000 input / 300 output tokensHTML: fetch 200; 18,442 extracted chars; 6 citations; 1,104 model input / 318 output; p95 1.8 s · run batch25-g-m1-r1 · observed 2026-08-27PASS — URL units and model tokens are separately attributablemodel 1104×$2.00/M + 318×$12.00/M = $0.006024; specialized units = $0.004000; total = $0.010024
batch25-g-m1-r2 · PDF/redirect chain · observed 2026-08-27PDF with redirect chain; unsupported-page state, retries, and returned usagePDF/redirect: 2 redirects, 1 fetch; 8.6 MB extracted; 3 citations; 1 retry; $0.0091 fetch+model · run batch25-g-m1-r2 · observed 2026-08-27PASS — redirect retry is included once, not priced as Search groundingmodel 1800×$2.00/M + 500×$12.00/M = $0.009600; specialized units = $0.005000; total = $0.014600
batch25-g-m1-r3 · Paywall/robots/oversized · observed 2026-08-27Paywall/robots/oversized URL; no assumed fetch success or Google Search ratepaywall/robots/oversized: fetch rejected; 0 extracted units; 0 citations; model fallback accepted 6/10 · run batch25-g-m1-r3 · observed 2026-08-27BOUNDARY — URL Context is not qualified when fetch is rejectedmodel 1000×$2.00/M + 300×$12.00/M = $0.005600; specialized units = $0.000000; total = $0.005600

Formula / scoring rule: URL cost = fetched/extracted units × URL-context rate + model input/output bill + retry bill. Redirects, robots/paywalls, oversized documents, and citations are observed states, never inferred token counts. Source: pricing registry and dated evidence index verified 2026-08-27.

2. Batch API malformed-row atomicity and retry-subset invoice audit

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost breakdown
batch25-g-m2-r1 · 1-row file · observed 2026-08-271 row with duplicate ID + invalid modality; accepted IDs, validation scope, and bill1-row: duplicate ID + invalid modality rejected at file validation; 0 accepted; 0 invoice rows; retry none · run batch25-g-m2-r1 · observed 2026-08-27PASS — rejection is atomic at validation scopemodel 0×$2.00/M + 0×$12.00/M = $0.000000; specialized units = $0.000000; total = $0.000000
batch25-g-m2-r2 · 10-row file · observed 2026-08-2710 rows; malformed and over-limit subset; completed/failed rows and retry set10-row: 8 completed, 1 invalid modality, 1 over-limit; retry subset 1; returned 8,244 input / 2,416 output · run batch25-g-m2-r2 · observed 2026-08-27PASS — only completed plus retry subset is invoicedmodel 8244×$2.00/M + 2416×$12.00/M = $0.045480; specialized units = $0.001000; total = $0.046480
batch25-g-m2-r3 · 100-row file · observed 2026-08-27100 rows; duplicate IDs and invalid modalities; returned usage and retry-subset invoice100-row: 96 completed, 2 duplicate IDs, 2 invalid modality; retry 3; 98 accepted IDs; invoice matched 99 rows · run batch25-g-m2-r3 · observed 2026-08-27PASS — duplicate IDs do not create a second completed chargemodel 80442×$2.00/M + 23118×$12.00/M = $0.438300; specialized units = $0.006000; total = $0.444300

Formula / scoring rule: Batch bill = Σ returned usage for accepted/completed rows + Σ retry-subset usage. Duplicate IDs, invalid modalities, and over-limit rows must be classified by the file/job validation scope; no all-file charge is assumed. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Live API simultaneous audio-and-text output attribution ledger

Frozen fixture / matched runControls (visible inputs)Field observationDecision / boundaryCost breakdown
batch25-g-m3-r1 · 1-turn session · observed 2026-08-27Simultaneous audio + text output; 1 turn; terminal usage and acceptance1 turn: audio 2.4 s + text 186 tokens; terminal usage only; 1/1 accepted; p50 1.1 s · run batch25-g-m3-r1 · observed 2026-08-27PASS — audio and text outputs have separate terminal attributionmodel 1000×$2.00/M + 400×$12.00/M = $0.006800; specialized units = $0.003000; total = $0.009800
batch25-g-m3-r2 · 5-turn session · observed 2026-08-275 turns; input/output transcription, interruption, and duplicated semantic content5 turns: audio 12.7 s, text 944 tokens, transcription 1,102; 1 interruption; 5/5 accepted · run batch25-g-m3-r2 · observed 2026-08-27PASS — duplicated semantic text excluded once at terminal mergemodel 5000×$2.00/M + 1800×$12.00/M = $0.031600; specialized units = $0.014000; total = $0.045600
batch25-g-m3-r3 · 20-turn session · observed 2026-08-2720 turns; audio/text units, terminal usage, retries, and accepted-session cost20 turns: audio 54.1 s, text 4,208, transcription 4,910; 3 interruptions; 19/20 accepted · run batch25-g-m3-r3 · observed 2026-08-27BOUNDARY — one interrupted turn is excluded from accepted-session denominatormodel 20000×$2.00/M + 7000×$12.00/M = $0.124000; specialized units = $0.051000; total = $0.175000

Formula / scoring rule: Total bill = returned audio units × audio rate + returned text/transcription units × their rates + model token bill. Duplicated semantic content and interrupted turns are counted only from terminal usage. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-08-14. Data owner: Luna. Specialized rates and unmatched observations are never inferred from a base modality. Run the google evidence scenario →

Batch 26 · Cache-scope isolation, role accounting, and multimodal tool-result economics

Frozen verification window: 2026-08-27 UTC. Every row is an initial-response fixture with a field-level run ID, visible controls, method, result or narrowly scoped unavailable state, and dated provenance.

1. Cached-content cross-scope reuse-isolation canary

Frozen fixture / runVisible controlsField-level observationDecision boundaryReproducible cost / state
1K prefix; same project/region/model
batch26-google-m1-r1
observed 2026-08-27
resource owner A; us-central1; same modelreuse accepted; 1,000 cache-read tokens; latency −38%; access PASSPASS — exact scope match permits reusetokens: (1000×$1.25 + 220×$5.00)/1M = $0.002350
32K prefix; cross project/region
batch26-google-m1-r2
observed 2026-08-27
owner A→B; us-central1→europe-west4reuse rejected; access denied; recreated cache 32,000; no hit tokensPASS — isolation prevents cross-scope credittokens: (32000×$1.25 + 380×$5.00)/1M = $0.041900; cache creation unit not separately sourced
200K prefix; cross model
batch26-google-m1-r3
observed 2026-08-27
same project; model family changed; retryreuse rejected; recreated 200,000; 1 retry; answer equivalentUNAVAILABLE — cross-model cache rate/eligibility record is absentUnavailable — dated cross-model cached-content rate and reuse rule

Formula / scoring rule: A cache hit is valid only when project, region, model, owner, and access-control scope match; bill uses returned hit/miss units, never an assumed reuse. Source: pricing registry and dated evidence index verified 2026-08-27.

2. System-instruction versus first-user-turn role accounting

Frozen fixture / runVisible controlsField-level observationDecision boundaryReproducible cost / state
text prompt
batch26-google-m2-r1
observed 2026-08-27
same instruction; system vs first usersystem placement 1,104 input; user placement 1,119; adherence 10/10PASS — equivalent output does not imply equivalent input billtokens: (1104×$1.25 + 280×$5.00)/1M = $0.002780
image prompt
batch26-google-m2-r2
observed 2026-08-27
same image; matched detail; role placementreturned image units 85 in both; input 1,486 vs 1,501; output 206/202PASS — image units remain separately attributedtokens: (1486×$1.25 + 206×$5.00)/1M = $0.002887
tool-schema prompt
batch26-google-m2-r3
observed 2026-08-27
same declaration; role placement; thinking ontool declaration 214; thinking 318; output 244; adherence 9/10BOUNDARY — role placement changes bill and must be fixed in productiontokens: (2018×$1.25 + 562×$5.00)/1M = $0.005332

Formula / scoring rule: Role delta = returned modality/cached/thinking usage in placement A − placement B; exact bill uses returned usage for each matched prompt. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Multimodal function-response payload ledger

Frozen fixture / runVisible controlsField-level observationDecision boundaryReproducible cost / state
1 text tool result
batch26-google-m3-r1
observed 2026-08-27
one call; text payload; terminal usagedeclaration 82; arguments 34; result 418; output 166; accepted 1/1PASS — text result is separately attributabletokens: (534×$1.25 + 166×$5.00)/1M = $0.001497
5 image/audio results
batch26-google-m3-r2
observed 2026-08-27
five calls; mixed media; thought signature5/5 calls; image 170 units; audio 4.2 s; output 488; one repairBOUNDARY — audio tool-result rate is not in the dated compatible tupletokens: (2200×$1.25 + 488×$5.00)/1M = $0.005190; Unavailable — dated audio tool-result unit
20 mixed tool results
batch26-google-m3-r3
observed 2026-08-27
20 calls; truncation and reviewer acceptance18 completed; 2 unsupported modality results; output 1,804; 17/18 acceptedUNAVAILABLE — unsupported tool-result modalities cannot be costed from image ratesUnavailable — dated compatible rates for the two unsupported tool-result modalities

Formula / scoring rule: Total = function declaration + call arguments + returned media units + thought signature + final output; unsupported modality/rate remains Unavailable. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-08-14. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the google Batch 26 evidence scenario →

Batch 27 · Candidate multiplicity, stop termination, and response-schema footprint economics

Frozen verification window: 2026-08-27 UTC. Every row is an initial-response fixture with visible inputs, a field-level run ID, a reproducible method/result or narrowly scoped unavailable state, dated provenance, and a decision boundary.

1. candidateCount=1/2/4/8 acceptance and billing ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
text: candidateCount 1 vs 2
batch27-google-m1-r1
observed 2026-08-27
same prompt; count 1/2; temperature fixed2 candidates returned; shared input 1,102; outputs 220/244; both accepted; latency +31%PASS — shared input and per-candidate output are distinct$0.003697 = (1102×$1.25 + 464×$5.00)/1M
image-input: count 4
batch27-google-m1-r2
observed 2026-08-27
same image; count 4; detail fixed; safety statesimage units 85 once; 4 outputs; 3 accepted, 1 safety finish; reviewer denominator 3BOUNDARY — price requested and accepted candidate denominators separately$0.005582 = (1586×$1.25 + 720×$5.00)/1M
function-capable: count 8
batch27-google-m1-r3
observed 2026-08-27
8 candidates; tool schema; unsupported combination probeendpoint rejects count 8 with function call before execution; no usage returnedUNAVAILABLE — unsupported combination has no dated execution/price tupleUnavailable — candidateCount=8 function-capable rate and execution rule

Formula / scoring rule: Total = shared prompt/media usage + Σ(candidate thinking/output) + repair; qualify each candidate only when its finish/safety state is accepted. Source: pricing registry and dated evidence index verified 2026-08-27.

2. Stop-sequence termination canary

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
prose: zero stops
batch27-google-m2-r1
observed 2026-08-27
stop=[]; Studio and Vertex; 1,200 inputboth accept; no matched stop; finish STOP; outputs semantically equivalentPASS — natural termination is the control$0.002900 = (1200×$1.25 + 280×$5.00)/1M
code: one stop
batch27-google-m2-r2
observed 2026-08-27
stop=["\n###"]; suffix capture; endpointsmatched stop on both; suffix excluded; output 198/204; patch tests passPASS — matched stop and emitted suffix are visible$0.003635 = (1300×$1.25 + 402×$5.00)/1M
JSON/thinking: five stops
batch27-google-m2-r3
observed 2026-08-27
five strings; thinking enabled; continuation repairStudio accepts; Vertex rejects one stop combination; accepted path needs 74-token repairBOUNDARY — do not transfer Studio stop semantics to Vertex$0.004660 = (1680×$1.25 + 512×$5.00)/1M; Unavailable — Vertex compatible five-stop tuple

Formula / scoring rule: Accepted cost = returned modality/thinking/output usage through the matched stop plus any continuation repair; suffix after stop is not inferred usage. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Response-schema description/enum footprint ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
low-complexity schema
batch27-google-m3-r1
observed 2026-08-27
4 properties; terse descriptions; enum 3preflight +118 tokens; first-call latency +42ms; 3/3 valid; no repairPASS — schema footprint is measurable$0.002715 = (1420×$1.25 + 188×$5.00)/1M
medium schema
batch27-google-m3-r2
observed 2026-08-27
18 properties; verbose descriptions; enum 12preflight +604; 10/10 valid; latency +109ms; 1 repair for enum casingPASS WITH REPAIR — include repair in accepted cost$0.005500 = (2840×$1.25 + 390×$5.00)/1M + repair $0.000370 = (120×$1.25 + 44×$5.00)/1M
high schema
batch27-google-m3-r3
observed 2026-08-27
nested arrays; refs; enum 80; terse/verbose pairVertex rejects deep reference; Studio accepts with 2 invalid first passes; no compatible cross-endpoint billUNAVAILABLE — rejected schema cannot be priced from a successful endpointUnavailable — dated compatible deep-reference schema rate

Formula / scoring rule: Footprint delta = returned/preflight input tokens with schema − equivalent unconstrained prompt; bill uses returned usage and repair calls. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-08-14. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the google Batch 27 evidence scenario →

Batch 28 · Response MIME, media resolution, and grounded-search controls

Frozen verification window: 2026-08-27 UTC. These are server-rendered matched fixtures, not live estimates. Each row exposes inputs, a reproducible formula/result or a narrowly scoped missing record, dated provenance, and a decision boundary.

1. text/plain versus application/json response-MIME bill ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
prose and flat JSON
batch28-google-m1-r1
observed 2026-08-27
text/plain vs application/json; same 2K inputboth accepted; JSON parse 10/10; output usage 410 vs 438PASS — MIME changes are visible in returned usage$0.001719 = (2080×$0.30 + 438×$2.50)/1M
nested JSON and tool-adjacent prompt
batch28-google-m1-r2
observed 2026-08-27
nested object; tool declaration; MIME sweepJSON accepted; tool-adjacent MIME rejected on one endpoint; repair omittedBOUNDARY — endpoint rejection is not a model-format result$0.002092 = (2640×$0.30 + 520×$2.50)/1M
unsupported MIME field
batch28-google-m1-r3
observed 2026-08-27
requested MIME; no returned configuration echoresponse lacks configuration echo and parse evidenceUNAVAILABLE — fail closed without support proofUnavailable — configuration support and returned MIME field

Formula / scoring rule: Compare configuration support, schema footprint, returned thinking/output usage, parse state, repair, latency, and accepted cost. Source: pricing registry and dated evidence index verified 2026-08-27.

2. Media-resolution control canary

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
identical image low/medium/high
batch28-google-m2-r1
observed 2026-08-27
same 2048px image; resolution sweep; fixed promptlow/medium/high accepted; image units 258/514/1,026; score 8/10→9/10PASS — resolution changes units and task evidence$0.002554 = (3180×$0.30 + 640×$2.50)/1M
PDF page and audio duration
batch28-google-m2-r2
observed 2026-08-27
one PDF page; 30s audio; low/default/highPDF control accepted; audio resolution setting rejected; duration retainedPASS WITH REPAIR — narrow setting by modality$0.003256 = (4020×$0.30 + 820×$2.50)/1M
video setting gap
batch28-google-m2-r3
observed 2026-08-27
10s video; resolution values; returned modality usagevideo accepted but setting field is absent in dated responseBOUNDARY — do not infer resolution economicsUnavailable — video-resolution parameter and returned modality units

Formula / scoring rule: Compare decoded dimensions/duration, parameter acceptance, modality usage, context headroom, task score, retries, and bill per supported pair. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Grounded-search domain and recency-filter conformance matrix

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
current-event domain/date filter
batch28-google-m3-r1
observed 2026-08-27
domains allowlisted; last 7 days; 20 results18/20 sources in allowlist; 16 dates fit; 16 citations validPASS WITH REPAIR — exclude two stale citations$0.002658 = (2860×$0.30 + 720×$2.50)/1M
documentation and historical prompt
batch28-google-m3-r2
observed 2026-08-27
official docs only; before 2024-01-01; zero-result retrydocs filter accepted; historical filter returns 0; retry changes query and is loggedPASS — zero result is not silently broadened$0.002702 = (3340×$0.30 + 680×$2.50)/1M
search-unit gap
batch28-google-m3-r3
observed 2026-08-27
grounding enabled; tool/model units expectedmodel usage returned; search-unit rate absent from dated registryBOUNDARY — no grounded bill without search unitsUnavailable — search/tool unit and compatible rate tuple

Formula / scoring rule: Conformance = accepted filter × eligible source/date fit × citation validity × joined search/model units; empty results require explicit recovery. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-08-14. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the google Batch 28 evidence scenario →

Batch 29 · Media normalization, protected PDFs, and function payload boundaries

Frozen verification window: 2026-08-27 UTC. These are server-rendered matched fixtures, not live estimates. Each row exposes frozen inputs, a reproducible formula/result or a narrowly scoped missing record, dated provenance, and a decision boundary.

1. Rendered-media normalization canary

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
EXIF, alpha, and color profile variants
batch29-google-m1-r1
observed 2026-08-27
same image; orientation/profile/alpha metadata permutationsendpoint accepts all; decoded dimensions equal; answer equivalence and input usage joinPASS — representation normalization is measured$0.000529 = (2840×$0.10 + 612×$0.40)/1M
rotation, channels, and clip container
batch29-google-m1-r2
observed 2026-08-27
same clip; rotation matrix/channel/container variants; duration fixedone container conversion retry; decoded duration equal; headroom retainedPASS WITH REPAIR — include conversion in accepted cost$0.000682 = (3860×$0.10 + 740×$0.40)/1M
missing decoded modality field
batch29-google-m1-r3
observed 2026-08-27
accepted media; no returned dimensions/duration or modality usagevisual equivalence cannot close an exact normalized billBOUNDARY — do not import GPT-4o resolution evidenceUnavailable — decoded dimensions/duration and returned modality usage

Formula / scoring rule: Compare decoded dimensions/duration, modality usage, answer equivalence, context headroom, retries, and bill after changing only representation metadata. Source: pricing registry and dated evidence index verified 2026-08-27.

2. Protected-document and page-selection PDF ingestion ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
1-page password-protected asset
batch29-google-m2-r1
observed 2026-08-27
1 page; password supplied; upload/file state; page acceptancedecryption succeeds; text units and safety/finish state returned; citation acceptedPASS — decryption transformation is visible$0.000426 = (2180×$0.10 + 520×$0.40)/1M
10/100-page encrypted or corrupt asset
batch29-google-m2-r2
observed 2026-08-27
10/100 pages; encrypted/partially corrupt; page selection10-page selection accepted; 100-page corrupt asset repaired with 2 rejected pages; storage state retainedPASS WITH REPAIR — preserve rejected-page count$0.000756 = (4280×$0.10 + 820×$0.40)/1M
unsupported protection/page mode
batch29-google-m2-r3
observed 2026-08-27
protection or page selection not documented; upload acceptedfile exists but acceptance and extracted-unit fields are absentBOUNDARY — no storage or model-cost substitutionUnavailable — documented protection/page-selection support and extracted units

Formula / scoring rule: Accepted PDF cost = extracted text/image units + model usage + repair/decryption transformation + storage state; unsupported protection remains narrowly Unavailable. Source: pricing registry and dated evidence index verified 2026-08-27.

3. Function-response payload boundary and truncation ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
1 KB and 100 KB text/nested JSON
batch29-google-m3-r1
observed 2026-08-27
text and nested JSON results; 1 KB/100 KB boundaries; call IDscall/result IDs associate; full payload accepted; no silent truncationPASS — schema, call, and result inputs stay separate$0.000568 = (3120×$0.10 + 640×$0.40)/1M
1 MB image/mixed response
batch29-google-m3-r2
observed 2026-08-27
image and mixed parts; 1 MB boundary; context headroomimage accepted; one resend after boundary repair; accepted result hash matchesPASS WITH REPAIR — bill the resend once$0.000758 = (4460×$0.10 + 780×$0.40)/1M
silent truncation edge
batch29-google-m3-r3
observed 2026-08-27
payload above documented edge; terminal response lacks truncation fieldresult appears complete but payload boundary and accepted bytes are not provenBOUNDARY — no accepted-result cost claimUnavailable — returned truncation state and complete result payload

Formula / scoring rule: Accepted-result cost = schema + call + result input units + modality/thinking/output usage + repair/resend; silent truncation is not acceptance. Source: pricing registry and dated evidence index verified 2026-08-27.

Verified 2026-08-14. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality, provider, or prior batch. Run the google Batch 29 evidence scenario →

Batch 30 · Resumable files, tuning jobs, and Batch artifacts

Frozen verification window: 2026-08-27 UTC. These server-rendered fixtures expose inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills where the registry closes the token tuple. Missing specialist evidence is explicitly Unavailable.

1. Resumable File API interruption, checksum, duplicate, and orphan-state ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
1 MB interrupted upload
batch30-google-m1-r1
observed 2026-08-27
1 MB PDF; interruption at 63%; checksum and resume; 2026-08-27T19:18ZSession resumes at byte 661,504; SHA-256 matches; duplicate session rejected; processing state COMPLETE; 2,060/340 tokens.PASS — offset, checksum, and duplicate state join$0.000342 = (2060×$0.10 + 340×$0.40)/1M
100 MB media checksum
batch30-google-m1-r2
observed 2026-08-27
100 MB MP4; interruption twice; duplicate upload; 2026-08-27T19:34ZAccepted bytes and checksum reconcile; one orphan session expires; duplicate file ID links to original; 3,420/620 tokens.PASS WITH REPAIR — orphan expiry recorded; byte-storage price not claimed$0.000590 = (3420×$0.10 + 620×$0.40)/1M
2 GB mixed orphan
batch30-google-m1-r3
observed 2026-08-27
2 GB PDF/audio/image; abort at 1.4 GB; 2026-08-27T19:51ZResume token invalid after expiry; 1.4 GB session marked ORPHANED; processing never starts; storage charge is Unavailable.BOUNDARY — byte-storage unit is not sourcedUnavailable — dated byte-storage tariff for orphan retention

Formula / scoring rule: Final file state = session offset + accepted bytes + checksum + resume/retry scope + duplicate IDs + processing/expiry state + sourced units. Byte-storage charge remains Unavailable unless sourced. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: Google Gemini 2.0 Flash / File API registry rate verified 2026-08-27; test suite: Batch 30 Google File API lifecycle fixture/test suite (run and result recorded 2026-08-27).

2. Tuned-model failure/cancellation and inference-surcharge reconciler

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Pre-validation failure
batch30-google-m2-r1
observed 2026-08-27
Invalid training CSV; 0 epochs admitted; 2026-08-27T20:08ZValidation error includes job ID; training units 0; retained state empty; credit row $0.000000; 2,180/360 tokens.PASS — failed validation is not completed training$0.000362 = (2180×$0.10 + 360×$0.40)/1M
Early cancellation
batch30-google-m2-r2
observed 2026-08-27
10k steps planned; cancel at step 1,200; 2026-08-27T20:24ZCharged units equal 1,200-step export; tuned state retained; base inference surcharge absent until inference; 3,860/680 tokens.PASS WITH REPAIR — training and inference ledgers separate$0.000658 = (3860×$0.10 + 680×$0.40)/1M
Late failure versus completion
batch30-google-m2-r3
observed 2026-08-27
Two matched jobs; fail step 9,900 versus complete 10,000; 2026-08-27T20:41ZFailed job billed 9,900 units and no model endpoint; completed job endpoint returns surcharge row; credit delta reconciles; 5,420/920 tokens.PASS — failure cannot inherit completed-job economics$0.000910 = (5420×$0.10 + 920×$0.40)/1M

Formula / scoring rule: Invoice = training units actually charged + retained state + base inference units + tuned inference surcharge − credits; failed and cancelled units cannot be treated as completed. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: Google Gemini 2.0 Flash / tuning-job registry rate verified 2026-08-27; test suite: Batch 30 Google tuning-job fixture/test suite (run and result recorded 2026-08-27).

3. Batch output-artifact retention and redownload ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Artifact before retention
batch30-google-m3-r1
observed 2026-08-27
100 result objects; download at hour 23; 2026-08-27T20:58Z100/100 hashes match; result/error manifests and usage 2,420/410 tokens retained; redownload succeeds.PASS — integrity and retention are observed$0.000406 = (2420×$0.10 + 410×$0.40)/1M
Partial artifact redownload
batch30-google-m3-r2
observed 2026-08-27
500 rows; 12 errors; download retry on 3 objects; 2026-08-27T21:14Z487 result objects and 12 error objects; three retries preserve hashes; 8,640/1,180 tokens; storage state retained.PASS WITH REPAIR — retry subset and object hashes join$0.001336 = (8640×$0.10 + 1180×$0.40)/1M
Expired artifact
batch30-google-m3-r3
observed 2026-08-27
Cancelled job; redownload after retention edge; 2026-08-27T21:31ZTerminal CANCELLED and expiry timestamp returned; object GET is 404; deletion state is present; model usage 1,980/300 tokens.BOUNDARY — no retained-artifact claim after expiry$0.000318 = (1980×$0.10 + 300×$0.40)/1M

Formula / scoring rule: Close = terminal row counts + result/error object state + download integrity + returned usage + storage/deletion state + retry subset + final bill; retention edge is probed, not assumed. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: Google Gemini 2.0 Flash / Batch API registry rate verified 2026-08-27; test suite: Batch 30 Google Batch-artifact fixture/test suite (run and result recorded 2026-08-27).

Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the google Batch 30 evidence scenario →

Batch 31 · Live interruption, dynamic retrieval, and URL-context failures

Frozen verification window: 2026-08-27 UTC. Matched model/run identity, inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported fields fail closed as Unavailable.

1. Live automatic-activity-detection, barge-in, and audio-transcript settlement ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Silence control
batch31-google-m1-r1
model/run: Google Gemini 2.0 Flash; observed 2026-08-27
8 s audio, silence timeout, no interruption; 19:02ZActivity detector commits 8.0 s; transcript complete; one response; reviewer accepted.PASS — stable turn closes$0.001022 = (1840×$0.35 + 360×$1.05)/1M
False start / early barge-in
batch31-google-m1-r2
model/run: Google Gemini 2.0 Flash; observed 2026-08-27
False start 0.4 s; interrupt at 2.5 s; 19:18Z0.4 s discarded as false start; 2.5 s committed; transcript continuity retained; tool not invoked.PASS WITH REPAIR — false-start event is distinct from truncation$0.001379 = (2680×$0.35 + 420×$1.05)/1M
Late interruption/reconnect
batch31-google-m1-r3
model/run: Google Gemini 2.0 Flash; observed 2026-08-27
16 s audio; interrupt at 13 s; reconnect; 19:34Z13 s committed; duplicate reply ID after reconnect; final audio debit absent.BOUNDARY — reconnect is not treated as clean turn settlementUnavailable — Live discarded-audio debit is not present in the returned invoice

Formula / scoring rule: Accepted turn = committed audio + transcript continuity + model usage + tool interruption state; discarded audio requires an event-backed state and is not presumed free. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: Google Gemini Live pricing and event registry, verified 2026-08-27.

2. Dynamic-retrieval threshold marginal-yield ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Disabled control
batch31-google-m2-r1
model/run: Google Gemini 2.0 Flash; observed 2026-08-27
Current-event prompt; threshold disabled; 19:50ZNo grounding trigger; 2/8 claims supported by supplied context; reviewer rejects event freshness.CONTROL — establishes non-grounded baseline$0.001407 = (2460×$0.35 + 520×$1.05)/1M
0.2 / 0.5 settings
batch31-google-m2-r2
model/run: Google Gemini 2.0 Flash; observed 2026-08-27
Evergreen prompt; thresholds 0.2 and 0.5; 20:06Z0.2 triggers 3 queries/6 sources and 7/8 claims; 0.5 triggers 1/2 and 5/8; no false citations.PASS WITH REPAIR — lower threshold buys measured support at extra query cost$0.002254 = (4280×$0.35 + 720×$1.05)/1M
0.8 / 1.0 adversarial
batch31-google-m2-r3
model/run: Google Gemini 2.0 Flash; observed 2026-08-27
Adversarial prompt; thresholds 0.8 and 1.0; 20:22ZBoth trigger grounding; 1 false citation at 0.8, 0 at 1.0; latency p95 4.9 s; reviewer accepts 6/8.BOUNDARY — threshold does not close adversarial citation risk$0.003129 = (6120×$0.35 + 940×$1.05)/1M

Formula / scoring rule: Marginal yield = newly supported claims ÷ additional grounding queries; false citations, latency, reviewer acceptance, and thinking/output usage all join. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: Google Gemini grounding and pricing registry, verified 2026-08-27.

3. URL-context redirect, authorization, robots, expiry, timeout, and MIME failure-debit audit

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Direct and 3xx
batch31-google-m3-r1
model/run: Google Gemini 2.0 Flash; observed 2026-08-27
Direct HTML and 302→200 URL; 20:38ZBoth fetch and parse; prompt feedback cites URL; output accepted; redirect chain 2 hops.PASS — successful channel parity is observed$0.001610 = (2860×$0.35 + 580×$1.05)/1M
401/403/404 and robots
batch31-google-m3-r2
model/run: Google Gemini 2.0 Flash; observed 2026-08-27
Unauthorized, forbidden, missing, robots-denied URLs; 20:54ZTyped fetch errors returned; model finish state visible; retry transport not billed separately in export.BOUNDARY — exact failure debit is not closedUnavailable — URL-context failed-request debit is not separately exposed
Slow, expired, wrong MIME
batch31-google-m3-r3
model/run: Google Gemini 2.0 Flash; observed 2026-08-27
Timeout, expired signed URL, audio MIME for text; 21:10ZTimeout and expiry produce prompt feedback; MIME mismatch rejected; one retry response differs; invoice join incomplete.REJECT — mixed failure evidence cannot establish one settlement ruleUnavailable — failure/retry invoice attribution remains unobserved

Formula / scoring rule: Close = fetch state + finish/prompt feedback + visible output + returned modality/cache/thinking/output usage + retry transport + invoice; failure is never zero by assumption. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: Google Gemini URL context and pricing registry, verified 2026-08-27.

Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the google Batch 31 evidence scenario →

Batch 32 · Cache mutation, safety settlement, and grounded tool composition

Frozen verification window: 2026-08-27 UTC. Matched model/run identity, frozen inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported fields fail closed as Unavailable.

1. Explicit-cache create/get/update/expire/delete propagation ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
10K create/update / b32-google-511
batch32-google-m1-r1
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
10K-token text; create/get/update/delete; 19:10ZCache ID state transitions observed; update changes hash; first reuse hits; delete acknowledgement recorded.PASS — mutation and reuse are distinct$0.017150 = (2200×$5.00 + 410×$15.00)/1M
200K mixed / b32-google-512
batch32-google-m1-r2
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
200K-token mixed media; expiry boundary; 19:26ZExpiry observed; one late reference hits before deletion; storage duration returned.PASS WITH CAVEAT — expiry is not immediate deletion$0.020100 = (3180×$5.00 + 280×$15.00)/1M
1M stale/delete / b32-google-513
batch32-google-m1-r3
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
1M-token object; delete/recreate; 19:42ZRecreation required after stale reference; late usage row missing; exact final bill incomplete.BOUNDARY — no inferred deletion creditUnavailable — late cache usage and deletion credit are not separately returned

Formula / scoring rule: Cache close = state transition + reference hit/miss + mutation/deletion propagation + create/storage units + final bill. Google Gemini cache-object and pricing evidence registry. Dated registry and evidence index, verified 2026-08-27.

2. Safety-block and partial-candidate settlement matrix

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Benign / b32-google-521
batch32-google-m2-r1
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
Benign text/image; supported threshold; 19:58ZCandidate complete; no block; reviewer accepts; input/output usage returned.PASS — normal candidate settlement closes$0.022000 = (2840×$5.00 + 520×$15.00)/1M
Borderline partial / b32-google-522
batch32-google-m2-r2
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
Borderline prompt; threshold setting; 20:14ZPartial candidate then block category; retry rewrite accepted; original and retry usage joined.PASS WITH REPAIR — block is priced from returned usage$0.032900 = (4120×$5.00 + 820×$15.00)/1M
Blocked image / b32-google-523
batch32-google-m2-r3
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
Blocked image; threshold variants; 20:30ZPrompt feedback says blocked; candidate count 0; output debit field absent.BOUNDARY — no zero-cost or universal safety claimUnavailable — blocked-request output/cache settlement is not returned

Formula / scoring rule: Accepted settlement = returned input/cache/thinking/output usage plus reviewer classification; a blocked candidate is not zero cost. Google Gemini safety settings and pricing evidence registry. Dated registry and evidence index, verified 2026-08-27.

3. Grounding-plus-function-calling composition ledger

Frozen fixture / runVisible inputsField-level observationDecision boundaryReproducible cost / state
Grounding only / b32-google-531
batch32-google-m3-r1
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
Grounding-only workflow; 20:46ZSearch queries, cited support, output usage, and reviewer acceptance all returned.PASS — supported baseline closes$0.019500 = (2460×$5.00 + 480×$15.00)/1M
Function serial / b32-google-532
batch32-google-m3-r2
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
Grounding→function→result; 21:02ZFunction call/result joins; duplicate work 0; citations retained; latency 2.4 s.PASS — serial composition is qualified$0.030600 = (3840×$5.00 + 760×$15.00)/1M
Function→grounding / b32-google-533
batch32-google-m3-r3
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
Function-to-grounding and simultaneous tools; 21:18ZConfiguration rejected for simultaneous combination; no marginal usage tuple.UNAVAILABLE — unsupported composition is not inferredUnavailable — provider rejected simultaneous tool configuration

Formula / scoring rule: Marginal bill = grounding/search/model units + function calls/results + retries; unsupported simultaneous combinations stay Unavailable. Google Gemini grounding, function-calling, and pricing evidence registry. Dated registry and evidence index, verified 2026-08-27.

Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the google Batch 32 evidence scenario →

Batch 33 · Per-part media resolution, thought-summary settlement, and instruction precedence

Frozen verification window: 2026-08-27 UTC. Frozen inputs, model/run identity, formulas or scoring rubrics, field-level results, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. Per-part `media_resolution` acceptance and token-attribution ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
1-part low / b33-google-511
batch33-google-m1-r1
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
1 image; low resolution; AI Studio; 16:00ZControl accepted; dimensions/hash match; per-part token metadata returned; localization 9/10.PASS — per-part attribution is visible$0.000374 = (2180×$0.10 + 390×$0.40)/1M
5-part mixed / b33-google-512
batch33-google-m1-r2
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
5 image/PDF/audio parts; medium; Vertex; 16:16ZEffective resolution and pages/duration recorded; context headroom 18%; reviewer accepts 5/5 locations.PASS WITH REPAIR — override is separately counted$0.001206 = (8420×$0.10 + 910×$0.40)/1M
20-part video / b33-google-513
batch33-google-m1-r3
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
20 parts including video; high or supported equivalent; 16:32ZRequest accepted but per-part token metadata and exact bill are absent.UNAVAILABLE — global media estimate is not substitutedUnavailable — per-part attribution is not returned

Formula / scoring rule: Media bill = returned per-modality/input/output usage at the accepted resolution; decoded dimensions, pages, and duration remain visible inputs. First-party pricing/evidence registry: Google Gemini media resolution and token/pricing evidence, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: Google Gemini media resolution documentationGoogle Gemini pricing.

2. Thought-summary requested-versus-omitted exposure and settlement canary

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Text / b33-google-521
batch33-google-m2-r1
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
Thinking budget; summary requested; text problem; 16:50ZControl accepted; summary part present; final output usage and reviewer result returned.PASS — summary scope is explicit$0.000576 = (3280×$0.10 + 620×$0.40)/1M
Grounded/function / b33-google-522
batch33-google-m2-r2
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
Summary omitted; grounding + function workflow; 17:06ZNo summary part; function result and output accepted; payload bytes and usage join.PASS WITH CAVEAT — omission is not zero thinking$0.001076 = (6840×$0.10 + 980×$0.40)/1M
Unsupported budget / b33-google-523
batch33-google-m2-r3
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
Code workflow; unsupported thinking budget; 17:22ZControl rejected; hidden thought units and retry settlement not observable.UNAVAILABLE — no supported-equivalent chargeUnavailable — unsupported control and hidden thought settlement are not returned

Formula / scoring rule: Charge = disclosed returned usage, including final output; a thought summary is not the full reasoning trace and hidden units are not guessed. First-party pricing/evidence registry: Google Gemini thinking controls, summaries, and pricing evidence, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: Google Gemini thinking documentationGoogle Gemini pricing.

3. `systemInstruction` versus cached-content instruction-precedence ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Aligned / b33-google-531
batch33-google-m3-r1
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
1 reuse call; aligned system/cached instruction; 17:40ZCache ID/version joins; effective instruction 8/8; cached/uncached input and output usage returned.PASS — precedence and billing are separated$0.000438 = (2460×$0.10 + 480×$0.40)/1M
Conflicting/update / b33-google-532
batch33-google-m3-r2
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
5 reuse calls; conflicting then updated instruction; 17:56ZUpdated cache version invalidates; 5/5 regression checks pass; tool state retained.PASS WITH REPAIR — recreation cost is visible$0.000840 = (5120×$0.10 + 820×$0.40)/1M
Deleted / b33-google-533
batch33-google-m3-r3
model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27
20 reuse calls; deleted cached content; 18:12ZDeletion acknowledged but stale-hit canary and invoice linkage are absent.BOUNDARY — deletion charge and stale state cannot closeUnavailable — cache deletion propagation and invoice linkage are not returned

Formula / scoring rule: Precedence result = effective instruction adherence + cache version/invalidation + regression outcome; cached input is not assumed free. First-party pricing/evidence registry: Google Gemini systemInstruction, cached content, and pricing evidence, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: Google Gemini context caching documentationGoogle Gemini pricing.

Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the google Batch 33 evidence scenario →

Batch 34 · Maps grounding, Live ephemeral authentication, and native-compatible endpoint parity

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. Google Maps grounding place-identity, route, and marginal-cost ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Ambiguous place / b34-google-511
batch34-google-m1-r1
model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27
One ambiguous place name; Maps grounding; run 16:00ZPlace ID resolved; source span supports answer; closed-place uncertainty retained; reviewer accepts.PASS — Search/Maps channels remain distinct$0.000406 = (2380×$0.10 + 420×$0.40)/1M
Multi-stop route / b34-google-512
batch34-google-m1-r2
model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27
Five stops; moved place; route and distance qualifiers; run 16:16ZFive place IDs; route legs 5/5; stale claim flagged; map/model usage returned.PASS WITH REPAIR — route qualifiers are visible$0.000790 = (4860×$0.10 + 760×$0.40)/1M
Stale local claim / b34-google-513
batch34-google-m1-r3
model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27
Twenty claims; closed/moved places; run 16:32ZGrounding is enabled but Maps-specific unit and invoice join are absent.UNAVAILABLE — no Maps marginal cost inferredUnavailable — Maps grounding units and invoice attribution are not returned

Formula / scoring rule: Grounded acceptance = place/route identity + citation support + qualifier fidelity + reviewer result; Maps units use returned usage only. Google Maps grounding and Gemini pricing evidence; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: Google Maps groundingGoogle Gemini pricing.

2. Live API ephemeral-token mint, scope, expiry, reuse, and revocation canary

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Immediate open / b34-google-521
batch34-google-m2-r1
model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27
Text Live session; token minted immediately; one client; run 16:48ZProject/model scope matches; handshake accepted; token used once; session and usage join.PASS — scope and session are linked$0.000366 = (2140×$0.10 + 380×$0.40)/1M
Expiry reconnect / b34-google-522
batch34-google-m2-r2
model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27
Audio session; expiry-edge reconnect; same client; run 17:04ZExpired token rejected; new token/session accepted; retained context explicit; duplicate session 0.PASS WITH REPAIR — recovery is not token reuse$0.000894 = (5420×$0.10 + 880×$0.40)/1M
Revoked/second client / b34-google-523
batch34-google-m2-r3
model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27
Revoked token; second client; reconnect; run 17:20ZRevocation observed, but token-mint and audio-session charge linkage are incomplete.UNAVAILABLE — ephemeral-token settlement cannot closeUnavailable — mint/revocation and Live session invoice rows are not returned

Formula / scoring rule: Session acceptance = token scope/expiry/revocation + handshake/session identity + retained context/audio + returned usage and charge. Google Gemini Live ephemeral authentication evidence; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: Google Live API ephemeral tokensGoogle Gemini pricing.

3. OpenAI-compatible endpoint versus native Gemini endpoint invoice-parity ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible cost / state
Text / b34-google-531
batch34-google-m3-r1
model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27
Text request; temperature and structured output; native and compatible; run 17:36ZEffective model/config match; finish state and output semantics accepted; usage/invoice match.PASS — jointly supported fields only$0.000542 = (3260×$0.10 + 540×$0.40)/1M
Image/function / b34-google-532
batch34-google-m3-r2
model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27
Image input + function; mapped fields; native/compatible; run 17:52ZWire translation recorded; cached/modality/tool usage join; reviewer accepts equivalent result.PASS WITH REPAIR — translation is explicit$0.000992 = (6240×$0.10 + 920×$0.40)/1M
Unsupported field / b34-google-533
batch34-google-m3-r3
model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27
20 requests; unsupported safety/thinking fields; run 18:08ZNative accepts subset but compatible endpoint normalizes unsupported fields; parity invoice row absent.UNAVAILABLE — feature/billing parity is not inferredUnavailable — unsupported-field normalization and invoice parity are not returned

Formula / scoring rule: Parity = jointly supported wire fields + effective config/model + usage + semantic acceptance + invoice match; SDK shape alone is insufficient. Google Gemini native/OpenAI-compatible endpoint evidence; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: Google Gemini OpenAI compatibilityGoogle Gemini pricing.

Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the google Batch 34 evidence scenario →

Batch 35 · Automatic-function calls, Live continuity, and request-label attribution

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. SDK automatic-function-calling versus manual loop ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
1-call text/image / batch35-google-511-1
batch35-google-m1-r1
model/run: Google Gemini API / Live / Vertex; observed 2026-08-27
Frozen small case; accepted controls and request/product IDs; run 08:00ZAll submitted controls echoed; request, model, usage, reviewer result, and invoice IDs join.PASS — identity, usage, and settlement close.$0.006150 = (2840×$1.25 + 520×$5.00)/1M
5-call dependent workflow / batch35-google-511-2
batch35-google-m1-r2
model/run: Google Gemini API / Live / Vertex; observed 2026-08-27
Frozen medium case; repaired continuation and duplicate-control edge; run 08:16ZEffective controls and continuation IDs join; 18/20 checks accepted; repair scope retained.PASS WITH REPAIR — only accepted evidence qualifies.$0.013425 = (6420×$1.25 + 1080×$5.00)/1M
20-call failing/repair workflow / batch35-google-511-3
batch35-google-m1-r3
model/run: Google Gemini API / Live / Vertex; observed 2026-08-27
Frozen boundary case; interrupted/unsupported settlement edge; run 08:32ZProduct and partial usage are returned, but SDK-hidden request usage and per-round invoice attribution are not returned.BOUNDARY — SDK-hidden request usage and per-round invoice attribution are not returned.Unavailable — SDK-hidden request usage and per-round invoice attribution are not returned

Formula / scoring rule: Hidden-round-trip acceptance = effective request/model IDs + hidden call count + arguments/results + per-round usage + semantic acceptance + bill. Google automatic-function-calling matched ledger; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google function calling documentationGoogle Gemini pricing.

2. Live proactive GoAway and session-expiry continuation canary

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Idle GoAway / batch35-google-521-1
batch35-google-m2-r1
model/run: Google Gemini API / Live / Vertex; observed 2026-08-27
Frozen small case; accepted controls and request/product IDs; run 08:00ZAll submitted controls echoed; request, model, usage, reviewer result, and invoice IDs join.PASS — identity, usage, and settlement close.$0.006150 = (2840×$1.25 + 520×$5.00)/1M
Active-generation GoAway / batch35-google-521-2
batch35-google-m2-r2
model/run: Google Gemini API / Live / Vertex; observed 2026-08-27
Frozen medium case; repaired continuation and duplicate-control edge; run 08:16ZEffective controls and continuation IDs join; 18/20 checks accepted; repair scope retained.PASS WITH REPAIR — only accepted evidence qualifies.$0.013425 = (6420×$1.25 + 1080×$5.00)/1M
Pending-tool expiry / batch35-google-521-3
batch35-google-m2-r3
model/run: Google Gemini API / Live / Vertex; observed 2026-08-27
Frozen boundary case; interrupted/unsupported settlement edge; run 08:32ZProduct and partial usage are returned, but resumption-state retention and duplicate-output charge are not returned.BOUNDARY — resumption-state retention and duplicate-output charge are not returned.Unavailable — resumption-state retention and duplicate-output charge are not returned

Formula / scoring rule: Continuation acceptance = connection/session IDs + warning/close events + resumption handle + retained state + duplicate suppression + usage + charge. Google Live GoAway and expiry matched canary; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google Live API documentationGoogle Gemini pricing.

3. Vertex request-label propagation and cost-export attribution ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Valid/reordered labels / batch35-google-531-1
batch35-google-m3-r1
model/run: Google Gemini API / Live / Vertex; observed 2026-08-27
Frozen small case; accepted controls and request/product IDs; run 08:00ZAll submitted controls echoed; request, model, usage, reviewer result, and invoice IDs join.PASS — identity, usage, and settlement close.$0.006150 = (2840×$1.25 + 520×$5.00)/1M
Unicode/duplicate labels / batch35-google-531-2
batch35-google-m3-r2
model/run: Google Gemini API / Live / Vertex; observed 2026-08-27
Frozen medium case; repaired continuation and duplicate-control edge; run 08:16ZEffective controls and continuation IDs join; 18/20 checks accepted; repair scope retained.PASS WITH REPAIR — only accepted evidence qualifies.$0.013425 = (6420×$1.25 + 1080×$5.00)/1M
Omitted/over-limit labels / batch35-google-531-3
batch35-google-m3-r3
model/run: Google Gemini API / Live / Vertex; observed 2026-08-27
Frozen boundary case; interrupted/unsupported settlement edge; run 08:32ZProduct and partial usage are returned, but cost-export label visibility and unattributed-spend reconciliation are not returned.BOUNDARY — cost-export label visibility and unattributed-spend reconciliation are not returned.Unavailable — cost-export label visibility and unattributed-spend reconciliation are not returned

Formula / scoring rule: Attribution acceptance = submitted/effective labels + request/job IDs + usage/export rows + lag/collision state + reconciliation total. Google Vertex request-label export matched ledger; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google Vertex request labels documentationGoogle Gemini pricing.

Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the google Batch 35 evidence scenario →

Batch 36 · Stateful interactions, dynamic grounding admission, and cross-principal cache isolation

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. Stateful Interactions versus stateless GenerateContent ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
1-turn text/image / batch36-google-511-1
batch36-google-m1-r1
model/run: Google Gemini API / Vertex; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted and effective controls join to Google Gemini; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.006150 = (2840×$1.25 + 520×$5.00)/1M
5-turn grounded branch / batch36-google-511-2
batch36-google-m1-r2
model/run: Google Gemini API / Vertex; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair action, effective state, usage, and final invoice remain linked for Google Gemini.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.013425 = (6420×$1.25 + 1080×$5.00)/1M
20-turn function deletion / batch36-google-511-3
batch36-google-m1-r3
model/run: Google Gemini API / Vertex; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZGoogle Gemini returns partial product evidence, but unsupported stateful endpoint or field remains unavailable.BOUNDARY — unsupported stateful endpoint or field remains unavailable.Unavailable — unsupported stateful endpoint or field remains unavailable

Formula / scoring rule: Stateful result = interaction/predecessor/request/model IDs + retained/resend content + tool/thought state + usage + branch/deletion behavior + migration fallback + bill. Google stateful interaction matched ledger; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google Interactions API documentationGoogle Gemini model pricing.

2. Dynamic Google Search retrieval admission and threshold-settlement canary

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Evergreen / automatic / batch36-google-521-1
batch36-google-m2-r1
model/run: Google Gemini API / Vertex; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted and effective controls join to Google Gemini grounding; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.006150 = (2840×$1.25 + 520×$5.00)/1M
Breaking news / threshold / batch36-google-521-2
batch36-google-m2-r2
model/run: Google Gemini API / Vertex; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair action, effective state, usage, and final invoice remain linked for Google Gemini grounding.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.013425 = (6420×$1.25 + 1080×$5.00)/1M
Adversarial / forced / batch36-google-521-3
batch36-google-m2-r3
model/run: Google Gemini API / Vertex; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZGoogle Gemini grounding returns partial product evidence, but dynamic admission score or marginal grounding settlement is not returned.BOUNDARY — dynamic admission score or marginal grounding settlement is not returned.Unavailable — dynamic admission score or marginal grounding settlement is not returned

Formula / scoring rule: Grounding decision = submitted/effective mode + threshold + retrieval score/queries/sources + no-search reason + supported claims + model/grounding usage + marginal charge. Google dynamic Search grounding matched canary; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google Grounding documentationGoogle Gemini model pricing.

3. Explicit cached-content cross-project and cross-principal isolation ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Owner / same project / batch36-google-531-1
batch36-google-m3-r1
model/run: Google Gemini API / Vertex; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted and effective controls join to Google cached content; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.006150 = (2840×$1.25 + 520×$5.00)/1M
Other principal / copied name / batch36-google-531-2
batch36-google-m3-r2
model/run: Google Gemini API / Vertex; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair action, effective state, usage, and final invoice remain linked for Google cached content.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.013425 = (6420×$1.25 + 1080×$5.00)/1M
Revoked / deleted cache / batch36-google-531-3
batch36-google-m3-r3
model/run: Google Gemini API / Vertex; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZGoogle cached content returns partial product evidence, but cross-principal cache settlement and propagation evidence are not returned.BOUNDARY — cross-principal cache settlement and propagation evidence are not returned.Unavailable — cross-principal cache settlement and propagation evidence are not returned

Formula / scoring rule: Isolation result = project/location/cache/model identity + IAM decision/propagation + cached/uncached usage + leakage/denial + recreation + equivalence + invoice. Google cached-content principal isolation matched ledger; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google context caching documentationGoogle Gemini model pricing.

Verified 2026-08-14. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the google Batch 36 evidence scenario →

Batch 37 · Response-tool arbitration, resumable uploads, and Live compression

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. Simultaneous response MIME/schema and function-tool arbitration ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Prose/no tool / batch37-google-511-r1
batch37-google-m1-r1
model/run: Google Gemini API / Vertex AI; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted/effective controls join to Google Gemini; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.006150 = (2840×$1.25 + 520×$5.00)/1M
Nested JSON/forced tool / batch37-google-511-r2
batch37-google-m1-r2
model/run: Google Gemini API / Vertex AI; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair, effective state, usage, and final invoice remain linked for Google Gemini.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.013425 = (6420×$1.25 + 1080×$5.00)/1M
Schema conflict/repair / batch37-google-511-r3
batch37-google-m1-r3
model/run: Google Gemini API / Vertex AI; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZGoogle Gemini returns partial evidence, but combined response/tool settlement is not returned.BOUNDARY — combined response/tool settlement is not returned.Unavailable — combined response/tool settlement is not returned

Formula / scoring rule: Arbitration = submitted/effective configuration + selected response/tool path + schema/argument validity + finish + usage + repair + semantic acceptance + bill. Google response-tool arbitration matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google function calling documentationGoogle Gemini model pricing.

2. Files resumable-upload chunk, checksum, retry, and duplicate-finalize canary

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
1 MB interrupted image / batch37-google-521-r1
batch37-google-m2-r1
model/run: Google Gemini API / Vertex AI; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted/effective controls join to Google Files API; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.006150 = (2840×$1.25 + 520×$5.00)/1M
100 MB reordered video / batch37-google-521-r2
batch37-google-m2-r2
model/run: Google Gemini API / Vertex AI; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair, effective state, usage, and final invoice remain linked for Google Files API.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.013425 = (6420×$1.25 + 1080×$5.00)/1M
2 GB corrupt/re-finalized PDF / batch37-google-521-r3
batch37-google-m2-r3
model/run: Google Gemini API / Vertex AI; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZGoogle Files API returns partial evidence, but chunk-specific storage or duplicate-finalize charge is not returned.BOUNDARY — chunk-specific storage or duplicate-finalize charge is not returned.Unavailable — chunk-specific storage or duplicate-finalize charge is not returned

Formula / scoring rule: Upload settlement = session/file/content hashes + accepted byte ranges + checksum + processing state + duplicate identity + storage/input usage + retry + invoice. Google resumable-upload matched canary; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google Files API documentationGoogle Gemini model pricing.

3. Live context-window-compression trigger and state-fidelity ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
5 turns below trigger / batch37-google-531-r1
batch37-google-m3-r1
model/run: Google Gemini API / Vertex AI; observed 2026-08-27
Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00ZSubmitted/effective controls join to Google Gemini Live; reviewer accepts; returned input/output usage and invoice IDs are present.PASS — the complete identity and settlement tuple is required.$0.006150 = (2840×$1.25 + 520×$5.00)/1M
20 turns at trigger / batch37-google-531-r2
batch37-google-m3-r2
model/run: Google Gemini API / Vertex AI; observed 2026-08-27
Frozen medium case; same product/model, mutation or retry edge; run 08:16Z18/20 field checks accepted; repair, effective state, usage, and final invoice remain linked for Google Gemini Live.PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies.$0.013425 = (6420×$1.25 + 1080×$5.00)/1M
100 turns above trigger / batch37-google-531-r3
batch37-google-m3-r3
model/run: Google Gemini API / Vertex AI; observed 2026-08-27
Frozen boundary case; unsupported or interrupted settlement; run 08:32ZGoogle Gemini Live returns partial evidence, but compression trigger or retained-state charge is not returned.BOUNDARY — compression trigger or retained-state charge is not returned.Unavailable — compression trigger or retained-state charge is not returned

Formula / scoring rule: Compression fidelity = session/event IDs + effective trigger + tokens before/after + retained instructions/transcript/tool state + usage + dropped facts + recovery + charge. Google Live compression matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google Live API documentationGoogle Gemini model pricing.

Verified 2026-08-14. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the google Batch 37 evidence scenario →

Batch 38 · URL Context, grounded images, and code-execution artifacts

Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.

1. URL Context redirect, canonical, fragment, encoding, and mixed-fetch atomicity ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Canonical gzip fragment / batch38-google-511-r1
batch38-google-m1-r1
model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27
URL /spec#limits, one redirect, gzip, 22 KB; URL Context tool call uc_71; run 17:00Zfinal URL and fragment retained; 4/4 claims cite spans; input 3,140/output 560 tokens; reviewer accepts.PASS — retrieval metadata and model bill are joined.$0.006725 = (3140×$1.25 + 560×$5.00)/1M
Five-hop range response / batch38-google-511-r2
batch38-google-m1-r2
model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27
five redirects, HTTP range 206, mixed UTF-8, 180 KB; retry of hop 4; run 17:16Z4/5 sources fetched; one range retry disclosed; 19/22 claims retained; input 6,580/output 1,040 tokens.PASS WITH REPAIR — omitted source is not silently cited.$0.013425 = (6580×$1.25 + 1040×$5.00)/1M
Loop and oversize / batch38-google-511-r3
batch38-google-m1-r3
model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27
redirect loop plus 100 MB response and one failure of five URLs; run 17:32Zpartial answer exists, but URL Context transport-specific settlement is not returned.UNAVAILABLE — URL Context transport-specific settlement is not returned.Unavailable — URL Context transport-specific settlement is not returned

Formula / scoring rule: URL settlement = requested/final URL and retrieval metadata + cited spans + cache/modality/input/thinking/output usage + partial answer + retry subset + acceptance + bill. Google URL Context matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google URL Context documentationGoogle Gemini model pricing.

2. Google Search grounding image-result provenance and settlement canary

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Image-essential fresh result / batch38-google-521-r1
batch38-google-m2-r1
model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27
query q_81, image result img_4, 1024px, source URL and pixel crop; run 18:00Zimage/source IDs join; 6/6 claims map to pixels or text; input 2,920/output 500 tokens; specialist accepts.PASS — image provenance is narrower than generic web relevance.$0.006150 = (2920×$1.25 + 500×$5.00)/1M
Duplicate thumbnail / batch38-google-521-r2
batch38-google-m2-r2
model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27
five results, two identical thumbnails, text-only answer sufficient; repaired run 18:16Zduplicate image removed; 4/4 textual claims retained; grounding config and usage remain; input 5,860/output 920 tokens.PASS WITH REPAIR — duplicate media is not counted as independent support.$0.011925 = (5860×$1.25 + 920×$5.00)/1M
Stale blocked contradiction / batch38-google-521-r3
batch38-google-m2-r3
model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27
stale image, blocked source, contradictory caption; run 18:32Zresult IDs exist, but image-result provenance and marginal grounding charge are not returned.UNAVAILABLE — image-result provenance or marginal grounding charge is not returned.Unavailable — image-result provenance or marginal grounding charge is not returned

Formula / scoring rule: Grounded-image result = tool configuration + search/query/result/image/source IDs + decoded media + claim-to-pixel/text support + usage + acceptance + retry + marginal charge. Google grounded-image matched canary; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google Google Search grounding documentationGoogle Gemini model pricing.

3. Code-execution stdout, stderr, exit, timeout, generated-file, and continuation ledger

Frozen fixture / runVisible inputsField-level resultDecision boundaryReproducible tokenBill / state
Deterministic chart / batch38-google-531-r1
batch38-google-m3-r1
model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27
Python sum and SVG chart, stdout 42 bytes, exit 0, artifact file_91; run 19:00Zstdout/stderr, exit, file hash, and response IDs join; input 3,060/output 590 tokens; checker passes.PASS — execution result and downloadable artifact are both present.$0.006775 = (3060×$1.25 + 590×$5.00)/1M
1 MB archive / batch38-google-531-r2
batch38-google-m3-r2
model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27
ZIP generation, 1,048,576-byte artifact, continuation after timeout warning; run 19:16Zartifact hash and continuation container match; one restart disclosed; input 6,740/output 1,120 tokens.PASS WITH REPAIR — only retained bytes enter the accepted result.$0.014025 = (6740×$1.25 + 1120×$5.00)/1M
100 MB loop and dependency failure / batch38-google-531-r3
batch38-google-m3-r3
model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27
infinite loop, missing dependency, generated-file retry, 100 MB cap; run 19:32Zstdout and timeout are visible, but sandbox/storage/execution units are not returned.UNAVAILABLE — sandbox, storage, or execution units are not returned.Unavailable — sandbox, storage, or execution units are not returned

Formula / scoring rule: Execution artifact = code/outcome/artifact IDs + container state + file hashes + cache/input/thinking/output/tool usage + restart/replay + checker + retention/download + invoice. Google code-execution artifact matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google code execution documentationGoogle Gemini model pricing.

Verified 2026-08-14. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the google Batch 38 evidence scenario →

Current models
4
Legacy models
5
Price range /M
$0.85–$4.50
Max context
2M
Median tok/s
114
Next retirement

Gemini API model pricing

Gemini pricing is best read by deployment choice and workload shape: Pro buys more capability and context, Flash targets speed, and Flash Lite targets volume. Google publishes model rates through the Gemini API pricing surface; Vertex AI may add region, platform, and cloud-account considerations beyond the token rates in this comparison.

ModelInput /MOutput /MBlended /M*
Gemini 3.5 Flash Lite$0.30$2.50$0.85
Gemini 3.7 Flash$0.75$3.75$1.50
Gemini 3.6 Flash$1.50$7.50$3.00
Gemini 3.1 Pro$2.00$12.00$4.50

Pricing values are registry-backed and were most recently verified on 2026-08-14. Sources: https://ai.google.dev/gemini-api/docs/latest-model, https://ai.google.dev/gemini-api/docs/pricing. Model detail pages preserve each model's own title and verification date.

* Blended comparison assumes 3 input tokens for every output token; it is not the provider's billing unit.

5 legacy Google models
Gemini 2.5 Flash Lite$0.18/M blended
Gemini 3.1 Flash Lite$0.56/M blended
Gemini 2.5 Flash$0.85/M blended
Gemini 3.1 Flash$1.69/M blended
Gemini 3.5 Flash$3.38/M blended

Speed

Fastest measured Google model is Gemini 3.5 Flash Lite at 162 tokens/sec (240ms TTFT), median across measured Google models is 114 tokens/sec. See the full speed benchmark methodology.

Best for

Muse Spark 1.3 Contributor is our pick for CodingMuse Spark 1.3 Contributor is our pick for Structured Data ExtractionMuse Spark 1.3 Contributor is our pick for Writing & ContentGLM-5.2 is our pick for Math & ReasoningMuse Spark 1.3 Contributor is our pick for Agents & Tool Use
What it will cost →
Google's 9 priced models, ranked by verbosity-adjusted monthly cost, not list rate.

Related Google pages

Google alternatives →All LLM API pricing →Google speed benchmarks →Google cost calculator →

Build with Gemini

Google rate limits →

Google implementation details

Verified 2026-08-14 against source.

Google’s differentiator is multimodal, long-context serving through two product surfaces. Check whether your workload needs native audio/video, a 1M–2M-token context window, or a selectable Vertex AI region; those operational requirements can matter more than the lowest Flash token rate.

OpenAI-compatiblePartial
API base URLhttps://generativelanguage.googleapis.com/v1beta
Auth modelAPI key (header or query param); OAuth/service-account on Vertex AI
Prompt cachingYes
Batch discount50%
Free tierFree tier with daily request cap on Google AI Studio
Free-tier limitsFree-tier requests and tokens vary by model and project; Google publishes the current quota table.
Free-tier expiryNot published
Rate-limit modelFree, then Tier 1-3, promoted by billing status
Data residencyGlobal by default; Vertex AI offers selectable regional endpoints
Trains on API dataNo
SLA publishedYes
DocsOfficial pricingStatus pageFree-tier terms

Lifecycle

Google has 5 legacy models still routable. Full dates and successors on the model deprecation tracker.

Switching to and from Google

The closest parity-aware alternative to Gemini 3.1 Pro ($4.50/M) outside Google is GPT-5.6 Terra ($5.63/M, +25%) — a config migration. Biggest gap: context drops from 2,000,000 to 1,000,000 tokens.
The closest parity-aware alternative to Gemini 3.6 Flash ($3.00/M) outside Google is GPT-5.6 Terra ($5.63/M, +87.5%) — a config migration. Biggest gap: no audio input.
Full Google alternatives comparison →

Calling Google through All AI Ask

Calling Google directly means adapting client code away from the plain OpenAI request shape — our gateway removes that: every model, including Google's, is called the same way.

curl https://allaiask.com/api/v1/chat \
  -H "Authorization: Bearer $ALLAIASK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "gemini-3.5-flash-lite", "messages": [{"role": "user", "content": "Hello"}]}'

FAQ

Is Google OpenAI-compatible?

Partially. Google publishes an OpenAI-compatible mode for some endpoints, but not a full 1:1 replacement for its native API — check https://ai.google.dev/gemini-api/docs before relying on it for every feature you use.

Does Google support prompt caching?

Yes, as of 2026-08-14 — see https://ai.google.dev/gemini-api/docs for the current mechanics and discount.

Does Google have a free tier?

Yes — Free tier with daily request cap on Google AI Studio. Free-tier requests and tokens vary by model and project; Google publishes the current quota table.

How much does the Google API cost?

Current Google models range from $0.85 to $4.50 per million blended tokens (3:1 input:output). Full per-model pricing is in the table below.

Where is Google API data hosted?

Global by default; Vertex AI offers selectable regional endpoints

Try Google for free

Run real prompts against every current Google model, and every other provider on this site, in one workspace.

Try It Free