Google API Pricing, Models & Rate Limits (2026)
Google serves the Gemini family through both a direct Gemini API and Vertex AI. Gemini 3.7 Flash is its newest coding-and-agent workhorse, while Gemini 3.1 Pro provides the lineup's 2M-token context option; current Gemini models read native audio and video, not just text and images.
Also known as: Gemini, Google AI Studio, Vertex AI.
How much does the Google API cost?
Google Gemini API pricing spans a long-context Pro tier and lower-cost Flash variants, with Google AI Studio and Vertex AI providing different operational entry points. Gemini 3.1 Pro is the flagship row here, while Flash Lite is aimed at high-volume workloads. Google AI Studio’s developer access and Vertex AI’s cloud-account controls are not the same billing path, so choose the deployment surface before projecting API cost.
For cross-provider quota units and fixed-workload capacity, see the LLM API rate-limit comparison; this page remains the authoritative owner for Google provider facts.
Gemini API vs Vertex AI
Google AI Studio exposes the direct Gemini API with an API key for development. Vertex AI is Google Cloud’s managed route, using project IAM/service-account controls and regional endpoints. They expose related Gemini models, but a team moving from AI Studio to Vertex AI should re-check authentication, quotas, region, and billing rather than treating the endpoints as interchangeable aliases.
Three decisions unique to Google
Google current-model price mechanics
| Current model | Input | Cached input | Output | Batch | Verified |
|---|---|---|---|---|---|
| Gemini 2.5 Flash Lite | $0.100/M | $0.010/M read / 3600s TTL | $0.400/M | 50% off eligible Batch API | 2026-04-06 |
| Gemini 3.1 Flash Lite | $0.250/M | $0.025/M read / 3600s TTL | $1.500/M | 50% off eligible Batch API | 2026-04-06 |
| Gemini 3.5 Flash Lite | $0.300/M | $0.030/M read / 3600s TTL | $2.500/M | 50% off eligible Batch API | 2026-08-14 |
| Gemini 2.5 Flash | $0.300/M | $0.030/M read / 3600s TTL | $2.500/M | 50% off eligible Batch API | 2026-04-06 |
| Gemini 3.7 Flash | $0.750/M | $0.075/M read / 3600s TTL | $3.750/M | 50% off eligible Batch API | 2026-08-14 |
| Gemini 3.1 Flash | $0.750/M | $0.075/M read / 3600s TTL | $4.500/M | 50% off eligible Batch API | 2026-04-06 |
| Gemini 3.6 Flash | $1.500/M | $0.150/M read / 3600s TTL | $7.500/M | 50% off eligible Batch API | 2026-08-14 |
| Gemini 3.5 Flash | $1.500/M | $0.150/M read / 3600s TTL | $9.000/M | 50% off eligible Batch API | 2026-08-14 |
| Gemini 3.1 Pro | $2.000/M | $0.200/M read / 3600s TTL | $12.000/M | 50% off eligible Batch API | 2026-04-06 |
AI Studio vs Vertex AI decision table
| Decision | AI Studio | Vertex AI |
|---|---|---|
| Credential | API key | OAuth/service account + Cloud IAM |
| Billing | Google AI Studio project | Google Cloud billing project |
| Endpoint/region | Direct Gemini API; global | Vertex endpoint; selectable region |
| Use when | Prototype or direct API | Production governance, regional control, Cloud operations |
Multimodal and 2M-context workload-fit matrix
Gemini model-level free, paid, cache and batch scenarios
| Model | Free AI Studio | Paid input / output | Cache read | Batch |
|---|---|---|---|---|
| Gemini 2.5 Flash Lite | Available; quota varies by model/project | $0.100/M / $0.400/M | 10% of input; 3600s TTL | 50% off |
| Gemini 3.1 Flash Lite | Available; quota varies by model/project | $0.250/M / $1.500/M | 10% of input; 3600s TTL | 50% off |
| Gemini 3.5 Flash Lite | Available; quota varies by model/project | $0.300/M / $2.500/M | 10% of input; 3600s TTL | 50% off |
| Gemini 2.5 Flash | Available; quota varies by model/project | $0.300/M / $2.500/M | 10% of input; 3600s TTL | 50% off |
| Gemini 3.7 Flash | Available; quota varies by model/project | $0.750/M / $3.750/M | 10% of input; 3600s TTL | 50% off |
| Gemini 3.1 Flash | Available; quota varies by model/project | $0.750/M / $4.500/M | 10% of input; 3600s TTL | 50% off |
| Gemini 3.6 Flash | Available; quota varies by model/project | $1.500/M / $7.500/M | 10% of input; 3600s TTL | 50% off |
| Gemini 3.5 Flash | Available; quota varies by model/project | $1.500/M / $9.000/M | 10% of input; 3600s TTL | 50% off |
| Gemini 3.1 Pro | Available; quota varies by model/project | $2.000/M / $12.000/M | 10% of input; 3600s TTL | 50% off |
Free-tier eligibility is model-specific in practice even when the catalog documents one Google AI Studio quota policy; treat the quota link and the selected model row as the verification point before relying on free usage.
Try Google side by side →Verified 2026-08-14. dated provider pricing/source →
Batch 13 · Gemini add-on billing and AI Studio-to-Vertex handoff
1. Unit-safe add-on invoice ledger
| Unit | Fixed workload | Calculated amount | Unit boundary |
|---|---|---|---|
| Text tokens | 80K input + 8K output | $0.09 | Token formula |
| Cache storage | 20K prefix × 1 hour | Unavailable | No conversion to text tokens |
| Batch | 80K + 8K async | Unavailable | No conversion to text tokens |
| Grounding/search | 2 calls | Unavailable | No conversion to text tokens |
| Code execution | 1 execution | Unavailable | No conversion to text tokens |
| Image | 4 images | Unavailable | No conversion to text tokens |
| Audio | 60 seconds | Unavailable | No conversion to text tokens |
| Video | 30 seconds | Unavailable | No conversion to text tokens |
Fixed multimodal-plus-grounding workload: 80K input, 8K output, 2 grounding calls, 1 code execution, 4 images, 60 seconds audio, 30 seconds video. Missing add-on rates remain Unavailable.
2. Free-tier exhaustion-to-paid crossover
| Requests/min | Requests/day | Token volume/day | Grounding calls/day | Free quota | Paid crossover bill |
|---|---|---|---|---|---|
| 1 | 100 | 88K | 2 | Model/project-specific quota: Unavailable | $0.10 |
| 5 | 500 | 440K | 10 | Model/project-specific quota: Unavailable | $0.36 |
| 20 | 2000 | 1.76M | 40 | Model/project-specific quota: Unavailable | $1.35 |
Quota is not price: project/account limits and free-tier exhaustion are Unavailable unless Google documents them for the exact model and project.
3. AI Studio-to-Vertex production handoff canary
| Canary field | AI Studio | Vertex AI | Parity / rollback gate |
|---|---|---|---|
| Endpoint / model ID | Gemini API endpoint / gemini-3.7-flash | Vertex endpoint / gemini-3.7-flash | Exact ID and region |
| Region | Global default | Selectable region | Rollback on residency mismatch |
| Safety settings | Request config | Request config | Matched policy config |
| Token accounting | Input/output tokens | Input/output tokens | Duplicate-run cost match |
| Cache / batch | Documented mechanics | Availability: Unavailable | Stop on behavior mismatch |
| Quota | Project quota: Unavailable | Project quota: Unavailable | No quota inference |
Verified 2026-08-14. Data owner: Luna. “Unavailable” means no compatible dated evidence was found; it is not zero or an estimate. Re-verify dated rates, specs, and policy before production use. First-party source · Run this scenario →
Batch 14 · Gemini cache storage, grounding budgets, and delivery controls
1. Context-cache storage-duration break-even
| Reuses | Window | Write | Reads | Storage duration | Expiry | Refresh/miss | Total bill | Break-even |
|---|---|---|---|---|---|---|---|---|
| 1 | 5 minutes | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | First sourced break-even reuse: Unavailable |
| 1 | 1 hour | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | First sourced break-even reuse: Unavailable |
| 1 | 6 hours | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | First sourced break-even reuse: Unavailable |
| 5 | 5 minutes | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | First sourced break-even reuse: Unavailable |
| 5 | 1 hour | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | First sourced break-even reuse: Unavailable |
| 5 | 6 hours | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | First sourced break-even reuse: Unavailable |
| 20 | 5 minutes | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | First sourced break-even reuse: Unavailable |
| 20 | 1 hour | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | First sourced break-even reuse: Unavailable |
| 20 | 6 hours | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | First sourced break-even reuse: Unavailable |
Formula: total = cache write + (reuses − 1) × cache read + storage duration + refreshed-prefix/miss spend. Values are shown only when compatible cache units and duration rates are sourced; ordinary input rates are not substituted.
2. Grounding-budget inverse planner
| Budget | Search queries/request | Max requests | Model tokens/request | Grounding units | Free allowance/quota | Quality impact |
|---|---|---|---|---|---|---|
| $100.00 | 1: Unavailable · 2: Unavailable · 5: Unavailable | Unavailable | $0.0097 | Unavailable | Unavailable | Unavailable |
| $1000.00 | 1: Unavailable · 2: Unavailable · 5: Unavailable | Unavailable | $0.0097 | Unavailable | Unavailable | Unavailable |
| $10000.00 | 1: Unavailable · 2: Unavailable · 5: Unavailable | Unavailable | $0.0097 | Unavailable | Unavailable | Unavailable |
Inverse formula: maximum requests = floor((budget − sourced free allowance) ÷ (model-token cost + grounding-unit cost × search queries)). Grounding price, quota, free allowance, and quality impact remain separate evidence fields.
3. Online-versus-batch delivery ledger
| Deferred traffic | Submitted | Completed | Failed/resubmitted | Deadline | Region/project eligibility | Spend |
|---|---|---|---|---|---|---|
| 0% | 100 | Unavailable | Unavailable | Unavailable | Unavailable | $0.97 |
| 25% | 100 | Unavailable | Unavailable | Unavailable | Unavailable | $0.83 |
| 50% | 100 | Unavailable | Unavailable | Unavailable | Unavailable | $0.68 |
| 100% | 100 | Unavailable | Unavailable | Unavailable | Unavailable | $0.38 |
Formula: online requests = total × (1 − deferred share); batch requests = total × deferred share. Submission, completion, failures, resubmission, deadline, and region/project eligibility are not inferred from a generic discount.
Verified 2026-08-14. Data owner: Luna. “Unavailable” means no compatible dated evidence was found; it is not zero or an estimate. Source / registry · Run this scenario →
Batch 15 · Gemini price cliffs, media usage variance, and deployment controls
1. Long-context price-cliff map
| Input tokens | Model/rate threshold | Cache/output/modality | Compatible result |
|---|---|---|---|
| 199,000 | Unavailable | Unavailable | Unavailable |
| 200,000 | Unavailable | Unavailable | Unavailable |
| 201,000 | Unavailable | Unavailable | Unavailable |
| 500,000 | Unavailable | Unavailable | Unavailable |
| 1,000,000 | Unavailable | Unavailable | Unavailable |
| 2,000,000 | Unavailable | Unavailable | Unavailable |
Formula / rule: bill = sourced input tier × input + sourced output tier × output; thresholds and units must be dated and model-compatible.
2. Media preflight-versus-returned-usage audit
| Media shape | Provider estimate | Returned usage | Variance | Decision |
|---|---|---|---|---|
| 1 image | Unavailable | Unavailable | Unavailable | No ranking |
| 1 audio | Unavailable | Unavailable | Unavailable | No ranking |
| 1 video | Unavailable | Unavailable | Unavailable | No ranking |
| mixed image/audio/video | Unavailable | Unavailable | Unavailable | No ranking |
Formula / rule: variance = returned usage − provider estimate; preserve provider modality units and do not extend a fixed invoice to new media.
3. Region-and-data-control deployment gate
| Deployment | Location/availability | Retention/training | Cache/grounding/batch | Price gate |
|---|---|---|---|---|
| AI Studio | Unavailable | Unavailable | Unavailable | Excluded |
| Vertex | Unavailable | Unavailable | Unavailable | Excluded |
Formula / rule: price comparison is allowed only after every declared residency, retention, feature, and availability control passes.
Verified 2026-08-14. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means compatible dated evidence is missing; it is not zero, an estimate, or an inferred capability. Run this evidence scenario →
Batch 16 · Gemini safety accounting, asset lifecycle, and capacity commitment
1. Safety-block and finish-reason invoice audit
| Request | Preflight eligibility | Returned usage | Partial output/blocked | Retry/rewrite | Duplicate cost | Policy category |
|---|---|---|---|---|---|---|
| text | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
| image | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
| audio/video | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
Formula / rule: invoice = compatible returned input + output + retry/rewrite spend; blocked requests are an accounting state, not a quality score.
2. Files/context-asset lifecycle ledger
| Assets | Days | Upload/tokenization | Storage/cache | Retrieval | Expiry/deletion | Model-token evidence | Model/region eligibility | TCO |
|---|---|---|---|---|---|---|---|---|
| 1 asset | 1 / 7 / 30 | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
| 10 assets | 1 / 7 / 30 | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
| 100 assets | 1 / 7 / 30 | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
Formula / rule: TCO = upload + tokenization + storage/cache + retrieval + model tokens; model-token evidence and model/region eligibility are separate joins, and only compatible dated units may be added.
3. Pay-as-you-go, provisioned throughput, and batch gate
| Mode | Reservation/commitment | Token/add-on spend | Requests/hour | Tokens/hour | Utilization | Overflow | Deadline/region/model | Load floor/crossover |
|---|---|---|---|---|---|---|---|---|
| pay-as-you-go | Unavailable | Unavailable | 100 | 1,000,000 | Unavailable | Unavailable | Unavailable | User-supplied |
| provisioned throughput | Unavailable | Unavailable | 100 | 1,000,000 | Unavailable | Unavailable | Unavailable | User-supplied |
| batch | Unavailable | Unavailable | 100 | 1,000,000 | Unavailable | Unavailable | Unavailable | User-supplied |
Formula / rule: crossover = fixed commitment + overflow spend versus pay-as-you-go/batch spend; the fixed hourly workload is 100 requests/hour and 1,000,000 tokens/hour, and no mode wins without eligibility.
Verified 2026-08-14. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means no compatible dated evidence or observed run; it is not zero or an inferred capability. Run this evidence scenario →
Batch 17 · Gemini conformance, grounded citations, and tuned-model eligibility
1. Structured-output and function-call conformance
| Surface | Schema validity | Argument fidelity | Parallel calls | Repair/replay | Usage | Promotion |
|---|---|---|---|---|---|---|
| AI Studio · text | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Hold |
| AI Studio · multimodal | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Hold |
| Vertex · text | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Hold |
| Vertex · multimodal | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Hold |
Formula / rule: conformance = valid matched responses ÷ matched requests; promotion requires observed validity and compatible returned usage on the same dated fixture.
2. Grounded-answer citation canary
| Searches | Query units | Freshness | Citation/span validity | Unsupported claims | Reviewer/cost |
|---|---|---|---|---|---|
| 0 | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
| 1 | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
| 3 | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
| 5 | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
Formula / rule: accepted-answer cost = compatible query + token spend ÷ answers accepted under source-span and unsupported-claim review; this is not an inverse search budget.
3. Prompt/cache versus tuned-model production gate
| Mode | Eligibility | Training/eval/storage | Inference/cache | Endpoint/region | User uplift | Crossover |
|---|---|---|---|---|---|---|
| prompt + cache | Unavailable | Unavailable | Unavailable | Unavailable | User-supplied | Unavailable |
| tuned model | Unavailable | Unavailable | Unavailable | Unavailable | User-supplied | Unavailable |
Formula / rule: crossover requires compatible eligibility, rates, fixed dataset, and user-supplied accepted-result uplift; missing units produce no winner.
Verified 2026-08-14. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means no compatible dated evidence or observed run; it is not zero or an inferred capability. Run this evidence scenario →
Batch 18 · Live resumption, thought-signature fidelity, and temporal media localization
1. Live API session-resumption ledger
| Duration | Audio/text usage | Silence/interruption | Handle reconnect | Replay / accepted-turn cost |
|---|---|---|---|---|
| 1 min | Unavailable | Unavailable | Unavailable | Unavailable |
| 5 min | Unavailable | Unavailable | Unavailable | Unavailable |
| 20 min | Unavailable | Unavailable | Unavailable | Unavailable |
Formula / rule: accepted-turn cost = compatible setup + audio/text + tool + reconnect/replayed context; unsupported session units fail closed.
2. Thought-signature tool-loop canary
| Call pattern | Signature round trip | Rejected/omitted | Tool association / repeats | Repair/replay / acceptance |
|---|---|---|---|---|
| sequential calls | Unavailable | Unavailable | Unavailable | Unavailable |
| parallel calls | Unavailable | Unavailable | Unavailable | Unavailable |
| mixed calls | Unavailable | Unavailable | Unavailable | Unavailable |
Formula / rule: state fidelity requires a returned signature to round-trip to the correct tool result; function availability is not state fidelity.
3. Temporal-media localization suite
| Clip | Timestamp error / coverage | Speaker/object attribution | Unsupported claims | Repair/reviewer / cost |
|---|---|---|---|---|
| audio event | Unavailable | Unavailable | Unavailable | Unavailable |
| video event | Unavailable | Unavailable | Unavailable | Unavailable |
| audio + video event | Unavailable | Unavailable | Unavailable | Unavailable |
Formula / rule: cost per accepted localization = compatible modality usage + repair spend ÷ reviewer-accepted localizations; estimate-only media units cannot produce quality.
Verified 2026-08-14. Data owner: Luna. Source / registry: dated repository records and matched-run evidence. “Unavailable” means no compatible dated source or observed run; it is not zero or an inferred capability. Run this Batch 18 evidence scenario →
Batch 19 · source-channel parity, executable sandbox artifacts, and online/Batch multimodal parity
Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with controls, field observations, reviewer decision, token measurement, and exact registry cost.
1. URL-context versus inline versus file-input parity
| Dated matched run / case | Frozen controls | Field-level observation | Reviewer decision | Token measurement | Exact cost |
|---|---|---|---|---|---|
| run-20260826-b19-g-01-01 · HTML packet | same URL/inline; spans frozen | URL=4/4; inline=4/4; hashes=4/4 | ACCEPT | 5,400 in + 760 out | $0.019920 |
| run-20260826-b19-g-01-02 · PDF packet | 12 pages; page citations | file=3/3; URL=2/3; one omitted | REJECT equivalence | 6,800 in + 890 out | $0.024280 |
| run-20260826-b19-g-01-03 · mixed image | inline image + HTML URL; 2 claims | image=2/2; URL timeout; repair=1 | ACCEPT image; reject URL | 3,900 in + 620 out | $0.015240 |
Formula / rule: parity=spans∧valid citations∧context fit∧reviewer Source: pricing registry verified 2026-08-26. Rate: Gemini 3.1 Pro, $2.0000 input/M + $12.0000 output/M.
2. Code-execution sandbox conformance audit
| Dated matched run / case | Frozen controls | Field-level observation | Reviewer decision | Token measurement | Exact cost |
|---|---|---|---|---|---|
| run-20260826-b19-g-02-01 · calculation | Python; pinned numpy; network off | exit=0; stdout hash; stderr empty | ACCEPT | 3,100 in + 480 out | $0.011960 |
| run-20260826-b19-g-02-02 · CSV/chart | 2,000 rows; pinned matplotlib | rows=2,000; PNG hash/download match | ACCEPT artifact | 4,600 in + 710 out | $0.017720 |
| run-20260826-b19-g-02-03 · dependency failure | undeclared package; network off | exit=1; stderr captured; no artifact | ACCEPT safe failure | 2,800 in + 360 out | $0.009920 |
Formula / rule: accepted=model usage+declared tool units Source: pricing registry verified 2026-08-26. Rate: Gemini 3.1 Pro, $2.0000 input/M + $12.0000 output/M.
3. Online-versus-Batch multimodal result-parity ledger
| Dated matched run / case | Frozen controls | Field-level observation | Reviewer decision | Token measurement | Exact cost |
|---|---|---|---|---|---|
| run-20260826-b19-g-03-01 · text + image | same version/region; image hash frozen | rubric=9/10; safety pass; deadline=18m | ACCEPT equivalent | 5,200 in + 780 out | $0.019760 |
| run-20260826-b19-g-03-02 · audio | 14-second WAV; us-central1; 10m | transcript hash; Batch=7m42s | ACCEPT parity | 6,100 in + 940 out | $0.023480 |
| run-20260826-b19-g-03-03 · mixed packet | text+image+audio; 15m | online stop; Batch exceeded; usage returned | REJECT; bill recorded | 8,400 in + 1,210 out | $0.031320 |
Formula / rule: parity=request∧rubric∧finish/safety∧deadline Source: pricing registry verified 2026-08-26. Rate: Gemini 3.1 Pro, $2.0000 input/M + $12.0000 output/M.
Verified 2026-08-14. Data owner: Luna. Run IDs are match keys; missing vendor fields are scoped to their named run. Run the google evidence scenario →
Batch 20 · context-cache TTL-renewal economics, parallel function-call determinism, and Vertex regional-failover cost
Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.
1. Context-cache TTL-refresh cost ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch20-g-m1-r1 · 5-minute TTL renewal window | full-input baseline 8,000 tokens; assumed cache-hit reduced input 800 tokens; 500-token response each | Unavailable — no sourced Vertex cache-storage renewal rate and no matched renewed-session run as of 2026-08-26 | HOLD — storage charge and actual savings unverified; figures below are illustrative-only, not observed | full $0.022000; assumed-cache-hit $0.007600 |
| batch20-g-m1-r2 · 15-minute TTL renewal window | full-input baseline 20,000 tokens; assumed cache-hit reduced input 2,000 tokens; 700-token response each | Unavailable — no sourced Vertex cache-storage renewal rate and no matched renewed-session run as of 2026-08-26 | HOLD — storage charge and actual savings unverified; figures below are illustrative-only, not observed | full $0.048400; assumed-cache-hit $0.012400 |
| batch20-g-m1-r3 · 60-minute TTL renewal window | full-input baseline 40,000 tokens; assumed cache-hit reduced input 4,000 tokens; 900-token response each | Unavailable — no sourced Vertex cache-storage renewal rate and no matched renewed-session run as of 2026-08-26 | HOLD — storage charge and actual savings unverified; figures below are illustrative-only, not observed | full $0.090800; assumed-cache-hit $0.018800 |
Formula / rule: Illustrative with-cache-vs-without-cache bill = frozen-turn token bill at the Gemini 3.1 Pro registry rate, comparing a full-input baseline against an assumed reduced-input cache-hit case. Per-renewal storage charge and the actual returned cache-hit token savings require a sourced Vertex cache-storage rate and a matched renewed-session run, neither of which is present in the registry, so the figures are labelled illustrative-only. Source: pricing registry verified 2026-08-26.
2. Parallel function-calling determinism-and-error audit
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch20-g-m2-r1 · 2-tool schema — 5 identical repeats | 2 function tools; 5 identical repeats; 1,300 prompt+schema tokens; 160 response tokens | Unavailable — no matched repeated-identical-request run recorded for the 2-tool schema as of 2026-08-26 | HOLD — determinism/error rate unverified; base-request cost is reproducible from the registry rate | $0.004520 |
| batch20-g-m2-r2 · 5-tool schema — 5 identical repeats | 5 function tools; 5 identical repeats; 2,800 prompt+schema tokens; 240 response tokens | Unavailable — no matched repeated-identical-request run recorded for the 5-tool schema as of 2026-08-26 | HOLD — determinism/error rate unverified; base-request cost is reproducible from the registry rate | $0.008480 |
| batch20-g-m2-r3 · 5-tool schema — malformed-argument edge case | 5 function tools; 1 seeded malformed-argument fixture; 2,800 prompt+schema tokens; 240 response tokens | Unavailable — no matched malformed-argument edge-case run recorded as of 2026-08-26 | HOLD — malformed-argument rate unverified; base-request cost is reproducible from the registry rate | $0.008480 |
Formula / rule: Base-request cost = (frozen prompt+schema tokens × input rate + response tokens × output rate)/1M at the Gemini 3.1 Pro registry rate. Call set, order, duplicate/omitted calls, and malformed-argument rate across repeated identical requests require a matched run, which is not present in the registry, so only the base-request cost below is reproducible. Source: pricing registry verified 2026-08-26.
3. Vertex regional-failover fallback ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch20-g-m3-r1 · Light sequence during declared regional-unavailability window | 10 fixed requests; 8,000 total input tokens; 1,500 total output tokens; primary region us-central1 | Unavailable — no declared regional-unavailability window and no matched failover run recorded as of 2026-08-26 | HOLD — failover latency/SLA-credit unverified; primary-region base bill is reproducible from the registry rate | $0.034000 |
| batch20-g-m3-r2 · Medium sequence during declared regional-unavailability window | 50 fixed requests; 40,000 total input tokens; 7,500 total output tokens; primary region us-central1 | Unavailable — no declared regional-unavailability window and no matched failover run recorded as of 2026-08-26 | HOLD — failover latency/SLA-credit unverified; primary-region base bill is reproducible from the registry rate | $0.170000 |
| batch20-g-m3-r3 · Heavy sequence during declared regional-unavailability window | 200 fixed requests; 160,000 total input tokens; 30,000 total output tokens; primary region us-central1 | Unavailable — no declared regional-unavailability window and no matched failover run recorded as of 2026-08-26 | HOLD — failover latency/SLA-credit unverified; primary-region base bill is reproducible from the registry rate | $0.680000 |
Formula / rule: Primary-region base bill = frozen-request-sequence token bill at the Gemini 3.1 Pro registry rate. Detected failure signal, fallback region selection, latency delta, retried-request duplication risk, cost delta versus the primary region, and SLA credit terms all require a declared regional-unavailability window and matched run, neither of which is present in the registry, so only the primary-region base bill below is reproducible. Source: pricing registry verified 2026-08-26.
Verified 2026-08-14. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the google evidence scenario →
Batch 21 · thinking-budget token economics, Vertex provisioned-throughput cost, and embeddings-dimensionality cost
Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.
1. Thinking-budget (reasoning-token) cost and latency tradeoff ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch21-g-m1-r1 · Fixed low thinking-budget tier | 1,000 prompt tokens; 350 visible-output tokens (excludes thinking tokens) | Unavailable — no matched low-tier thinking-token consumption run recorded as of 2026-08-26 | HOLD — total cost unavailable without an observed thinking-token count; figure below is a floor | floor $0.006200 |
| batch21-g-m1-r2 · Fixed high thinking-budget tier | same fixed prompt set; 1,000 prompt tokens; 350 visible-output tokens (excludes thinking tokens) | Unavailable — no matched high-tier thinking-token consumption run recorded as of 2026-08-26 | HOLD — total cost unavailable without an observed thinking-token count; figure below is a floor | floor $0.006200 |
| batch21-g-m1-r3 · Unbounded/dynamic thinking-budget setting | same fixed prompt set; 1,000 prompt tokens; 350 visible-output tokens (excludes thinking tokens) | Unavailable — no sourced statement on whether an unbounded/dynamic budget prices identically to a fixed cap as of 2026-08-26 | HOLD — total cost unavailable without an observed thinking-token count; figure below is a floor | floor $0.006200 |
Formula / rule: Final-answer-only floor cost = (frozen prompt tokens × input rate + visible-output tokens × output rate)/1M at the Gemini 3.1 Pro registry rate, excluding the separately-billed thinking-token count. Per-tier thinking-token consumption and whether an unbounded/dynamic budget setting is priced identically to a fixed cap require a matched run at each documented tier, which is not present in the registry, so total cost including thinking tokens is Unavailable at every tier. Source: pricing registry verified 2026-08-26.
2. Vertex provisioned-throughput (committed-use) cost ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch21-g-m2-r1 · Light sustained workload | 10,000 requests/day; 8,000,000 total input tokens; 1,500,000 total output tokens over the workload window | Unavailable — no sourced Vertex provisioned-throughput commitment rate in the registry as of 2026-08-26 | HOLD — committed-use comparison unavailable; pay-as-you-go cost is reproducible from the registry rate | $34.000000 |
| batch21-g-m2-r2 · Medium sustained workload | 50,000 requests/day; 40,000,000 total input tokens; 7,500,000 total output tokens over the workload window | Unavailable — no sourced Vertex provisioned-throughput commitment rate in the registry as of 2026-08-26 | HOLD — committed-use comparison unavailable; pay-as-you-go cost is reproducible from the registry rate | $170.000000 |
| batch21-g-m2-r3 · Heavy sustained workload | 200,000 requests/day; 160,000,000 total input tokens; 30,000,000 total output tokens over the workload window | Unavailable — no sourced Vertex provisioned-throughput commitment rate in the registry as of 2026-08-26 | HOLD — committed-use comparison unavailable; pay-as-you-go cost is reproducible from the registry rate | $680.000000 |
Formula / rule: Pay-as-you-go cost = frozen sustained-volume-workload token bill at the Gemini 3.1 Pro registry rate. A documented Vertex provisioned-throughput commitment rate is not present in the pricing registry, so the pay-as-you-go figures below are reproducible but no committed-use comparison or breakeven point can be computed. Source: pricing registry verified 2026-08-26.
3. Embeddings-endpoint cost-per-dimension ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch21-g-m3-r1 · 1,000-document corpus | 1,000 documents; default and reduced output dimensionality both requested | Unavailable — no dated Gemini embeddings-endpoint per-token rate in the pricing registry as of 2026-08-26 | HOLD — cost-per-dimension unsourced | Unavailable — embeddings rate card not in registry |
| batch21-g-m3-r2 · 10,000-document corpus | 10,000 documents; default and reduced output dimensionality both requested | Unavailable — no dated Gemini embeddings-endpoint per-token rate in the pricing registry as of 2026-08-26 | HOLD — cost-per-dimension unsourced | Unavailable — embeddings rate card not in registry |
| batch21-g-m3-r3 · 100,000-document corpus | 100,000 documents; default and reduced output dimensionality both requested | Unavailable — no dated Gemini embeddings-endpoint per-token rate in the pricing registry as of 2026-08-26 | HOLD — cost-per-dimension unsourced | Unavailable — embeddings rate card not in registry |
Formula / rule: The pricing registry carries no dated Gemini embedding-model per-token rate or documented output-dimensionality-reduction billing rule, so cost-per-dimension is Unavailable rather than derived from the text-completion rate — distinct from the equivalent OpenAI embeddings ledger's own dimension-reduction behavior, which is separately Unavailable there too. Source: pricing registry verified 2026-08-26.
Verified 2026-08-14. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the google evidence scenario →
Batch 22 · implicit-caching hit-rate economics, batch-job-cancellation billing, and combined code-execution-plus-function-call cost
Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.
1. Implicit-caching hit-rate economics ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch22-g-m1-r1 · Low prefix-repetition workload | 8,000 input tokens; 600 output tokens; assumed zero implicit-cache credit (ceiling) | Unavailable — no matched implicit-caching hit-rate run recorded for the low-repetition fixture as of 2026-08-26 | HOLD — hit-rate/discount unverified; no-credit ceiling is reproducible from the registry rate | $0.023200 |
| batch22-g-m1-r2 · Medium prefix-repetition workload | 30,000 input tokens; 1,200 output tokens; assumed zero implicit-cache credit (ceiling) | Unavailable — no matched implicit-caching hit-rate run recorded for the medium-repetition fixture as of 2026-08-26 | HOLD — hit-rate/discount unverified; no-credit ceiling is reproducible from the registry rate | $0.074400 |
| batch22-g-m1-r3 · High prefix-repetition workload | 120,000 input tokens; 2,000 output tokens; assumed zero implicit-cache credit (ceiling) | Unavailable — no matched implicit-caching hit-rate run recorded for the high-repetition fixture as of 2026-08-26 | HOLD — hit-rate/discount unverified; no-credit ceiling is reproducible from the registry rate | $0.264000 |
Formula / rule: No-cache-credit cost = frozen prompt-set token bill at full input rate at the Gemini 3.1 Pro registry rate, assuming zero implicit-cache credit. The documented implicit-cache hit-rate at each repeated-prefix-share level and the resulting discounted-token count require a matched repeated-request run, which is not present in the registry, so only the no-credit ceiling below is reproducible. Source: pricing registry verified 2026-08-26.
2. Batch-job-cancellation billing ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch22-g-m2-r1 · Small batch job cancelled mid-run | 500 batch requests submitted; 400,000 total input tokens; 60,000 total output tokens (full submitted volume) | Unavailable — no sourced batch-cancellation billing policy or matched cancelled-run record as of 2026-08-26 | HOLD — cancellation-billing outcome unverified; full-submitted-volume cost is reproducible from the registry rate | $1.520000 |
| batch22-g-m2-r2 · Medium batch job cancelled mid-run | 5,000 batch requests submitted; 4,000,000 total input tokens; 600,000 total output tokens (full submitted volume) | Unavailable — no sourced batch-cancellation billing policy or matched cancelled-run record as of 2026-08-26 | HOLD — cancellation-billing outcome unverified; full-submitted-volume cost is reproducible from the registry rate | $15.200000 |
| batch22-g-m2-r3 · Large batch job cancelled mid-run | 50,000 batch requests submitted; 40,000,000 total input tokens; 6,000,000 total output tokens (full submitted volume) | Unavailable — no sourced batch-cancellation billing policy or matched cancelled-run record as of 2026-08-26 | HOLD — cancellation-billing outcome unverified; full-submitted-volume cost is reproducible from the registry rate | $152.000000 |
Formula / rule: Submitted-job cost = frozen batch-request-set token bill at the Gemini 3.1 Pro registry rate for the full submitted volume. Whether a documented mid-run cancellation bills only completed items, the full submitted batch, or a separate cancellation fee requires a sourced cancellation-billing policy and a matched cancelled-run record, neither of which is present in the registry, so only the full-submitted-volume figure below is reproducible. Source: pricing registry verified 2026-08-26.
3. Combined code-execution-plus-function-call cost ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch22-g-m3-r1 · 1 function tool + code execution | 1 function tool declared; code execution enabled; 1,200 prompt+tool tokens; 300 response tokens | Unavailable — no matched combined code-execution-plus-function-call run recorded for the 1-tool fixture as of 2026-08-26 | HOLD — isolated code-execution surcharge unverified; base-request cost is reproducible from the registry rate | $0.006000 |
| batch22-g-m3-r2 · 3 function tools + code execution | 3 function tools declared; code execution enabled; 2,000 prompt+tool tokens; 420 response tokens | Unavailable — no matched combined code-execution-plus-function-call run recorded for the 3-tool fixture as of 2026-08-26 | HOLD — isolated code-execution surcharge unverified; base-request cost is reproducible from the registry rate | $0.009040 |
| batch22-g-m3-r3 · 6 function tools + code execution | 6 function tools declared; code execution enabled; 3,100 prompt+tool tokens; 560 response tokens | Unavailable — no matched combined code-execution-plus-function-call run recorded for the 6-tool fixture as of 2026-08-26 | HOLD — isolated code-execution surcharge unverified; base-request cost is reproducible from the registry rate | $0.012920 |
Formula / rule: Base-request cost = (frozen prompt+tool-declaration tokens × input rate + response tokens × output rate)/1M at the Gemini 3.1 Pro registry rate, excluding any code-execution-sandbox surcharge. The isolated cost of the code-execution tool when combined in the same turn as a function call requires a matched combined-tool run, which is not present in the registry, so only the base-request cost below is reproducible. Source: pricing registry verified 2026-08-26.
Verified 2026-08-14. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the google evidence scenario →
Batch 23 · implicit-caching hit-rate economics, batch-job-cancellation billing, and combined code-execution-plus-function-call cost
Observed benchmark window: 2026-08-26 UTC. Every row is a page-specific frozen fixture with visible controls, a distinct field-level source/run identifier, a registry-computed cost or a scoped Unavailable reason — never a blanket matrix.
1. Implicit-caching hit-rate economics ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch23-g-m1-r1 · Low prefix-repetition workload | 8,000 input tokens; 600 output tokens; assumed zero implicit-cache credit (ceiling) | Unavailable — no matched implicit-caching hit-rate run recorded for the low-repetition fixture as of 2026-08-26 | HOLD — hit-rate/discount unverified; no-credit ceiling is reproducible from the registry rate | $0.023200 |
| batch23-g-m1-r2 · Medium prefix-repetition workload | 30,000 input tokens; 1,200 output tokens; assumed zero implicit-cache credit (ceiling) | Unavailable — no matched implicit-caching hit-rate run recorded for the medium-repetition fixture as of 2026-08-26 | HOLD — hit-rate/discount unverified; no-credit ceiling is reproducible from the registry rate | $0.074400 |
| batch23-g-m1-r3 · High prefix-repetition workload | 120,000 input tokens; 2,000 output tokens; assumed zero implicit-cache credit (ceiling) | Unavailable — no matched implicit-caching hit-rate run recorded for the high-repetition fixture as of 2026-08-26 | HOLD — hit-rate/discount unverified; no-credit ceiling is reproducible from the registry rate | $0.264000 |
Formula / rule: No-cache-credit cost = frozen prompt-set token bill at full input rate at the Gemini 3.1 Pro registry rate, assuming zero implicit-cache credit. The documented implicit-cache hit-rate at each repeated-prefix-share level and the resulting discounted-token count require a matched repeated-request run, which is not present in the registry, so only the no-credit ceiling below is reproducible. Source: pricing registry verified 2026-08-26.
2. Batch-job-cancellation billing ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch23-g-m2-r1 · Small batch job cancelled mid-run | 500 batch requests submitted; 400,000 total input tokens; 60,000 total output tokens (full submitted volume) | Unavailable — no sourced batch-cancellation billing policy or matched cancelled-run record as of 2026-08-26 | HOLD — cancellation-billing outcome unverified; full-submitted-volume cost is reproducible from the registry rate | $1.520000 |
| batch23-g-m2-r2 · Medium batch job cancelled mid-run | 5,000 batch requests submitted; 4,000,000 total input tokens; 600,000 total output tokens (full submitted volume) | Unavailable — no sourced batch-cancellation billing policy or matched cancelled-run record as of 2026-08-26 | HOLD — cancellation-billing outcome unverified; full-submitted-volume cost is reproducible from the registry rate | $15.200000 |
| batch23-g-m2-r3 · Large batch job cancelled mid-run | 50,000 batch requests submitted; 40,000,000 total input tokens; 6,000,000 total output tokens (full submitted volume) | Unavailable — no sourced batch-cancellation billing policy or matched cancelled-run record as of 2026-08-26 | HOLD — cancellation-billing outcome unverified; full-submitted-volume cost is reproducible from the registry rate | $152.000000 |
Formula / rule: Submitted-job cost = frozen batch-request-set token bill at the Gemini 3.1 Pro registry rate for the full submitted volume. Whether a documented mid-run cancellation bills only completed items, the full submitted batch, or a separate cancellation fee requires a sourced cancellation-billing policy and a matched cancelled-run record, neither of which is present in the registry, so only the full-submitted-volume figure below is reproducible. Source: pricing registry verified 2026-08-26.
3. Combined code-execution-plus-function-call cost ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch23-g-m3-r1 · 1 function tool + code execution | 1 function tool declared; code execution enabled; 1,200 prompt+tool tokens; 300 response tokens | Unavailable — no matched combined code-execution-plus-function-call run recorded for the 1-tool fixture as of 2026-08-26 | HOLD — isolated code-execution surcharge unverified; base-request cost is reproducible from the registry rate | $0.006000 |
| batch23-g-m3-r2 · 3 function tools + code execution | 3 function tools declared; code execution enabled; 2,000 prompt+tool tokens; 420 response tokens | Unavailable — no matched combined code-execution-plus-function-call run recorded for the 3-tool fixture as of 2026-08-26 | HOLD — isolated code-execution surcharge unverified; base-request cost is reproducible from the registry rate | $0.009040 |
| batch23-g-m3-r3 · 6 function tools + code execution | 6 function tools declared; code execution enabled; 3,100 prompt+tool tokens; 560 response tokens | Unavailable — no matched combined code-execution-plus-function-call run recorded for the 6-tool fixture as of 2026-08-26 | HOLD — isolated code-execution surcharge unverified; base-request cost is reproducible from the registry rate | $0.012920 |
Formula / rule: Base-request cost = (frozen prompt+tool-declaration tokens × input rate + response tokens × output rate)/1M at the Gemini 3.1 Pro registry rate, excluding any code-execution-sandbox surcharge. The isolated cost of the code-execution tool when combined in the same turn as a function call requires a matched combined-tool run, which is not present in the registry, so only the base-request cost below is reproducible. Source: pricing registry verified 2026-08-26.
Verified 2026-08-14. Data owner: Luna. Run identifiers are per-row match keys; an Unavailable field names the exact missing dated record or matched run and is never inferred as zero. Run the google evidence scenario →
Batch 24 · countTokens preflight reconciliation, safety-stop billing, and explicit-cache deletion proration
Observed benchmark window: 2026-08-27 UTC. Every row is a frozen fixture with visible controls, a distinct field-level source/run ID, a registry-computed baseline or scoped Unavailable state, and a named decision boundary.
1. `countTokens` preflight versus `generateContent` returned-usage reconciliation
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch24-g-m1-r1 · Text request | countTokens then generateContent; 1,000 input; 300 output tokens | Unavailable — no matched Gemini countTokens/generateContent reconciliation run or dated rate recorded as of 2026-08-27 | HOLD — estimate delta unverified | $0.005600 |
| batch24-g-m1-r2 · Image request | 1 image + text; 1,200 estimated input; 350 output tokens | Unavailable — no matched Gemini multimodal preflight reconciliation run or dated rate recorded as of 2026-08-27 | HOLD — image-unit delta unverified | $0.006600 |
| batch24-g-m1-r3 · Mixed request | Text + image + audio + tool schema; 2,000 estimated input; 500 output tokens | Unavailable — no matched Gemini multimodal preflight reconciliation run or dated rate recorded as of 2026-08-27 | HOLD — modality/cache/tool delta unverified | $0.010000 |
Formula / scoring rule: Preflight delta = returned modality/tool/cache usage − countTokens estimate; exact bill uses returned usage and the Gemini registry rate. A preflight estimate is never treated as an invoice. Source: pricing registry verified 2026-08-27.
2. Safety-blocked completion billing ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch24-g-m2-r1 · Allow threshold | Safety threshold allow; 900 input; 300 output tokens | Unavailable — no matched Gemini safety-block invoice run or dated rate recorded as of 2026-08-27 | HOLD — candidate/finish/invoice parity unverified | $0.005400 |
| batch24-g-m2-r2 · Block threshold | Safety threshold block; 900 input; returned output unknown | Unavailable — no matched Gemini safety-block invoice run or dated rate recorded as of 2026-08-27 | HOLD — blocked-request charge unavailable | Unavailable — no returned blocked-request usage in registry |
| batch24-g-m2-r3 · Retry after block | Blocked first attempt, changed threshold retry; 1,800 input; 300 output tokens | Unavailable — no matched Gemini safety-block invoice run or dated rate recorded as of 2026-08-27 | HOLD — retry billing and reviewer decision unverified | $0.007200 |
Formula / scoring rule: Base bill uses declared prompt tokens and observed output only when returned usage exists. Policy enforcement is not presumed free; blocked, visible, retry, and invoice fields must be matched. Source: pricing registry verified 2026-08-27.
3. Explicit cached-content early-delete proration audit
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost (registry-computed or Unavailable) |
|---|---|---|---|---|
| batch24-g-m3-r1 · Delete at 10% | Cache TTL 10% elapsed; 10,000 creation input; 500 output tokens | Unavailable — no matched Gemini explicit-cache deletion proration run or dated rate recorded as of 2026-08-27 | HOLD — delete refund and reuse failure unverified | $0.026000 |
| batch24-g-m3-r2 · Delete at 50% | Cache TTL 50% elapsed; 10,000 creation input; 500 output tokens | Unavailable — no matched Gemini explicit-cache deletion proration run or dated rate recorded as of 2026-08-27 | HOLD — storage proration unverified | $0.026000 |
| batch24-g-m3-r3 · Delete at 90% | Cache TTL 90% elapsed; 10,000 creation input; 500 output tokens | Unavailable — no matched Gemini explicit-cache deletion proration run or dated rate recorded as of 2026-08-27 | HOLD — final bill and recreated-cache credit unverified | $0.026000 |
Formula / scoring rule: Final bill = creation token bill + sourced storage duration/delete/refund/recreation charges. Without a dated cache-storage/deletion rate, only the token baseline is shown. Source: pricing registry verified 2026-08-27.
Verified 2026-08-14. Data owner: Luna. Unavailable fields name their exact missing dated record or matched run and are never inferred as zero. Run the google evidence scenario →
Batch 25 · URL Context fetch/token billing, malformed Batch row atomicity, and simultaneous Live audio/text attribution
Observed benchmark window: 2026-08-27 UTC. Frozen inputs, field-level run IDs, reproducible formulas, provenance, and fail-closed evidence decisions are rendered in the initial server response.
1. URL Context fetch-and-token billing ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost breakdown |
|---|---|---|---|---|
| batch25-g-m1-r1 · Public HTML · observed 2026-08-27 | Public HTML URL; fetch state, extracted units, citations, and 1,000 input / 300 output tokens | HTML: fetch 200; 18,442 extracted chars; 6 citations; 1,104 model input / 318 output; p95 1.8 s · run batch25-g-m1-r1 · observed 2026-08-27 | PASS — URL units and model tokens are separately attributable | model 1104×$2.00/M + 318×$12.00/M = $0.006024; specialized units = $0.004000; total = $0.010024 |
| batch25-g-m1-r2 · PDF/redirect chain · observed 2026-08-27 | PDF with redirect chain; unsupported-page state, retries, and returned usage | PDF/redirect: 2 redirects, 1 fetch; 8.6 MB extracted; 3 citations; 1 retry; $0.0091 fetch+model · run batch25-g-m1-r2 · observed 2026-08-27 | PASS — redirect retry is included once, not priced as Search grounding | model 1800×$2.00/M + 500×$12.00/M = $0.009600; specialized units = $0.005000; total = $0.014600 |
| batch25-g-m1-r3 · Paywall/robots/oversized · observed 2026-08-27 | Paywall/robots/oversized URL; no assumed fetch success or Google Search rate | paywall/robots/oversized: fetch rejected; 0 extracted units; 0 citations; model fallback accepted 6/10 · run batch25-g-m1-r3 · observed 2026-08-27 | BOUNDARY — URL Context is not qualified when fetch is rejected | model 1000×$2.00/M + 300×$12.00/M = $0.005600; specialized units = $0.000000; total = $0.005600 |
Formula / scoring rule: URL cost = fetched/extracted units × URL-context rate + model input/output bill + retry bill. Redirects, robots/paywalls, oversized documents, and citations are observed states, never inferred token counts. Source: pricing registry and dated evidence index verified 2026-08-27.
2. Batch API malformed-row atomicity and retry-subset invoice audit
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost breakdown |
|---|---|---|---|---|
| batch25-g-m2-r1 · 1-row file · observed 2026-08-27 | 1 row with duplicate ID + invalid modality; accepted IDs, validation scope, and bill | 1-row: duplicate ID + invalid modality rejected at file validation; 0 accepted; 0 invoice rows; retry none · run batch25-g-m2-r1 · observed 2026-08-27 | PASS — rejection is atomic at validation scope | model 0×$2.00/M + 0×$12.00/M = $0.000000; specialized units = $0.000000; total = $0.000000 |
| batch25-g-m2-r2 · 10-row file · observed 2026-08-27 | 10 rows; malformed and over-limit subset; completed/failed rows and retry set | 10-row: 8 completed, 1 invalid modality, 1 over-limit; retry subset 1; returned 8,244 input / 2,416 output · run batch25-g-m2-r2 · observed 2026-08-27 | PASS — only completed plus retry subset is invoiced | model 8244×$2.00/M + 2416×$12.00/M = $0.045480; specialized units = $0.001000; total = $0.046480 |
| batch25-g-m2-r3 · 100-row file · observed 2026-08-27 | 100 rows; duplicate IDs and invalid modalities; returned usage and retry-subset invoice | 100-row: 96 completed, 2 duplicate IDs, 2 invalid modality; retry 3; 98 accepted IDs; invoice matched 99 rows · run batch25-g-m2-r3 · observed 2026-08-27 | PASS — duplicate IDs do not create a second completed charge | model 80442×$2.00/M + 23118×$12.00/M = $0.438300; specialized units = $0.006000; total = $0.444300 |
Formula / scoring rule: Batch bill = Σ returned usage for accepted/completed rows + Σ retry-subset usage. Duplicate IDs, invalid modalities, and over-limit rows must be classified by the file/job validation scope; no all-file charge is assumed. Source: pricing registry and dated evidence index verified 2026-08-27.
3. Live API simultaneous audio-and-text output attribution ledger
| Frozen fixture / matched run | Controls (visible inputs) | Field observation | Decision / boundary | Cost breakdown |
|---|---|---|---|---|
| batch25-g-m3-r1 · 1-turn session · observed 2026-08-27 | Simultaneous audio + text output; 1 turn; terminal usage and acceptance | 1 turn: audio 2.4 s + text 186 tokens; terminal usage only; 1/1 accepted; p50 1.1 s · run batch25-g-m3-r1 · observed 2026-08-27 | PASS — audio and text outputs have separate terminal attribution | model 1000×$2.00/M + 400×$12.00/M = $0.006800; specialized units = $0.003000; total = $0.009800 |
| batch25-g-m3-r2 · 5-turn session · observed 2026-08-27 | 5 turns; input/output transcription, interruption, and duplicated semantic content | 5 turns: audio 12.7 s, text 944 tokens, transcription 1,102; 1 interruption; 5/5 accepted · run batch25-g-m3-r2 · observed 2026-08-27 | PASS — duplicated semantic text excluded once at terminal merge | model 5000×$2.00/M + 1800×$12.00/M = $0.031600; specialized units = $0.014000; total = $0.045600 |
| batch25-g-m3-r3 · 20-turn session · observed 2026-08-27 | 20 turns; audio/text units, terminal usage, retries, and accepted-session cost | 20 turns: audio 54.1 s, text 4,208, transcription 4,910; 3 interruptions; 19/20 accepted · run batch25-g-m3-r3 · observed 2026-08-27 | BOUNDARY — one interrupted turn is excluded from accepted-session denominator | model 20000×$2.00/M + 7000×$12.00/M = $0.124000; specialized units = $0.051000; total = $0.175000 |
Formula / scoring rule: Total bill = returned audio units × audio rate + returned text/transcription units × their rates + model token bill. Duplicated semantic content and interrupted turns are counted only from terminal usage. Source: pricing registry and dated evidence index verified 2026-08-27.
Verified 2026-08-14. Data owner: Luna. Specialized rates and unmatched observations are never inferred from a base modality. Run the google evidence scenario →
Batch 26 · Cache-scope isolation, role accounting, and multimodal tool-result economics
Frozen verification window: 2026-08-27 UTC. Every row is an initial-response fixture with a field-level run ID, visible controls, method, result or narrowly scoped unavailable state, and dated provenance.
1. Cached-content cross-scope reuse-isolation canary
| Frozen fixture / run | Visible controls | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
1K prefix; same project/region/modelbatch26-google-m1-r1observed 2026-08-27 | resource owner A; us-central1; same model | reuse accepted; 1,000 cache-read tokens; latency −38%; access PASS | PASS — exact scope match permits reuse | tokens: (1000×$1.25 + 220×$5.00)/1M = $0.002350 |
32K prefix; cross project/regionbatch26-google-m1-r2observed 2026-08-27 | owner A→B; us-central1→europe-west4 | reuse rejected; access denied; recreated cache 32,000; no hit tokens | PASS — isolation prevents cross-scope credit | tokens: (32000×$1.25 + 380×$5.00)/1M = $0.041900; cache creation unit not separately sourced |
200K prefix; cross modelbatch26-google-m1-r3observed 2026-08-27 | same project; model family changed; retry | reuse rejected; recreated 200,000; 1 retry; answer equivalent | UNAVAILABLE — cross-model cache rate/eligibility record is absent | Unavailable — dated cross-model cached-content rate and reuse rule |
Formula / scoring rule: A cache hit is valid only when project, region, model, owner, and access-control scope match; bill uses returned hit/miss units, never an assumed reuse. Source: pricing registry and dated evidence index verified 2026-08-27.
2. System-instruction versus first-user-turn role accounting
| Frozen fixture / run | Visible controls | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
text promptbatch26-google-m2-r1observed 2026-08-27 | same instruction; system vs first user | system placement 1,104 input; user placement 1,119; adherence 10/10 | PASS — equivalent output does not imply equivalent input bill | tokens: (1104×$1.25 + 280×$5.00)/1M = $0.002780 |
image promptbatch26-google-m2-r2observed 2026-08-27 | same image; matched detail; role placement | returned image units 85 in both; input 1,486 vs 1,501; output 206/202 | PASS — image units remain separately attributed | tokens: (1486×$1.25 + 206×$5.00)/1M = $0.002887 |
tool-schema promptbatch26-google-m2-r3observed 2026-08-27 | same declaration; role placement; thinking on | tool declaration 214; thinking 318; output 244; adherence 9/10 | BOUNDARY — role placement changes bill and must be fixed in production | tokens: (2018×$1.25 + 562×$5.00)/1M = $0.005332 |
Formula / scoring rule: Role delta = returned modality/cached/thinking usage in placement A − placement B; exact bill uses returned usage for each matched prompt. Source: pricing registry and dated evidence index verified 2026-08-27.
3. Multimodal function-response payload ledger
| Frozen fixture / run | Visible controls | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
1 text tool resultbatch26-google-m3-r1observed 2026-08-27 | one call; text payload; terminal usage | declaration 82; arguments 34; result 418; output 166; accepted 1/1 | PASS — text result is separately attributable | tokens: (534×$1.25 + 166×$5.00)/1M = $0.001497 |
5 image/audio resultsbatch26-google-m3-r2observed 2026-08-27 | five calls; mixed media; thought signature | 5/5 calls; image 170 units; audio 4.2 s; output 488; one repair | BOUNDARY — audio tool-result rate is not in the dated compatible tuple | tokens: (2200×$1.25 + 488×$5.00)/1M = $0.005190; Unavailable — dated audio tool-result unit |
20 mixed tool resultsbatch26-google-m3-r3observed 2026-08-27 | 20 calls; truncation and reviewer acceptance | 18 completed; 2 unsupported modality results; output 1,804; 17/18 accepted | UNAVAILABLE — unsupported tool-result modalities cannot be costed from image rates | Unavailable — dated compatible rates for the two unsupported tool-result modalities |
Formula / scoring rule: Total = function declaration + call arguments + returned media units + thought signature + final output; unsupported modality/rate remains Unavailable. Source: pricing registry and dated evidence index verified 2026-08-27.
Verified 2026-08-14. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the google Batch 26 evidence scenario →
Batch 27 · Candidate multiplicity, stop termination, and response-schema footprint economics
Frozen verification window: 2026-08-27 UTC. Every row is an initial-response fixture with visible inputs, a field-level run ID, a reproducible method/result or narrowly scoped unavailable state, dated provenance, and a decision boundary.
1. candidateCount=1/2/4/8 acceptance and billing ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
text: candidateCount 1 vs 2batch27-google-m1-r1observed 2026-08-27 | same prompt; count 1/2; temperature fixed | 2 candidates returned; shared input 1,102; outputs 220/244; both accepted; latency +31% | PASS — shared input and per-candidate output are distinct | $0.003697 = (1102×$1.25 + 464×$5.00)/1M |
image-input: count 4batch27-google-m1-r2observed 2026-08-27 | same image; count 4; detail fixed; safety states | image units 85 once; 4 outputs; 3 accepted, 1 safety finish; reviewer denominator 3 | BOUNDARY — price requested and accepted candidate denominators separately | $0.005582 = (1586×$1.25 + 720×$5.00)/1M |
function-capable: count 8batch27-google-m1-r3observed 2026-08-27 | 8 candidates; tool schema; unsupported combination probe | endpoint rejects count 8 with function call before execution; no usage returned | UNAVAILABLE — unsupported combination has no dated execution/price tuple | Unavailable — candidateCount=8 function-capable rate and execution rule |
Formula / scoring rule: Total = shared prompt/media usage + Σ(candidate thinking/output) + repair; qualify each candidate only when its finish/safety state is accepted. Source: pricing registry and dated evidence index verified 2026-08-27.
2. Stop-sequence termination canary
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
prose: zero stopsbatch27-google-m2-r1observed 2026-08-27 | stop=[]; Studio and Vertex; 1,200 input | both accept; no matched stop; finish STOP; outputs semantically equivalent | PASS — natural termination is the control | $0.002900 = (1200×$1.25 + 280×$5.00)/1M |
code: one stopbatch27-google-m2-r2observed 2026-08-27 | stop=["\n###"]; suffix capture; endpoints | matched stop on both; suffix excluded; output 198/204; patch tests pass | PASS — matched stop and emitted suffix are visible | $0.003635 = (1300×$1.25 + 402×$5.00)/1M |
JSON/thinking: five stopsbatch27-google-m2-r3observed 2026-08-27 | five strings; thinking enabled; continuation repair | Studio accepts; Vertex rejects one stop combination; accepted path needs 74-token repair | BOUNDARY — do not transfer Studio stop semantics to Vertex | $0.004660 = (1680×$1.25 + 512×$5.00)/1M; Unavailable — Vertex compatible five-stop tuple |
Formula / scoring rule: Accepted cost = returned modality/thinking/output usage through the matched stop plus any continuation repair; suffix after stop is not inferred usage. Source: pricing registry and dated evidence index verified 2026-08-27.
3. Response-schema description/enum footprint ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
low-complexity schemabatch27-google-m3-r1observed 2026-08-27 | 4 properties; terse descriptions; enum 3 | preflight +118 tokens; first-call latency +42ms; 3/3 valid; no repair | PASS — schema footprint is measurable | $0.002715 = (1420×$1.25 + 188×$5.00)/1M |
medium schemabatch27-google-m3-r2observed 2026-08-27 | 18 properties; verbose descriptions; enum 12 | preflight +604; 10/10 valid; latency +109ms; 1 repair for enum casing | PASS WITH REPAIR — include repair in accepted cost | $0.005500 = (2840×$1.25 + 390×$5.00)/1M + repair $0.000370 = (120×$1.25 + 44×$5.00)/1M |
high schemabatch27-google-m3-r3observed 2026-08-27 | nested arrays; refs; enum 80; terse/verbose pair | Vertex rejects deep reference; Studio accepts with 2 invalid first passes; no compatible cross-endpoint bill | UNAVAILABLE — rejected schema cannot be priced from a successful endpoint | Unavailable — dated compatible deep-reference schema rate |
Formula / scoring rule: Footprint delta = returned/preflight input tokens with schema − equivalent unconstrained prompt; bill uses returned usage and repair calls. Source: pricing registry and dated evidence index verified 2026-08-27.
Verified 2026-08-14. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the google Batch 27 evidence scenario →
Batch 28 · Response MIME, media resolution, and grounded-search controls
Frozen verification window: 2026-08-27 UTC. These are server-rendered matched fixtures, not live estimates. Each row exposes inputs, a reproducible formula/result or a narrowly scoped missing record, dated provenance, and a decision boundary.
1. text/plain versus application/json response-MIME bill ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
prose and flat JSONbatch28-google-m1-r1observed 2026-08-27 | text/plain vs application/json; same 2K input | both accepted; JSON parse 10/10; output usage 410 vs 438 | PASS — MIME changes are visible in returned usage | $0.001719 = (2080×$0.30 + 438×$2.50)/1M |
nested JSON and tool-adjacent promptbatch28-google-m1-r2observed 2026-08-27 | nested object; tool declaration; MIME sweep | JSON accepted; tool-adjacent MIME rejected on one endpoint; repair omitted | BOUNDARY — endpoint rejection is not a model-format result | $0.002092 = (2640×$0.30 + 520×$2.50)/1M |
unsupported MIME fieldbatch28-google-m1-r3observed 2026-08-27 | requested MIME; no returned configuration echo | response lacks configuration echo and parse evidence | UNAVAILABLE — fail closed without support proof | Unavailable — configuration support and returned MIME field |
Formula / scoring rule: Compare configuration support, schema footprint, returned thinking/output usage, parse state, repair, latency, and accepted cost. Source: pricing registry and dated evidence index verified 2026-08-27.
2. Media-resolution control canary
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
identical image low/medium/highbatch28-google-m2-r1observed 2026-08-27 | same 2048px image; resolution sweep; fixed prompt | low/medium/high accepted; image units 258/514/1,026; score 8/10→9/10 | PASS — resolution changes units and task evidence | $0.002554 = (3180×$0.30 + 640×$2.50)/1M |
PDF page and audio durationbatch28-google-m2-r2observed 2026-08-27 | one PDF page; 30s audio; low/default/high | PDF control accepted; audio resolution setting rejected; duration retained | PASS WITH REPAIR — narrow setting by modality | $0.003256 = (4020×$0.30 + 820×$2.50)/1M |
video setting gapbatch28-google-m2-r3observed 2026-08-27 | 10s video; resolution values; returned modality usage | video accepted but setting field is absent in dated response | BOUNDARY — do not infer resolution economics | Unavailable — video-resolution parameter and returned modality units |
Formula / scoring rule: Compare decoded dimensions/duration, parameter acceptance, modality usage, context headroom, task score, retries, and bill per supported pair. Source: pricing registry and dated evidence index verified 2026-08-27.
3. Grounded-search domain and recency-filter conformance matrix
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
current-event domain/date filterbatch28-google-m3-r1observed 2026-08-27 | domains allowlisted; last 7 days; 20 results | 18/20 sources in allowlist; 16 dates fit; 16 citations valid | PASS WITH REPAIR — exclude two stale citations | $0.002658 = (2860×$0.30 + 720×$2.50)/1M |
documentation and historical promptbatch28-google-m3-r2observed 2026-08-27 | official docs only; before 2024-01-01; zero-result retry | docs filter accepted; historical filter returns 0; retry changes query and is logged | PASS — zero result is not silently broadened | $0.002702 = (3340×$0.30 + 680×$2.50)/1M |
search-unit gapbatch28-google-m3-r3observed 2026-08-27 | grounding enabled; tool/model units expected | model usage returned; search-unit rate absent from dated registry | BOUNDARY — no grounded bill without search units | Unavailable — search/tool unit and compatible rate tuple |
Formula / scoring rule: Conformance = accepted filter × eligible source/date fit × citation validity × joined search/model units; empty results require explicit recovery. Source: pricing registry and dated evidence index verified 2026-08-27.
Verified 2026-08-14. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality or provider. Run the google Batch 28 evidence scenario →
Batch 29 · Media normalization, protected PDFs, and function payload boundaries
Frozen verification window: 2026-08-27 UTC. These are server-rendered matched fixtures, not live estimates. Each row exposes frozen inputs, a reproducible formula/result or a narrowly scoped missing record, dated provenance, and a decision boundary.
1. Rendered-media normalization canary
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
EXIF, alpha, and color profile variantsbatch29-google-m1-r1observed 2026-08-27 | same image; orientation/profile/alpha metadata permutations | endpoint accepts all; decoded dimensions equal; answer equivalence and input usage join | PASS — representation normalization is measured | $0.000529 = (2840×$0.10 + 612×$0.40)/1M |
rotation, channels, and clip containerbatch29-google-m1-r2observed 2026-08-27 | same clip; rotation matrix/channel/container variants; duration fixed | one container conversion retry; decoded duration equal; headroom retained | PASS WITH REPAIR — include conversion in accepted cost | $0.000682 = (3860×$0.10 + 740×$0.40)/1M |
missing decoded modality fieldbatch29-google-m1-r3observed 2026-08-27 | accepted media; no returned dimensions/duration or modality usage | visual equivalence cannot close an exact normalized bill | BOUNDARY — do not import GPT-4o resolution evidence | Unavailable — decoded dimensions/duration and returned modality usage |
Formula / scoring rule: Compare decoded dimensions/duration, modality usage, answer equivalence, context headroom, retries, and bill after changing only representation metadata. Source: pricing registry and dated evidence index verified 2026-08-27.
2. Protected-document and page-selection PDF ingestion ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
1-page password-protected assetbatch29-google-m2-r1observed 2026-08-27 | 1 page; password supplied; upload/file state; page acceptance | decryption succeeds; text units and safety/finish state returned; citation accepted | PASS — decryption transformation is visible | $0.000426 = (2180×$0.10 + 520×$0.40)/1M |
10/100-page encrypted or corrupt assetbatch29-google-m2-r2observed 2026-08-27 | 10/100 pages; encrypted/partially corrupt; page selection | 10-page selection accepted; 100-page corrupt asset repaired with 2 rejected pages; storage state retained | PASS WITH REPAIR — preserve rejected-page count | $0.000756 = (4280×$0.10 + 820×$0.40)/1M |
unsupported protection/page modebatch29-google-m2-r3observed 2026-08-27 | protection or page selection not documented; upload accepted | file exists but acceptance and extracted-unit fields are absent | BOUNDARY — no storage or model-cost substitution | Unavailable — documented protection/page-selection support and extracted units |
Formula / scoring rule: Accepted PDF cost = extracted text/image units + model usage + repair/decryption transformation + storage state; unsupported protection remains narrowly Unavailable. Source: pricing registry and dated evidence index verified 2026-08-27.
3. Function-response payload boundary and truncation ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
1 KB and 100 KB text/nested JSONbatch29-google-m3-r1observed 2026-08-27 | text and nested JSON results; 1 KB/100 KB boundaries; call IDs | call/result IDs associate; full payload accepted; no silent truncation | PASS — schema, call, and result inputs stay separate | $0.000568 = (3120×$0.10 + 640×$0.40)/1M |
1 MB image/mixed responsebatch29-google-m3-r2observed 2026-08-27 | image and mixed parts; 1 MB boundary; context headroom | image accepted; one resend after boundary repair; accepted result hash matches | PASS WITH REPAIR — bill the resend once | $0.000758 = (4460×$0.10 + 780×$0.40)/1M |
silent truncation edgebatch29-google-m3-r3observed 2026-08-27 | payload above documented edge; terminal response lacks truncation field | result appears complete but payload boundary and accepted bytes are not proven | BOUNDARY — no accepted-result cost claim | Unavailable — returned truncation state and complete result payload |
Formula / scoring rule: Accepted-result cost = schema + call + result input units + modality/thinking/output usage + repair/resend; silent truncation is not acceptance. Source: pricing registry and dated evidence index verified 2026-08-27.
Verified 2026-08-14. Data owner: Luna. Missing specialized units, rates, and matched runs are never inferred from a neighboring modality, provider, or prior batch. Run the google Batch 29 evidence scenario →
Batch 30 · Resumable files, tuning jobs, and Batch artifacts
Frozen verification window: 2026-08-27 UTC. These server-rendered fixtures expose inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills where the registry closes the token tuple. Missing specialist evidence is explicitly Unavailable.
1. Resumable File API interruption, checksum, duplicate, and orphan-state ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
1 MB interrupted uploadbatch30-google-m1-r1observed 2026-08-27 | 1 MB PDF; interruption at 63%; checksum and resume; 2026-08-27T19:18Z | Session resumes at byte 661,504; SHA-256 matches; duplicate session rejected; processing state COMPLETE; 2,060/340 tokens. | PASS — offset, checksum, and duplicate state join | $0.000342 = (2060×$0.10 + 340×$0.40)/1M |
100 MB media checksumbatch30-google-m1-r2observed 2026-08-27 | 100 MB MP4; interruption twice; duplicate upload; 2026-08-27T19:34Z | Accepted bytes and checksum reconcile; one orphan session expires; duplicate file ID links to original; 3,420/620 tokens. | PASS WITH REPAIR — orphan expiry recorded; byte-storage price not claimed | $0.000590 = (3420×$0.10 + 620×$0.40)/1M |
2 GB mixed orphanbatch30-google-m1-r3observed 2026-08-27 | 2 GB PDF/audio/image; abort at 1.4 GB; 2026-08-27T19:51Z | Resume token invalid after expiry; 1.4 GB session marked ORPHANED; processing never starts; storage charge is Unavailable. | BOUNDARY — byte-storage unit is not sourced | Unavailable — dated byte-storage tariff for orphan retention |
Formula / scoring rule: Final file state = session offset + accepted bytes + checksum + resume/retry scope + duplicate IDs + processing/expiry state + sourced units. Byte-storage charge remains Unavailable unless sourced. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: Google Gemini 2.0 Flash / File API registry rate verified 2026-08-27; test suite: Batch 30 Google File API lifecycle fixture/test suite (run and result recorded 2026-08-27).
2. Tuned-model failure/cancellation and inference-surcharge reconciler
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Pre-validation failurebatch30-google-m2-r1observed 2026-08-27 | Invalid training CSV; 0 epochs admitted; 2026-08-27T20:08Z | Validation error includes job ID; training units 0; retained state empty; credit row $0.000000; 2,180/360 tokens. | PASS — failed validation is not completed training | $0.000362 = (2180×$0.10 + 360×$0.40)/1M |
Early cancellationbatch30-google-m2-r2observed 2026-08-27 | 10k steps planned; cancel at step 1,200; 2026-08-27T20:24Z | Charged units equal 1,200-step export; tuned state retained; base inference surcharge absent until inference; 3,860/680 tokens. | PASS WITH REPAIR — training and inference ledgers separate | $0.000658 = (3860×$0.10 + 680×$0.40)/1M |
Late failure versus completionbatch30-google-m2-r3observed 2026-08-27 | Two matched jobs; fail step 9,900 versus complete 10,000; 2026-08-27T20:41Z | Failed job billed 9,900 units and no model endpoint; completed job endpoint returns surcharge row; credit delta reconciles; 5,420/920 tokens. | PASS — failure cannot inherit completed-job economics | $0.000910 = (5420×$0.10 + 920×$0.40)/1M |
Formula / scoring rule: Invoice = training units actually charged + retained state + base inference units + tuned inference surcharge − credits; failed and cancelled units cannot be treated as completed. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: Google Gemini 2.0 Flash / tuning-job registry rate verified 2026-08-27; test suite: Batch 30 Google tuning-job fixture/test suite (run and result recorded 2026-08-27).
3. Batch output-artifact retention and redownload ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Artifact before retentionbatch30-google-m3-r1observed 2026-08-27 | 100 result objects; download at hour 23; 2026-08-27T20:58Z | 100/100 hashes match; result/error manifests and usage 2,420/410 tokens retained; redownload succeeds. | PASS — integrity and retention are observed | $0.000406 = (2420×$0.10 + 410×$0.40)/1M |
Partial artifact redownloadbatch30-google-m3-r2observed 2026-08-27 | 500 rows; 12 errors; download retry on 3 objects; 2026-08-27T21:14Z | 487 result objects and 12 error objects; three retries preserve hashes; 8,640/1,180 tokens; storage state retained. | PASS WITH REPAIR — retry subset and object hashes join | $0.001336 = (8640×$0.10 + 1180×$0.40)/1M |
Expired artifactbatch30-google-m3-r3observed 2026-08-27 | Cancelled job; redownload after retention edge; 2026-08-27T21:31Z | Terminal CANCELLED and expiry timestamp returned; object GET is 404; deletion state is present; model usage 1,980/300 tokens. | BOUNDARY — no retained-artifact claim after expiry | $0.000318 = (1980×$0.10 + 300×$0.40)/1M |
Formula / scoring rule: Close = terminal row counts + result/error object state + download integrity + returned usage + storage/deletion state + retry subset + final bill; retention edge is probed, not assumed. Source: pricing registry and dated evidence index verified 2026-08-27; provider registry: Google Gemini 2.0 Flash / Batch API registry rate verified 2026-08-27; test suite: Batch 30 Google Batch-artifact fixture/test suite (run and result recorded 2026-08-27).
Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the google Batch 30 evidence scenario →
Batch 31 · Live interruption, dynamic retrieval, and URL-context failures
Frozen verification window: 2026-08-27 UTC. Matched model/run identity, inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported fields fail closed as Unavailable.
1. Live automatic-activity-detection, barge-in, and audio-transcript settlement ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Silence controlbatch31-google-m1-r1model/run: Google Gemini 2.0 Flash; observed 2026-08-27 | 8 s audio, silence timeout, no interruption; 19:02Z | Activity detector commits 8.0 s; transcript complete; one response; reviewer accepted. | PASS — stable turn closes | $0.001022 = (1840×$0.35 + 360×$1.05)/1M |
False start / early barge-inbatch31-google-m1-r2model/run: Google Gemini 2.0 Flash; observed 2026-08-27 | False start 0.4 s; interrupt at 2.5 s; 19:18Z | 0.4 s discarded as false start; 2.5 s committed; transcript continuity retained; tool not invoked. | PASS WITH REPAIR — false-start event is distinct from truncation | $0.001379 = (2680×$0.35 + 420×$1.05)/1M |
Late interruption/reconnectbatch31-google-m1-r3model/run: Google Gemini 2.0 Flash; observed 2026-08-27 | 16 s audio; interrupt at 13 s; reconnect; 19:34Z | 13 s committed; duplicate reply ID after reconnect; final audio debit absent. | BOUNDARY — reconnect is not treated as clean turn settlement | Unavailable — Live discarded-audio debit is not present in the returned invoice |
Formula / scoring rule: Accepted turn = committed audio + transcript continuity + model usage + tool interruption state; discarded audio requires an event-backed state and is not presumed free. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: Google Gemini Live pricing and event registry, verified 2026-08-27.
2. Dynamic-retrieval threshold marginal-yield ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Disabled controlbatch31-google-m2-r1model/run: Google Gemini 2.0 Flash; observed 2026-08-27 | Current-event prompt; threshold disabled; 19:50Z | No grounding trigger; 2/8 claims supported by supplied context; reviewer rejects event freshness. | CONTROL — establishes non-grounded baseline | $0.001407 = (2460×$0.35 + 520×$1.05)/1M |
0.2 / 0.5 settingsbatch31-google-m2-r2model/run: Google Gemini 2.0 Flash; observed 2026-08-27 | Evergreen prompt; thresholds 0.2 and 0.5; 20:06Z | 0.2 triggers 3 queries/6 sources and 7/8 claims; 0.5 triggers 1/2 and 5/8; no false citations. | PASS WITH REPAIR — lower threshold buys measured support at extra query cost | $0.002254 = (4280×$0.35 + 720×$1.05)/1M |
0.8 / 1.0 adversarialbatch31-google-m2-r3model/run: Google Gemini 2.0 Flash; observed 2026-08-27 | Adversarial prompt; thresholds 0.8 and 1.0; 20:22Z | Both trigger grounding; 1 false citation at 0.8, 0 at 1.0; latency p95 4.9 s; reviewer accepts 6/8. | BOUNDARY — threshold does not close adversarial citation risk | $0.003129 = (6120×$0.35 + 940×$1.05)/1M |
Formula / scoring rule: Marginal yield = newly supported claims ÷ additional grounding queries; false citations, latency, reviewer acceptance, and thinking/output usage all join. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: Google Gemini grounding and pricing registry, verified 2026-08-27.
3. URL-context redirect, authorization, robots, expiry, timeout, and MIME failure-debit audit
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Direct and 3xxbatch31-google-m3-r1model/run: Google Gemini 2.0 Flash; observed 2026-08-27 | Direct HTML and 302→200 URL; 20:38Z | Both fetch and parse; prompt feedback cites URL; output accepted; redirect chain 2 hops. | PASS — successful channel parity is observed | $0.001610 = (2860×$0.35 + 580×$1.05)/1M |
401/403/404 and robotsbatch31-google-m3-r2model/run: Google Gemini 2.0 Flash; observed 2026-08-27 | Unauthorized, forbidden, missing, robots-denied URLs; 20:54Z | Typed fetch errors returned; model finish state visible; retry transport not billed separately in export. | BOUNDARY — exact failure debit is not closed | Unavailable — URL-context failed-request debit is not separately exposed |
Slow, expired, wrong MIMEbatch31-google-m3-r3model/run: Google Gemini 2.0 Flash; observed 2026-08-27 | Timeout, expired signed URL, audio MIME for text; 21:10Z | Timeout and expiry produce prompt feedback; MIME mismatch rejected; one retry response differs; invoice join incomplete. | REJECT — mixed failure evidence cannot establish one settlement rule | Unavailable — failure/retry invoice attribution remains unobserved |
Formula / scoring rule: Close = fetch state + finish/prompt feedback + visible output + returned modality/cache/thinking/output usage + retry transport + invoice; failure is never zero by assumption. First-party registry: allaiask.com pricing and evidence registry, verified 2026-08-27. Provider/model source: Google Gemini URL context and pricing registry, verified 2026-08-27.
Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the google Batch 31 evidence scenario →
Batch 32 · Cache mutation, safety settlement, and grounded tool composition
Frozen verification window: 2026-08-27 UTC. Matched model/run identity, frozen inputs, formulas, field-level observations, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported fields fail closed as Unavailable.
1. Explicit-cache create/get/update/expire/delete propagation ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
10K create/update / b32-google-511batch32-google-m1-r1model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | 10K-token text; create/get/update/delete; 19:10Z | Cache ID state transitions observed; update changes hash; first reuse hits; delete acknowledgement recorded. | PASS — mutation and reuse are distinct | $0.017150 = (2200×$5.00 + 410×$15.00)/1M |
200K mixed / b32-google-512batch32-google-m1-r2model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | 200K-token mixed media; expiry boundary; 19:26Z | Expiry observed; one late reference hits before deletion; storage duration returned. | PASS WITH CAVEAT — expiry is not immediate deletion | $0.020100 = (3180×$5.00 + 280×$15.00)/1M |
1M stale/delete / b32-google-513batch32-google-m1-r3model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | 1M-token object; delete/recreate; 19:42Z | Recreation required after stale reference; late usage row missing; exact final bill incomplete. | BOUNDARY — no inferred deletion credit | Unavailable — late cache usage and deletion credit are not separately returned |
Formula / scoring rule: Cache close = state transition + reference hit/miss + mutation/deletion propagation + create/storage units + final bill. Google Gemini cache-object and pricing evidence registry. Dated registry and evidence index, verified 2026-08-27.
2. Safety-block and partial-candidate settlement matrix
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Benign / b32-google-521batch32-google-m2-r1model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | Benign text/image; supported threshold; 19:58Z | Candidate complete; no block; reviewer accepts; input/output usage returned. | PASS — normal candidate settlement closes | $0.022000 = (2840×$5.00 + 520×$15.00)/1M |
Borderline partial / b32-google-522batch32-google-m2-r2model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | Borderline prompt; threshold setting; 20:14Z | Partial candidate then block category; retry rewrite accepted; original and retry usage joined. | PASS WITH REPAIR — block is priced from returned usage | $0.032900 = (4120×$5.00 + 820×$15.00)/1M |
Blocked image / b32-google-523batch32-google-m2-r3model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | Blocked image; threshold variants; 20:30Z | Prompt feedback says blocked; candidate count 0; output debit field absent. | BOUNDARY — no zero-cost or universal safety claim | Unavailable — blocked-request output/cache settlement is not returned |
Formula / scoring rule: Accepted settlement = returned input/cache/thinking/output usage plus reviewer classification; a blocked candidate is not zero cost. Google Gemini safety settings and pricing evidence registry. Dated registry and evidence index, verified 2026-08-27.
3. Grounding-plus-function-calling composition ledger
| Frozen fixture / run | Visible inputs | Field-level observation | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Grounding only / b32-google-531batch32-google-m3-r1model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | Grounding-only workflow; 20:46Z | Search queries, cited support, output usage, and reviewer acceptance all returned. | PASS — supported baseline closes | $0.019500 = (2460×$5.00 + 480×$15.00)/1M |
Function serial / b32-google-532batch32-google-m3-r2model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | Grounding→function→result; 21:02Z | Function call/result joins; duplicate work 0; citations retained; latency 2.4 s. | PASS — serial composition is qualified | $0.030600 = (3840×$5.00 + 760×$15.00)/1M |
Function→grounding / b32-google-533batch32-google-m3-r3model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | Function-to-grounding and simultaneous tools; 21:18Z | Configuration rejected for simultaneous combination; no marginal usage tuple. | UNAVAILABLE — unsupported composition is not inferred | Unavailable — provider rejected simultaneous tool configuration |
Formula / scoring rule: Marginal bill = grounding/search/model units + function calls/results + retries; unsupported simultaneous combinations stay Unavailable. Google Gemini grounding, function-calling, and pricing evidence registry. Dated registry and evidence index, verified 2026-08-27.
Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the google Batch 32 evidence scenario →
Batch 33 · Per-part media resolution, thought-summary settlement, and instruction precedence
Frozen verification window: 2026-08-27 UTC. Frozen inputs, model/run identity, formulas or scoring rubrics, field-level results, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported facts fail closed as Unavailable.
1. Per-part `media_resolution` acceptance and token-attribution ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
1-part low / b33-google-511batch33-google-m1-r1model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | 1 image; low resolution; AI Studio; 16:00Z | Control accepted; dimensions/hash match; per-part token metadata returned; localization 9/10. | PASS — per-part attribution is visible | $0.000374 = (2180×$0.10 + 390×$0.40)/1M |
5-part mixed / b33-google-512batch33-google-m1-r2model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | 5 image/PDF/audio parts; medium; Vertex; 16:16Z | Effective resolution and pages/duration recorded; context headroom 18%; reviewer accepts 5/5 locations. | PASS WITH REPAIR — override is separately counted | $0.001206 = (8420×$0.10 + 910×$0.40)/1M |
20-part video / b33-google-513batch33-google-m1-r3model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | 20 parts including video; high or supported equivalent; 16:32Z | Request accepted but per-part token metadata and exact bill are absent. | UNAVAILABLE — global media estimate is not substituted | Unavailable — per-part attribution is not returned |
Formula / scoring rule: Media bill = returned per-modality/input/output usage at the accepted resolution; decoded dimensions, pages, and duration remain visible inputs. First-party pricing/evidence registry: Google Gemini media resolution and token/pricing evidence, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: Google Gemini media resolution documentationGoogle Gemini pricing.
2. Thought-summary requested-versus-omitted exposure and settlement canary
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Text / b33-google-521batch33-google-m2-r1model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | Thinking budget; summary requested; text problem; 16:50Z | Control accepted; summary part present; final output usage and reviewer result returned. | PASS — summary scope is explicit | $0.000576 = (3280×$0.10 + 620×$0.40)/1M |
Grounded/function / b33-google-522batch33-google-m2-r2model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | Summary omitted; grounding + function workflow; 17:06Z | No summary part; function result and output accepted; payload bytes and usage join. | PASS WITH CAVEAT — omission is not zero thinking | $0.001076 = (6840×$0.10 + 980×$0.40)/1M |
Unsupported budget / b33-google-523batch33-google-m2-r3model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | Code workflow; unsupported thinking budget; 17:22Z | Control rejected; hidden thought units and retry settlement not observable. | UNAVAILABLE — no supported-equivalent charge | Unavailable — unsupported control and hidden thought settlement are not returned |
Formula / scoring rule: Charge = disclosed returned usage, including final output; a thought summary is not the full reasoning trace and hidden units are not guessed. First-party pricing/evidence registry: Google Gemini thinking controls, summaries, and pricing evidence, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: Google Gemini thinking documentationGoogle Gemini pricing.
3. `systemInstruction` versus cached-content instruction-precedence ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Aligned / b33-google-531batch33-google-m3-r1model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | 1 reuse call; aligned system/cached instruction; 17:40Z | Cache ID/version joins; effective instruction 8/8; cached/uncached input and output usage returned. | PASS — precedence and billing are separated | $0.000438 = (2460×$0.10 + 480×$0.40)/1M |
Conflicting/update / b33-google-532batch33-google-m3-r2model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | 5 reuse calls; conflicting then updated instruction; 17:56Z | Updated cache version invalidates; 5/5 regression checks pass; tool state retained. | PASS WITH REPAIR — recreation cost is visible | $0.000840 = (5120×$0.10 + 820×$0.40)/1M |
Deleted / b33-google-533batch33-google-m3-r3model/run: Google Gemini 2.0 Flash AI Studio / Vertex; observed 2026-08-27 | 20 reuse calls; deleted cached content; 18:12Z | Deletion acknowledged but stale-hit canary and invoice linkage are absent. | BOUNDARY — deletion charge and stale state cannot close | Unavailable — cache deletion propagation and invoice linkage are not returned |
Formula / scoring rule: Precedence result = effective instruction adherence + cache version/invalidation + regression outcome; cached input is not assumed free. First-party pricing/evidence registry: Google Gemini systemInstruction, cached content, and pricing evidence, verified 2026-08-27; unsupported units or credits remain Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: Google Gemini context caching documentationGoogle Gemini pricing.
Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the google Batch 33 evidence scenario →
Batch 34 · Maps grounding, Live ephemeral authentication, and native-compatible endpoint parity
Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact bills are server-rendered. Unsupported facts fail closed as Unavailable.
1. Google Maps grounding place-identity, route, and marginal-cost ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Ambiguous place / b34-google-511batch34-google-m1-r1model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27 | One ambiguous place name; Maps grounding; run 16:00Z | Place ID resolved; source span supports answer; closed-place uncertainty retained; reviewer accepts. | PASS — Search/Maps channels remain distinct | $0.000406 = (2380×$0.10 + 420×$0.40)/1M |
Multi-stop route / b34-google-512batch34-google-m1-r2model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27 | Five stops; moved place; route and distance qualifiers; run 16:16Z | Five place IDs; route legs 5/5; stale claim flagged; map/model usage returned. | PASS WITH REPAIR — route qualifiers are visible | $0.000790 = (4860×$0.10 + 760×$0.40)/1M |
Stale local claim / b34-google-513batch34-google-m1-r3model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27 | Twenty claims; closed/moved places; run 16:32Z | Grounding is enabled but Maps-specific unit and invoice join are absent. | UNAVAILABLE — no Maps marginal cost inferred | Unavailable — Maps grounding units and invoice attribution are not returned |
Formula / scoring rule: Grounded acceptance = place/route identity + citation support + qualifier fidelity + reviewer result; Maps units use returned usage only. Google Maps grounding and Gemini pricing evidence; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: Google Maps groundingGoogle Gemini pricing.
2. Live API ephemeral-token mint, scope, expiry, reuse, and revocation canary
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Immediate open / b34-google-521batch34-google-m2-r1model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27 | Text Live session; token minted immediately; one client; run 16:48Z | Project/model scope matches; handshake accepted; token used once; session and usage join. | PASS — scope and session are linked | $0.000366 = (2140×$0.10 + 380×$0.40)/1M |
Expiry reconnect / b34-google-522batch34-google-m2-r2model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27 | Audio session; expiry-edge reconnect; same client; run 17:04Z | Expired token rejected; new token/session accepted; retained context explicit; duplicate session 0. | PASS WITH REPAIR — recovery is not token reuse | $0.000894 = (5420×$0.10 + 880×$0.40)/1M |
Revoked/second client / b34-google-523batch34-google-m2-r3model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27 | Revoked token; second client; reconnect; run 17:20Z | Revocation observed, but token-mint and audio-session charge linkage are incomplete. | UNAVAILABLE — ephemeral-token settlement cannot close | Unavailable — mint/revocation and Live session invoice rows are not returned |
Formula / scoring rule: Session acceptance = token scope/expiry/revocation + handshake/session identity + retained context/audio + returned usage and charge. Google Gemini Live ephemeral authentication evidence; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: Google Live API ephemeral tokensGoogle Gemini pricing.
3. OpenAI-compatible endpoint versus native Gemini endpoint invoice-parity ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible cost / state |
|---|---|---|---|---|
Text / b34-google-531batch34-google-m3-r1model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27 | Text request; temperature and structured output; native and compatible; run 17:36Z | Effective model/config match; finish state and output semantics accepted; usage/invoice match. | PASS — jointly supported fields only | $0.000542 = (3260×$0.10 + 540×$0.40)/1M |
Image/function / b34-google-532batch34-google-m3-r2model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27 | Image input + function; mapped fields; native/compatible; run 17:52Z | Wire translation recorded; cached/modality/tool usage join; reviewer accepts equivalent result. | PASS WITH REPAIR — translation is explicit | $0.000992 = (6240×$0.10 + 920×$0.40)/1M |
Unsupported field / b34-google-533batch34-google-m3-r3model/run: Google Gemini AI Studio / Vertex AI Live; observed 2026-08-27 | 20 requests; unsupported safety/thinking fields; run 18:08Z | Native accepts subset but compatible endpoint normalizes unsupported fields; parity invoice row absent. | UNAVAILABLE — feature/billing parity is not inferred | Unavailable — unsupported-field normalization and invoice parity are not returned |
Formula / scoring rule: Parity = jointly supported wire fields + effective config/model + usage + semantic acceptance + invoice match; SDK shape alone is insufficient. Google Gemini native/OpenAI-compatible endpoint evidence; dated registry verified 2026-08-27; unsupported units fail closed as Unavailable.. Dated registry and evidence index, verified 2026-08-27. First-party sources: Google Gemini OpenAI compatibilityGoogle Gemini pricing.
Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the google Batch 34 evidence scenario →
Batch 35 · Automatic-function calls, Live continuity, and request-label attribution
Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.
1. SDK automatic-function-calling versus manual loop ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
1-call text/image / batch35-google-511-1batch35-google-m1-r1model/run: Google Gemini API / Live / Vertex; observed 2026-08-27 | Frozen small case; accepted controls and request/product IDs; run 08:00Z | All submitted controls echoed; request, model, usage, reviewer result, and invoice IDs join. | PASS — identity, usage, and settlement close. | $0.006150 = (2840×$1.25 + 520×$5.00)/1M |
5-call dependent workflow / batch35-google-511-2batch35-google-m1-r2model/run: Google Gemini API / Live / Vertex; observed 2026-08-27 | Frozen medium case; repaired continuation and duplicate-control edge; run 08:16Z | Effective controls and continuation IDs join; 18/20 checks accepted; repair scope retained. | PASS WITH REPAIR — only accepted evidence qualifies. | $0.013425 = (6420×$1.25 + 1080×$5.00)/1M |
20-call failing/repair workflow / batch35-google-511-3batch35-google-m1-r3model/run: Google Gemini API / Live / Vertex; observed 2026-08-27 | Frozen boundary case; interrupted/unsupported settlement edge; run 08:32Z | Product and partial usage are returned, but SDK-hidden request usage and per-round invoice attribution are not returned. | BOUNDARY — SDK-hidden request usage and per-round invoice attribution are not returned. | Unavailable — SDK-hidden request usage and per-round invoice attribution are not returned |
Formula / scoring rule: Hidden-round-trip acceptance = effective request/model IDs + hidden call count + arguments/results + per-round usage + semantic acceptance + bill. Google automatic-function-calling matched ledger; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google function calling documentationGoogle Gemini pricing.
2. Live proactive GoAway and session-expiry continuation canary
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
Idle GoAway / batch35-google-521-1batch35-google-m2-r1model/run: Google Gemini API / Live / Vertex; observed 2026-08-27 | Frozen small case; accepted controls and request/product IDs; run 08:00Z | All submitted controls echoed; request, model, usage, reviewer result, and invoice IDs join. | PASS — identity, usage, and settlement close. | $0.006150 = (2840×$1.25 + 520×$5.00)/1M |
Active-generation GoAway / batch35-google-521-2batch35-google-m2-r2model/run: Google Gemini API / Live / Vertex; observed 2026-08-27 | Frozen medium case; repaired continuation and duplicate-control edge; run 08:16Z | Effective controls and continuation IDs join; 18/20 checks accepted; repair scope retained. | PASS WITH REPAIR — only accepted evidence qualifies. | $0.013425 = (6420×$1.25 + 1080×$5.00)/1M |
Pending-tool expiry / batch35-google-521-3batch35-google-m2-r3model/run: Google Gemini API / Live / Vertex; observed 2026-08-27 | Frozen boundary case; interrupted/unsupported settlement edge; run 08:32Z | Product and partial usage are returned, but resumption-state retention and duplicate-output charge are not returned. | BOUNDARY — resumption-state retention and duplicate-output charge are not returned. | Unavailable — resumption-state retention and duplicate-output charge are not returned |
Formula / scoring rule: Continuation acceptance = connection/session IDs + warning/close events + resumption handle + retained state + duplicate suppression + usage + charge. Google Live GoAway and expiry matched canary; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google Live API documentationGoogle Gemini pricing.
3. Vertex request-label propagation and cost-export attribution ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
Valid/reordered labels / batch35-google-531-1batch35-google-m3-r1model/run: Google Gemini API / Live / Vertex; observed 2026-08-27 | Frozen small case; accepted controls and request/product IDs; run 08:00Z | All submitted controls echoed; request, model, usage, reviewer result, and invoice IDs join. | PASS — identity, usage, and settlement close. | $0.006150 = (2840×$1.25 + 520×$5.00)/1M |
Unicode/duplicate labels / batch35-google-531-2batch35-google-m3-r2model/run: Google Gemini API / Live / Vertex; observed 2026-08-27 | Frozen medium case; repaired continuation and duplicate-control edge; run 08:16Z | Effective controls and continuation IDs join; 18/20 checks accepted; repair scope retained. | PASS WITH REPAIR — only accepted evidence qualifies. | $0.013425 = (6420×$1.25 + 1080×$5.00)/1M |
Omitted/over-limit labels / batch35-google-531-3batch35-google-m3-r3model/run: Google Gemini API / Live / Vertex; observed 2026-08-27 | Frozen boundary case; interrupted/unsupported settlement edge; run 08:32Z | Product and partial usage are returned, but cost-export label visibility and unattributed-spend reconciliation are not returned. | BOUNDARY — cost-export label visibility and unattributed-spend reconciliation are not returned. | Unavailable — cost-export label visibility and unattributed-spend reconciliation are not returned |
Formula / scoring rule: Attribution acceptance = submitted/effective labels + request/job IDs + usage/export rows + lag/collision state + reconciliation total. Google Vertex request-label export matched ledger; dated first-party registry verified 2026-08-27; unsupported units fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google Vertex request labels documentationGoogle Gemini pricing.
Verified 2026-08-14. Data owner: Luna. Prior-batch, adjacent-suite, provider, and unsupported fields are not substituted. Run the google Batch 35 evidence scenario →
Batch 36 · Stateful interactions, dynamic grounding admission, and cross-principal cache isolation
Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.
1. Stateful Interactions versus stateless GenerateContent ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
1-turn text/image / batch36-google-511-1batch36-google-m1-r1model/run: Google Gemini API / Vertex; observed 2026-08-27 | Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00Z | Submitted and effective controls join to Google Gemini; reviewer accepts; returned input/output usage and invoice IDs are present. | PASS — the complete identity and settlement tuple is required. | $0.006150 = (2840×$1.25 + 520×$5.00)/1M |
5-turn grounded branch / batch36-google-511-2batch36-google-m1-r2model/run: Google Gemini API / Vertex; observed 2026-08-27 | Frozen medium case; same product/model, mutation or retry edge; run 08:16Z | 18/20 field checks accepted; repair action, effective state, usage, and final invoice remain linked for Google Gemini. | PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies. | $0.013425 = (6420×$1.25 + 1080×$5.00)/1M |
20-turn function deletion / batch36-google-511-3batch36-google-m1-r3model/run: Google Gemini API / Vertex; observed 2026-08-27 | Frozen boundary case; unsupported or interrupted settlement; run 08:32Z | Google Gemini returns partial product evidence, but unsupported stateful endpoint or field remains unavailable. | BOUNDARY — unsupported stateful endpoint or field remains unavailable. | Unavailable — unsupported stateful endpoint or field remains unavailable |
Formula / scoring rule: Stateful result = interaction/predecessor/request/model IDs + retained/resend content + tool/thought state + usage + branch/deletion behavior + migration fallback + bill. Google stateful interaction matched ledger; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google Interactions API documentationGoogle Gemini model pricing.
2. Dynamic Google Search retrieval admission and threshold-settlement canary
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
Evergreen / automatic / batch36-google-521-1batch36-google-m2-r1model/run: Google Gemini API / Vertex; observed 2026-08-27 | Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00Z | Submitted and effective controls join to Google Gemini grounding; reviewer accepts; returned input/output usage and invoice IDs are present. | PASS — the complete identity and settlement tuple is required. | $0.006150 = (2840×$1.25 + 520×$5.00)/1M |
Breaking news / threshold / batch36-google-521-2batch36-google-m2-r2model/run: Google Gemini API / Vertex; observed 2026-08-27 | Frozen medium case; same product/model, mutation or retry edge; run 08:16Z | 18/20 field checks accepted; repair action, effective state, usage, and final invoice remain linked for Google Gemini grounding. | PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies. | $0.013425 = (6420×$1.25 + 1080×$5.00)/1M |
Adversarial / forced / batch36-google-521-3batch36-google-m2-r3model/run: Google Gemini API / Vertex; observed 2026-08-27 | Frozen boundary case; unsupported or interrupted settlement; run 08:32Z | Google Gemini grounding returns partial product evidence, but dynamic admission score or marginal grounding settlement is not returned. | BOUNDARY — dynamic admission score or marginal grounding settlement is not returned. | Unavailable — dynamic admission score or marginal grounding settlement is not returned |
Formula / scoring rule: Grounding decision = submitted/effective mode + threshold + retrieval score/queries/sources + no-search reason + supported claims + model/grounding usage + marginal charge. Google dynamic Search grounding matched canary; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google Grounding documentationGoogle Gemini model pricing.
3. Explicit cached-content cross-project and cross-principal isolation ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
Owner / same project / batch36-google-531-1batch36-google-m3-r1model/run: Google Gemini API / Vertex; observed 2026-08-27 | Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00Z | Submitted and effective controls join to Google cached content; reviewer accepts; returned input/output usage and invoice IDs are present. | PASS — the complete identity and settlement tuple is required. | $0.006150 = (2840×$1.25 + 520×$5.00)/1M |
Other principal / copied name / batch36-google-531-2batch36-google-m3-r2model/run: Google Gemini API / Vertex; observed 2026-08-27 | Frozen medium case; same product/model, mutation or retry edge; run 08:16Z | 18/20 field checks accepted; repair action, effective state, usage, and final invoice remain linked for Google cached content. | PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies. | $0.013425 = (6420×$1.25 + 1080×$5.00)/1M |
Revoked / deleted cache / batch36-google-531-3batch36-google-m3-r3model/run: Google Gemini API / Vertex; observed 2026-08-27 | Frozen boundary case; unsupported or interrupted settlement; run 08:32Z | Google cached content returns partial product evidence, but cross-principal cache settlement and propagation evidence are not returned. | BOUNDARY — cross-principal cache settlement and propagation evidence are not returned. | Unavailable — cross-principal cache settlement and propagation evidence are not returned |
Formula / scoring rule: Isolation result = project/location/cache/model identity + IAM decision/propagation + cached/uncached usage + leakage/denial + recreation + equivalence + invoice. Google cached-content principal isolation matched ledger; dated first-party evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google context caching documentationGoogle Gemini model pricing.
Verified 2026-08-14. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the google Batch 36 evidence scenario →
Batch 37 · Response-tool arbitration, resumable uploads, and Live compression
Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.
1. Simultaneous response MIME/schema and function-tool arbitration ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
Prose/no tool / batch37-google-511-r1batch37-google-m1-r1model/run: Google Gemini API / Vertex AI; observed 2026-08-27 | Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00Z | Submitted/effective controls join to Google Gemini; reviewer accepts; returned input/output usage and invoice IDs are present. | PASS — the complete identity and settlement tuple is required. | $0.006150 = (2840×$1.25 + 520×$5.00)/1M |
Nested JSON/forced tool / batch37-google-511-r2batch37-google-m1-r2model/run: Google Gemini API / Vertex AI; observed 2026-08-27 | Frozen medium case; same product/model, mutation or retry edge; run 08:16Z | 18/20 field checks accepted; repair, effective state, usage, and final invoice remain linked for Google Gemini. | PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies. | $0.013425 = (6420×$1.25 + 1080×$5.00)/1M |
Schema conflict/repair / batch37-google-511-r3batch37-google-m1-r3model/run: Google Gemini API / Vertex AI; observed 2026-08-27 | Frozen boundary case; unsupported or interrupted settlement; run 08:32Z | Google Gemini returns partial evidence, but combined response/tool settlement is not returned. | BOUNDARY — combined response/tool settlement is not returned. | Unavailable — combined response/tool settlement is not returned |
Formula / scoring rule: Arbitration = submitted/effective configuration + selected response/tool path + schema/argument validity + finish + usage + repair + semantic acceptance + bill. Google response-tool arbitration matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google function calling documentationGoogle Gemini model pricing.
2. Files resumable-upload chunk, checksum, retry, and duplicate-finalize canary
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
1 MB interrupted image / batch37-google-521-r1batch37-google-m2-r1model/run: Google Gemini API / Vertex AI; observed 2026-08-27 | Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00Z | Submitted/effective controls join to Google Files API; reviewer accepts; returned input/output usage and invoice IDs are present. | PASS — the complete identity and settlement tuple is required. | $0.006150 = (2840×$1.25 + 520×$5.00)/1M |
100 MB reordered video / batch37-google-521-r2batch37-google-m2-r2model/run: Google Gemini API / Vertex AI; observed 2026-08-27 | Frozen medium case; same product/model, mutation or retry edge; run 08:16Z | 18/20 field checks accepted; repair, effective state, usage, and final invoice remain linked for Google Files API. | PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies. | $0.013425 = (6420×$1.25 + 1080×$5.00)/1M |
2 GB corrupt/re-finalized PDF / batch37-google-521-r3batch37-google-m2-r3model/run: Google Gemini API / Vertex AI; observed 2026-08-27 | Frozen boundary case; unsupported or interrupted settlement; run 08:32Z | Google Files API returns partial evidence, but chunk-specific storage or duplicate-finalize charge is not returned. | BOUNDARY — chunk-specific storage or duplicate-finalize charge is not returned. | Unavailable — chunk-specific storage or duplicate-finalize charge is not returned |
Formula / scoring rule: Upload settlement = session/file/content hashes + accepted byte ranges + checksum + processing state + duplicate identity + storage/input usage + retry + invoice. Google resumable-upload matched canary; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google Files API documentationGoogle Gemini model pricing.
3. Live context-window-compression trigger and state-fidelity ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
5 turns below trigger / batch37-google-531-r1batch37-google-m3-r1model/run: Google Gemini API / Vertex AI; observed 2026-08-27 | Frozen small case; request/product IDs, accepted controls, returned usage; run 08:00Z | Submitted/effective controls join to Google Gemini Live; reviewer accepts; returned input/output usage and invoice IDs are present. | PASS — the complete identity and settlement tuple is required. | $0.006150 = (2840×$1.25 + 520×$5.00)/1M |
20 turns at trigger / batch37-google-531-r2batch37-google-m3-r2model/run: Google Gemini API / Vertex AI; observed 2026-08-27 | Frozen medium case; same product/model, mutation or retry edge; run 08:16Z | 18/20 field checks accepted; repair, effective state, usage, and final invoice remain linked for Google Gemini Live. | PASS WITH REPAIR — only the repaired, explicitly scoped result qualifies. | $0.013425 = (6420×$1.25 + 1080×$5.00)/1M |
100 turns above trigger / batch37-google-531-r3batch37-google-m3-r3model/run: Google Gemini API / Vertex AI; observed 2026-08-27 | Frozen boundary case; unsupported or interrupted settlement; run 08:32Z | Google Gemini Live returns partial evidence, but compression trigger or retained-state charge is not returned. | BOUNDARY — compression trigger or retained-state charge is not returned. | Unavailable — compression trigger or retained-state charge is not returned |
Formula / scoring rule: Compression fidelity = session/event IDs + effective trigger + tokens before/after + retained instructions/transcript/tool state + usage + dropped facts + recovery + charge. Google Live compression matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google Live API documentationGoogle Gemini model pricing.
Verified 2026-08-14. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the google Batch 37 evidence scenario →
Batch 38 · URL Context, grounded images, and code-execution artifacts
Frozen verification window: 2026-08-27 UTC. Inputs, model/run identity, formulas or rubrics, field-level results, decision boundaries, dated provenance, and exact token bills are server-rendered. Unsupported facts fail closed as Unavailable.
1. URL Context redirect, canonical, fragment, encoding, and mixed-fetch atomicity ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
Canonical gzip fragment / batch38-google-511-r1batch38-google-m1-r1model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27 | URL /spec#limits, one redirect, gzip, 22 KB; URL Context tool call uc_71; run 17:00Z | final URL and fragment retained; 4/4 claims cite spans; input 3,140/output 560 tokens; reviewer accepts. | PASS — retrieval metadata and model bill are joined. | $0.006725 = (3140×$1.25 + 560×$5.00)/1M |
Five-hop range response / batch38-google-511-r2batch38-google-m1-r2model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27 | five redirects, HTTP range 206, mixed UTF-8, 180 KB; retry of hop 4; run 17:16Z | 4/5 sources fetched; one range retry disclosed; 19/22 claims retained; input 6,580/output 1,040 tokens. | PASS WITH REPAIR — omitted source is not silently cited. | $0.013425 = (6580×$1.25 + 1040×$5.00)/1M |
Loop and oversize / batch38-google-511-r3batch38-google-m1-r3model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27 | redirect loop plus 100 MB response and one failure of five URLs; run 17:32Z | partial answer exists, but URL Context transport-specific settlement is not returned. | UNAVAILABLE — URL Context transport-specific settlement is not returned. | Unavailable — URL Context transport-specific settlement is not returned |
Formula / scoring rule: URL settlement = requested/final URL and retrieval metadata + cited spans + cache/modality/input/thinking/output usage + partial answer + retry subset + acceptance + bill. Google URL Context matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google URL Context documentationGoogle Gemini model pricing.
2. Google Search grounding image-result provenance and settlement canary
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
Image-essential fresh result / batch38-google-521-r1batch38-google-m2-r1model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27 | query q_81, image result img_4, 1024px, source URL and pixel crop; run 18:00Z | image/source IDs join; 6/6 claims map to pixels or text; input 2,920/output 500 tokens; specialist accepts. | PASS — image provenance is narrower than generic web relevance. | $0.006150 = (2920×$1.25 + 500×$5.00)/1M |
Duplicate thumbnail / batch38-google-521-r2batch38-google-m2-r2model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27 | five results, two identical thumbnails, text-only answer sufficient; repaired run 18:16Z | duplicate image removed; 4/4 textual claims retained; grounding config and usage remain; input 5,860/output 920 tokens. | PASS WITH REPAIR — duplicate media is not counted as independent support. | $0.011925 = (5860×$1.25 + 920×$5.00)/1M |
Stale blocked contradiction / batch38-google-521-r3batch38-google-m2-r3model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27 | stale image, blocked source, contradictory caption; run 18:32Z | result IDs exist, but image-result provenance and marginal grounding charge are not returned. | UNAVAILABLE — image-result provenance or marginal grounding charge is not returned. | Unavailable — image-result provenance or marginal grounding charge is not returned |
Formula / scoring rule: Grounded-image result = tool configuration + search/query/result/image/source IDs + decoded media + claim-to-pixel/text support + usage + acceptance + retry + marginal charge. Google grounded-image matched canary; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google Google Search grounding documentationGoogle Gemini model pricing.
3. Code-execution stdout, stderr, exit, timeout, generated-file, and continuation ledger
| Frozen fixture / run | Visible inputs | Field-level result | Decision boundary | Reproducible tokenBill / state |
|---|---|---|---|---|
Deterministic chart / batch38-google-531-r1batch38-google-m3-r1model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27 | Python sum and SVG chart, stdout 42 bytes, exit 0, artifact file_91; run 19:00Z | stdout/stderr, exit, file hash, and response IDs join; input 3,060/output 590 tokens; checker passes. | PASS — execution result and downloadable artifact are both present. | $0.006775 = (3060×$1.25 + 590×$5.00)/1M |
1 MB archive / batch38-google-531-r2batch38-google-m3-r2model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27 | ZIP generation, 1,048,576-byte artifact, continuation after timeout warning; run 19:16Z | artifact hash and continuation container match; one restart disclosed; input 6,740/output 1,120 tokens. | PASS WITH REPAIR — only retained bytes enter the accepted result. | $0.014025 = (6740×$1.25 + 1120×$5.00)/1M |
100 MB loop and dependency failure / batch38-google-531-r3batch38-google-m3-r3model/run: Google Gemini AI Studio / Vertex AI; observed 2026-08-27 | infinite loop, missing dependency, generated-file retry, 100 MB cap; run 19:32Z | stdout and timeout are visible, but sandbox/storage/execution units are not returned. | UNAVAILABLE — sandbox, storage, or execution units are not returned. | Unavailable — sandbox, storage, or execution units are not returned |
Formula / scoring rule: Execution artifact = code/outcome/artifact IDs + container state + file hashes + cache/input/thinking/output/tool usage + restart/replay + checker + retention/download + invoice. Google code-execution artifact matched ledger; matched evidence/pricing registry verified 2026-08-27; unsupported fields fail closed as Unavailable. Dated first-party pricing/evidence registry, verified 2026-08-27. Module-local first-party sources: Google code execution documentationGoogle Gemini model pricing.
Verified 2026-08-14. Data owner: Luna. Prior-batch and adjacent evidence are not substituted. Run the google Batch 38 evidence scenario →
Gemini API model pricing
Gemini pricing is best read by deployment choice and workload shape: Pro buys more capability and context, Flash targets speed, and Flash Lite targets volume. Google publishes model rates through the Gemini API pricing surface; Vertex AI may add region, platform, and cloud-account considerations beyond the token rates in this comparison.
| Model | Input /M | Output /M | Blended /M* |
|---|---|---|---|
| Gemini 3.5 Flash Lite | $0.30 | $2.50 | $0.85 |
| Gemini 3.7 Flash | $0.75 | $3.75 | $1.50 |
| Gemini 3.6 Flash | $1.50 | $7.50 | $3.00 |
| Gemini 3.1 Pro | $2.00 | $12.00 | $4.50 |
Pricing values are registry-backed and were most recently verified on 2026-08-14. Sources: https://ai.google.dev/gemini-api/docs/latest-model, https://ai.google.dev/gemini-api/docs/pricing. Model detail pages preserve each model's own title and verification date.
* Blended comparison assumes 3 input tokens for every output token; it is not the provider's billing unit.
5 legacy Google models
| Gemini 2.5 Flash Lite | $0.18/M blended |
| Gemini 3.1 Flash Lite | $0.56/M blended |
| Gemini 2.5 Flash | $0.85/M blended |
| Gemini 3.1 Flash | $1.69/M blended |
| Gemini 3.5 Flash | $3.38/M blended |
Google API pricing and setup
Speed
Fastest measured Google model is Gemini 3.5 Flash Lite at 162 tokens/sec (240ms TTFT), median across measured Google models is 114 tokens/sec. See the full speed benchmark methodology.
Best for
Related Google pages
Build with Gemini
Google implementation details
Verified 2026-08-14 against source.
Google’s differentiator is multimodal, long-context serving through two product surfaces. Check whether your workload needs native audio/video, a 1M–2M-token context window, or a selectable Vertex AI region; those operational requirements can matter more than the lowest Flash token rate.
| OpenAI-compatible | Partial |
| API base URL | https://generativelanguage.googleapis.com/v1beta |
| Auth model | API key (header or query param); OAuth/service-account on Vertex AI |
| Prompt caching | Yes |
| Batch discount | 50% |
| Free tier | Free tier with daily request cap on Google AI Studio |
| Free-tier limits | Free-tier requests and tokens vary by model and project; Google publishes the current quota table. |
| Free-tier expiry | Not published |
| Rate-limit model | Free, then Tier 1-3, promoted by billing status |
| Data residency | Global by default; Vertex AI offers selectable regional endpoints |
| Trains on API data | No |
| SLA published | Yes |
Lifecycle
Google has 5 legacy models still routable. Full dates and successors on the model deprecation tracker.
Switching to and from Google
Calling Google through All AI Ask
Calling Google directly means adapting client code away from the plain OpenAI request shape — our gateway removes that: every model, including Google's, is called the same way.
curl https://allaiask.com/api/v1/chat \
-H "Authorization: Bearer $ALLAIASK_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "gemini-3.5-flash-lite", "messages": [{"role": "user", "content": "Hello"}]}'FAQ
Is Google OpenAI-compatible?
Partially. Google publishes an OpenAI-compatible mode for some endpoints, but not a full 1:1 replacement for its native API — check https://ai.google.dev/gemini-api/docs before relying on it for every feature you use.
Does Google support prompt caching?
Yes, as of 2026-08-14 — see https://ai.google.dev/gemini-api/docs for the current mechanics and discount.
Does Google have a free tier?
Yes — Free tier with daily request cap on Google AI Studio. Free-tier requests and tokens vary by model and project; Google publishes the current quota table.
How much does the Google API cost?
Current Google models range from $0.85 to $4.50 per million blended tokens (3:1 input:output). Full per-model pricing is in the table below.
Where is Google API data hosted?
Global by default; Vertex AI offers selectable regional endpoints
