Google API Cost Calculator
How much does the Google API cost per month?
For Gemini 2.5 Flash Lite at 200K requests/month, 2,400 input tokens and 350 output tokens per request costs about $31.52 per month after 30% cache use and 100% batch share. Across Google's 9 priced models, the cheapest default ranking is Gemini 2.5 Flash Lite at $76.00 per month.
Pricing data as of June 2026. Sources: Google pricing and model documentation. Parameters are shareable in this URL.
Gemini 2.5 Flash Lite: estimated monthly cost
Formula: provider-specific cached input multiplier × cache rate, plus uncached input and verbosity-adjusted output; cache writes are amortised over the provider TTL, then batch savings are applied.
| Per request | $0.000 |
|---|---|
| Per day | $1.05 |
| Per month | $31.52 |
| Per year | $378 |
OpenAI scenario sensitivity
The default is a server-rendered estimate. Change the cache and batch shares to see when a cheaper qualified model overtakes the selected model; the URL is shareable and preserves the inputs.
| Default | $31.52/month |
|---|---|
| No caching or batch | $76.00/month |
| 50% cache, 50% batch | $40.80/month |
| 100% cache, 100% batch | $16.40/month |
Batch 11 Google workload decision depth
1. Long-context billing-cliff audit
| Input tokens | Model | Whole-request bill | Rule |
|---|---|---|---|
| 127,999 | Gemini 2.5 Flash Lite | $0.0047 | Threshold treatment: dated model rule or Unavailable |
| 127,999 | Gemini 3.1 Flash Lite | $0.01 | Threshold treatment: dated model rule or Unavailable |
| 127,999 | Gemini 3.5 Flash Lite | $0.01 | Threshold treatment: dated model rule or Unavailable |
| 128,000 | Gemini 2.5 Flash Lite | $0.0047 | Threshold treatment: dated model rule or Unavailable |
| 128,000 | Gemini 3.1 Flash Lite | $0.01 | Threshold treatment: dated model rule or Unavailable |
| 128,000 | Gemini 3.5 Flash Lite | $0.01 | Threshold treatment: dated model rule or Unavailable |
| 128,001 | Gemini 2.5 Flash Lite | $0.0047 | Threshold treatment: dated model rule or Unavailable |
| 128,001 | Gemini 3.1 Flash Lite | $0.01 | Threshold treatment: dated model rule or Unavailable |
| 128,001 | Gemini 3.5 Flash Lite | $0.01 | Threshold treatment: dated model rule or Unavailable |
| 500,000 | Gemini 2.5 Flash Lite | $0.02 | Threshold treatment: dated model rule or Unavailable |
| 500,000 | Gemini 3.1 Flash Lite | $0.05 | Threshold treatment: dated model rule or Unavailable |
| 500,000 | Gemini 3.5 Flash Lite | $0.06 | Threshold treatment: dated model rule or Unavailable |
| 500,001 | Gemini 2.5 Flash Lite | $0.02 | Threshold treatment: dated model rule or Unavailable |
| 500,001 | Gemini 3.1 Flash Lite | $0.05 | Threshold treatment: dated model rule or Unavailable |
| 500,001 | Gemini 3.5 Flash Lite | $0.06 | Threshold treatment: dated model rule or Unavailable |
| 1,000,000 | Gemini 2.5 Flash Lite | $0.04 | Threshold treatment: dated model rule or Unavailable |
| 1,000,000 | Gemini 3.1 Flash Lite | $0.09 | Threshold treatment: dated model rule or Unavailable |
| 1,000,000 | Gemini 3.5 Flash Lite | $0.11 | Threshold treatment: dated model rule or Unavailable |
2. 128K/500K/1M fit and cost matrix
| Input shape | Gemini 2.5 Flash Lite | Gemini 3.1 Flash Lite | Gemini 3.5 Flash Lite | Gemini 2.5 Flash |
|---|---|---|---|---|
| 128K | $0.01 | $0.03 | $0.04 | $0.04 |
| 500K | $0.05 | $0.13 | $0.15 | $0.15 |
| 1M | $0.10 | $0.25 | $0.30 | $0.30 |
Formula: ($/M input × input + $/M output × output) ÷ 1,000,000. Image/audio units are Unavailable when not text-compatible.
3. AI Studio versus Vertex reconciliation
| Endpoint | Same workload bill | Rate/cap source | Unmatched facts |
|---|---|---|---|
| AI Studio | $31.52 | Dated model record only | Region, free tier, endpoint caps: Unavailable |
| Vertex | $31.52 | Dated model record only | Region, free tier, endpoint caps: Unavailable |
Provenance: Batch 11 Google threshold/endpoint reconciliation module; selected model Gemini 2.5 Flash Lite; inputs are 2,400 input tokens, 350 output tokens, 200,000 calls/month, cache rate 30%, batch share 100%. Verified 2026-06-21. Provider source: https://ai.google.dev/gemini-api/docs/pricing · Run this scenario →
Batch 59 · server-rendered evidence boards · verified 2026-09-07
Intent answer: Google Gemini API bills vary by context window threshold (<=128K vs >128K tokens), cached prompt storage/read rates, and 50% batch discounts. Interactive estimates must join exact model IDs, tier crossovers, and multimodal tokens. Verified 2026-09-07.
Demand evidence: Qualitative demand: dedicated Gemini cost calculators reviewed 2026-09-07; exact US monthly volume is unavailable.
Scope boundary: Calculate and forecast Google Gemini API monthly expenses across 2.5 Flash, 3.5 Flash Lite, 3.6 Flash, 3.7 Flash, and 3.1 Pro with context tiers, caching, and batch. Exact joins required; unresolved joins render Unavailable.
Gemini context tier crossover evaluator
Deterministic formula / rule: rate = (tokens <= 128000) ? base_rate : long_context_rate; caching discount = 75% on cache reads.
Boundary: Owns multi-tier Gemini cost forecasting; single model pricing cards remain on model pages.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch59-calc-google-m1-r1short chat (<32K) | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$short chat (<32K); input tokens; output tokens; context threshold; base vs tier rate; cached write/read rates; deterministic monthly bill; verified=2026-09-07 | Unavailable — frozen calc-google fixture requires an exact source, identity, and result receipt | UNTESTED — module 1 rule is reproducible but no production observation is claimed |
batch59-calc-google-m1-r2standard extraction (64K) | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$standard extraction (64K); input tokens; output tokens; context threshold; base vs tier rate; cached write/read rates; deterministic monthly bill; verified=2026-09-07 | Unavailable — frozen calc-google fixture requires an exact source, identity, and result receipt | UNTESTED — module 1 rule is reproducible but no production observation is claimed |
batch59-calc-google-m1-r3borderline tier (128K) | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$borderline tier (128K); input tokens; output tokens; context threshold; base vs tier rate; cached write/read rates; deterministic monthly bill; verified=2026-09-07 | Unavailable — frozen calc-google fixture requires an exact source, identity, and result receipt | UNTESTED — module 1 rule is reproducible but no production observation is claimed |
batch59-calc-google-m1-r4large repository analysis (250K) | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$large repository analysis (250K); input tokens; output tokens; context threshold; base vs tier rate; cached write/read rates; deterministic monthly bill; verified=2026-09-07 | Unavailable — frozen calc-google fixture requires an exact source, identity, and result receipt | UNTESTED — module 1 rule is reproducible but no production observation is claimed |
batch59-calc-google-m1-r5massive multimodal context (1M) | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$massive multimodal context (1M); input tokens; output tokens; context threshold; base vs tier rate; cached write/read rates; deterministic monthly bill; verified=2026-09-07 | Unavailable — frozen calc-google fixture requires an exact source, identity, and result receipt | UNTESTED — module 1 rule is reproducible but no production observation is claimed |
batch59-calc-google-m1-r6unsupported context threshold | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$unsupported context threshold; input tokens; output tokens; context threshold; base vs tier rate; cached write/read rates; deterministic monthly bill; verified=2026-09-07 | Unavailable — exact calc-google evidence join is not closed for "unsupported context threshold" | FAIL CLOSED — manual, probe, or source evidence required |
First-party citation: Google Gemini API official pricing. Verified 2026-09-07; missing or conflicting joins fail closed.
AI Studio vs Vertex AI cost & feature ledger
Deterministic formula / rule: vertex_delta = enterprise SLA + private VPC + custom quotas; base model token rates match published API tariff.
Boundary: Owns enterprise routing cost comparison across Google serving surfaces.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch59-calc-google-m2-r1standard developer API | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$standard developer API; endpoint realm; tool execution surcharges; quota tier; search grounding add-on; monthly baseline estimate; fail-closed notice; verified=2026-09-07 | Unavailable — frozen calc-google fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch59-calc-google-m2-r2Vertex AI enterprise project | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$Vertex AI enterprise project; endpoint realm; tool execution surcharges; quota tier; search grounding add-on; monthly baseline estimate; fail-closed notice; verified=2026-09-07 | Unavailable — frozen calc-google fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch59-calc-google-m2-r3grounding with Google Search | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$grounding with Google Search; endpoint realm; tool execution surcharges; quota tier; search grounding add-on; monthly baseline estimate; fail-closed notice; verified=2026-09-07 | Unavailable — frozen calc-google fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch59-calc-google-m2-r4code execution tool run | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$code execution tool run; endpoint realm; tool execution surcharges; quota tier; search grounding add-on; monthly baseline estimate; fail-closed notice; verified=2026-09-07 | Unavailable — frozen calc-google fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch59-calc-google-m2-r5provisioned throughput | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$provisioned throughput; endpoint realm; tool execution surcharges; quota tier; search grounding add-on; monthly baseline estimate; fail-closed notice; verified=2026-09-07 | Unavailable — frozen calc-google fixture requires an exact source, identity, and result receipt | UNTESTED — module 2 rule is reproducible but no production observation is claimed |
batch59-calc-google-m2-r6unresolved billing tier | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$unresolved billing tier; endpoint realm; tool execution surcharges; quota tier; search grounding add-on; monthly baseline estimate; fail-closed notice; verified=2026-09-07 | Unavailable — exact calc-google evidence join is not closed for "unresolved billing tier" | FAIL CLOSED — manual, probe, or source evidence required |
First-party citation: Google Gemini API official pricing. Verified 2026-09-07; missing or conflicting joins fail closed.
Gemini batch vs synchronous economics board
Deterministic formula / rule: batch_savings = standard_cost * 0.50 for 24h asynchronous queue jobs.
Boundary: Owns Gemini batch cost optimization calculations.
| Frozen scenario / field ID | Exact identity and evidence fields | Result | State |
|---|---|---|---|
batch59-calc-google-m3-r1classification batch (100K calls) | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$classification batch (100K calls); batch job tokens; 50% discount multiplier; turnaround SLA (24h); estimated net savings; reconciliation rule; verified=2026-09-07 | Unavailable — frozen calc-google fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch59-calc-google-m3-r2document summarization (10K docs) | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$document summarization (10K docs); batch job tokens; 50% discount multiplier; turnaround SLA (24h); estimated net savings; reconciliation rule; verified=2026-09-07 | Unavailable — frozen calc-google fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch59-calc-google-m3-r3mixed text-multimodal batch | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$mixed text-multimodal batch; batch job tokens; 50% discount multiplier; turnaround SLA (24h); estimated net savings; reconciliation rule; verified=2026-09-07 | Unavailable — frozen calc-google fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch59-calc-google-m3-r4urgent SLA bypass | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$urgent SLA bypass; batch job tokens; 50% discount multiplier; turnaround SLA (24h); estimated net savings; reconciliation rule; verified=2026-09-07 | Unavailable — exact calc-google evidence join is not closed for "urgent SLA bypass" | FAIL CLOSED — manual, probe, or source evidence required |
batch59-calc-google-m3-r5expired batch fallback | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$expired batch fallback; batch job tokens; 50% discount multiplier; turnaround SLA (24h); estimated net savings; reconciliation rule; verified=2026-09-07 | Unavailable — frozen calc-google fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
batch59-calc-google-m3-r6cancelled batch reconciliation | route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$cancelled batch reconciliation; batch job tokens; 50% discount multiplier; turnaround SLA (24h); estimated net savings; reconciliation rule; verified=2026-09-07 | Unavailable — frozen calc-google fixture requires an exact source, identity, and result receipt | UNTESTED — module 3 rule is reproducible but no production observation is claimed |
First-party citation: Google Gemini API official pricing. Verified 2026-09-07; missing or conflicting joins fail closed.
Method and limitations: this board exposes deterministic rules, first-party citations, and dated evidence identities. It does not invent volume, coverage, entitlement, retention, residency, quota, capacity, feature support, quality, price, reliability, or legal conclusions. Run the calc-google evidence flow →
Ranked cost — 200K requests/month
| Model | Provider | List monthly | Effective monthly | Verbosity | Rank Δ |
|---|---|---|---|---|---|
| Gemini 2.5 Flash Litebudgetlegacy | $76.00 | $63.09 | — | — | |
| Gemini 3.1 Flash Litebudgetlegacy | $225 | $179 | 0.87× | — | |
| Gemini 3.5 Flash Litebudget | $319 | $280 | — | — | |
| Gemini 2.5 Flashbudgetlegacy | $319 | $280 | — | — | |
| Gemini 3.7 Flashmid | $623 | $526 | — | — | |
| Gemini 3.1 Flashmidlegacy | $675 | $578 | — | — | |
| Gemini 3.6 Flashmid | $1,245 | $1,051 | — | — | |
| Gemini 3.5 Flashmidlegacy | $1,350 | $1,156 | — | — | |
| Gemini 3.1 Promid | $1,800 | $1,231 | 0.63× | — |
Levers live on Google
The billing gotcha
Gemini splits pricing by context length on some models — a call that crosses the long-context threshold bills at a higher per-token rate for the whole request, not just the tokens past the line. Google AI Studio's free tier also shares a daily request cap across every model you test, which does not show up in a per-model monthly projection.
