Google API Cost Calculator

How much does the Google API cost per month?

For Gemini 2.5 Flash Lite at 200K requests/month, 2,400 input tokens and 350 output tokens per request costs about $31.52 per month after 30% cache use and 100% batch share. Across Google's 9 priced models, the cheapest default ranking is Gemini 2.5 Flash Lite at $76.00 per month.

Verified 2026-06-21

Pricing data as of June 2026. Sources: Google pricing and model documentation. Parameters are shareable in this URL.

Gemini 2.5 Flash Lite: estimated monthly cost

$31.52

Formula: provider-specific cached input multiplier × cache rate, plus uncached input and verbosity-adjusted output; cache writes are amortised over the provider TTL, then batch savings are applied.

Per request$0.000
Per day$1.05
Per month$31.52
Per year$378

OpenAI scenario sensitivity

The default is a server-rendered estimate. Change the cache and batch shares to see when a cheaper qualified model overtakes the selected model; the URL is shareable and preserves the inputs.

Default$31.52/month
No caching or batch$76.00/month
50% cache, 50% batch$40.80/month
100% cache, 100% batch$16.40/month

Batch 11 Google workload decision depth

1. Long-context billing-cliff audit

Input tokensModelWhole-request billRule
127,999Gemini 2.5 Flash Lite$0.0047Threshold treatment: dated model rule or Unavailable
127,999Gemini 3.1 Flash Lite$0.01Threshold treatment: dated model rule or Unavailable
127,999Gemini 3.5 Flash Lite$0.01Threshold treatment: dated model rule or Unavailable
128,000Gemini 2.5 Flash Lite$0.0047Threshold treatment: dated model rule or Unavailable
128,000Gemini 3.1 Flash Lite$0.01Threshold treatment: dated model rule or Unavailable
128,000Gemini 3.5 Flash Lite$0.01Threshold treatment: dated model rule or Unavailable
128,001Gemini 2.5 Flash Lite$0.0047Threshold treatment: dated model rule or Unavailable
128,001Gemini 3.1 Flash Lite$0.01Threshold treatment: dated model rule or Unavailable
128,001Gemini 3.5 Flash Lite$0.01Threshold treatment: dated model rule or Unavailable
500,000Gemini 2.5 Flash Lite$0.02Threshold treatment: dated model rule or Unavailable
500,000Gemini 3.1 Flash Lite$0.05Threshold treatment: dated model rule or Unavailable
500,000Gemini 3.5 Flash Lite$0.06Threshold treatment: dated model rule or Unavailable
500,001Gemini 2.5 Flash Lite$0.02Threshold treatment: dated model rule or Unavailable
500,001Gemini 3.1 Flash Lite$0.05Threshold treatment: dated model rule or Unavailable
500,001Gemini 3.5 Flash Lite$0.06Threshold treatment: dated model rule or Unavailable
1,000,000Gemini 2.5 Flash Lite$0.04Threshold treatment: dated model rule or Unavailable
1,000,000Gemini 3.1 Flash Lite$0.09Threshold treatment: dated model rule or Unavailable
1,000,000Gemini 3.5 Flash Lite$0.11Threshold treatment: dated model rule or Unavailable

2. 128K/500K/1M fit and cost matrix

Input shapeGemini 2.5 Flash LiteGemini 3.1 Flash LiteGemini 3.5 Flash LiteGemini 2.5 Flash
128K$0.01$0.03$0.04$0.04
500K$0.05$0.13$0.15$0.15
1M$0.10$0.25$0.30$0.30

Formula: ($/M input × input + $/M output × output) ÷ 1,000,000. Image/audio units are Unavailable when not text-compatible.

3. AI Studio versus Vertex reconciliation

EndpointSame workload billRate/cap sourceUnmatched facts
AI Studio$31.52Dated model record onlyRegion, free tier, endpoint caps: Unavailable
Vertex$31.52Dated model record onlyRegion, free tier, endpoint caps: Unavailable

Provenance: Batch 11 Google threshold/endpoint reconciliation module; selected model Gemini 2.5 Flash Lite; inputs are 2,400 input tokens, 350 output tokens, 200,000 calls/month, cache rate 30%, batch share 100%. Verified 2026-06-21. Provider source: https://ai.google.dev/gemini-api/docs/pricing · Run this scenario →

Batch 59 · server-rendered evidence boards · verified 2026-09-07

Intent answer: Google Gemini API bills vary by context window threshold (<=128K vs >128K tokens), cached prompt storage/read rates, and 50% batch discounts. Interactive estimates must join exact model IDs, tier crossovers, and multimodal tokens. Verified 2026-09-07.

Demand evidence: Qualitative demand: dedicated Gemini cost calculators reviewed 2026-09-07; exact US monthly volume is unavailable.

Scope boundary: Calculate and forecast Google Gemini API monthly expenses across 2.5 Flash, 3.5 Flash Lite, 3.6 Flash, 3.7 Flash, and 3.1 Pro with context tiers, caching, and batch. Exact joins required; unresolved joins render Unavailable.

Gemini context tier crossover evaluator

Deterministic formula / rule: rate = (tokens <= 128000) ? base_rate : long_context_rate; caching discount = 75% on cache reads.

Boundary: Owns multi-tier Gemini cost forecasting; single model pricing cards remain on model pages.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch59-calc-google-m1-r1
short chat (<32K)
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$short chat (<32K); input tokens; output tokens; context threshold; base vs tier rate; cached write/read rates; deterministic monthly bill; verified=2026-09-07Unavailable — frozen calc-google fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-calc-google-m1-r2
standard extraction (64K)
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$standard extraction (64K); input tokens; output tokens; context threshold; base vs tier rate; cached write/read rates; deterministic monthly bill; verified=2026-09-07Unavailable — frozen calc-google fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-calc-google-m1-r3
borderline tier (128K)
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$borderline tier (128K); input tokens; output tokens; context threshold; base vs tier rate; cached write/read rates; deterministic monthly bill; verified=2026-09-07Unavailable — frozen calc-google fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-calc-google-m1-r4
large repository analysis (250K)
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$large repository analysis (250K); input tokens; output tokens; context threshold; base vs tier rate; cached write/read rates; deterministic monthly bill; verified=2026-09-07Unavailable — frozen calc-google fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-calc-google-m1-r5
massive multimodal context (1M)
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$massive multimodal context (1M); input tokens; output tokens; context threshold; base vs tier rate; cached write/read rates; deterministic monthly bill; verified=2026-09-07Unavailable — frozen calc-google fixture requires an exact source, identity, and result receiptUNTESTED — module 1 rule is reproducible but no production observation is claimed
batch59-calc-google-m1-r6
unsupported context threshold
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$unsupported context threshold; input tokens; output tokens; context threshold; base vs tier rate; cached write/read rates; deterministic monthly bill; verified=2026-09-07Unavailable — exact calc-google evidence join is not closed for "unsupported context threshold"FAIL CLOSED — manual, probe, or source evidence required

First-party citation: Google Gemini API official pricing. Verified 2026-09-07; missing or conflicting joins fail closed.

AI Studio vs Vertex AI cost & feature ledger

Deterministic formula / rule: vertex_delta = enterprise SLA + private VPC + custom quotas; base model token rates match published API tariff.

Boundary: Owns enterprise routing cost comparison across Google serving surfaces.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch59-calc-google-m2-r1
standard developer API
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$standard developer API; endpoint realm; tool execution surcharges; quota tier; search grounding add-on; monthly baseline estimate; fail-closed notice; verified=2026-09-07Unavailable — frozen calc-google fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-calc-google-m2-r2
Vertex AI enterprise project
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$Vertex AI enterprise project; endpoint realm; tool execution surcharges; quota tier; search grounding add-on; monthly baseline estimate; fail-closed notice; verified=2026-09-07Unavailable — frozen calc-google fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-calc-google-m2-r3
grounding with Google Search
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$grounding with Google Search; endpoint realm; tool execution surcharges; quota tier; search grounding add-on; monthly baseline estimate; fail-closed notice; verified=2026-09-07Unavailable — frozen calc-google fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-calc-google-m2-r4
code execution tool run
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$code execution tool run; endpoint realm; tool execution surcharges; quota tier; search grounding add-on; monthly baseline estimate; fail-closed notice; verified=2026-09-07Unavailable — frozen calc-google fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-calc-google-m2-r5
provisioned throughput
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$provisioned throughput; endpoint realm; tool execution surcharges; quota tier; search grounding add-on; monthly baseline estimate; fail-closed notice; verified=2026-09-07Unavailable — frozen calc-google fixture requires an exact source, identity, and result receiptUNTESTED — module 2 rule is reproducible but no production observation is claimed
batch59-calc-google-m2-r6
unresolved billing tier
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$unresolved billing tier; endpoint realm; tool execution surcharges; quota tier; search grounding add-on; monthly baseline estimate; fail-closed notice; verified=2026-09-07Unavailable — exact calc-google evidence join is not closed for "unresolved billing tier"FAIL CLOSED — manual, probe, or source evidence required

First-party citation: Google Gemini API official pricing. Verified 2026-09-07; missing or conflicting joins fail closed.

Gemini batch vs synchronous economics board

Deterministic formula / rule: batch_savings = standard_cost * 0.50 for 24h asynchronous queue jobs.

Boundary: Owns Gemini batch cost optimization calculations.

Frozen scenario / field IDExact identity and evidence fieldsResultState
batch59-calc-google-m3-r1
classification batch (100K calls)
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$classification batch (100K calls); batch job tokens; 50% discount multiplier; turnaround SLA (24h); estimated net savings; reconciliation rule; verified=2026-09-07Unavailable — frozen calc-google fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-calc-google-m3-r2
document summarization (10K docs)
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$document summarization (10K docs); batch job tokens; 50% discount multiplier; turnaround SLA (24h); estimated net savings; reconciliation rule; verified=2026-09-07Unavailable — frozen calc-google fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-calc-google-m3-r3
mixed text-multimodal batch
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$mixed text-multimodal batch; batch job tokens; 50% discount multiplier; turnaround SLA (24h); estimated net savings; reconciliation rule; verified=2026-09-07Unavailable — frozen calc-google fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-calc-google-m3-r4
urgent SLA bypass
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$urgent SLA bypass; batch job tokens; 50% discount multiplier; turnaround SLA (24h); estimated net savings; reconciliation rule; verified=2026-09-07Unavailable — exact calc-google evidence join is not closed for "urgent SLA bypass"FAIL CLOSED — manual, probe, or source evidence required
batch59-calc-google-m3-r5
expired batch fallback
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$expired batch fallback; batch job tokens; 50% discount multiplier; turnaround SLA (24h); estimated net savings; reconciliation rule; verified=2026-09-07Unavailable — frozen calc-google fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed
batch59-calc-google-m3-r6
cancelled batch reconciliation
route=$/llm-cost-calculator/google; owner=$calc-google; scenario=$cancelled batch reconciliation; batch job tokens; 50% discount multiplier; turnaround SLA (24h); estimated net savings; reconciliation rule; verified=2026-09-07Unavailable — frozen calc-google fixture requires an exact source, identity, and result receiptUNTESTED — module 3 rule is reproducible but no production observation is claimed

First-party citation: Google Gemini API official pricing. Verified 2026-09-07; missing or conflicting joins fail closed.

Method and limitations: this board exposes deterministic rules, first-party citations, and dated evidence identities. It does not invent volume, coverage, entitlement, retention, residency, quota, capacity, feature support, quality, price, reliability, or legal conclusions. Run the calc-google evidence flow →

Ranked cost — 200K requests/month

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
Gemini 2.5 Flash LitebudgetlegacyGoogle$76.00$63.09
Gemini 3.1 Flash LitebudgetlegacyGoogle$225$1790.87×
Gemini 3.5 Flash LitebudgetGoogle$319$280
Gemini 2.5 FlashbudgetlegacyGoogle$319$280
Gemini 3.7 FlashmidGoogle$623$526
Gemini 3.1 FlashmidlegacyGoogle$675$578
Gemini 3.6 FlashmidGoogle$1,245$1,051
Gemini 3.5 FlashmidlegacyGoogle$1,350$1,156
Gemini 3.1 PromidGoogle$1,800$1,2310.63×

Levers live on Google

Prompt caching — available, guide coming soon
Batch API: up to 50%Model verbosity: up to 95%Context trimming: up to 47%

The billing gotcha

Gemini splits pricing by context length on some models — a call that crosses the long-context threshold bills at a higher per-token rate for the whole request, not just the tokens past the line. Google AI Studio's free tier also shares a daily request cap across every model you test, which does not show up in a per-model monthly projection.

Related

Google provider profile →Compare against another provider →Global cost calculator →How to reduce LLM API costs →