OpenAI API Cost Calculator

How much does the OpenAI API cost per month?

For GPT-5 Nano at 200K requests/month, 2,400 input tokens and 350 output tokens per request costs about $22.76 per month after 30% cache use and 100% batch share. Across OpenAI's 15 priced models, the cheapest default ranking is GPT-5 Nano at $52.00 per month.

Verified 2026-06-21

Pricing data as of June 2026. Sources: OpenAI pricing and model documentation. Parameters are shareable in this URL.

GPT-5 Nano: estimated monthly cost

$22.76

Formula: provider-specific cached input multiplier × cache rate, plus uncached input and verbosity-adjusted output; cache writes are amortised over the provider TTL, then batch savings are applied.

Per request$0.000
Per day$0.759
Per month$22.76
Per year$273

OpenAI scenario sensitivity

The default is a server-rendered estimate. Change the cache and batch shares to see when a cheaper qualified model overtakes the selected model; the URL is shareable and preserves the inputs.

Default$22.76/month
No caching or batch$52.00/month
50% cache, 50% batch$30.90/month
100% cache, 100% batch$15.20/month

OpenAI calculator defaults, sensitivity, and crossover

Selected model / assumptionsRequests/dayRequests/monthRequests/yearMonthly / annual cost
GPT-5.6 Terra: 2,400 in + 350 out; 30% cache; 100% batch6,667200,0002,400,000$963.00 / $11556.00
ScenarioCache hitBatch shareMonthly cost
No optimization0%0%$2250.00
Default30%100%$963.00
High cache, sync80%0%$1386.00
High cache + batch80%100%$693.00

Cheaper qualified alternatives

AlternativeDefault monthlyOvertake condition
Amazon Nova Micro$12.12Amazon Nova Micro is cheaper at the default 30% cache / 100% batch mix ($12.12 < $963.00)
Amazon Nova Lite$22.04Amazon Nova Lite is cheaper at the default 30% cache / 100% batch mix ($22.04 < $963.00)
Gemini 2.5 Flash Lite$31.52Gemini 2.5 Flash Lite is cheaper at the default 30% cache / 100% batch mix ($31.52 < $963.00)

The shareable calculation is parameterized by in, out, calls, cache, and batch query values on this page; the static default above is the reproducible landing state. Overtake is recalculated from the same formula as the interactive calculator, not a prose estimate.

Verified 2026-06-21. shareable parameterized scenario

Batch 14 · OpenAI forecast bands, prompt churn, and invoice reconciliation

1. User-entered p50/p95 forecast band

ScenarioInputOutputRetriesCache hitBatch shareForecastPercentile evidence
p502,400350030%100%$18.20Statistical percentile: Unavailable unless observed distribution supplied
p952,400350030%100%$18.20Statistical percentile: Unavailable unless observed distribution supplied

Retries are the user-entered retry rate for both displayed scenarios. Formula: forecast = base token bill × (1 + retries) × (1 − cache-hit share) × (1 − 0.5 × batch share). The percentile labels preserve your entered scenario; they are not statistical claims without an observed distribution.

2. Cache warm-up and prompt-churn curve

Prefix changeCallsMiss inputWrite costRead costOutput costTotal
0%10UnavailableUnavailable$0.0001Unavailable
0%50UnavailableUnavailable$0.0007Unavailable
0%200UnavailableUnavailable$0.0028Unavailable
0%1000UnavailableUnavailable$0.01Unavailable
1%124UnavailableUnavailable$0.0001Unavailable
1%524UnavailableUnavailable$0.0007Unavailable
1%2024UnavailableUnavailable$0.0028Unavailable
1%10024UnavailableUnavailable$0.01Unavailable
5%1120UnavailableUnavailable$0.0001Unavailable
5%5120UnavailableUnavailable$0.0007Unavailable
5%20120UnavailableUnavailable$0.0028Unavailable
5%100120UnavailableUnavailable$0.01Unavailable
20%1480UnavailableUnavailable$0.0001Unavailable
20%5480UnavailableUnavailable$0.0007Unavailable
20%20480UnavailableUnavailable$0.0028Unavailable
20%100480UnavailableUnavailable$0.01Unavailable

Formula: miss input = entered input × changed-prefix share; total = miss cost + write cost + read cost + output cost. Cache write/read prices, TTL, and first sourced break-even call are Unavailable unless the selected registry row provides compatible evidence.

3. Forecast-to-invoice reconciliation

Invoice fieldPredicted parameterObserved bill fieldVariance / mapping
input tokens2,400UnavailableUnavailable
cached inputUnavailableUnavailableUnavailable
output/reasoning output350UnavailableUnavailable
service classUnavailableUnavailableUnavailable
toolsUnavailableUnavailableUnavailable
failed requestsUnavailableUnavailableUnavailable
roundingUnavailable$240.00 fixed observed billFixed observed bill for reconciliation; unexplained variance: Unavailable

Reconciliation formula: observed invoice = input + cached input + output/reasoning output + service class + tools + failed requests, after documented rounding. The fixed observed bill for this example is $240.00; unsupported invoice fields and unexplained variance are Unavailable.

Verified 2026-06-21. Data owner: Luna. “Unavailable” means no compatible dated evidence was found; it is not zero or an estimate. Source / registry · Run this scenario →

Batch 16 · user-entered capacity, environment allocation, and phased rollout

1. Inverse budget-and-capacity solver

BudgetShapeCache/batch/toolsRetries/service classTarget modelMax callsBinding constraint
10002,400 in / 350 out30.0% / 100.0% / $0.00000% / standardGPT-5 Nano8,787,346Monthly budget

Formula / rule: max calls = floor(user-entered budget ÷ (((cached input + output) + tool cost) × (1 − batch share × batch discount)) × (1 + retry rate)); cache and batch discounts apply before the per-call solver math.

2. Development/evaluation/shadow/production allocation forecast

EnvironmentAllocation shareDuplicate shareModel/callsToken/tool spendBudget varianceSuccess probability
development10.0%10.0%GPT-5 Nano / 20,000 / standard$2.5036$997.4964Unavailable
evaluation10.0%20.0%GPT-5 Nano / 20,000 / standard$2.7312$997.2688Unavailable
shadow10.0%5.0%GPT-5 Nano / 20,000 / standard$2.3898$997.6102Unavailable
production70.0%0%GPT-5 Nano / 140,000 / standard$15.9320$984.0680Unavailable

Formula / rule: environment spend = allocation share × calls × discounted token-plus-tool bill × (1 + user-entered duplicate share); service class and target model remain visible inputs, and success probability is never invented.

3. Within-OpenAI phased-rollout forecast

Traffic stageSource/targetAllocation shareDuplicate shareRollback retained stateTools/service classSpendQuality/defect/promotion
1%GPT-5 Nano → GPT-5 Nano1%0%Unavailable$0.0000 / standard$0.2276User-supplied
5.0%GPT-5 Nano → GPT-5 Nano5.0%0%Unavailable$0.0000 / standard$1.1380User-supplied
10.0%GPT-5 Nano → GPT-5 Nano10.0%0%Unavailable$0.0000 / standard$2.2760User-supplied
25.0%GPT-5 Nano → GPT-5 Nano25.0%0%Unavailable$0.0000 / standard$5.6900User-supplied
50.0%GPT-5 Nano → GPT-5 Nano50.0%0%Unavailable$0.0000 / standard$11.3800User-supplied
100.0%GPT-5 Nano → GPT-5 Nano100.0%0%Unavailable$0.0000 / standard$22.7600User-supplied

Formula / rule: rollout spend = allocation share × primary discounted bill + duplicate share × target discounted bill + rollback retained-state/tool spend; target model, tool cost, and service class are user-entered, while quality, defect loss, and promotion thresholds are user-supplied.

Verified 2026-06-21. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means no compatible dated evidence or observed run; it is not zero or an inferred capability. Run this evidence scenario →

Batch 18 · multimodal mix, contract credits, and spend anomaly guardrails

1. Multimodal workload-mix forecast

User-entered unitVolume / compatible rateCache / retry treatmentMonthly subtotalDecision
text tokens2400 in + 350 out30% cache; 0% retryUnavailableUse compatible rate only
imagesUnavailableUnavailableUnavailableUnavailable
audioUnavailableUnavailableUnavailableUnavailable
RealtimeUnavailableUnavailableUnavailableUnavailable
built-in toolsUnavailableUnavailableUnavailableUnavailable
storageUnavailableUnavailableUnavailableUnavailable
accepted outputsUnavailableUnavailableUnavailableUnavailable

Formula / rule: monthly total = Σ(user-entered compatible unit volume × compatible rate × cache/batch/retry treatment); unsupported cross-modal conversions remain Unavailable.

2. Contract-discount and credit burn-down planner

User-entered controlValueGross / eligible discountCredit / overageExhaustion / decision
tier / commit spend$1000.00UnavailableUnavailableUnavailable
discount / creditUnavailableUnavailableUnavailableUnavailable
expiry / overage rateUnavailableUnavailableUnavailableUnavailable
service classstandardUnavailableUnavailableUnavailable

Formula / rule: net spend = gross workload − eligible discount − credit consumed + overage; private terms are user-supplied, not OpenAI facts, and expiry is required for an exhaustion month.

3. Spend-anomaly guardrail planner

User-entered observationValueBaseline / deviation ruleProjected month-endFirst breach / review cost
daily distributionUnavailableUnavailableUnavailableUnavailable
seasonality windowUnavailableUnavailableUnavailableUnavailable
retry/tool mixUnavailableUnavailableUnavailableUnavailable
alert budgetUnavailableUnavailableUnavailableUnavailable

Formula / rule: baseline uses the entered observations and rule; projected month-end = observed baseline × remaining days (plus entered mix). Current tool-only floor is $0.00 for 200000 calls; p95/probability/anomaly claims are Unavailable without observations.

Verified 2026-06-21. Data owner: Luna. Source / registry: dated repository records and matched-run evidence. “Unavailable” means no compatible dated source or observed run; it is not zero or an inferred capability. Run this Batch 18 evidence scenario →

Ranked cost — 200K requests/month

ModelProviderList monthlyEffective monthlyVerbosityRank Δ
GPT-5 NanobudgetlegacyOpenAI$52.00$45.58
GPT-4o MinibudgetlegacyOpenAI$114$94.74
GPT-5.4 NanobudgetlegacyOpenAI$184$1390.78×
GPT-5 MinibudgetlegacyOpenAI$260$228
GPT-5.4 MinimidlegacyOpenAI$675$579
o3-MinimidlegacyOpenAI$836$695
GPT-5.6 LunamidOpenAI$900$772
GPT-5midlegacyOpenAI$1,300$1,139
GPT-4.1midlegacyOpenAI$1,520$1,263
GPT-4omidlegacyOpenAI$1,900$1,579
GPT-5.6 TerramidOpenAI$2,250$1,929
GPT-5.4midlegacyOpenAI$2,250$1,929
GPT-5.6 SolmidOpenAI$3,320$2,806
GPT-4 TurbofrontierlegacyOpenAI$6,900$5,616
GPT-5.4 ProfrontierlegacyOpenAI$27,000$23,6521.04×

Levers live on OpenAI

Prompt caching — available, guide coming soon
Batch API: up to 50%Model verbosity: up to 95%Context trimming: up to 47%

The billing gotcha

OpenAI's usage tiers auto-promote on cumulative spend and account age, not a plan you pick — a new key starts on a low rate-limit tier regardless of the monthly total this calculator projects, so the number below is what you'll pay once you're promoted, not necessarily what you can call today. Batch API halves the price but returns results within 24 hours, not synchronously, so it only helps workloads that can tolerate that delay.

Related

OpenAI provider profile →Try this OpenAI scenario →Compare against another provider →Global cost calculator →How to reduce LLM API costs →