OpenAI API Cost Calculator
How much does the OpenAI API cost per month?
For GPT-5 Nano at 200K requests/month, 2,400 input tokens and 350 output tokens per request costs about $22.76 per month after 30% cache use and 100% batch share. Across OpenAI's 15 priced models, the cheapest default ranking is GPT-5 Nano at $52.00 per month.
Pricing data as of June 2026. Sources: OpenAI pricing and model documentation. Parameters are shareable in this URL.
GPT-5 Nano: estimated monthly cost
Formula: provider-specific cached input multiplier × cache rate, plus uncached input and verbosity-adjusted output; cache writes are amortised over the provider TTL, then batch savings are applied.
| Per request | $0.000 |
|---|---|
| Per day | $0.759 |
| Per month | $22.76 |
| Per year | $273 |
OpenAI scenario sensitivity
The default is a server-rendered estimate. Change the cache and batch shares to see when a cheaper qualified model overtakes the selected model; the URL is shareable and preserves the inputs.
| Default | $22.76/month |
|---|---|
| No caching or batch | $52.00/month |
| 50% cache, 50% batch | $30.90/month |
| 100% cache, 100% batch | $15.20/month |
OpenAI calculator defaults, sensitivity, and crossover
| Selected model / assumptions | Requests/day | Requests/month | Requests/year | Monthly / annual cost |
|---|---|---|---|---|
| GPT-5.6 Terra: 2,400 in + 350 out; 30% cache; 100% batch | 6,667 | 200,000 | 2,400,000 | $963.00 / $11556.00 |
| Scenario | Cache hit | Batch share | Monthly cost |
|---|---|---|---|
| No optimization | 0% | 0% | $2250.00 |
| Default | 30% | 100% | $963.00 |
| High cache, sync | 80% | 0% | $1386.00 |
| High cache + batch | 80% | 100% | $693.00 |
Cheaper qualified alternatives
| Alternative | Default monthly | Overtake condition |
|---|---|---|
| Amazon Nova Micro | $12.12 | Amazon Nova Micro is cheaper at the default 30% cache / 100% batch mix ($12.12 < $963.00) |
| Amazon Nova Lite | $22.04 | Amazon Nova Lite is cheaper at the default 30% cache / 100% batch mix ($22.04 < $963.00) |
| Gemini 2.5 Flash Lite | $31.52 | Gemini 2.5 Flash Lite is cheaper at the default 30% cache / 100% batch mix ($31.52 < $963.00) |
The shareable calculation is parameterized by in, out, calls, cache, and batch query values on this page; the static default above is the reproducible landing state. Overtake is recalculated from the same formula as the interactive calculator, not a prose estimate.
Verified 2026-06-21. shareable parameterized scenario →
Batch 14 · OpenAI forecast bands, prompt churn, and invoice reconciliation
1. User-entered p50/p95 forecast band
| Scenario | Input | Output | Retries | Cache hit | Batch share | Forecast | Percentile evidence |
|---|---|---|---|---|---|---|---|
| p50 | 2,400 | 350 | 0 | 30% | 100% | $18.20 | Statistical percentile: Unavailable unless observed distribution supplied |
| p95 | 2,400 | 350 | 0 | 30% | 100% | $18.20 | Statistical percentile: Unavailable unless observed distribution supplied |
Retries are the user-entered retry rate for both displayed scenarios. Formula: forecast = base token bill × (1 + retries) × (1 − cache-hit share) × (1 − 0.5 × batch share). The percentile labels preserve your entered scenario; they are not statistical claims without an observed distribution.
2. Cache warm-up and prompt-churn curve
| Prefix change | Calls | Miss input | Write cost | Read cost | Output cost | Total |
|---|---|---|---|---|---|---|
| 0% | 1 | 0 | Unavailable | Unavailable | $0.0001 | Unavailable |
| 0% | 5 | 0 | Unavailable | Unavailable | $0.0007 | Unavailable |
| 0% | 20 | 0 | Unavailable | Unavailable | $0.0028 | Unavailable |
| 0% | 100 | 0 | Unavailable | Unavailable | $0.01 | Unavailable |
| 1% | 1 | 24 | Unavailable | Unavailable | $0.0001 | Unavailable |
| 1% | 5 | 24 | Unavailable | Unavailable | $0.0007 | Unavailable |
| 1% | 20 | 24 | Unavailable | Unavailable | $0.0028 | Unavailable |
| 1% | 100 | 24 | Unavailable | Unavailable | $0.01 | Unavailable |
| 5% | 1 | 120 | Unavailable | Unavailable | $0.0001 | Unavailable |
| 5% | 5 | 120 | Unavailable | Unavailable | $0.0007 | Unavailable |
| 5% | 20 | 120 | Unavailable | Unavailable | $0.0028 | Unavailable |
| 5% | 100 | 120 | Unavailable | Unavailable | $0.01 | Unavailable |
| 20% | 1 | 480 | Unavailable | Unavailable | $0.0001 | Unavailable |
| 20% | 5 | 480 | Unavailable | Unavailable | $0.0007 | Unavailable |
| 20% | 20 | 480 | Unavailable | Unavailable | $0.0028 | Unavailable |
| 20% | 100 | 480 | Unavailable | Unavailable | $0.01 | Unavailable |
Formula: miss input = entered input × changed-prefix share; total = miss cost + write cost + read cost + output cost. Cache write/read prices, TTL, and first sourced break-even call are Unavailable unless the selected registry row provides compatible evidence.
3. Forecast-to-invoice reconciliation
| Invoice field | Predicted parameter | Observed bill field | Variance / mapping |
|---|---|---|---|
| input tokens | 2,400 | Unavailable | Unavailable |
| cached input | Unavailable | Unavailable | Unavailable |
| output/reasoning output | 350 | Unavailable | Unavailable |
| service class | Unavailable | Unavailable | Unavailable |
| tools | Unavailable | Unavailable | Unavailable |
| failed requests | Unavailable | Unavailable | Unavailable |
| rounding | Unavailable | $240.00 fixed observed bill | Fixed observed bill for reconciliation; unexplained variance: Unavailable |
Reconciliation formula: observed invoice = input + cached input + output/reasoning output + service class + tools + failed requests, after documented rounding. The fixed observed bill for this example is $240.00; unsupported invoice fields and unexplained variance are Unavailable.
Verified 2026-06-21. Data owner: Luna. “Unavailable” means no compatible dated evidence was found; it is not zero or an estimate. Source / registry · Run this scenario →
Batch 16 · user-entered capacity, environment allocation, and phased rollout
1. Inverse budget-and-capacity solver
| Budget | Shape | Cache/batch/tools | Retries/service class | Target model | Max calls | Binding constraint |
|---|---|---|---|---|---|---|
| 1000 | 2,400 in / 350 out | 30.0% / 100.0% / $0.0000 | 0% / standard | GPT-5 Nano | 8,787,346 | Monthly budget |
Formula / rule: max calls = floor(user-entered budget ÷ (((cached input + output) + tool cost) × (1 − batch share × batch discount)) × (1 + retry rate)); cache and batch discounts apply before the per-call solver math.
2. Development/evaluation/shadow/production allocation forecast
| Environment | Allocation share | Duplicate share | Model/calls | Token/tool spend | Budget variance | Success probability |
|---|---|---|---|---|---|---|
| development | 10.0% | 10.0% | GPT-5 Nano / 20,000 / standard | $2.5036 | $997.4964 | Unavailable |
| evaluation | 10.0% | 20.0% | GPT-5 Nano / 20,000 / standard | $2.7312 | $997.2688 | Unavailable |
| shadow | 10.0% | 5.0% | GPT-5 Nano / 20,000 / standard | $2.3898 | $997.6102 | Unavailable |
| production | 70.0% | 0% | GPT-5 Nano / 140,000 / standard | $15.9320 | $984.0680 | Unavailable |
Formula / rule: environment spend = allocation share × calls × discounted token-plus-tool bill × (1 + user-entered duplicate share); service class and target model remain visible inputs, and success probability is never invented.
3. Within-OpenAI phased-rollout forecast
| Traffic stage | Source/target | Allocation share | Duplicate share | Rollback retained state | Tools/service class | Spend | Quality/defect/promotion |
|---|---|---|---|---|---|---|---|
| 1% | GPT-5 Nano → GPT-5 Nano | 1% | 0% | Unavailable | $0.0000 / standard | $0.2276 | User-supplied |
| 5.0% | GPT-5 Nano → GPT-5 Nano | 5.0% | 0% | Unavailable | $0.0000 / standard | $1.1380 | User-supplied |
| 10.0% | GPT-5 Nano → GPT-5 Nano | 10.0% | 0% | Unavailable | $0.0000 / standard | $2.2760 | User-supplied |
| 25.0% | GPT-5 Nano → GPT-5 Nano | 25.0% | 0% | Unavailable | $0.0000 / standard | $5.6900 | User-supplied |
| 50.0% | GPT-5 Nano → GPT-5 Nano | 50.0% | 0% | Unavailable | $0.0000 / standard | $11.3800 | User-supplied |
| 100.0% | GPT-5 Nano → GPT-5 Nano | 100.0% | 0% | Unavailable | $0.0000 / standard | $22.7600 | User-supplied |
Formula / rule: rollout spend = allocation share × primary discounted bill + duplicate share × target discounted bill + rollback retained-state/tool spend; target model, tool cost, and service class are user-entered, while quality, defect loss, and promotion thresholds are user-supplied.
Verified 2026-06-21. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means no compatible dated evidence or observed run; it is not zero or an inferred capability. Run this evidence scenario →
Batch 18 · multimodal mix, contract credits, and spend anomaly guardrails
1. Multimodal workload-mix forecast
| User-entered unit | Volume / compatible rate | Cache / retry treatment | Monthly subtotal | Decision |
|---|---|---|---|---|
| text tokens | 2400 in + 350 out | 30% cache; 0% retry | Unavailable | Use compatible rate only |
| images | Unavailable | Unavailable | Unavailable | Unavailable |
| audio | Unavailable | Unavailable | Unavailable | Unavailable |
| Realtime | Unavailable | Unavailable | Unavailable | Unavailable |
| built-in tools | Unavailable | Unavailable | Unavailable | Unavailable |
| storage | Unavailable | Unavailable | Unavailable | Unavailable |
| accepted outputs | Unavailable | Unavailable | Unavailable | Unavailable |
Formula / rule: monthly total = Σ(user-entered compatible unit volume × compatible rate × cache/batch/retry treatment); unsupported cross-modal conversions remain Unavailable.
2. Contract-discount and credit burn-down planner
| User-entered control | Value | Gross / eligible discount | Credit / overage | Exhaustion / decision |
|---|---|---|---|---|
| tier / commit spend | $1000.00 | Unavailable | Unavailable | Unavailable |
| discount / credit | Unavailable | Unavailable | Unavailable | Unavailable |
| expiry / overage rate | Unavailable | Unavailable | Unavailable | Unavailable |
| service class | standard | Unavailable | Unavailable | Unavailable |
Formula / rule: net spend = gross workload − eligible discount − credit consumed + overage; private terms are user-supplied, not OpenAI facts, and expiry is required for an exhaustion month.
3. Spend-anomaly guardrail planner
| User-entered observation | Value | Baseline / deviation rule | Projected month-end | First breach / review cost |
|---|---|---|---|---|
| daily distribution | Unavailable | Unavailable | Unavailable | Unavailable |
| seasonality window | Unavailable | Unavailable | Unavailable | Unavailable |
| retry/tool mix | Unavailable | Unavailable | Unavailable | Unavailable |
| alert budget | Unavailable | Unavailable | Unavailable | Unavailable |
Formula / rule: baseline uses the entered observations and rule; projected month-end = observed baseline × remaining days (plus entered mix). Current tool-only floor is $0.00 for 200000 calls; p95/probability/anomaly claims are Unavailable without observations.
Verified 2026-06-21. Data owner: Luna. Source / registry: dated repository records and matched-run evidence. “Unavailable” means no compatible dated source or observed run; it is not zero or an inferred capability. Run this Batch 18 evidence scenario →
Ranked cost — 200K requests/month
| Model | Provider | List monthly | Effective monthly | Verbosity | Rank Δ |
|---|---|---|---|---|---|
| GPT-5 Nanobudgetlegacy | OpenAI | $52.00 | $45.58 | — | — |
| GPT-4o Minibudgetlegacy | OpenAI | $114 | $94.74 | — | — |
| GPT-5.4 Nanobudgetlegacy | OpenAI | $184 | $139 | 0.78× | — |
| GPT-5 Minibudgetlegacy | OpenAI | $260 | $228 | — | — |
| GPT-5.4 Minimidlegacy | OpenAI | $675 | $579 | — | — |
| o3-Minimidlegacy | OpenAI | $836 | $695 | — | — |
| GPT-5.6 Lunamid | OpenAI | $900 | $772 | — | — |
| GPT-5midlegacy | OpenAI | $1,300 | $1,139 | — | — |
| GPT-4.1midlegacy | OpenAI | $1,520 | $1,263 | — | — |
| GPT-4omidlegacy | OpenAI | $1,900 | $1,579 | — | — |
| GPT-5.6 Terramid | OpenAI | $2,250 | $1,929 | — | — |
| GPT-5.4midlegacy | OpenAI | $2,250 | $1,929 | — | — |
| GPT-5.6 Solmid | OpenAI | $3,320 | $2,806 | — | — |
| GPT-4 Turbofrontierlegacy | OpenAI | $6,900 | $5,616 | — | — |
| GPT-5.4 Profrontierlegacy | OpenAI | $27,000 | $23,652 | 1.04× | — |
Levers live on OpenAI
The billing gotcha
OpenAI's usage tiers auto-promote on cumulative spend and account age, not a plan you pick — a new key starts on a low rate-limit tier regardless of the monthly total this calculator projects, so the number below is what you'll pay once you're promoted, not necessarily what you can call today. Batch API halves the price but returns results within 24 hours, not synchronously, so it only helps workloads that can tolerate that delay.
