GPT-5 API Pricing: Flagship Reasoning and Agentic Workflows
Comprehensive GPT-5 API pricing analysis ($1.25/M input, $10.00/M output), reasoning depth benchmarks, prompt caching breaks, and enterprise ROI modeling.
How much does GPT-5 cost per million tokens?
GPT-5 costs $1.25 per million input tokens and $10.00 per million output tokens ($3.4375/M blended at 3:1). Provides frontier cognitive performance across advanced coding, analysis, and autonomous agent orchestration. Verified 2026-09-08.
How much does GPT-5 cost per 1,000 requests?
Computed from generated token pricing. Each row assumes the listed input and output tokens per request; this model has no measured verbosity factor, so the unadjusted output estimate is shown.
| Request shape | Input tokens | Output tokens | Cost / 1,000 requests |
|---|---|---|---|
| Short | 100 | 50 | $0.6250 |
| Medium | 1,000 | 500 | $6.2500 |
| Long | 4,000 | 2,000 | $25.0000 |
Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Unadjusted — no measured verbosity factor is available.
Three model-specific legacy pricing decisions
GPT-5 owns the historical reasoning, coding, and agent bill. GPT-5.6 Sol is the designated successor input, not a broad family verdict.
1. Fixed model-specific workload bills
| Workload | Input / output | 100K requests | Evidence boundary |
|---|---|---|---|
| Reasoning | 8,000 / 1,200 | $3400.00 | 2× output expansion is explicit |
| Coding | 4,000 / 1,000 | $1500.00 | Text tokens |
| Agent | 16,000 / 2,000 | $4000.00 | Text tokens |
Formula: requests × (input tokens × input $/M + output tokens × output $/M × output expansion) ÷ 1,000,000. Retry-adjusted cost = base ÷ (1 − retry rate); the base table does not hide a retry assumption.
2. GPT-5 → Sol accepted-result crossover
| Fixed shape | GPT-5 | GPT-5.6 Sol | Numeric decision boundary |
|---|---|---|---|
| Reasoning · 8,000 / 1,200 | $3777.78 | $5957.45 | 58% accepted-result uplift required after fixed retry assumptions (10% → 6%) |
| Coding · 4,000 / 1,000 | $1666.67 | $3829.79 | 58% accepted-result uplift required after fixed retry assumptions (10% → 6%) |
This is a cost-per-accepted-result threshold, not a measured quality claim. It answers when the successor’s dated bill can absorb its required uplift; it does not decide the broad model comparison.
3. Staged migration budget
| Traffic / evidence | Result | Safe treatment |
|---|---|---|
| 100K requests | $3777.78 | Fixed first workload; text-token units |
| 1M requests | $34000.00 | Linear token spend only; no volume discount inferred |
| 10M requests | $340000.00 | Budget exposure; quota and latency remain unavailable |
| Cache and batch: Unavailable | Unavailable | Do not infer, zero-price, or import a neighboring model’s mechanic |
| Shutdown date: Unavailable | Unavailable | Do not infer, zero-price, or import a neighboring model’s mechanic |
| Benchmark coverage: Unavailable for this exact historical row | Unavailable | Do not infer, zero-price, or import a neighboring model’s mechanic |
| Lifecycle | legacy; no sourced announcement date | No sourced shutdown date; revalidate before current claims |
Staged traffic allocation budget
| GPT-5 / Sol traffic | Monthly blended cost | Budget delta vs 0% Sol | Calculation |
|---|---|---|---|
| 100% / 0% | $34000.00 | $0.0000 vs 0% transfer | $34000.00 legacy + $80000.00 successor monthly bills |
| 75% / 25% | $45500.00 | $11500.00 vs 0% transfer | $34000.00 legacy + $80000.00 successor monthly bills |
| 50% / 50% | $57000.00 | $23000.00 vs 0% transfer | $34000.00 legacy + $80000.00 successor monthly bills |
| 0% / 100% | $80000.00 | $46000.00 vs 0% transfer | $34000.00 legacy + $80000.00 successor monthly bills |
Monthly budget uses 1,000,000 requests of the fixed reasoning shape, including its explicit 2× output expansion. Cache, batch, shutdown, and benchmark evidence are unavailable.
Price verified 2026-04-06; lifecycle verified 2026-08-14. Luna is the data owner. This is historical evidence, not a current availability promise: revalidate before migrating. “Unavailable” means no compatible dated evidence was found; it is never treated as zero. First-party price source · lifecycle source · Test this model in All AI Ask.
Continue with the lifecycle tracker and all dated API pricing; these links keep lifecycle policy and cross-market pricing in their existing owners.
All three Batch 6 contributions are server-rendered for GPT-5; fixed inputs, formulas, dated provenance, successor boundary, lifecycle state, and missing-data treatment remain visible.
gpt-5GPT-5 API Pricing: Flagship Reasoning and Agentic Workflows
GPT-5 costs $1.25 per million input tokens and $10.00 per million output tokens ($3.4375/M blended at 3:1). Provides frontier cognitive performance across advanced coding, analysis, and autonomous agent orchestration. Verified 2026-09-08.
GPT-5 balances top-tier analytical reasoning with competitive token pricing for serious development.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Complex multi-file refactoring prompt (16K in, 4K out): $0.06000 per run |
| Scenario 2 | Financial earnings report deep analysis (32K in, 2K out): $0.06000 per report |
| Scenario 3 | Autonomous research agent loop (64K in, 8K out): $0.16000 per research cycle |
| Scenario 4 | Full-stack feature implementation (128K in, 16K out): $0.32000 per feature PR |
| Scenario 5 | Competitive intelligence landscape audit (48K in, 6K out): $0.12000 per dossier |
| Scenario 6 | Monthly enterprise developer tier (100M blended tokens): $343.75 per seat allocation |
Prompt caching dramatically expands sustainable context window depth without compounding query budgets.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Monorepo AST cache reuse (32K prefix, 4K new in): 42.8% input cost savings |
| Scenario 2 | Enterprise knowledge base context (64K prefix, 8K new in): 44.4% input cost savings |
| Scenario 3 | Interactive multi-turn coding session (10 turns cached): 46.1% cumulative input savings |
| Scenario 4 | API documentation schema cache: $0.02000 savings per developer interaction turn |
| Scenario 5 | Long-form conversation state caching: amortizes heavy multi-turn context costs |
| Scenario 6 | Break-even reached on turn 2 of identical agent tool definitions |
Frontier reasoning spend is heavily offset by the elimination of manual debugging and deployment rollbacks.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Production bug prevention in CI/CD pipeline: $1,200 saved per caught regression |
| Scenario 2 | Automated test suite generation across 50 services: $48.00 API cost vs 20 dev hours |
| Scenario 3 | Database schema migration validation: zero downtime deployment insurance |
| Scenario 4 | Security vulnerability detection in third-party dependencies: immediate compliance ROI |
| Scenario 5 | 1,000 pull request review runs ($160 total spend): caught 14 critical edge-case bugs |
| Scenario 6 | Net positive software engineering ROI estimated at >15x API expenditures |
How fast is GPT-5?
How much does GPT-5 cost at scale?
| Tokens / month | Est. cost (blended 3:1) |
|---|---|
| 100,000 | $0.34 |
| 1,000,000 | $3.44 |
| 10,000,000 | $34.38 |
| 100,000,000 | $343.75 |
How does GPT-5 compare with other models?
What should you explore next for GPT-5?
What are common questions about GPT-5?
Is GPT-5 cheaper than Gemini 3.5 Flash?
GPT-5 costs $3.44/M blended tokens, Gemini 3.5 Flash costs $3.38/M — Gemini 3.5 Flash is cheaper.
How much does 1 million tokens cost with GPT-5?
At a 3:1 input:output ratio, 1 million blended tokens costs approximately $3.44. Pure input costs $1.25/M; pure output costs $10.00/M.
What does GPT-5 cost at high volume?
At 100 million blended tokens a month, GPT-5 costs approximately $343.75. See the cost-at-scale table below for other volumes.
