GPT-4 Turbo API Pricing: Legacy Frontier Economics and Migration Guide
Comprehensive GPT-4 Turbo API pricing analysis ($10.00/M input, $30.00/M output), 128K context economics, and clear migration guidance to modern GPT-5 models.
How much does GPT-4 Turbo cost per million tokens?
GPT-4 Turbo costs $10.00 per million input tokens and $30.00 per million output tokens ($15.00/M blended at 3:1). While a pioneering frontier model with 128K context, modern alternatives offer higher intelligence at lower costs. Verified 2026-09-08.
How much does GPT-4 Turbo cost per 1,000 requests?
Computed from generated token pricing. Each row assumes the listed input and output tokens per request; this model has no measured verbosity factor, so the unadjusted output estimate is shown.
| Request shape | Input tokens | Output tokens | Cost / 1,000 requests |
|---|---|---|---|
| Short | 100 | 50 | $2.5000 |
| Medium | 1,000 | 500 | $25.0000 |
| Long | 4,000 | 2,000 | $100.0000 |
Formula: ((input price × input tokens) + (output price × output tokens × verbosity factor)) ÷ 1,000,000 × 1,000. Assumptions: short 100/50, medium 1,000/500, long 4,000/2,000 input/output tokens per request. Unadjusted — no measured verbosity factor is available.
Volume ladder, output-cost sensitivity, and migration ledger for GPT-4 Turbo
1. Fixed-shape monthly cost ladder
| Workload | Input / output tokens | 100K requests | 1M requests | 10M requests |
|---|---|---|---|---|
| Short chat | 500 / 150 | $950.00 | $9500.00 | $95000.00 |
| Code review | 4,000 / 800 | $6400.00 | $64000.00 | $640000.00 |
| Document summary | 16,000 / 2,000 | $22000.00 | $220000.00 | $2200000.00 |
Cost = requests × (input tokens × $10.00/M + output tokens × $30.00/M) ÷ 1,000,000, at three fixed text-token workload shapes. No volume discount, cache, or batch rate is inferred — GPT-4 Turbo's registry entry does not price those tiers.
2. Output-token cost-sensitivity band
| Workload | Input cost (1 request) | Output cost (1 request) | Output share of spend |
|---|---|---|---|
| Short chat | $0.0050 | $0.0045 | 47.4% |
| Code review | $0.04 | $0.02 | 37.5% |
| Document summary | $0.16 | $0.06 | 27.3% |
Output share = output-token spend ÷ (input-token spend + output-token spend) for a single request at each shape. Across these three shapes, GPT-4 Turbo's output share spans 27.3% to 47.4% — a 20.1%-point swing driven entirely by workload shape, not by any change in the $10.00/$30.00 per-1M rates.
3. Successor rate delta and lifecycle ledger
| Rate | GPT-4 Turbo | GPT-5.6 Terra | Delta ($ / %) |
|---|---|---|---|
| Input $/M | $10.00 | $2.50 | $-7.5000 (-75.0%) |
| Output $/M | $30.00 | $15.00 | $-15.0000 (-50.0%) |
Delta = successor rate − GPT-4 Turbo rate, both taken from the current dated pricing registry. Cache, batch, and context-tier pricing are not substituted across models; each figure is model-specific or marked Unavailable.
| Field | Recorded value | Note |
|---|---|---|
| Lifecycle status | legacy | Verified 2026-08-14 |
| Deprecation announced | Unavailable | No inference beyond the dated record |
| Shutdown date | Unavailable | Null/unavailable is not a promise of indefinite availability |
| Successor | GPT-5.6 Terra | gpt-5-6-terra priced in this registry |
| Context window | 128K tokens | Verified 2026-08-14 |
Price verified 2026-04-06; lifecycle verified 2026-08-14. Luna is the data owner. "Unavailable" means no compatible dated evidence was found for that field; it is never treated as zero. Dated price source · lifecycle source · Test GPT-4 Turbo in All AI Ask.
gpt-4-turboGPT-4 Turbo API Pricing: Legacy Frontier Economics and Migration Guide
GPT-4 Turbo costs $10.00 per million input tokens and $30.00 per million output tokens ($15.00/M blended at 3:1). While a pioneering frontier model with 128K context, modern alternatives offer higher intelligence at lower costs. Verified 2026-09-08.
GPT-4 Turbo remains supported but carries a notable cost premium compared to newer architectures.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Standard document extraction (8K in, 1K out): $0.11000 per document |
| Scenario 2 | Legal clause comparison prompt (16K in, 2K out): $0.22000 per analysis |
| Scenario 3 | Customer service multi-turn session (4K in, 500 out): $0.05500 per interaction |
| Scenario 4 | Code generation and refactoring turn (12K in, 2K out): $0.18000 per snippet |
| Scenario 5 | Enterprise monthly quota (10M blended tokens): $150.00 legacy maintenance cost |
| Scenario 6 | High-volume production tier (100M blended tokens): $1,500.00 monthly baseline spend |
Migrating off GPT-4 Turbo delivers dramatic token cost reductions alongside major speed improvements.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Migrating 50M tokens/mo to GPT-5.4 ($5.625/M blended): saves $468.75/mo (62.5% reduction) |
| Scenario 2 | Migrating 50M tokens/mo to GPT-5.4 Mini ($0.6875/M blended): saves $715.63/mo (95.4% reduction) |
| Scenario 3 | Migrating 100M tokens/mo to GPT-5.6 Terra: saves $1,000.00/mo while upgrading intelligence |
| Scenario 4 | Prompt compatibility: drop-in replacement with minimal system prompt adjustments required |
| Scenario 5 | Latency improvements: modern models execute up to 3.2x faster on first-token response |
| Scenario 6 | Payback period on migration engineering: typically under 3 weeks for active production fleets |
Structured shadow validation prevents unexpected edge-case regressions during architecture transitions.
| Scenario | Rendered Evidence & Bounds |
|---|---|
| Scenario 1 | Run golden eval dataset (100 representative queries) against candidate target model |
| Scenario 2 | Verify strict JSON schema adherence and field name matching on structured endpoints |
| Scenario 3 | Check temperature sensitivity: modern models often perform best at lower temperature (0.0–0.2) |
| Scenario 4 | Implement dual-run shadow routing for 72 hours to validate real-world production outputs |
| Scenario 5 | Compare token count distributions: modern tokenizers may consume slightly fewer tokens |
| Scenario 6 | Execute 100% traffic cutover once output parity and latency benchmarks are confirmed |
How fast is GPT-4 Turbo?
How much does GPT-4 Turbo cost at scale?
| Tokens / month | Est. cost (blended 3:1) |
|---|---|
| 100,000 | $1.50 |
| 1,000,000 | $15.00 |
| 10,000,000 | $150.00 |
| 100,000,000 | $1500.00 |
How does GPT-4 Turbo compare with other models?
What is GPT-4 Turbo best for?
What should you explore next for GPT-4 Turbo?
Which GPT-4 Turbo head-to-head comparisons are available?
What are common questions about GPT-4 Turbo?
Is GPT-4 Turbo cheaper than Claude Opus 4.8?
GPT-4 Turbo costs $15.00/M blended tokens, Claude Opus 4.8 costs $10.00/M — Claude Opus 4.8 is cheaper.
How much does 1 million tokens cost with GPT-4 Turbo?
At a 3:1 input:output ratio, 1 million blended tokens costs approximately $15.00. Pure input costs $10.00/M; pure output costs $30.00/M.
What does GPT-4 Turbo cost at high volume?
At 100 million blended tokens a month, GPT-4 Turbo costs approximately $1500.00. See the cost-at-scale table below for other volumes.
