LLM API Pricing Comparison — 68 Models, Updated April 2026
We compare standard API rates across 68 models, from Amazon Nova Micro at $0.061 per million blended tokens to the most expensive model at GPT-5.4 Pro at $67.50 per million — a 1102x spread. Every price below is sourced directly from the provider's official pricing page and dated with the last verification. Use the calculator to estimate your own monthly cost across all 68 models.
Quick answer
How should I compare LLM API prices?
Compare input, cached-input, output, blended workload cost, context, and freshness in one table. For the single lowest-cost recommendation, use the cheapest LLM API ranking; this page remains the source comparison for the underlying prices.
| Model | Provider | $ / M input | $ / M output | Blended ↓ | Your cost / mo |
|---|---|---|---|---|---|
| Amazon Nova Micro | Amazon | $0.03 | $0.14 | $0.06 | $0.07 |
| Amazon Nova Lite | Amazon | $0.06 | $0.24 | $0.11 | $0.12 |
| Muse Spark 1.3 Contributor | Meta | $0.10 | $0.20 | $0.13 | $0.15 |
| GPT-OSS 20B | Groq | $0.07 | $0.30 | $0.13 | $0.15 |
| Ministral 8B | Mistral | $0.15 | $0.15 | $0.15 | $0.19 |
| GPT-OSS 120B | Groq | $0.15 | $0.60 | $0.26 | $0.30 |
| Mistral Small 3.1 | Mistral | $0.15 | $0.60 | $0.26 | $0.30 |
| Codestral | Mistral | $0.30 | $0.90 | $0.45 | $0.53 |
| GPT-OSS 120B (Cerebras) | Cerebras | $0.35 | $0.75 | $0.45 | $0.54 |
| DeepSeek V4 Flash | DeepSeek | $0.44 | $1.32 | $0.66 | $0.77 |
| Mistral Large 3 | Mistral | $0.50 | $1.50 | $0.75 | $0.88 |
| Gemini 3.5 Flash Lite | $0.30 | $2.50 | $0.85 | $0.93 | |
| Qwen 3.7 Plus | Qwen | $0.80 | $2.00 | $1.10 | $1.30 |
| Qwen 3.8 30B | Groq | $0.60 | $3.00 | $1.20 | $1.35 |
| Amazon Nova Pro | Amazon | $0.80 | $3.20 | $1.40 | $1.60 |
| Gemini 3.7 Flash | $0.75 | $3.75 | $1.50 | $1.69 | |
| Grok 4.3 | xAI | $1.25 | $2.50 | $1.56 | $1.88 |
| DeepSeek V4 Pro | DeepSeek | $1.32 | $3.96 | $1.98 | $2.31 |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | $2.00 | $2.25 |
| Muse Spark 1.3 | Meta | $1.25 | $4.25 | $2.00 | $2.31 |
| GLM-5.2 | Z.ai | $1.40 | $4.40 | $2.15 | $2.50 |
| GPT-5.6 Luna | OpenAI | $1.00 | $6.00 | $2.25 | $2.50 |
| GLM 4.7 (Cerebras) | Cerebras | $2.25 | $2.75 | $2.38 | $2.94 |
| Qwen 3.8 Max | Qwen | $1.60 | $6.40 | $2.80 | $3.20 |
| Qwen 3.7 Max | Qwen | $1.60 | $6.40 | $2.80 | $3.20 |
| Grok-4.20 Reasoning | xAI | $2.00 | $6.00 | $3.00 | $3.50 |
| Grok-4.20 | xAI | $2.00 | $6.00 | $3.00 | $3.50 |
| Grok 4.6 | xAI | $2.00 | $6.00 | $3.00 | $3.50 |
| Grok 4.5 | xAI | $2.00 | $6.00 | $3.00 | $3.50 |
| Gemini 3.6 Flash | $1.50 | $7.50 | $3.00 | $3.38 | |
| Mistral Medium 3 | Mistral | $1.50 | $7.50 | $3.00 | $3.38 |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | $4.00 | $4.50 |
| Gemini 3.1 Pro | $2.00 | $12.00 | $4.50 | $5.00 | |
| GPT-5.6 Terra | OpenAI | $2.50 | $15.00 | $5.63 | $6.25 |
| Claude Sonnet 4.6 | Anthropic | $3.00 | $15.00 | $6.00 | $6.75 |
| GPT-5.6 Sol | OpenAI | $4.00 | $20.00 | $8.00 | $9.00 |
| Claude Opus 4.8 | Anthropic | $5.00 | $25.00 | $10.00 | $11.25 |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | $20.00 | $22.50 |
| Claude Opus 5 | Anthropic | $15.00 | $75.00 | $30.00 | $33.75 |
Three reproducible pricing decisions
Every result is normalized to 100,000 requests. List, prompt-cache, and batch are separate applicable-cost columns; cache uses the provider read multiplier and batch uses a documented 50% scenario only for providers with a reusable rule.
1. Workload-normalized list / cache / batch cost
| Model | Workload | List | Prompt-cache read | Batch scenario |
|---|---|---|---|---|
| GPT-5 Nano | Short chat | $12.00 | $8.40 | $8.00 |
| GPT-5 Nano | Document extraction | $72.00 | $36.00 | $56.00 |
| GPT-5 Nano | Output-heavy drafting | $165.00 | $160.50 | $85.00 |
| Gemini 2.5 Flash Lite | Short chat | $16.00 | $8.80 | $12.00 |
| Gemini 2.5 Flash Lite | Document extraction | $112.00 | $40.00 | $96.00 |
| Gemini 2.5 Flash Lite | Output-heavy drafting | $170.00 | $161.00 | $90.00 |
| GPT-4o Mini | Short chat | $24.00 | $13.20 | $18.00 |
| GPT-4o Mini | Document extraction | $168.00 | $60.00 | $144.00 |
| GPT-4o Mini | Output-heavy drafting | $255.00 | $241.50 | $135.00 |
| Grok-3 Mini | Short chat | $24.00 | Unavailable | Unavailable |
| Grok-3 Mini | Document extraction | $168.00 | Unavailable | Unavailable |
| Grok-3 Mini | Output-heavy drafting | $255.00 | Unavailable | Unavailable |
| GPT-5.4 Nano | Short chat | $41.00 | $26.60 | $28.50 |
| GPT-5.4 Nano | Document extraction | $260.00 | $116.00 | $210.00 |
| GPT-5.4 Nano | Output-heavy drafting | $520.00 | $502.00 | $270.00 |
| Gemini 3.1 Flash Lite | Short chat | $50.00 | $32.00 | $35.00 |
| Gemini 3.1 Flash Lite | Document extraction | $320.00 | $140.00 | $260.00 |
| Gemini 3.1 Flash Lite | Output-heavy drafting | $625.00 | $602.50 | $325.00 |
| DeepSeek V4 Flash | Short chat | $61.60 | Unavailable | Unavailable |
| DeepSeek V4 Flash | Document extraction | $457.60 | Unavailable | Unavailable |
| DeepSeek V4 Flash | Output-heavy drafting | $572.00 | Unavailable | Unavailable |
| GPT-5 Mini | Short chat | $60.00 | $42.00 | $40.00 |
| GPT-5 Mini | Document extraction | $360.00 | $180.00 | $280.00 |
| GPT-5 Mini | Output-heavy drafting | $825.00 | $802.50 | $425.00 |
| Gemini 3.5 Flash Lite | Short chat | $74.00 | $52.40 | $49.00 |
| Gemini 3.5 Flash Lite | Document extraction | $440.00 | $224.00 | $340.00 |
| Gemini 3.5 Flash Lite | Output-heavy drafting | $1030.00 | $1003.00 | $530.00 |
| Gemini 2.5 Flash | Short chat | $74.00 | $52.40 | $49.00 |
| Gemini 2.5 Flash | Document extraction | $440.00 | $224.00 | $340.00 |
| Gemini 2.5 Flash | Output-heavy drafting | $1030.00 | $1003.00 | $530.00 |
| Gemini 3.7 Flash | Short chat | $135.00 | $81.00 | $97.50 |
| Gemini 3.7 Flash | Document extraction | $900.00 | $360.00 | $750.00 |
| Gemini 3.7 Flash | Output-heavy drafting | $1575.00 | $1507.50 | $825.00 |
| Grok 4.3 | Short chat | $150.00 | Unavailable | Unavailable |
| Grok 4.3 | Document extraction | $1200.00 | Unavailable | Unavailable |
| Grok 4.3 | Output-heavy drafting | $1125.00 | Unavailable | Unavailable |
| GPT-5.4 Mini | Short chat | $150.00 | $96.00 | $105.00 |
| GPT-5.4 Mini | Document extraction | $960.00 | $420.00 | $780.00 |
| GPT-5.4 Mini | Output-heavy drafting | $1875.00 | $1807.50 | $975.00 |
| Gemini 3.1 Flash | Short chat | $150.00 | $96.00 | $105.00 |
| Gemini 3.1 Flash | Document extraction | $960.00 | $420.00 | $780.00 |
| Gemini 3.1 Flash | Output-heavy drafting | $1875.00 | $1807.50 | $975.00 |
| o3-Mini | Short chat | $176.00 | $96.80 | $132.00 |
| o3-Mini | Document extraction | $1232.00 | $440.00 | $1056.00 |
| o3-Mini | Output-heavy drafting | $1870.00 | $1771.00 | $990.00 |
| DeepSeek V4 Pro | Short chat | $184.80 | Unavailable | Unavailable |
| DeepSeek V4 Pro | Document extraction | $1372.80 | Unavailable | Unavailable |
| DeepSeek V4 Pro | Output-heavy drafting | $1716.00 | Unavailable | Unavailable |
| Claude Haiku 4.5 | Short chat | $180.00 | $108.00 | $130.00 |
| Claude Haiku 4.5 | Document extraction | $1200.00 | $480.00 | $1000.00 |
| Claude Haiku 4.5 | Output-heavy drafting | $2100.00 | $2010.00 | $1100.00 |
| GPT-5.6 Luna | Short chat | $200.00 | $128.00 | $140.00 |
| GPT-5.6 Luna | Document extraction | $1280.00 | $560.00 | $1040.00 |
| GPT-5.6 Luna | Output-heavy drafting | $2500.00 | $2410.00 | $1300.00 |
| Grok-3 | Short chat | $240.00 | Unavailable | Unavailable |
| Grok-3 | Document extraction | $1920.00 | Unavailable | Unavailable |
| Grok-3 | Output-heavy drafting | $1800.00 | Unavailable | Unavailable |
| Grok-4.20 Reasoning | Short chat | $280.00 | Unavailable | Unavailable |
| Grok-4.20 Reasoning | Document extraction | $2080.00 | Unavailable | Unavailable |
| Grok-4.20 Reasoning | Output-heavy drafting | $2600.00 | Unavailable | Unavailable |
| Grok-4.20 | Short chat | $280.00 | Unavailable | Unavailable |
| Grok-4.20 | Document extraction | $2080.00 | Unavailable | Unavailable |
| Grok-4.20 | Output-heavy drafting | $2600.00 | Unavailable | Unavailable |
| Grok 4.6 | Short chat | $280.00 | Unavailable | Unavailable |
| Grok 4.6 | Document extraction | $2080.00 | Unavailable | Unavailable |
| Grok 4.6 | Output-heavy drafting | $2600.00 | Unavailable | Unavailable |
| Grok 4.5 | Short chat | $280.00 | Unavailable | Unavailable |
| Grok 4.5 | Document extraction | $2080.00 | Unavailable | Unavailable |
| Grok 4.5 | Output-heavy drafting | $2600.00 | Unavailable | Unavailable |
| Gemini 3.6 Flash | Short chat | $270.00 | $162.00 | $195.00 |
| Gemini 3.6 Flash | Document extraction | $1800.00 | $720.00 | $1500.00 |
| Gemini 3.6 Flash | Output-heavy drafting | $3150.00 | $3015.00 | $1650.00 |
| Gemini 3.5 Flash | Short chat | $300.00 | $192.00 | $210.00 |
| Gemini 3.5 Flash | Document extraction | $1920.00 | $840.00 | $1560.00 |
| Gemini 3.5 Flash | Output-heavy drafting | $3750.00 | $3615.00 | $1950.00 |
| GPT-5 | Short chat | $300.00 | $210.00 | $200.00 |
| GPT-5 | Document extraction | $1800.00 | $900.00 | $1400.00 |
| GPT-5 | Output-heavy drafting | $4125.00 | $4012.50 | $2125.00 |
| GPT-4.1 | Short chat | $320.00 | $176.00 | $240.00 |
| GPT-4.1 | Document extraction | $2240.00 | $800.00 | $1920.00 |
| GPT-4.1 | Output-heavy drafting | $3400.00 | $3220.00 | $1800.00 |
| Claude Sonnet 5 | Short chat | $360.00 | $216.00 | $260.00 |
| Claude Sonnet 5 | Document extraction | $2400.00 | $960.00 | $2000.00 |
| Claude Sonnet 5 | Output-heavy drafting | $4200.00 | $4020.00 | $2200.00 |
| GPT-4o | Short chat | $400.00 | $220.00 | $300.00 |
| GPT-4o | Document extraction | $2800.00 | $1000.00 | $2400.00 |
| GPT-4o | Output-heavy drafting | $4250.00 | $4025.00 | $2250.00 |
| Gemini 3.1 Pro | Short chat | $400.00 | $256.00 | $280.00 |
| Gemini 3.1 Pro | Document extraction | $2560.00 | $1120.00 | $2080.00 |
| Gemini 3.1 Pro | Output-heavy drafting | $5000.00 | $4820.00 | $2600.00 |
| GPT-5.6 Terra | Short chat | $500.00 | $320.00 | $350.00 |
| GPT-5.6 Terra | Document extraction | $3200.00 | $1400.00 | $2600.00 |
| GPT-5.6 Terra | Output-heavy drafting | $6250.00 | $6025.00 | $3250.00 |
| GPT-5.4 | Short chat | $500.00 | $320.00 | $350.00 |
| GPT-5.4 | Document extraction | $3200.00 | $1400.00 | $2600.00 |
| GPT-5.4 | Output-heavy drafting | $6250.00 | $6025.00 | $3250.00 |
| Claude Sonnet 4.6 | Short chat | $540.00 | $324.00 | $390.00 |
| Claude Sonnet 4.6 | Document extraction | $3600.00 | $1440.00 | $3000.00 |
| Claude Sonnet 4.6 | Output-heavy drafting | $6300.00 | $6030.00 | $3300.00 |
| Claude Sonnet 4.5 | Short chat | $540.00 | $324.00 | $390.00 |
| Claude Sonnet 4.5 | Document extraction | $3600.00 | $1440.00 | $3000.00 |
| Claude Sonnet 4.5 | Output-heavy drafting | $6300.00 | $6030.00 | $3300.00 |
| Claude Sonnet 4 | Short chat | $540.00 | $324.00 | $390.00 |
| Claude Sonnet 4 | Document extraction | $3600.00 | $1440.00 | $3000.00 |
| Claude Sonnet 4 | Output-heavy drafting | $6300.00 | $6030.00 | $3300.00 |
| GPT-5.6 Sol | Short chat | $720.00 | $432.00 | $520.00 |
| GPT-5.6 Sol | Document extraction | $4800.00 | $1920.00 | $4000.00 |
| GPT-5.6 Sol | Output-heavy drafting | $8400.00 | $8040.00 | $4400.00 |
| Claude Opus 4.8 | Short chat | $900.00 | $540.00 | $650.00 |
| Claude Opus 4.8 | Document extraction | $6000.00 | $2400.00 | $5000.00 |
| Claude Opus 4.8 | Output-heavy drafting | $10500.00 | $10050.00 | $5500.00 |
| Claude Opus 4.7 | Short chat | $900.00 | $540.00 | $650.00 |
| Claude Opus 4.7 | Document extraction | $6000.00 | $2400.00 | $5000.00 |
| Claude Opus 4.7 | Output-heavy drafting | $10500.00 | $10050.00 | $5500.00 |
| Claude Opus 4.6 | Short chat | $900.00 | $540.00 | $650.00 |
| Claude Opus 4.6 | Document extraction | $6000.00 | $2400.00 | $5000.00 |
| Claude Opus 4.6 | Output-heavy drafting | $10500.00 | $10050.00 | $5500.00 |
| Claude Opus 4.5 | Short chat | $900.00 | $540.00 | $650.00 |
| Claude Opus 4.5 | Document extraction | $6000.00 | $2400.00 | $5000.00 |
| Claude Opus 4.5 | Output-heavy drafting | $10500.00 | $10050.00 | $5500.00 |
| GPT-4 Turbo | Short chat | $1400.00 | $680.00 | $1100.00 |
| GPT-4 Turbo | Document extraction | $10400.00 | $3200.00 | $9200.00 |
| GPT-4 Turbo | Output-heavy drafting | $13000.00 | $12100.00 | $7000.00 |
| Claude Fable 5 | Short chat | $1800.00 | $1080.00 | $1300.00 |
| Claude Fable 5 | Document extraction | $12000.00 | $4800.00 | $10000.00 |
| Claude Fable 5 | Output-heavy drafting | $21000.00 | $20100.00 | $11000.00 |
| Claude Opus 5 | Short chat | $2700.00 | $1620.00 | $1950.00 |
| Claude Opus 5 | Document extraction | $18000.00 | $7200.00 | $15000.00 |
| Claude Opus 5 | Output-heavy drafting | $31500.00 | $30150.00 | $16500.00 |
| Claude Opus 4.1 | Short chat | $2700.00 | $1620.00 | $1950.00 |
| Claude Opus 4.1 | Document extraction | $18000.00 | $7200.00 | $15000.00 |
| Claude Opus 4.1 | Output-heavy drafting | $31500.00 | $30150.00 | $16500.00 |
| Claude Opus 4 | Short chat | $2700.00 | $1620.00 | $1950.00 |
| Claude Opus 4 | Document extraction | $18000.00 | $7200.00 | $15000.00 |
| Claude Opus 4 | Output-heavy drafting | $31500.00 | $30150.00 | $16500.00 |
| GPT-5.4 Pro | Short chat | $6000.00 | $3840.00 | $4200.00 |
| GPT-5.4 Pro | Document extraction | $38400.00 | $16800.00 | $31200.00 |
| GPT-5.4 Pro | Output-heavy drafting | $75000.00 | $72300.00 | $39000.00 |
2. Workload crossover
| Workload | Input / output | List winner | List cost / 100K | Applicability note |
|---|---|---|---|---|
| Short chat | 800 / 200 | GPT-5 Nano | $12.00 | Cache/batch columns below retain unavailable states per model |
| Document extraction | 8,000 / 800 | GPT-5 Nano | $72.00 | Cache/batch columns below retain unavailable states per model |
| Output-heavy drafting | 1,000 / 4,000 | GPT-5 Nano | $165.00 | Cache/batch columns below retain unavailable states per model |
Formula: 100,000 × (input tokens × input $/M + output tokens × output $/M) ÷ 1,000,000. A missing cache or batch rule is shown as unavailable, never as free.
3. Dated price-change ledger
| Model | Verified | Before / after delta | Source |
|---|---|---|---|
| GPT-5 Nano | 2026-04-06 | Before value unavailable in current registry | First-party price record |
| Gemini 2.5 Flash Lite | 2026-04-06 | Before value unavailable in current registry | First-party price record |
| GPT-4o Mini | 2026-04-06 | Before value unavailable in current registry | First-party price record |
| Grok-3 Mini | 2026-04-06 | Before value unavailable in current registry | First-party price record |
| GPT-5.4 Nano | 2026-04-06 | Before value unavailable in current registry | First-party price record |
| Gemini 3.1 Flash Lite | 2026-04-06 | Before value unavailable in current registry | First-party price record |
| DeepSeek V4 Flash | 2026-08-14 | Before value unavailable in current registry | First-party price record |
| GPT-5 Mini | 2026-04-06 | Before value unavailable in current registry | First-party price record |
Historical deltas stay unavailable when the repository has only the current value.
Run an attributed cross-model comparison →Verified 2026-08-14. Formulas use the dated registry; unavailable means no model-specific dated source supports a numeric claim. dated normalized pricing registry.
Batch 14 · rate freshness, unit normalization, and market shock governance
1. Rate-freshness control board
| Age band | Rows | Source reachability | Lifecycle | Unit completeness | Comparison treatment |
|---|---|---|---|---|---|
| 0–7 days | 20 | Registry source URL present | Registry lifecycle filter applied | Text input/output rows only | Eligible only if units match |
| 8–30 days | 2 | Registry source URL present | Registry lifecycle filter applied | Text input/output rows only | Eligible only if units match |
| 31–60 days | 6 | Registry source URL present | Registry lifecycle filter applied | Text input/output rows only | Eligible only if units match |
| 61–90 days | 8 | Registry source URL present | Registry lifecycle filter applied | Text input/output rows only | Eligible only if units match |
| 90+ days | 32 | Registry source URL present | Registry lifecycle filter applied | Text input/output rows only | Exclude all stale rows |
Formula: age = verification date − source date. Rows are excluded when stale, unreachable, lifecycle-incompatible, or unit-incomplete; no stale row is silently carried forward.
2. Cross-provider unit-normalization audit
| Unit family | Registry fields | Compatible comparison | Blended price | Verdict |
|---|---|---|---|---|
| text input/output | inputPerM + outputPerM | Like-for-like token shape | Never blend unlike units | Compatible |
| cached tokens | Unavailable | Unavailable | Unavailable | Unavailable — no cross-unit blend |
| image/audio/video | Unavailable | Unavailable | Unavailable | Unavailable — no cross-unit blend |
| tools/search | Unavailable | Unavailable | Unavailable | Unavailable — no cross-unit blend |
| storage | Unavailable | Unavailable | Unavailable | Unavailable — no cross-unit blend |
| batch | Unavailable | Unavailable | Unavailable | Unavailable — no cross-unit blend |
| long-context tiers | Unavailable | Unavailable | Unavailable | Unavailable — no cross-unit blend |
3. Neutral portfolio-shock table
| Basket | Affected models/providers | Base bill | 0% shock | 10% shock | 25% shock | 50% shock | Concentration |
|---|---|---|---|---|---|---|---|
| text-heavy | OpenAI/GPT-5 Nano · Anthropic/Claude Haiku 4.5 · Google/Gemini 2.5 Flash Lite · xAI/Grok-3 Mini | $0.02 | $0.02 | $0.02 | $0.03 | $0.03 | 77.3% top-model share |
| balanced | OpenAI/GPT-5 Nano · Anthropic/Claude Haiku 4.5 · Google/Gemini 2.5 Flash Lite · xAI/Grok-3 Mini | $0.03 | $0.03 | $0.03 | $0.04 | $0.04 | 77.7% top-model share |
| output-heavy | OpenAI/GPT-5 Nano · Anthropic/Claude Haiku 4.5 · Google/Gemini 2.5 Flash Lite · xAI/Grok-3 Mini | $0.06 | $0.06 | $0.06 | $0.07 | $0.08 | 78.0% top-model share |
Formula: shocked bill = base compatible bill × (1 + rate shock). This table reports affected rows and spend risk; it does not name a universal cheapest provider.
Verified 2026-08-14. Data owner: Luna. “Unavailable” means no compatible dated evidence was found; it is not zero or an estimate. Source / registry · Run this scenario →
Batch 16 · neutral channel parity, tariff transitions, and missing-unit sensitivity
1. Direct-provider versus cloud-marketplace parity
| Model/version | Region | Context tier | Input/output/cache/batch/tool | Currency/date | Missing cells | Verdict |
|---|---|---|---|---|---|---|
| Amazon Nova Micro | Unavailable | Unavailable | $0.034999999999999996/M · $0.13999999999999999/M · Unavailable | USD / 2026-06-14 | Unavailable | No parity verdict |
| Amazon Nova Lite | Unavailable | Unavailable | $0.060000000000000005/M · $0.24000000000000002/M · Unavailable | USD / 2026-06-14 | Unavailable | No parity verdict |
| Muse Spark 1.3 Contributor | Unavailable | Unavailable | $0.1/M · $0.2/M · Unavailable | USD / 2026-08-14 | Unavailable | No parity verdict |
| GPT-OSS 20B | Unavailable | Unavailable | $0.075/M · $0.3/M · Unavailable | USD / 2026-04-06 | Unavailable | No parity verdict |
| GPT-5 Nano | Unavailable | Unavailable | $0.05/M · $0.4/M · Unavailable | USD / 2026-04-06 | Unavailable | No parity verdict |
| Ministral 8B | Unavailable | Unavailable | $0.15/M · $0.15/M · Unavailable | USD / 2026-08-14 | Unavailable | No parity verdict |
Formula / rule: parity requires exact model/version, region, tier, units, currency, and effective date; no cheapest-channel verdict is issued.
2. Tariff effective-date transition board
| Model/alias | Old row | New row | Announcement | Effective | Observed | Basket delta | Stale quarantine |
|---|---|---|---|---|---|---|---|
| Amazon Nova Micro | Unavailable | Unavailable | Unavailable | Unavailable | 2026-06-14 | Unavailable | Quarantine until joined |
| Amazon Nova Lite | Unavailable | Unavailable | Unavailable | Unavailable | 2026-06-14 | Unavailable | Quarantine until joined |
| Muse Spark 1.3 Contributor | Unavailable | Unavailable | Unavailable | Unavailable | 2026-08-14 | Unavailable | Quarantine until joined |
| GPT-OSS 20B | Unavailable | Unavailable | Unavailable | Unavailable | 2026-04-06 | Unavailable | Quarantine until joined |
| GPT-5 Nano | Unavailable | Unavailable | Unavailable | Unavailable | 2026-04-06 | Unavailable | Quarantine until joined |
Formula / rule: eligible row = observed date ≥ effective date and not superseded; future, stale, and alias-unmatched rows are quarantined before comparison.
3. Missing-unit sensitivity frontier
| Attempt | Known bill | Unknown unit | Switching threshold | Neutral result |
|---|---|---|---|---|
| text | Unavailable | Unavailable | Unavailable | Unavailable |
| media | Unavailable | Unavailable | Unavailable | Unavailable |
| tool/search | Unavailable | Unavailable | Unavailable | Unavailable |
| storage | Unavailable | Unavailable | Unavailable | Unavailable |
| long-context | Unavailable | Unavailable | Unavailable | Unavailable |
Formula / rule: threshold = (competing known bill − known bill) ÷ unknown quantity; if the quantity or compatible unit is absent, result is Unavailable rather than blended.
Verified 2026-08-14. Data owner: Luna. Source / registry: dated repository pricing and provider records. “Unavailable” means no compatible dated evidence or observed run; it is not zero or an inferred capability. Run this evidence scenario →
Batch 18 · billing increments, neutral normalization, and source conflict quarantine
1. Billing-granularity and rounding audit
| Workload | Token rounding / minimum | Free allowance | Invoice precision | Theoretical / sourced bill |
|---|---|---|---|---|
| micro request | Unavailable | Unavailable | Unavailable | Unavailable |
| short request | Unavailable | Unavailable | Unavailable | Unavailable |
| long request | Unavailable | Unavailable | Unavailable | Unavailable |
Formula / rule: effective cost is Unavailable when billing increment, minimum, allowance, or invoice precision is undocumented; theoretical bills never replace sourced bills.
2. Region-currency-tax comparability board
| Model/version | Region / currency | Unit basis / FX date | Tax / surcharge / effective | Pretax USD / eligibility |
|---|---|---|---|---|
| compatible row A | Unavailable | Unavailable | Unavailable | Unavailable |
| compatible row B | Unavailable | Unavailable | Unavailable | Unavailable |
| compatible row C | Unavailable | Unavailable | Unavailable | Unavailable |
Formula / rule: normalized pretax USD = quoted compatible unit × dated FX, excluding tax; every required field must match before inclusion, and no cheapest provider is named.
3. Price-source conflict resolver
| Model/version | Provider docs / console | Marketplace / registry | Disagreement / recency | Quarantine / precedence |
|---|---|---|---|---|
| row A | Unavailable | Unavailable | Unavailable | Unavailable |
| row B | Unavailable | Unavailable | Unavailable | Unavailable |
| row C | Unavailable | Unavailable | Unavailable | Unavailable |
Formula / rule: join by model/version and effective date; quarantine disagreement until deterministic precedence (first-party effective documentation, then dated export, then marketplace/registry) is satisfied.
Verified 2026-08-14. Data owner: Luna. Source / registry: dated repository records and matched-run evidence. “Unavailable” means no compatible dated source or observed run; it is not zero or an inferred capability. Run this Batch 18 evidence scenario →
For cross-dataset findings and a citation-ready snapshot, read the latest LLM pricing and performance report.
Cheapest by tier
Popular head-to-heads
Methodology & freshness
Prices are hand-verified against each provider's official pricing page and stored with a source URL and verification date. "Blended cost" assumes 3 input tokens per 1 output token. The oldest verified price in this dataset is from 2026-04-06; the newest is from 2026-08-14. Legacy models remain listed (toggle above) because they still see real API traffic — each links to its recommended successor. Cite this: All AI Ask LLM API Pricing Dataset, retrieved 2026-08-14.
FAQ
What does this LLM API pricing comparison include?
This hub compares standard input, cached-input where available, and output prices across 68 models, with verification dates and a consistent blended-cost calculation. For a ranked answer to the separate cheapest-LLM-API question, see the dedicated cost ranking.
Why is output more expensive than input?
Generating a token requires a full forward pass through the model, while processing an input token can be batched and cached. Providers price output 3–6x higher than input to reflect this compute cost.
Do these prices include caching discounts?
No — these are the standard, non-cached list prices published by each provider. Prompt caching (where available) can cut input costs significantly for repeated context.
How often are these prices updated?
We verify prices against each provider's official pricing page and re-check on every model addition. The most recently verified price in this table is dated 2026-08-14; the oldest is 2026-04-06.
What does "blended cost" mean?
Blended cost assumes 3 input tokens for every 1 output token — a rough approximation of typical chat/agent usage — computed as (3 × input price + 1 × output price) / 4.
