Monthly State of LLM Pricing release · Q3 2026
State of LLM API Pricing and Performance
A reproducible snapshot of verified API pricing and controlled gateway performance. The report is descriptive: it does not declare a universal best model.
Data as of 2026-08-14. Permanent edition URL: https://allaiask.com/reports/llm-state-q3-2026. Release archive: /reports
Key findings
- The pricing table covers 68 models. The cheapest blended price is $0.06/M for Amazon Nova Micro; the highest is $67.50/M for GPT-5.4 Pro. [P]
- The median blended price in the published pricing rows is $2.15/M. [P]
- GPT-OSS 120B (Cerebras) is the fastest measured model at 2450 tokens/sec, while Claude Fable 5 is at 41 tokens/sec: a 59.8× spread within this snapshot. [S]
- Throughput and price are separate dimensions: the speed table publishes TTFT, median throughput, p95 throughput, blended price, and cost per throughput unit so readers can make a workload-specific choice. [P][S]
Accessible tables
These tables are the accessible, machine-readable counterpart to any visual comparison. They contain the full rows used for the findings above; use the downloads for analysis.
Pricing snapshot
| Model | Provider | Input / M | Output / M | Blended / M | Verified |
|---|---|---|---|---|---|
| Amazon Nova Micro | Amazon | $0.03 | $0.14 | $0.06 | 2026-06-14 |
| Amazon Nova Lite | Amazon | $0.06 | $0.24 | $0.11 | 2026-06-14 |
| Muse Spark 1.3 Contributor | Meta | $0.10 | $0.20 | $0.13 | 2026-08-14 |
| GPT-OSS 20B | Groq | $0.07 | $0.30 | $0.13 | 2026-04-06 |
| GPT-5 Nano | OpenAI | $0.05 | $0.40 | $0.14 | 2026-04-06 |
| Ministral 8B | Mistral | $0.15 | $0.15 | $0.15 | 2026-08-14 |
| Gemini 2.5 Flash Lite | $0.10 | $0.40 | $0.18 | 2026-04-06 | |
| GPT-4o Mini | OpenAI | $0.15 | $0.60 | $0.26 | 2026-04-06 |
| Grok-3 Mini | xAI | $0.15 | $0.60 | $0.26 | 2026-04-06 |
| GPT-OSS 120B | Groq | $0.15 | $0.60 | $0.26 | 2026-04-06 |
| Mistral Small 3.1 | Mistral | $0.15 | $0.60 | $0.26 | 2026-08-14 |
| Llama 4 Maverick | Groq | $0.20 | $0.60 | $0.30 | 2026-07-10 |
| Codestral | Mistral | $0.30 | $0.90 | $0.45 | 2026-06-14 |
| GPT-OSS 120B (Cerebras) | Cerebras | $0.35 | $0.75 | $0.45 | 2026-06-14 |
| GPT-5.4 Nano | OpenAI | $0.20 | $1.25 | $0.46 | 2026-04-06 |
| Gemini 3.1 Flash Lite | $0.25 | $1.50 | $0.56 | 2026-04-06 | |
| DeepSeek V4 Flash | DeepSeek | $0.44 | $1.32 | $0.66 | 2026-08-14 |
| GPT-5 Mini | OpenAI | $0.25 | $2.00 | $0.69 | 2026-04-06 |
| Mistral Large 3 | Mistral | $0.50 | $1.50 | $0.75 | 2026-08-14 |
| Gemini 3.5 Flash Lite | $0.30 | $2.50 | $0.85 | 2026-08-14 | |
| Gemini 2.5 Flash | $0.30 | $2.50 | $0.85 | 2026-04-06 | |
| GLM-5.1 | Z.ai | $0.60 | $2.20 | $1.00 | 2026-06-19 |
| Qwen 3.7 Plus | Qwen | $0.80 | $2.00 | $1.10 | 2026-07-23 |
| Qwen 3.8 30B | Groq | $0.60 | $3.00 | $1.20 | 2026-07-10 |
| Qwen 3.6 27B | Groq | $0.60 | $3.00 | $1.20 | 2026-06-19 |
| Amazon Nova Pro | Amazon | $0.80 | $3.20 | $1.40 | 2026-06-14 |
| Gemini 3.7 Flash | $0.75 | $3.75 | $1.50 | 2026-08-14 | |
| Grok 4.3 | xAI | $1.25 | $2.50 | $1.56 | 2026-05-19 |
| GPT-5.4 Mini | OpenAI | $0.75 | $4.50 | $1.69 | 2026-04-06 |
| Gemini 3.1 Flash | $0.75 | $4.50 | $1.69 | 2026-04-06 | |
| o3-Mini | OpenAI | $1.10 | $4.40 | $1.93 | 2026-04-06 |
| DeepSeek V4 Pro | DeepSeek | $1.32 | $3.96 | $1.98 | 2026-08-14 |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | $2.00 | 2026-04-06 |
| Muse Spark 1.3 | Meta | $1.25 | $4.25 | $2.00 | 2026-08-14 |
| GLM-5.2 | Z.ai | $1.40 | $4.40 | $2.15 | 2026-06-19 |
| GPT-5.6 Luna | OpenAI | $1.00 | $6.00 | $2.25 | 2026-08-14 |
| GLM 4.7 (Cerebras) | Cerebras | $2.25 | $2.75 | $2.38 | 2026-06-14 |
| Grok-3 | xAI | $2.00 | $4.00 | $2.50 | 2026-04-06 |
| Qwen 3.8 Max | Qwen | $1.60 | $6.40 | $2.80 | 2026-07-10 |
| Qwen 3.7 Max | Qwen | $1.60 | $6.40 | $2.80 | 2026-07-23 |
| Grok-4.20 Reasoning | xAI | $2.00 | $6.00 | $3.00 | 2026-04-06 |
| Grok-4.20 | xAI | $2.00 | $6.00 | $3.00 | 2026-04-06 |
| Grok 4.6 | xAI | $2.00 | $6.00 | $3.00 | 2026-08-14 |
| Grok 4.5 | xAI | $2.00 | $6.00 | $3.00 | 2026-08-14 |
| Gemini 3.6 Flash | $1.50 | $7.50 | $3.00 | 2026-08-14 | |
| Mistral Medium 3 | Mistral | $1.50 | $7.50 | $3.00 | 2026-08-14 |
| Gemini 3.5 Flash | $1.50 | $9.00 | $3.38 | 2026-08-14 | |
| GPT-5 | OpenAI | $1.25 | $10.00 | $3.44 | 2026-04-06 |
| GPT-4.1 | OpenAI | $2.00 | $8.00 | $3.50 | 2026-04-06 |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | $4.00 | 2026-08-14 |
| GPT-4o | OpenAI | $2.50 | $10.00 | $4.38 | 2026-04-06 |
| Gemini 3.1 Pro | $2.00 | $12.00 | $4.50 | 2026-04-06 | |
| GPT-5.6 Terra | OpenAI | $2.50 | $15.00 | $5.63 | 2026-08-14 |
| GPT-5.4 | OpenAI | $2.50 | $15.00 | $5.63 | 2026-04-06 |
| Claude Sonnet 4.6 | Anthropic | $3.00 | $15.00 | $6.00 | 2026-04-06 |
| Claude Sonnet 4.5 | Anthropic | $3.00 | $15.00 | $6.00 | 2026-04-06 |
| Claude Sonnet 4 | Anthropic | $3.00 | $15.00 | $6.00 | 2026-04-06 |
| GPT-5.6 Sol | OpenAI | $4.00 | $20.00 | $8.00 | 2026-08-14 |
| Claude Opus 4.8 | Anthropic | $5.00 | $25.00 | $10.00 | 2026-06-07 |
| Claude Opus 4.7 | Anthropic | $5.00 | $25.00 | $10.00 | 2026-04-06 |
| Claude Opus 4.6 | Anthropic | $5.00 | $25.00 | $10.00 | 2026-04-06 |
| Claude Opus 4.5 | Anthropic | $5.00 | $25.00 | $10.00 | 2026-04-06 |
| GPT-4 Turbo | OpenAI | $10.00 | $30.00 | $15.00 | 2026-04-06 |
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | $20.00 | 2026-08-14 |
| Claude Opus 5 | Anthropic | $15.00 | $75.00 | $30.00 | 2026-08-09 |
| Claude Opus 4.1 | Anthropic | $15.00 | $75.00 | $30.00 | 2026-04-06 |
| Claude Opus 4 | Anthropic | $15.00 | $75.00 | $30.00 | 2026-04-06 |
| GPT-5.4 Pro | OpenAI | $30.00 | $180.00 | $67.50 | 2026-04-06 |
Performance and cost-per-speed snapshot
| Rank | Model | Provider | TTFT ms | Tokens/sec | p95 tokens/sec | Blended / M | $/M ÷ t/s |
|---|---|---|---|---|---|---|---|
| 1 | GPT-OSS 120B (Cerebras) | Cerebras | 90 | 2450 | 2082.5 | $0.45 | $0.0002 |
| 2 | GLM 4.7 (Cerebras) | Cerebras | 110 | 1980 | 1702.8 | $2.38 | $0.0012 |
| 3 | GPT-OSS 20B | Groq | 140 | 1120 | 974.4 | $0.13 | $0.0001 |
| 4 | GPT-OSS 120B | Groq | 160 | 780 | 639.6 | $0.26 | $0.0003 |
| 5 | Qwen 3.8 30B | Groq | 150 | 690 | 634.8 | $1.20 | $0.0017 |
| 6 | Amazon Nova Micro | Amazon | 220 | 168 | 142.8 | $0.06 | $0.0004 |
| 7 | Gemini 3.5 Flash Lite | 240 | 162 | 129.6 | $0.85 | $0.0052 | |
| 8 | Ministral 8B | Mistral | 210 | 158 | 131.14 | $0.15 | $0.0009 |
| 9 | Claude Haiku 4.5 | Anthropic | 260 | 148 | 118.4 | $2.00 | $0.0135 |
| 10 | DeepSeek V4 Flash | DeepSeek | 280 | 132 | 118.8 | $0.66 | $0.0050 |
| 11 | GPT-5.6 Luna | OpenAI | 300 | 126 | 108.36 | $2.25 | $0.0179 |
| 12 | Mistral Small 3.1 | Mistral | 260 | 121 | 110.11 | $0.26 | $0.0022 |
| 13 | Codestral | Mistral | 270 | 118 | 99.12 | $0.45 | $0.0038 |
| 14 | Gemini 3.6 Flash | 310 | 114 | 93.48 | $3.00 | $0.0263 | |
| 15 | Amazon Nova Lite | Amazon | 250 | 108 | 95.04 | $0.11 | $0.0010 |
| 16 | Grok-4.20 | xAI | 290 | 104 | 91.52 | $3.00 | $0.0288 |
| 17 | Grok 4.3 | xAI | 320 | 98 | 82.32 | $1.56 | $0.0159 |
| 18 | Mistral Medium 3 | Mistral | 320 | 92 | 73.6 | $3.00 | $0.0326 |
| 19 | Qwen 3.7 Plus | Qwen | 340 | 84 | 75.6 | $1.10 | $0.0131 |
| 20 | GPT-5.6 Terra | OpenAI | 380 | 78 | 63.96 | $5.63 | $0.0721 |
| 21 | Claude Sonnet 4.6 | Anthropic | 360 | 76 | 69.92 | $6.00 | $0.0789 |
| 22 | DeepSeek V4 Pro | DeepSeek | 480 | 68 | 62.56 | $1.98 | $0.0291 |
| 23 | Amazon Nova Pro | Amazon | 350 | 64 | 55.04 | $1.40 | $0.0219 |
| 24 | Mistral Large 3 | Mistral | 400 | 61 | 50.02 | $0.75 | $0.0123 |
| 25 | Claude Opus 4.8 | Anthropic | 470 | 58 | 53.36 | $10.00 | $0.1724 |
| 26 | Gemini 3.1 Pro | 420 | 55 | 50.6 | $4.50 | $0.0818 | |
| 27 | Grok-4.20 Reasoning | xAI | 540 | 52 | 46.8 | $3.00 | $0.0577 |
| 28 | Qwen 3.7 Max | Qwen | 460 | 49 | 45.08 | $2.80 | $0.0571 |
| 29 | Qwen 3.8 Max | Qwen | 470 | 47 | 37.6 | $2.80 | $0.0596 |
| 30 | GPT-5.6 Sol | OpenAI | 560 | 44 | 35.64 | $8.00 | $0.1818 |
| 31 | Claude Fable 5 | Anthropic | 610 | 41 | 35.67 | $20.00 | $0.4878 |
Methodology
Pricing is read from the versioned provider registry, converted from $/1K to $/1M, and blended as (3 × input + output) ÷ 4. The performance snapshot uses one fixed prompt, 5 runs per model, measured through the All AI Ask gateway from us-east-1. TTFT is reported separately from throughput. Rankings exclude estimated rows.
Limitations
- List prices and one gateway route do not predict your total bill or end-to-end latency.
- Five runs per model are a snapshot, not a confidence interval or a sustained-load test.
- Provider regions, batching, caching, context length, output length, and model availability can change results.
- No historical quarterly baseline is included in this repository, so “key changes” here means notable findings in this version, not a claimed quarter-over-quarter delta.
Download and cite
Download JSON · Download CSV · Read dataset metadata
Citation: All AI Ask. State of LLM API Pricing and Performance — Q3 2026. Data as of 2026-08-14. https://allaiask.com/reports/llm-state-q3-2026. Cite dataset IDs all-ai-ask-pricing-registry-2026-08-08 and all-ai-ask-speed-benchmark-2026-08-08.
Release changelog
- 2026-08-14: Initial Q3 2026 pricing and performance snapshot published.
[P] Versioned pricing registry · [S] Versioned speed benchmark
Batch 49 · q3 decision and evidence contributions. Surface verification: 2026-08-14. These are route-local, server-rendered fixtures; unavailable values are not inferred.
Q3 dataset coverage-and-exclusion funnel
Frozen Batch 49 fixture board. Formula / decision rule: coverage = included / input × 100; excluded rows cannot influence aggregates Boundary: Identity-conflict or unresolved rows are excluded, never silently normalized.
| Frozen fixture / field ID | Joined inputs and observation | Calculated result | State |
|---|---|---|---|
batch49-q3-m1-r1pricing-registry rows · verified price rows | dataset=pricing-q3-v1; input=42; verified=39; excluded=3; reason=missing source/date; join=model+provider+version 39 rows have complete price and identity joins. | coverage=39/42×100=92.86%; aggregate cohort=39 | PASS — exclusions explicit. |
batch49-q3-m1-r2benchmark candidates · completed measured rows · price-speed joinable rows | dataset=speed-q3-v1; candidates=28; measured=20; joinable=18; excluded=8; reasons=timeout/identity conflict Only 18 rows join both measurement and pricing identities. | measurement coverage=20/28=71.43%; price-speed=18/28=64.29% | PASS WITH LIMIT — joinable cohort only. |
batch49-q3-m1-r3identity-conflict rows | provider=ambiguous; model=alias-x; versions=v1/v2; input=4; included=0; reason=conflicting owner/version Conflicting provider/model/version fields cannot enter a published aggregate. | included=0/4; contribution=0 | EXCLUDED — unresolved identity. |
Provenance: Batch 49 q3 module 1 first-party evidence and surface verification date 2026-08-14. All AI Ask Q3 2026 report. Missing joins fail closed.
Frozen token-mix sensitivity table
Frozen Batch 49 fixture board. Formula / decision rule: blended = (input price × input share) + (output price × output share) Boundary: Fixed mixes describe this snapshot; they are not user-specific forecasts or universal recommendations.
| Frozen fixture / field ID | Joined inputs and observation | Calculated result | State |
|---|---|---|---|
batch49-q3-m2-r190:10 input:output mix | cohort=complete price rows; mix=(0.90,0.10); formula=0.90×Pin+0.10×Pout; missing-price=2 The complete cohort is recalculated under a token-heavy input mix; missing-price rows stay out. | median/quartiles=descriptive snapshot; unavailable=2 | AVAILABLE — fixed sensitivity only. |
batch49-q3-m2-r275:25 input:output mix | cohort=complete price rows; mix=(0.75,0.25); formula=0.75×Pin+0.25×Pout; alias-conflict=1 Alias-conflict row is excluded rather than assigned to a neighboring model. | rank-change=reported only if joined; alias row=Unavailable | AVAILABLE WITH EXCLUSION. |
batch49-q3-m2-r350:50 input:output mix | cohort=complete price rows; mix=(0.50,0.50); formula=0.50×Pin+0.50×Pout; forecast=No Equal mix is a descriptive check on the frozen cohort, not a forecast for any workload. | extrema=descriptive only; user-specific cost=Unsupported | PASS WITH BOUNDARY — no recommendation. |
Provenance: Batch 49 q3 module 2 first-party evidence and surface verification date 2026-08-14. All AI Ask Q3 2026 report. Missing joins fail closed.
Finding-level reproducibility receipt board
Frozen Batch 49 fixture board. Formula / decision rule: reproducible = source checksum + filter + formula + row count + published-text match Boundary: Historical claims without a baseline or raw measurement are Unsupported.
| Frozen fixture / field ID | Joined inputs and observation | Calculated result | State |
|---|---|---|---|
batch49-q3-m3-r1cheapest/highest price · median blended price | finding=P-01/P-02; source=pricing-q3-v1; filter=verified; formula=blended; rows=39; checksum=joined Extrema and median recompute from the declared frozen price cohort. | published-text match=Yes; state=Reproducible | REPRODUCIBLE — price receipt. |
batch49-q3-m3-r2fastest/slowest measured throughput · price-speed join | finding=S-01/S-02; source=speed-q3-v1; filter=measured; formula=median t/s; rows=20/18; checksum=joined Speed extrema and joined cost-per-speed use only completed and joinable rows. | published-text match=Yes; joined metric coverage=18/20 | REPRODUCIBLE WITH LIMIT. |
batch49-q3-m3-r3unsupported quarter-over-quarter change claim | finding=Δ-Q2→Q3; baseline=missing; source=Q3 only; formula=delta; rows=0 baseline No historical baseline exists in the declared artifacts. | reproducibility=Unsupported; published change=not claimed | UNSUPPORTED — baseline absent. |
Provenance: Batch 49 q3 module 3 first-party evidence and surface verification date 2026-08-14. All AI Ask Q3 2026 report. Missing joins fail closed.
