LLM API Providers Compared — 12 Providers, 39 Models, Updated August 2026
Raw dataset: data.json. Cite this: All AI Ask LLM API Provider Capability Dataset, retrieved 2026-08-14.
Amazon is the cheapest provider by entry-level current-model price. Cerebras is the fastest by median measured throughput, at 2215 tokens/sec. Google has the widest context window, at 2M tokens.
How to read this: "lab" providers train the weights they serve; "host" providers (Groq, Cerebras) serve someone else's open-weight models on their own inference hardware — their differentiator is speed, not model quality.
| Provider | Kind | Models | Price range /M | Max context | Median tok/s | OpenAI-compatible | Prompt caching | Batch discount | Free tier / limits | Data residency | Next retirement |
|---|---|---|---|---|---|---|---|---|---|---|---|
| OpenAI | lab | 3 (+12 legacy) | $2.25–$8.00 | 1M | 78 | Yes | Yes | 50% | No free tier No free API tier published; API usage is billed under the account's usage tier. Official terms | US by default; EU data residency available on enterprise agreements | — |
| Anthropic | lab | 6 (+7 legacy) | $2.00–$30.00 | 1M | 67 | Partial | Yes | 50% | No free tier No free API tier published; API usage requires an enabled billing account. Official terms | Not documented | — |
| lab | 4 (+5 legacy) | $0.85–$4.50 | 2M | 114 | Partial | Yes | 50% | Free tier with daily request cap on Google AI Studio Free-tier requests and tokens vary by model and project; Google publishes the current quota table. Official terms | Global by default; Vertex AI offers selectable regional endpoints | — | |
| xAI | lab | 5 (+2 legacy) | $1.56–$3.00 | 1M | 98 | Yes | Not documented | Not documented | Free starting credits for new accounts Promotional credits are limited to eligible new accounts; amount and expiry vary by account. Official terms | Not documented | — |
| DeepSeek | lab | 2 | $0.66–$1.98 | 1M | 100 | Yes | Not documented | Not documented | No free tier No free API tier published; API usage is billed at the listed token rates. Official terms | Not documented | — |
| Mistral | lab | 5 | $0.15–$3.00 | 256K | 118 | Yes | Not documented | 50% | Free tier with rate-limited experimentation Experiment access is rate-limited; current limits depend on account and model. Official terms | EU-hosted by default | — |
| Groq | host | 3 (+2 legacy) | $0.13–$1.20 | 131K | 780 | Yes | Not documented | 50% | Free tier with per-minute and per-day token caps Per-minute and per-day request/token caps vary by model and account tier. Official terms | US | — |
| Cerebras | host | 2 | $0.45–$2.38 | 200K | 2215 | Yes | Not documented | Not documented | Free tier with a daily token cap Free access is subject to a daily token cap and account quota. Official terms | US | — |
| Qwen | lab | 3 | $1.10–$2.80 | 256K | 49 | Partial | Not documented | Not documented | Free quota for new Alibaba Cloud accounts Quota is limited to eligible new accounts and varies by model, region, and account. Official terms | Singapore/international region via DashScope Intl; mainland China served from a separate region | — |
| Amazon | lab | 3 | $0.06–$1.40 | 300K | 108 | No | Not documented | 50% | No free tier No provider-wide Nova API free tier is published; AWS account promotions, if any, are separate. Official terms | Selectable AWS region | — |
| Z.ai | lab | 1 (+1 legacy) | $2.15–$2.15 | 1M | — | Yes | Not documented | Not documented | Free trial credits for new accounts Trial credits are limited to eligible new accounts; amount and expiry vary by account. Official terms | Not documented | — |
| Meta | lab | 2 | $0.13–$2.00 | 1.0M | — | Yes | Not documented | Not documented | No free tier No published free tier for the Model API; the Contributor tier is a steep discount, not a free tier. Official terms | Not documented | — |
Every cell is a sourced fact or an explicit "Not documented" — never an inferred value. Full sourcing and verification dates on each provider's hub.
Provider hubs
One key, every provider
You don't need each provider's SDK, key, billing relationship, or rate-limit tier — the same call reaches all 12 of them.
Try It FreeFAQ
Which LLM provider is cheapest?
By entry-level current-model price, Amazon is the cheapest provider on this site, with Amazon Nova Micro at $0.06/M blended tokens.
Which providers are OpenAI-compatible?
OpenAI, xAI, DeepSeek, Mistral, Groq, Cerebras, Z.ai, Meta expose a fully OpenAI-compatible endpoint. Anthropic, Google, Qwen offer a partial/beta compatibility layer. Amazon do not.
Do I need a separate API key for each provider?
Yes, if you call each provider directly — each has its own key, billing relationship, and rate-limit tier. Routing every model through All AI Ask removes that: one key reaches all of them.
Which providers offer a free tier?
Google, xAI, Mistral, Groq, Cerebras, Qwen, Z.ai publish a free tier. The rest explicitly publish no free tier as of their verification date.
Which provider has the largest context window?
Google has the widest context window among current models, at 2M tokens.
Provider custody taxonomy, hard-gate shortlist, and concentration planner
Batch 47 decision and evidence surface · verified 2026-08-14 · exact route allowlist match: /llm-providers.
Provider-class and custody normalizer
Frozen Batch 47 provider-directory fixture — Sitemap providers classified by model owner, API seller, serving host, and support party; first-party evidence checked 2026-08-14; unresolved joins render Unavailable.
Method: every calculation uses only the joined fields shown in its row; missing joins fail closed. First-party source: OpenAI API documentation. Verified: 2026-08-14.
| Fixture / field ID | Named inputs | Observation | Formula / boundary | Decision |
|---|---|---|---|---|
lab-trained / direct APIbatch47-provider-directory-m1-r1 | model owner=OpenAI; API seller=OpenAI; serving host=OpenAI; endpoint/region=joined; artifact=Unavailable; billing/support=OpenAI | Owner, seller, and host align; artifact visibility remains Unavailable for the managed surface. | class = lab-trained/direct when owner = seller = serving host Boundary: Managed artifact opacity is not a model defect. | PASS — exact layer retained. |
lab-trained / cloud-servedbatch47-provider-directory-m1-r2 | model owner=Meta; API seller=cloud platform; serving host=cloud platform; region=eu-west-1; revision=joined; billing=seller | Model owner differs from API seller and host; support follows the seller. | class = lab-trained/cloud-served when model owner ≠ API seller = host Boundary: Provider counts must not collapse lab and host identities. | PASS — split custody displayed. |
third-party host + private artifactbatch47-provider-directory-m1-r3 | model owner=Mistral/unknown; API seller=inference host/private team; host=host/private; hash=joined/private; precision=Unavailable; support=split | Third-party host and private artifact have different support and evidence boundaries. | class = exact declared layer; missing precision remains Unavailable Boundary: A private artifact cannot inherit the host provider promise. | UNAVAILABLE — precision/support edge unresolved. |
Hard-requirement shortlist compiler
Frozen Batch 47 provider-directory fixture — Frozen workload gates with pass/fail/unknown outcomes before ranking; first-party evidence checked 2026-08-14; unresolved joins render Unavailable.
Method: every calculation uses only the joined fields shown in its row; missing joins fail closed. First-party source: OpenAI API documentation. Verified: 2026-08-14.
| Fixture / field ID | Named inputs | Observation | Formula / boundary | Decision |
|---|---|---|---|---|
text extraction + multimodal RAGbatch47-provider-directory-m2-r1 | gates=schema/context/citations/media; provider evidence=3/4; eligible set=Unavailable; disqualifier=media custody | Text extraction passes; multimodal RAG has unknown media custody, so the set is not eligible. | eligible = every hard gate passes; unknown ≠ pass Boundary: Price and speed are not ranked before every gate passes. | UNAVAILABLE — media custody gate. |
five-tool agent + EU-controlled workloadbatch47-provider-directory-m2-r2 | gates=5 tools/EU region/logging; pass=2; fail=1; unknown=1; eligible=none; disqualifier=tool effect key | Tool effect-key failure disqualifies the candidate; EU logging remains unknown. | shortlist = providers where fail=0 ∧ unknown=0 across hard gates Boundary: A model-quality pass cannot waive an effect-control failure. | FAIL — no eligible provider. |
500K synthesis + high-burst open-weight + AWS-native appbatch47-provider-directory-m2-r3 | gates=context/burst/artifact/IAM; context=unknown; burst=pass; artifact=pass; IAM=pass; eligible=Unavailable | Three gates pass, but 500K context evidence is unknown; no provider is promoted. | gate coverage = pass / declared = 3/4; eligibility requires 4/4 Boundary: Partial gate coverage cannot produce a winner. | UNAVAILABLE — context gate unresolved. |
Portfolio concentration and fallback planner
Frozen Batch 47 provider-directory fixture — Workload shares, owner/host concentration, correlated controls, and fallback qualification; first-party evidence checked 2026-08-14; unresolved joins render Unavailable.
Method: every calculation uses only the joined fields shown in its row; missing joins fail closed. First-party source: OpenAI API documentation. Verified: 2026-08-14.
| Fixture / field ID | Named inputs | Observation | Formula / boundary | Decision |
|---|---|---|---|---|
single-provider architecture · six workload rolesbatch47-provider-directory-m3-r1 | shares=text 40%; RAG 20%; code 15%; tools 10%; media 10%; batch 5%; owner concentration=100%; host=100% | All roles share one model owner and serving provider; fallback is not qualified. | concentration = Σ(share²); owner=1.00; serving provider=1.00 Boundary: A single-provider plan is a concentration measure, not a quality verdict. | FAIL — fallback absent. |
lab-plus-host + two-lab architecturebatch47-provider-directory-m3-r2 | owner shares=60/40; serving shares=50/30/20; region=US/EU; common control plane=partial; fallback=2/6 roles | Owner and serving-provider concentration differ; only two roles have replay-qualified fallbacks. | uncovered roles = declared roles − qualified fallback roles = 6 − 2 = 4 Boundary: Provider diversity does not guarantee workload fallback. | UNAVAILABLE — four roles uncovered. |
region-diverse + private-fallback architecturebatch47-provider-directory-m3-r3 | roles=6; regions=US/EU/private; owner shares=45/35/20; serving shares=40/30/20/10; correlated control=none joined; fallback=6/6 | All six roles have a declared fallback; correlated-control evidence is still only partially joined. | fallback coverage = 6/6 = 100%; concentration reported separately by owner and host Boundary: Coverage does not prove independence when common controls are unknown. | PASS WITH UNAVAILABLE EDGE — independence review remains. |
Route owner CTA: test this decision path at All AI Ask. This surface does not transfer evidence to another host, model, region, tier, artifact, or scenario.
