← All providers

Mistral API Pricing, Models & Rate Limits (2026)

Mistral is the largest EU-based lab on this site, serving the open-weight Large 3 flagship, Medium 3.5, Small 4, edge-class Ministral 3, and code-specialised Codestral. EU hosting is the headline operational differentiator for teams with data-residency constraints.

How much does the Mistral API cost?

Mistral API pricing is usage-based and varies by model and token volume. Use the current model table and provider documentation to estimate a workload before committing to production.

Verified 2026-08-14 source

For cross-provider quota units and fixed-workload capacity, see the LLM API rate-limit comparison; this page remains the authoritative owner for Mistral provider facts.

Mistral product vs API

Mistral consumer access and API billing are separate surfaces; check the provider documentation for current account terms.

Three decisions unique to Mistral

Mistral current-model price mechanics

Current modelInputCached inputOutputBatchVerified
Ministral 8B$0.150/MUnavailable — no model cache rate$0.150/M50% off eligible Batch API2026-08-14
Mistral Small 3.1$0.150/MUnavailable — no model cache rate$0.600/M50% off eligible Batch API2026-08-14
Codestral$0.300/MUnavailable — no model cache rate$0.900/M50% off eligible Batch API2026-06-14
Mistral Large 3$0.500/MUnavailable — no model cache rate$1.500/M50% off eligible Batch API2026-08-14
Mistral Medium 3$1.500/MUnavailable — no model cache rate$7.500/M50% off eligible Batch API2026-08-14

Messages API vs OpenAI compatibility map

ChoiceDecision ruleEvidence
BillingChatGPT plan never includes API creditsSeparate metered API account
Input/cached/outputToken prices are model rowsUse calculator for workload totals
Limits/authTier-based, promoted by spend · Bearer API keyVerify before production

Adoption map: what is documented versus unavailable

DimensionRecorded valueDecision consequence
Try Mistral side by side →

Verified 2026-08-14. dated provider pricing/source

Current models
5
Legacy models
0
Price range /M
$0.15–$3.00
Max context
256K
Median tok/s
118
Next retirement

Mistral model pricing

Compare Mistral models by input, output, and blended token cost below.

ModelInput /MOutput /MBlended /M
Ministral 8B$0.15$0.15$0.15
Mistral Small 3.1$0.15$0.60$0.26
Codestral$0.30$0.90$0.45
Mistral Large 3$0.50$1.50$0.75
Mistral Medium 3$1.50$7.50$3.00

Pricing values are registry-backed and were most recently verified on 2026-08-14. Sources: https://docs.mistral.ai/models/model-cards/ministral-3-8b-25-12, https://docs.mistral.ai/models/model-cards/mistral-small-4-0-26-03, https://mistral.ai/pricing, https://docs.mistral.ai/models/model-cards/mistral-large-3-25-12, https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04. Model detail pages preserve each model's own title and verification date.

* Blended comparison assumes 3 input tokens for every output token; it is not the provider's billing unit.

Speed

Fastest measured Mistral model is Ministral 8B at 158 tokens/sec (210ms TTFT), median across measured Mistral models is 118 tokens/sec. See the full speed benchmark methodology.

Best for

Muse Spark 1.3 Contributor is our pick for CodingMuse Spark 1.3 Contributor is our pick for Structured Data ExtractionMuse Spark 1.3 Contributor is our pick for Writing & ContentGLM-5.2 is our pick for Math & ReasoningMuse Spark 1.3 Contributor is our pick for Agents & Tool Use
What it will cost →
Mistral's 5 priced models, ranked by verbosity-adjusted monthly cost, not list rate.

Related Mistral pages

Mistral alternatives →All LLM API pricing →Mistral speed benchmarks →Mistral cost calculator →

Build with Mistral

Mistral rate limits →

Mistral implementation details

Verified 2026-08-14 against source.

Mistral publishes its current authentication, limits, and data-handling details in the linked documentation.

OpenAI-compatibleYes
API base URLhttps://api.mistral.ai/v1
Auth modelBearer API key
Prompt cachingNot documented
Batch discount50%
Free tierFree tier with rate-limited experimentation
Free-tier limitsExperiment access is rate-limited; current limits depend on account and model.
Free-tier expiryNot published
Rate-limit modelTier-based, promoted by spend
Data residencyEU-hosted by default
Trains on API dataNo
SLA publishedYes
DocsOfficial pricingStatus pageFree-tier terms

Switching to and from Mistral

The closest parity-aware alternative to Mistral Medium 3 ($3.00/M) outside Mistral is Gemini 3.7 Flash ($1.50/M, -50%) — a code-change migration.
The closest parity-aware alternative to Mistral Large 3 ($0.75/M) outside Mistral is Gemini 3.7 Flash ($1.50/M, +100%) — a code-change migration.
Full Mistral alternatives comparison →

Calling Mistral through All AI Ask

curl https://allaiask.com/api/v1/chat \
  -H "Authorization: Bearer $ALLAIASK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "ministral-8b", "messages": [{"role": "user", "content": "Hello"}]}'

FAQ

Is Mistral OpenAI-compatible?

Yes — Mistral's API base (https://api.mistral.ai/v1) accepts the OpenAI SDK request/response shape, so existing OpenAI client code works with a base-URL and key swap.

Does Mistral support prompt caching?

Not documented as of 2026-08-14 — we did not find a published prompt-caching feature for Mistral. If that changes, this page updates.

Does Mistral have a free tier?

Yes — Free tier with rate-limited experimentation. Experiment access is rate-limited; current limits depend on account and model.

How much does the Mistral API cost?

Current Mistral models range from $0.15 to $3.00 per million blended tokens (3:1 input:output). Full per-model pricing is in the table below.

Where is Mistral API data hosted?

EU-hosted by default

Mistral role registry, European controls, and artifact provenance

Batch 47 decision and evidence surface · verified 2026-08-14 · exact route allowlist match: /llm-providers/mistral.

Mistral model-role and surface registry

Frozen Batch 47 mistral fixture — Declared Mistral roles with exact surface identity and lifecycle fields; first-party evidence checked 2026-08-14; unresolved joins render Unavailable.

Method: every calculation uses only the joined fields shown in its row; missing joins fail closed. First-party source: Mistral AI documentation. Verified: 2026-08-14.

Fixture / field IDNamed inputsObservationFormula / boundaryDecision
general chat + low-latency text
batch47-mistral-m1-r1
role=general/fast; endpoint=api.mistral.ai; model=Large/Small; managed=Yes; context=joined; tools=supported/UnavailableGeneral model identity joins; low-latency tool support remains model-specific and is not filled from the sibling.role admission = exact endpoint/model ∧ declared context/modality/tool fields
Boundary: A sibling model cannot fill an undocumented capability cell.
UNAVAILABLE — fast-role tool field missing.
edge/local + coding + embeddings
batch47-mistral-m1-r2
role=edge/coding/embed; artifact=weights for edge/coding; endpoint=embed; license=joined; lifecycle=currentArtifact and endpoint roles are separate; coding and embedding evidence does not merge.registry row = one exact role identity; no sibling substitution
Boundary: Open-weight availability does not imply managed API equivalence.
PASS WITH SPLIT — role boundaries preserved.
OCR/document + audio + moderation
batch47-mistral-m1-r3
role=OCR/audio/moderation; endpoint=Unavailable for audio; schema=OCR joined; lifecycle=Unavailable; source date=2026-08-14OCR schema is evidenced; audio endpoint and moderation lifecycle are not documented in this fixture.documented fields only; missing field = Unavailable
Boundary: No modality is inferred from a general model page.
UNAVAILABLE — modality evidence incomplete.

European-control evidence chain

Frozen Batch 47 mistral fixture — Service scope, location, logging, support, training, deletion, and incident claims; first-party evidence checked 2026-08-14; unresolved joins render Unavailable.

Method: every calculation uses only the joined fields shown in its row; missing joins fail closed. First-party source: Mistral AI documentation. Verified: 2026-08-14.

Fixture / field IDNamed inputsObservationFormula / boundaryDecision
managed EU API + customer cloud
batch47-mistral-m2-r1
claim=EU processing; scope=managed API/customer cloud; primary evidence=service terms; region=EU; owner=legalEU endpoint is joined; customer-cloud control is a separate scope and not implied by managed API text.claim pass = exact service scope ∧ primary evidence ∧ verified date ∧ owner
Boundary: Headquarters or endpoint geography alone cannot prove compliance.
PASS WITH SCOPE — managed API only.
on-premises + non-EU consumer + logging/support access
batch47-mistral-m2-r2
scope=on-prem/non-EU; logging=deployer; support access=Unavailable; training use=terms; conflict=noneOn-prem logging is deployer-owned; support-access evidence is missing for the non-EU consumer surface.control status = pass | conflict | unknown per exact scope
Boundary: One compliant scope cannot cover another product surface.
UNAVAILABLE — support access unknown.
training use + deletion + incident response
batch47-mistral-m2-r3
claim=training/deletion/incident; evidence=terms + security page; verified=2026-08-14; owner=legal/security; result=partialTraining and deletion claims are sourced; incident-response timing is not stated for the declared artifact.evidence chain pass only when every declared claim has exact scope and owner
Boundary: Partial policy evidence cannot be promoted to a full workload guarantee.
PASS WITH UNAVAILABLE EDGE — incident timing missing.

Managed-to-artifact provenance ledger

Frozen Batch 47 mistral fixture — Managed model, weights, license, tokenizer, quantizations, runtimes, and drift ownership; first-party evidence checked 2026-08-14; unresolved joins render Unavailable.

Method: every calculation uses only the joined fields shown in its row; missing joins fail closed. First-party source: Mistral AI documentation. Verified: 2026-08-14.

Fixture / field IDNamed inputsObservationFormula / boundaryDecision
official API model + pinned reference weights
batch47-mistral-m3-r1
model=Large; API endpoint=joined; weights=reference; license=revision 3; artifact hash=sha256:31a0…; context=joinedAPI and reference weights are linked by declared model name; operational equivalence still requires replay.provenance complete = endpoint ∧ model ∧ license ∧ hash ∧ tokenizer/template ∧ context
Boundary: Provider promise ends at the managed API boundary.
PASS WITH DRIFT GATE — replay required.
full precision + two named quantizations
batch47-mistral-m3-r2
fp16=sha256:31a0…; INT8=sha256:88bd…; INT4=sha256:71ce…; context=fp16 128K/quantized Unavailable; tools=UnavailableThree hashes are distinct; quantized context and tool availability are not assumed from fp16.equivalence field = documented per artifact; otherwise Unavailable
Boundary: A quantized artifact cannot inherit full-precision controls.
UNAVAILABLE — quantized envelope incomplete.
two runtimes + tokenizer/chat template
batch47-mistral-m3-r3
runtime=vLLM/TGI; tokenizer=joined; template=v1/Unavailable; drift owner=ML platform; checker=JSON passTokenizer joins across runtimes; TGI template version is missing, so drift ownership remains active.drift test required = runtime ∧ tokenizer ∧ template ∧ checker
Boundary: JSON success does not establish conversational template parity.
PASS WITH UNAVAILABLE EDGE — template missing.

Route owner CTA: test this decision path at All AI Ask. This surface does not transfer evidence to another host, model, region, tier, artifact, or scenario.

Try Mistral for free

Run real prompts against every current Mistral model, and every other provider on this site, in one workspace.

Try It Free