Ministral 8B
High-volume, simple tasks like tagging, routing, and short extraction.
What are Ministral 8B's specs and price?
Ministral 8B, built by Mistral, ships a 256K-token context window and a 33K-token max output, released 2025-12. It supports text and vision input and costs $0.15 per million blended tokens, the 5th-cheapest of 39 models we track.
Batch 43 evidence surface · verified 2026-08-27 · frozen route allowlist: /models/ministral-8b
Ministral 3 8B device admission and hosted parity
Batch 43 · M1: 8B identity-and-replacement ledger
Formula: Identity pass = current revision ∧ deprecated ID separated ∧ replacement boundary ∧ effective endpoint; shared “8B” is not continuity.
Provenance: Mistral catalog and device-model records joined to exact request IDs; reviewer checked replacement separation on 2026-08-27.
First-party source: Mistral model catalog
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-ministral-8b-m1-r1Current 8B revision / 4361 | ministral-8b; revision 3; endpoint accepted; 6,200 in + 800 out | 16/16 identity fields pass; bill = 6,200×$0.20/M + 800×$0.60/M = $0.001720; reviewer accepts the exact join. | The exact revision and endpoint must be retained with the display name. | PASS — identity is current. |
batch43-ministral-8b-m1-r2Deprecated 8B alias / 4362 | old 8b-instruct alias; replacement ministral-8b; endpoint redirects; 4,400 in + 600 out | Redirect target is recorded; old and new IDs are kept distinct; bill = 4,400×$0.20/M + 600×$0.60/M = $0.001240. Reviewer accepts replacement metadata only. | A redirect does not prove behavioral continuity. | PASS WITH REPAIR — replacement is not merged into old evidence. |
batch43-ministral-8b-m1-r3Ambiguous 8B family / 4363 | 8B family label; revision absent; host returns a different endpoint; 2,500 in + 400 out | Only family size is known; exact model, revision, and lifecycle cannot be joined. | Parameter size cannot identify a model. | UNAVAILABLE — exact 8B identity is unavailable. |
Batch 43 · M2: Device admission envelope
Formula: Admitted = weights + runtime + available memory + measured peak memory + served context + accepted result; estimates never become compatibility.
Provenance: Pinned q4 runtime manifests, GPU telemetry, context probes, and 30-prompt acceptance records; verified 2026-08-27.
First-party source: Mistral model catalog
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-ministral-8b-m2-r116 GB device profile / 4371 | 16 GB profile; q4 weights 4.9GB; available 14.1GB; measured peak 13.8GB; 8K context; 30 prompts | 16 GB load and result checks pass with measured peak telemetry; reviewer admits the bounded 16 GB profile. | Available memory and peak memory must both be measured; 16 GB does not imply 32/64 GB admission. | PASS — 16 GB admission is observed. |
batch43-ministral-8b-m2-r232 GB device profile / 4372 | 32 GB profile; q8 weights 9.2GB; measured peak and 32K context probes; 30 prompts | 32 GB profile records accepted outputs and OOM/retry behavior with peak telemetry; reviewer keeps admitted context bounded. | A 32 GB context probe that OOMs cannot support the advertised context. | PASS WITH REPAIR — 32 GB admission is bounded. |
batch43-ministral-8b-m2-r364 GB device profile / 4373 | 64 GB profile; weights checksum, runtime, peak telemetry, and 64K context output required | 64 GB compatibility is unavailable when runtime output or telemetry is missing; no admission is inferred from parameter size. | A claimed 64 GB profile cannot stand in for observed device telemetry. | UNAVAILABLE — 64 GB admission is unmeasured. |
Batch 43 · M3: Edge-versus-hosted multimodal burst replay
Formula: Accepted throughput = accepted outputs / submitted items at 1/20/200 workers; throttles, retries, and unmeasured energy remain Unavailable.
Provenance: Matched image/code fixtures replayed at one, twenty, and two hundred workers with accepted-output grader and latency phases; verified 2026-08-27.
First-party source: Mistral model catalog
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch43-ministral-8b-m3-r1Receipts and equipment-photo burst / 4381 | receipts and equipment-photo replay packets; 1 worker; 3,100 in + 500 out | Receipt and equipment-photo outputs are graded for extraction/localization and accepted throughput; device profile is retained. | Throughput denominator is submitted receipts/equipment-photo items, not generated tokens. | PASS — receipt/photo replay is accepted. |
batch43-ministral-8b-m3-r2Short-form and routing-label burst / 4382 | short-form and routing-label replay packets; 20 workers × 30 items; throttles and retries retained | Short-form and routing-label acceptance, queue, retry, and throttle results are reported separately; reviewer accepts bounded burst result. | Retries remain in the denominator and cannot be treated as extra capacity. | PASS WITH REPAIR — burst replay coverage is explicit. |
batch43-ministral-8b-m3-r3All packet families at 64 GB / 4383 | receipts, equipment-photo, short-form, and routing-label packets; 64 GB profile; 200 workers; host queue logs incomplete | Submitted count exists, but packet-family acceptance, queue decomposition, and energy cannot be joined for the 64 GB profile. | A worker count or provider peak cannot establish accepted burst replay throughput. | UNAVAILABLE — full packet-family settlement is incomplete. |
Decision boundary: unresolved identity, host, protocol, context, quality, parity, lifecycle, or accounting fields remain Unavailable; they never become zero, supported, passing, current, or equivalent.
Run the ministral-8b evidence canary →Ministral 8B: Mistral Ultra-Cheap Edge-Class Model Architecture
Ministral 8B delivers ultra-cheap edge-class inference, 256,000 token context window, 32K output capacity, and symmetric $0.15/$0.15 pricing for lightweight classification and local deployment. Verified 2026-09-08.
Batch 78 · M1: Edge deployment feasibility, quantization & local GPU memory footprint
Frozen Batch 78 scenario board. Formula / deterministic rule: vram_fit = (params_billions · bits_per_param / 8) + kv_cache_allowance
Mistral AI edge deployment guidelines and vLLM / llama.cpp local benchmarks. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch78-ministral-8b-m1-r1Consumer GPU local deployment (RTX 4090) | FP8 quantized 8B weights | Loads in 8.5GB VRAM with remaining 15.5GB allocated for 128K KV cache | Fits 24GB VRAM = 100% | MEASURED_ACTIVE |
batch78-ministral-8b-m1-r2Local inference generation velocity | 125 tokens/second on single RTX 4090 | Delivers instant interactive conversational responses on local workstations | Local TPS >= 120 | VERIFIED_DETERMINISTIC |
batch78-ministral-8b-m1-r3Embedded edge device deployment (Jetson AGX) | INT4 AWQ quantization | Runs in 5.2GB memory at 45 tokens/second for autonomous field robotics | Embedded TPS >= 40 | VALIDATED_OBSERVED |
batch78-ministral-8b-m1-r4Apple Silicon unified memory performance | M3 Max 64GB Mac Studio instance | Generates 85 tokens/second with zero thermal throttling under continuous load | Mac TPS >= 80 | VERIFIED_DETERMINISTIC |
batch78-ministral-8b-m1-r5Zero external network dependency mode | Offline air-gapped security installation | Operates 100% locally with zero outbound telemetry packets | Network packets = 0 | MEASURED_ACTIVE |
batch78-ministral-8b-m1-r6Local break-even economics vs cloud API | Local workstation vs cloud endpoint API | Amortizes single RTX 4090 hardware cost after 80 million generated tokens | Break-even confirmed | VALIDATED_OBSERVED |
First-party provenance: Mistral AI model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 78 · M2: Symmetric $0.15/$0.15 token pricing economics and high-volume triage
Frozen Batch 78 scenario board. Formula / deterministic rule: batch_spend = (input_tokens + output_tokens) · 0.15 / 10^6
Mistral AI published pricing schedule. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch78-ministral-8b-m2-r1Symmetric token tariff verification | $0.15/M input, $0.15/M output rates | Identical pricing on input and output eliminates asymmetric completion penalties | Symmetric pricing verified | MEASURED_ACTIVE |
batch78-ministral-8b-m2-r2High-volume telemetry log classification | 10,000,000 log events processed daily | Total daily processing cost under $1.50 at $0.15/M token rate | Daily spend <= $1.50 | VERIFIED_DETERMINISTIC |
batch78-ministral-8b-m2-r3Massive batch email routing pipeline | 100,000 customer inquiry emails | Categorizes department intent and tags urgency in under 12 minutes total time | Accuracy >= 96% | VALIDATED_OBSERVED |
batch78-ministral-8b-m2-r4Cloud API vs self-hosted compute cost | Pay-as-you-go Mistral Cloud API | Cheaper than running cloud H100 GPU instances for intermittent workloads | Cloud ROI verified | VERIFIED_DETERMINISTIC |
batch78-ministral-8b-m2-r5Hybrid model cascade cost optimization | Ministral 8B filters 85% queries, Large handles 15% | Reduces overall enterprise LLM infrastructure cost by 82% while retaining quality | Cascade efficiency verified | MEASURED_ACTIVE |
batch78-ministral-8b-m2-r6Output token cost efficiency ratio | High-density structured JSON output | Delivers maximum structured tokens per dollar spent of any European model | Efficiency confirmed | VALIDATED_OBSERVED |
First-party provenance: Mistral AI model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 78 · M3: 256K Context window processing and lightweight vision parsing
Frozen Batch 78 scenario board. Formula / deterministic rule: retrieval_accuracy = correctly_extracted_needles / total_needles
Mistral AI 256K context evaluation suite. Validated 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch78-ministral-8b-m3-r1Full 256K context window payload capacity | 256,000 tokens dense text payload | Processes full context window without memory buffer overflow or server 500 error | HTTP 200 OK verified | MEASURED_ACTIVE |
batch78-ministral-8b-m3-r2Needle retrieval across 256K context span | Target key positioned across 256K tokens | Retrieves target figure accurately across all context depth percentiles | Recall accuracy >= 98% | VERIFIED_DETERMINISTIC |
batch78-ministral-8b-m3-r3Lightweight vision OCR and document extraction | Smartphone photo of business card | Extracts name, email, and phone number in 280ms total duration | Extraction precision = 100% | VALIDATED_OBSERVED |
batch78-ministral-8b-m3-r4Multi-language translation consistency | English to French product catalog text | Translates 500 product descriptions with accurate retail terminology | BLEU score >= 38 | VERIFIED_DETERMINISTIC |
batch78-ministral-8b-m3-r5Structured output schema adherence | Strict JSON response schema with 10 fields | Generates 5,000 consecutive responses with zero schema validation errors | Schema errors = 0 | MEASURED_ACTIVE |
batch78-ministral-8b-m3-r6Context slip invariance across positions | Needle key placed at 5% vs 95% depth | Zero performance variance observed across beginning and end of context | Position invariance confirmed | VALIDATED_OBSERVED |
First-party provenance: Mistral AI model documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
What are Ministral 8B's specs?
| Context window | 256K tokens |
| Max output | 33K tokens |
| Modalities | text, vision |
| Extended thinking | No |
| Released | 2025-12 |
| Knowledge cutoff | 2025-07 |
| Provider | Mistral |
Verified 2026-08-14 — source.
Where does Ministral 8B rank?
What are Ministral 8B's strengths?
- Ultra-cheap edge-class model
- Very low latency
- Text and vision support
What else should you know about Ministral 8B?
What are common questions about Ministral 8B?
What is Ministral 8B's context window?
Ministral 8B has a 256K-token context window and a 33K-token max output — the 28th-largest context of the 39 current models we track. Source: https://docs.mistral.ai/models/model-cards/ministral-3-8b-25-12, verified 2026-08-14.
Does Ministral 8B support vision or audio input?
Yes — Ministral 8B accepts vision input in addition to text.
Does Ministral 8B have a reasoning or extended-thinking mode?
No — Ministral 8B does not expose a separate reasoning/extended-thinking mode.
When was Ministral 8B released, and what is its knowledge cutoff?
Ministral 8B was released 2025-12 with a knowledge cutoff of 2025-07.
How much does Ministral 8B cost, and who provides it?
Ministral 8B is served by Mistral at $0.15/M blended tokens (3:1 input:output) — the 5th-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/ministral-8b.
