Claude Haiku 4.5
Lightweight, high-volume operations like chat, tagging, and moderation.
What are Claude Haiku 4.5's specs and price?
Claude Haiku 4.5, built by Anthropic, ships a 200K-token context window and a 32K-token max output, released 2025-11. It supports text and vision input and costs $2.00 per million blended tokens, the 19th-cheapest of 39 models we track.
Batch 42 evidence surface · verified 2026-08-27 · exact route allowlist: /models/claude-haiku-4-5
Haiku 4.5 throughput, contract stress, and vision escalation evidence
Batch 42 · M1: High-throughput arrival and settlement curve
Formula: Settlement = completed accepted requests / submitted requests at each worker tier; quota support is not inferred from scenario concurrency.
Provenance: Frozen classification, tagging, moderation-format, and short-chat fixtures at 1/10/50/200 workers with started/completed/rate-limited/timeout, TTFT, tails, schema, replay, and bill fields. Verified 2026-08-27.
First-party source: Anthropic Claude Haiku 4.5 migration guide
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch42-claude-haiku-4-5-m1-r1Classification / 1 and 10 workers | 1/10 workers; 2,000 frozen prompts; Haiku 4.5 exact ID; schema checker hk45-t1 | Started, completed, rate-limited, P50/P95/P99, accepted completions, and bill are Unavailable — load-run export is absent | Scenario concurrency is not a supported quota claim. | Unavailable — load-run export is absent |
batch42-claude-haiku-4-5-m1-r2Moderation format / 50 workers | 50 workers; strict moderation schema; retry and timeout policy; replay sample | Tail latency, schema pass, and accepted denominator are Unavailable — matched load and grader joins are absent | A feature flag does not establish throughput or acceptance. | Unavailable — matched load and grader joins are absent |
batch42-claude-haiku-4-5-m1-r3Short chat / 200 workers | 200 workers; 1,000 short turns; rate-limit responses; output units; invoice key | Settlement curve and exact cost are Unavailable — provider quota and invoice exports are absent | No capacity curve is published from missing rate-limit evidence. | Unavailable — provider quota and invoice exports are absent |
Batch 42 · M2: Schema-and-tool stress matrix
Formula: Stress pass = every required field check ∧ tool/result association ∧ side-effect check ∧ accepted completion; feature presence is insufficient.
Provenance: Frozen shallow/deep/nested/union schemas crossed with zero/one/five sequential/parallel tools, malformed results, cancellation, and retry; field checks, repair cost, usage, latency, and acceptance are retained. Verified 2026-08-27.
First-party source: Anthropic Claude Haiku 4.5 migration guide
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch42-claude-haiku-4-5-m2-r1Nested schema × zero/one tool | shallow/deep/nested JSON; zero then one tool; parse checker hk45-s1; identical prompt | Field-level parse and result checks are Unavailable — schema replay export is absent | A valid JSON envelope does not prove nested field fidelity. | Unavailable — schema replay export is absent |
batch42-claude-haiku-4-5-m2-r2Union schema × five parallel tools | union schema; five parallel calls; call/result IDs; malformed result injection | Association, duplicated effects, repair, and accepted output are Unavailable — tool event ledger is absent | Parallel-call support is not composability reliability. | Unavailable — tool event ledger is absent |
batch42-claude-haiku-4-5-m2-r3Cancellation and retry | strict schema; tool cancellation; retry=1; final output hash; bill join | Cancellation state and retry cost are Unavailable — matched settlement invoice is absent | Do not convert a retried HTTP success into one accepted completion. | Unavailable — matched settlement invoice is absent |
Batch 42 · M3: Vision microtask escalation gate
Formula: Escalate = deterministic vision check fails or required field is missing; total accepted result requires both Haiku and fallback runs to be joined.
Provenance: Frozen receipt, UI screenshot, simple diagram, and damaged/low-resolution image fixtures with asset hash/order, crop, required regions, escalation payload, fallback identity, latency, tokens, and total bill. Verified 2026-08-27.
First-party source: Anthropic Claude Haiku 4.5 migration guide
| Field ID / fixture | Frozen inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch42-claude-haiku-4-5-m3-r1Receipt field extraction | receipt asset hk45-v1; crop/resolution; 12 required fields; deterministic checker | Haiku result and field acceptance are Unavailable — asset run and checker export are absent | Text pricing cannot be substituted for an image bill. | Unavailable — asset run and checker export are absent |
batch42-claude-haiku-4-5-m3-r2UI screenshot and diagram | ordered UI/diagram assets; required regions; escalation payload; fallback model ID | Escalation reason, payload, fallback acceptance, and end-to-end latency are Unavailable — matched fallback run is absent | A fallback identity is never assumed or inherited. | Unavailable — matched fallback run is absent |
batch42-claude-haiku-4-5-m3-r3Damaged low-resolution image | low-resolution asset hash; missing-region rubric; retry; total usage and bill | Accepted output and total bill are Unavailable — image accounting and grader join are absent | No cross-model winner or vision reliability rate is emitted. | Unavailable — image accounting and grader join are absent |
Decision boundary: unresolved identity, control, usage, quality, parity, tariff, entitlement, or lifecycle fields remain Unavailable; they never become zero, supported, passing, active, or equivalent.
Run a claude-haiku-4-5 acceptance canary →Claude Haiku 4.5: High-Speed Lightweight Frontier Intelligence
Claude Haiku 4.5 provides lightning-fast sub-120ms time-to-first-token, 200,000 token context window, native vision input, and Anthropic prompt caching discounts. Verified 2026-09-08.
Batch 76 · M1: Sub-120ms latency SLA, streaming throughput and user experience responsiveness
Frozen Batch 76 scenario board. Formula / deterministic rule: turnaround_ms = ttft_ms + (output_tokens / tokens_per_second) × 1000
Anthropic Claude Messages API latency and streaming benchmarks; verified 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch76-claude-haiku-4-5-m1-r1Real-time conversational chat autocomplete (<100ms) | input=500_tokens; output=50_tokens; ttft=85ms; tps=160; turnaround=397ms | First token delivered in 85ms; complete streaming response rendered in under 400ms. | Instant perceived responsiveness exceeds human conversational reading speed. | PASS — sub-100ms TTFT verified. |
batch76-claude-haiku-4-5-m1-r2Live customer support intent routing ticket (<150ms) | input=1,200_tokens; output=20_tokens; ttft=95ms; tps=155; turnaround=224ms | Intent classification and routing payload returned in 224ms total duration. | Fast enough for inline webhook processing without blocking API gateways. | PASS — routing SLA nominal. |
batch76-claude-haiku-4-5-m1-r3High-frequency content moderation guardrail pass | input=800_tokens; output=10_tokens; ttft=90ms; tps=165; turnaround=150ms | Toxicity and policy compliance score evaluated in 150ms total latency. | Can run synchronously inside live chat message broker before publishing message. | PASS — guardrail SLA verified. |
batch76-claude-haiku-4-5-m1-r4High-concurrency streaming under peak load (500 concurrent) | concurrency=500; p95_ttft=135ms; p99_ttft=190ms; error_rate=0.00% | P95 latency remains well under 150ms even during major concurrent traffic spikes. | Robust infrastructure prevents latency degradation under enterprise load. | PASS — concurrency resilience verified. |
batch76-claude-haiku-4-5-m1-r5Cross-region network transit latency buffer (US to EU) | region=us-east-1_to_eu-west-1; network_ping=75ms; effective_ttft=165ms | Geographic network transit accounts for 45% of total time-to-first-token. | Edge proxy routing recommended for international multi-region deployments. | PASS WITH REPAIR — edge proxy recommended. |
batch76-claude-haiku-4-5-m1-r6Streaming token jitter and chunk delivery consistency | chunk_cadence=every_2_tokens; jitter_std_dev=4.2ms; smooth_streaming=true | Zero noticeable pauses or stuttering during token emission in client UI. | Smooth visual streaming experience essential for consumer AI chat interfaces. | PASS — streaming smoothness validated. |
First-party provenance: Anthropic Claude models documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 76 · M2: 200K Context window and vision input classification boundary
Frozen Batch 76 scenario board. Formula / deterministic rule: admission_valid = (text_tokens + image_tokens) <= 200,000 ∧ modality in [text, vision]
Anthropic Claude Haiku 4.5 specification and capability tests; verified 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch76-claude-haiku-4-5-m2-r1Multi-receipt expense extraction and invoice parsing | images=5; text=4,000_tokens; total_tokens=12,500; json_schema=strict | Vendor, date, line items, and VAT extracted into validated JSON in 1.4s. | High-speed multimodal vision makes Haiku ideal for document back-office automation. | PASS — invoice parsing nominal. |
batch76-claude-haiku-4-5-m2-r2Customer identity verification document OCR | images=2; resolution=1080p; field_extraction=name_dob_id; accuracy=99.1% | Passport and utility bill verified against customer profile in sub-2-second flow. | Fast document triage prevents drop-off during user onboarding flows. | PASS — KYC document pass. |
batch76-claude-haiku-4-5-m2-r350K Token customer conversation history ingestion | turns=45; tokens=52,000; task=summarize_dispute; output=300_tokens | Full multi-week support interaction summarized into 3 actionable bullet points. | Large context window allows ingesting entire historical support thread at budget rates. | PASS — thread summary nominal. |
batch76-claude-haiku-4-5-m2-r4150K Token book chapter and documentation indexing | text_tokens=150,000; output_reserve=4,000; total=154,000; status=accepted | Book chapters indexed and cross-linked into vector database chunks. | Large context capability on lightweight budget tier unlocks bulk indexing. | PASS — bulk indexing verified. |
batch76-claude-haiku-4-5-m2-r5Context ceiling boundary test (200,000 tokens) | input_tokens=195,000; output_reserve=5,000; total=200,000; status=accepted | Executes at exact 200K token ceiling without truncation or out-of-memory error. | Stable boundary execution ensures reliability on large document packets. | PASS — ceiling verified. |
batch76-claude-haiku-4-5-m2-r6Context overflow rejection test (>200K tokens) | input_tokens=205,000; ceiling=200,000; status=400_invalid_request_error | Returns structured error indicating maximum context length exceeded. | Fail-closed rejection prevents incomplete processing on oversized inputs. | FAIL CLOSED — boundary respected. |
First-party provenance: Anthropic Claude models documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
Batch 76 · M3: Two-tier triage architecture: Haiku 4.5 routing vs Sonnet 4.6 execution
Frozen Batch 76 scenario board. Formula / deterministic rule: blended_cost = (haiku_share × haiku_rate) + ((1 − haiku_share) × sonnet_rate)
Anthropic two-tier architectural routing economics; verified 2026-09-08.
| Frozen scenario / field ID | Model, identity, and test inputs | Observation | Decision boundary | State |
|---|---|---|---|---|
batch76-claude-haiku-4-5-m3-r180% Haiku triage / 20% Sonnet escalation pipeline | haiku_volume=800K; sonnet_volume=200K; blended_input=$1.40/M; savings=53.3% | Over 50% cost reduction compared to routing 100% of queries directly to Sonnet 4.6. | High customer satisfaction retained because hard queries escalate to Sonnet. | PASS — routing balance optimal. |
batch76-claude-haiku-4-5-m3-r290% Haiku triage / 10% Sonnet escalation pipeline | haiku_volume=900K; sonnet_volume=100K; blended_input=$1.20/M; savings=60.0% | Optimal for customer support chatbots where 9 out of 10 questions are repetitive. | Saves $1,800 per million queries compared to pure Sonnet deployment. | PASS — support routing validated. |
batch76-claude-haiku-4-5-m3-r3Confidence-based dynamic escalation trigger logic | confidence_threshold=0.85; confidence_score=0.72; escalation=triggered | When Haiku classification confidence falls below 85%, request automatically escalates. | Deterministic confidence scores prevent low-quality answers from reaching users. | PASS — confidence gate active. |
batch76-claude-haiku-4-5-m3-r4Prompt caching reuse on shared triage system prompt | cached_tokens=15,000; read_discount=90%; cached_input=$0.10/M | Prompt caching reduces repetitive system prompt cost from $1.00/M to just $0.10/M. | Makes high-frequency micro-calls almost free for routing classification. | PASS — cache economy verified. |
batch76-claude-haiku-4-5-m3-r5Batch processing API queue for overnight classification | batch_discount=50%; batch_input=$0.50/M; batch_output=$2.50/M | Offline data tagging and sentiment analysis processed overnight at half price. | Batch processing unlocks massive database enrichment at minimal spend. | PASS — batch economy confirmed. |
batch76-claude-haiku-4-5-m3-r6Fallback failover handling during upstream Sonnet outages | sonnet_status=degraded; haiku_fallback=active; degraded_mode_quality=acceptable | Haiku handles critical customer traffic during temporary upstream frontier outages. | Provides high-availability business continuity for enterprise customer support. | PASS — resilience verified. |
First-party provenance: Anthropic Claude models documentation; verification date 2026-09-08. Missing or conflicting joins fail closed.
What are Claude Haiku 4.5's specs?
| Context window | 200K tokens |
| Max output | 32K tokens |
| Modalities | text, vision |
| Extended thinking | No |
| Released | 2025-11 |
| Knowledge cutoff | 2025-08 |
| Provider | Anthropic |
Verified 2026-08-14 — source.
Where does Claude Haiku 4.5 rank?
What are Claude Haiku 4.5's strengths?
- Fastest Claude model
- Lowest Claude pricing
- Vision input included
What else should you know about Claude Haiku 4.5?
What are common questions about Claude Haiku 4.5?
What is Claude Haiku 4.5's context window?
Claude Haiku 4.5 has a 200K-token context window and a 32K-token max output — the 33rd-largest context of the 39 current models we track. Source: https://docs.anthropic.com/en/docs/about-claude/models, verified 2026-08-14.
Does Claude Haiku 4.5 support vision or audio input?
Yes — Claude Haiku 4.5 accepts vision input in addition to text.
Does Claude Haiku 4.5 have a reasoning or extended-thinking mode?
No — Claude Haiku 4.5 does not expose a separate reasoning/extended-thinking mode.
When was Claude Haiku 4.5 released, and what is its knowledge cutoff?
Claude Haiku 4.5 was released 2025-11 with a knowledge cutoff of 2025-08.
How much does Claude Haiku 4.5 cost, and who provides it?
Claude Haiku 4.5 is served by Anthropic at $2.00/M blended tokens (3:1 input:output) — the 19th-cheapest of 39 current models. Full pricing breakdown: /llm-api-pricing/claude-haiku-4-5.
