Claude Haiku 5.5
Claude Haiku 5.5 API: Anthropic's fastest, cheapest and most capable small model yet, with a 1M-token context, the first Haiku with effort control, and list prices up to 90% below Haiku 4.5
Claude Haiku 5.5 is the new Haiku from Anthropic, released on October 7, 2026 as the successor to Claude Haiku 4.5. Anthropic calls it the cheapest, fastest and most capable small model it has ever released, and it has the lowest latency in the current Claude lineup. It is built for high-volume work where speed and cost matter: classification, routing, extraction, summaries and context compaction, live customer support, and subagent work inside coding agents led by Claude Opus 5.5 or Sonnet 5.5. It has a 1M-token context window and up to 128K output tokens, runs adaptive thinking by default, and is the first Haiku with an effort setting (five levels from low to max; default medium). For prompts up to 100K tokens, the list price is $0.10 input / $0.50 output per million tokens, 90% below Haiku 4.5. We are integrating it now: leave your email and we will tell you the day it goes live.
Released by the vendor: 2026-10-07 · not yet callable on NezhaGate
Highlights
- Anthropic's cheapest, fastest and most capable small model yet, with the lowest latency in the current Claude lineup
- Much lower list prices: for prompts up to 100K tokens, $0.10 input, $0.50 output and $0.01 cache reads per million tokens, 90% below Haiku 4.5; over 100K tokens, $0.50 / $2.50, still 50% lower; the Batch API saves another 50%
- 1M-token context (Haiku 4.5: 200K) and up to 128K output tokens per request (was 64K); image input; reliable knowledge cutoff of June 2026
- The first Haiku with effort control: low, medium, high, xhigh and max, default medium; adaptive thinking is on by default, and simple requests can be answered without thinking
- A big jump in agentic skills: OSWorld 2.1 computer use from 15.7% to 72.4% and Terminal-Bench 4.0 command-line tasks from 0.0% to 39.2%, both ahead of GPT-6 Luna (48.9% and 16.4%)
- Full tool support: tool use (including forcing a specific tool), structured outputs, prompt caching, the Batch API, the computer use toolset and the new browser use tool
- A pairing Anthropic recommends: inside coding agents led by Opus 5.5 or Sonnet 5.5, Haiku 5.5 runs as a subagent on the many small jobs such as searching, reading and summarizing, cutting overall cost and latency
- Early customer results on Anthropic's launch page: Box measured scores 11 points higher than Haiku 4.5 at about half the latency; Asana cut task-completion latency by more than 30%
How Haiku 5.5 improves on Haiku 4.5
Five times the context, twice the output, effort control for the first time and much stronger agentic scores, at a lower price.
| Claude Haiku 4.5 | Claude Haiku 5.5 | |
|---|---|---|
| Released | 2025-10-15 | 2026-10-07 |
| Context window | 200K tokens | 1M tokens |
| Max output per request | 64K tokens | 128K tokens |
| Thinking | Manual budget (budget_tokens) | Adaptive thinking, on by default |
| Effort control | Not supported | low / medium / high / xhigh / max, default medium |
| Reliable knowledge cutoff | February 2025 | June 2026 |
| Browser use tool | Not supported | Supported |
| List price: input (per 1M tokens) | $1 | $0.10 (prompts up to 100K tokens) / $0.50 (over 100K) |
| List price: output (per 1M tokens) | $5 | $0.50 / $2.50 |
| List price: cache reads (per 1M tokens) | $0.10 | $0.01 / $0.05 |
| OSWorld 2.1, offline subset (computer use) | 15.7% | 72.4% |
| Terminal-Bench 4.0 (command-line agent tasks) | 0.0% | 39.2% |
| Humanity's Last Exam (no tools / with tools) | 10.2% / 18.7% | 45.9% / 57.4% |
| GDPval-AA v2.1 (real-world work tasks, Elo) | 735 | 1620 |
| AA-Briefcase v1.1 (agentic knowledge work, Elo) | 614 | 1578 |
Specs and prices are from Anthropic's documentation; benchmark scores are from Anthropic's launch announcement of October 7, 2026. Prices in this table are Anthropic's list prices, not ours. Haiku 5.5 uses a newer tokenizer that counts the same text as roughly 30% more tokens; Anthropic says it still costs around 75% less to run than Haiku 4.5 on average.
Use cases
FAQ
What improved in Claude Haiku 5.5 compared with Haiku 4.5?
What is the official price of Claude Haiku 5.5?
Does Haiku 5.5 use more tokens?
Which effort level should I use, and can thinking be turned off?
What changes when migrating from Haiku 4.5?
When can I call Claude Haiku 5.5 here?
How will I call it, and does it work with Claude Code?
Is there a Claude model I can use today?
Available now
These models are live on NezhaGate today:
Claude Sonnet 5.5
Anthropic's newest Sonnet · the best mix of speed and intelligence
Claude Opus 5.5
Anthropic's next Opus · 1M context · adaptive thinking
Claude Sonnet 5
Anthropic's new-generation Sonnet · near-Opus coding
Gemini 3.8 Flash
Latest Flash · adaptive thinking
DeepSeek V4.1 Flash
DeepSeek's current Flash model - thinking on or off
GPT-6 Luna
The most efficient GPT-6, built for high-volume work