NezhaGateNezhaGate
Anthropic ● Coming soon

Claude Haiku 5.5

claude-haiku-5-5

Claude Haiku 5.5 API: Anthropic's fastest, cheapest and most capable small model yet, with a 1M-token context, the first Haiku with effort control, and list prices up to 90% below Haiku 4.5

Claude Haiku 5.5 is the new Haiku from Anthropic, released on October 7, 2026 as the successor to Claude Haiku 4.5. Anthropic calls it the cheapest, fastest and most capable small model it has ever released, and it has the lowest latency in the current Claude lineup. It is built for high-volume work where speed and cost matter: classification, routing, extraction, summaries and context compaction, live customer support, and subagent work inside coding agents led by Claude Opus 5.5 or Sonnet 5.5. It has a 1M-token context window and up to 128K output tokens, runs adaptive thinking by default, and is the first Haiku with an effort setting (five levels from low to max; default medium). For prompts up to 100K tokens, the list price is $0.10 input / $0.50 output per million tokens, 90% below Haiku 4.5. We are integrating it now: leave your email and we will tell you the day it goes live.

Released by the vendor: 2026-10-07 · not yet callable on NezhaGate

Highlights

How Haiku 5.5 improves on Haiku 4.5

Five times the context, twice the output, effort control for the first time and much stronger agentic scores, at a lower price.

Claude Haiku 4.5Claude Haiku 5.5
Released2025-10-152026-10-07
Context window200K tokens1M tokens
Max output per request64K tokens128K tokens
ThinkingManual budget (budget_tokens)Adaptive thinking, on by default
Effort controlNot supportedlow / medium / high / xhigh / max, default medium
Reliable knowledge cutoffFebruary 2025June 2026
Browser use toolNot supportedSupported
List price: input (per 1M tokens)$1$0.10 (prompts up to 100K tokens) / $0.50 (over 100K)
List price: output (per 1M tokens)$5$0.50 / $2.50
List price: cache reads (per 1M tokens)$0.10$0.01 / $0.05
OSWorld 2.1, offline subset (computer use)15.7%72.4%
Terminal-Bench 4.0 (command-line agent tasks)0.0%39.2%
Humanity's Last Exam (no tools / with tools)10.2% / 18.7%45.9% / 57.4%
GDPval-AA v2.1 (real-world work tasks, Elo)7351620
AA-Briefcase v1.1 (agentic knowledge work, Elo)6141578

Specs and prices are from Anthropic's documentation; benchmark scores are from Anthropic's launch announcement of October 7, 2026. Prices in this table are Anthropic's list prices, not ours. Haiku 5.5 uses a newer tokenizer that counts the same text as roughly 30% more tokens; Anthropic says it still costs around 75% less to run than Haiku 4.5 on average.

Use cases

High-volume classification, tagging, intent detection and request routing
Extraction: pull fields out of contracts, web pages and chat logs into structured JSON
Summaries and context compaction for long documents and long conversations
Subagents in coding agents: searching code, reading files, running commands and collecting results for the main model
Live customer support, chatbots and database-query style Q&A
Browser and computer use agents: filling in forms, clicking through pages, gathering information across sites

FAQ

What improved in Claude Haiku 5.5 compared with Haiku 4.5?
Five things stand out. (1) The context window grows from 200K to 1M tokens and the output cap from 64K to 128K tokens. (2) Thinking moves from a manual budget (budget_tokens) to adaptive thinking, and Haiku gets effort control for the first time. (3) Agentic skills jump: OSWorld 2.1 from 15.7% to 72.4%, Terminal-Bench 4.0 from 0.0% to 39.2%, and Humanity's Last Exam without tools from 10.2% to 45.9%. (4) The reliable knowledge cutoff moves from February 2025 to June 2026, and the browser use tool is new. (5) Prices drop: 90% lower per token for prompts up to 100K tokens and 50% lower above that. Anthropic also reports major improvements over Haiku 4.5 on almost all of its alignment evaluations.
What is the official price of Claude Haiku 5.5?
Anthropic prices it in two tiers by prompt length (per million tokens). Prompts up to 100K tokens: $0.10 input, $0.50 output, $0.125 for 5-minute cache writes, $0.20 for 1-hour cache writes and $0.01 for cache reads. Prompts over 100K tokens: every rate is five times higher, so $0.50 input, $2.50 output and $0.05 for cache reads. The Batch API saves another 50%. For comparison, Haiku 4.5 is $1 input / $5 output. These are Anthropic's list prices; our own price will be published at launch.
Does Haiku 5.5 use more tokens?
Two things to know. First, it uses the newer tokenizer introduced with Claude Opus 4.7, so the same text counts as roughly 30% more tokens than on Haiku 4.5. Second, thinking is on by default, and thinking tokens are billed as output and count toward max_tokens; the higher the effort, the more it thinks. The independent benchmarker Artificial Analysis measured about 162K output tokens per Intelligence Index task at max effort. For high-volume simple requests, low or the default medium is usually enough. Net of all this, Anthropic says Haiku 5.5 costs around 75% less to run than Haiku 4.5 on average.
Which effort level should I use, and can thinking be turned off?
Anthropic's guidance: start with the default, medium, for most work, including agentic coding. Use low, the fastest and cheapest level, for chat, short tool tasks and simple high-volume requests; in long agent prompts the model is more likely to skip a search or stop early at low. Use high for knowledge work, longer agent tasks and strict instruction following. Use xhigh or max only where your evals show a quality gain, and compare them with Sonnet 5.5 on quality, cost and speed. Thinking can be switched off with thinking: disabled at high effort or below, but not at xhigh or max. Thinking text is omitted by default; set thinking.display to summarized to get a summary.
What changes when migrating from Haiku 4.5?
The main ones: the model ID becomes claude-haiku-5-5 (no date suffix); a thinking budget (budget_tokens) now returns a 400, so use adaptive thinking with effort; drop temperature, top_p and top_k (top_k, or a non-default temperature or top_p, returns a 400); assistant prefill is no longer accepted, so use structured outputs or state the format in the system prompt; a response can start with thinking blocks, so pick text blocks by type rather than by position; the same text counts as more tokens, so recheck max_tokens and cost estimates; computer use moves to the computer_toolset_20260801 toolset; and handle stop_reason refusal, since safety classifiers can decline a request.
When can I call Claude Haiku 5.5 here?
We are integrating and testing it now. The moment it is reliably callable, this page switches to live and shows the price. Leave your email and we will write the day it goes live.
How will I call it, and does it work with Claude Code?
Both APIs will work: the OpenAI-compatible API (/v1/chat/completions) and the native Anthropic Messages API (/anthropic), with the model set to claude-haiku-5-5. Claude Code sends many small background tasks to its Haiku model; once Haiku 5.5 is live, set Claude Code's Haiku model to claude-haiku-5-5. The Claude Code page in our docs shows how.
Is there a Claude model I can use today?
Yes. Claude Sonnet 5.5, Claude Opus 5.5 and Claude Sonnet 5 are live and callable right now. If you want a cheap, fast small model in the meantime, Gemini 3.8 Flash, DeepSeek V4.1 Flash and GPT-6 Luna are available too.

Available now

These models are live on NezhaGate today:

View all →