Anthropic released Claude Haiku 5.5 on 7 October with sharply lower token prices and a new control over how much effort it spends reasoning. The appeal is straightforward: a cheaper model for repeated, narrowly defined jobs. Early independent evidence adds an important wrinkle. A low token price does not guarantee a small bill—or a quick answer.
Based on 8 primary sources · See sources
A small model with a bigger range
Anthropic’s launch announcement positions Haiku for summaries, classification and other frequent tasks, including support and browser use. It is available through the Claude Platform and major cloud platforms. Anthropic calls it its fastest model at standard speed; that qualification matters when comparing it with premium fast modes.
Anthropic introduces Haiku 5.5 and its claimed average cost reduction.
View the original post on X ↗The model documentation lists text and image input, text output and a one-million-token context window—the space available for instructions and supplied material. Adaptive thinking is on by default, with medium effort. The effort setting lets a developer trade reasoning depth against time and token consumption.
For European developers, the supported API regions include Belgium, France, Germany and the UK. That establishes access, rather than a guarantee about a particular cloud region or data-residency arrangement.
The headline price has a 100,000-token boundary
The current price list puts Haiku 5.5’s standard input and output rates at one-tenth of Haiku 4.5’s for prompts up to 100,000 tokens. Above that boundary, the new rates rise fivefold. The larger context window is useful, but filling it changes the economics.
These are rates per million tokens, not prices per completed job. Tokens are the units into which text is split for processing. Anthropic’s migration guide says the same text counts as approximately 30% more tokens with the newer tokenizer; the exact change depends on the content. Old token counts therefore make poor new cost estimates.
Haiku’s standard token tariffs
USD per million tokens. Prompt length selects the Haiku 5.5 tier.
| Charge | 5.5 · up to 100k | 5.5 · over 100k | 4.5 |
|---|---|---|---|
| Input | $0.10 | $0.50 | $1.00 |
| Output | $0.50 | $2.50 | $5.00 |
| Cache reads | $0.01 | $0.05 | $0.10 |
The independent gains come with heavier token use
Artificial Analysis’s launch evaluation gives Haiku 5.5 at maximum effort 43 on its Intelligence Index, a combined measure across ten tests. It reports 38 for GPT-6 Luna at maximum effort and 17 for the previous Haiku. This is evidence of a broad capability gain, rather than a test of brand voice or customer satisfaction.
The same evaluator reports about 162,000 output tokens per task at maximum effort, roughly three times Luna’s. Moving from extra-high to maximum effort adds two index points while using about 1.8 times as many tokens. That is the price of the extra reasoning, even when each token is inexpensive.
There is also a billing caveat: Artificial Analysis says its provisional cost figures did not yet account for Haiku’s higher tariff above 100,000 input tokens. Its published cost-per-task estimates therefore do not yet capture long-prompt bills. The evaluator also found excessive safety refusals in its pre-release automation test; a rerun remains pending.
Broad capability rises beyond the previous Haiku
Artificial Analysis Intelligence Index, launch snapshot of 7 October. Higher is better.
One pelican makes the waiting visible
Developer Simon Willison tested the same bicycle-riding-pelican drawing prompt at five effort levels and published the outputs and usage records. His low-effort run took about seven seconds and cost 0.0936 cents. Maximum effort took five minutes and nine seconds, at 3.3826 cents.
“I got a good bicycle frame for everything beyond low.”
These are single drawing attempts, not a success-rate study, but they make the difference between fast token generation and fast task completion easy to see. A model can produce tokens quickly while spending a long time deciding what to produce. Willison also notes that the model recognised the familiar drawing test and tried a less familiar prompt. The pelicans illustrate a trade-off, rather than a general quality ranking.
Choose the effort for the job
For a repeated job such as sorting incoming feedback, the useful question is what effort level preserves the required accuracy at an acceptable wait and cost. Maximum effort is an option, not a sensible default simply because its benchmark bar is taller.
The migration documentation also describes changes to thinking configuration, output handling and sampling parameters. Existing integrations need more than a model-name swap. And Haiku remains a hosted, proprietary model: Mistral’s preview and planned downloadable weights offer a different route to control, rather than an interchangeable price comparison.
Haiku 5.5 makes narrowly scoped automation more interesting to test. The strongest early lesson is to measure a finished task: what it got right, how long it took and all the tokens it consumed.
Sources
- Haiku 5.5 launch and related updates
Anthropic · Announcement ·
- Model capabilities, pricing tiers and defaults
Anthropic · Documentation
- Current token tariffs and long-prompt pricing
Anthropic · Documentation
- Migration requirements and token counting
Anthropic · Documentation
- Independent launch evaluation and cost-model caveat
Artificial Analysis · Research ·
- Five-effort drawing test with saved outputs
Simon Willison · Research ·
- Original model announcement
Claude / Anthropic · Announcement ·
- Commercial API supported regions
Anthropic · Documentation
Written with AI assistance from the sources above. Analysis and practical implications are our interpretation. Our editorial approach.



