Claude Haiku 5.5 Costs $0.10 Per Million Input Tokens, a Tenth of Haiku 4.5, and Scores 39.2% on Terminal-Bench 4.0
A.I. / news
Claude Haiku 5.5 Costs $0.10 Per Million Input Tokens, a Tenth of Haiku 4.5, and Scores 39.2% on Terminal-Bench 4.0
Anthropic's smallest model is also its first Haiku with an effort setting, and its own table shows Sonnet 5.5 still ahead on every listed test.

Anthropic released Claude Haiku 5.5 on Wednesday, pricing it at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. That is one tenth of what Haiku 4.5 charges for input, according to the launch page.
The model ID is claude-haiku-5-5. It is available now on the Claude Platform, Amazon Web Services, Google Cloud and Microsoft Azure, Anthropic said.
Claude Haiku 5.5 pricing against Haiku 4.5 and Sonnet 5.5
Anthropic splits Haiku 5.5 pricing at a 100,000-token prompt. Above that length the rates rise to $0.50 input and $2.50 output. Haiku 4.5 had a single rate of $1.00 input and $5.00 output, a figure Artificial Analysis also lists, along with a 200,000-token context window.
| Per million tokens | Haiku 5.5 (up to 100k / over 100k) | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|
| Input | $0.10 / $0.50 | $1.00 | $2.00 |
| Output | $0.50 / $2.50 | $5.00 | $10.00 |
| Cache reads | $0.01 / $0.05 | $0.10 | $0.10 |
Anthropic said Haiku 5.5 costs about 75% less to run than Haiku 4.5 on average. The gap is smaller than the 90% list-price cut because the model uses a new tokenizer, which the company said consumes slightly more tokens per task.
The company also halved cache-read pricing on Sonnet 5.5, from $0.20 to $0.10 per million tokens. Anthropic said that cuts Sonnet's cost on most agentic tasks by about 20%.
The input price now matches OpenAI's Decisions API, which bills $0.10 per million input tokens and nothing for output. That service returns a decision rather than text, so the two are not interchangeable. At the open end, Google's EmbeddingGemma 2 ships free weights under Apache 2.0 but only produces embeddings.
What the Haiku 5.5 benchmark table shows
Anthropic compared Haiku 5.5 with Haiku 4.5, OpenAI's GPT-6 Luna and Sonnet 5.5. The figures are vendor-supplied. The page credits Artificial Analysis with running one of them, GDPval-AA v2.1, and does not say who ran the rest. It points to the Haiku 5.5 system card for methods.

- Sonnet 5.51840 Elo
- Haiku 5.51620 Elo
- GPT-6 Luna1437 Elo
- Haiku 4.5735 Elo
Source: Anthropic, Claude Haiku 5.5 launch page, accessed 2026-10-07; benchmark credited to Artificial Analysis
The largest jump is in agentic coding. The page lists Haiku 4.5 at 0.0% on Terminal-Bench 4.0 and Haiku 5.5 at 39.2%, against 70.6% for Sonnet 5.5 and 16.4% for GPT-6 Luna.
On OSWorld 2.1, an offline computer-use subset, Haiku 5.5 scores 72.4% against 15.7% for Haiku 4.5 and 83.9% for Sonnet 5.5. On Humanity's Last Exam without tools it scores 45.9%, against 10.2% for Haiku 4.5.
Sonnet 5.5 beats Haiku 5.5 on every listed benchmark. Anthropic said Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding, and positioned Haiku 5.5 for summaries, compaction, database queries and classification.
The first Haiku with an effort setting
Haiku 5.5 is the first Haiku-class model with an adjustable effort setting, running from Low through Med, High and Xhigh to Max. Anthropic said it is the fastest model it has released at standard speed, and slower than the Opus models in Fast Mode.

Customer figures on the page are Anthropic's selection. Aaron Vinh, a Staff Software Engineer at Asana, reported over 30% lower latency on task completions. Yashodha Bhavnani, VP of AI Products at Box, reported a score 11 points above Haiku 4.5 at about half the latency.
Limits Anthropic set on Haiku 5.5
The page does not state a context window or an output limit for Haiku 5.5. Its cyber safeguards are stricter than Haiku 4.5's but looser than Sonnet 5.5's, and still block penetration testing and similar attacker-oriented techniques. Organisations that need wider access can apply to the Cyber Verification Program.
Anthropic said monthly API credits for subscribers roll out this week: $100 on Max 5x, $200 on Max 20x, and up to $500 pooled on Team. The credits work on any model.
Sources
More in A.I.
- 01LTX-2.5's Free Commercial Licence Stops at $10 Million in Group Revenue, and Its Card Asks for Contact DetailsLightricks' open-weight video and audio model is free for production use below that line, but the revenue test counts parent companies and affiliates, and the card publishes no benchmark scores.
- 02Gemini 4 Argon Costs $1.99 a Task at Promo Price Against $0.72 for GPT-6.1 Sol, and Is Not Yet on SaleGoogle's introductory $2 and $10 rates match OpenAI's Sol per token, but Artificial Analysis figures cited by eesel show Argon writing 62,000 output tokens a task where GPT-6 Astra writes 27,000.
- 03EmbeddingGemma 2 Embeds Text, Images, Audio and Video in 567MB of RAM on a Pixel 11 ProGoogle DeepMind's Apache 2.0 embedding model has 740M parameters in three modular pieces, and the only benchmark number its launch post prints is a 9.92-point gain on MTEB Code.
- 04Kolibri-1 Is Apache 2.0 and Fits on One B200, but Trails Qwen3.8 27B by 9.1 Points in GermanAleph Alpha's 78B-parameter mixture-of-experts model activates 3.46B parameters per token, and its own model card shows a larger dense Qwen ahead on every headline benchmark.