# Claude Haiku 5.5 cuts token prices but changes the migration math
> Anthropic's new Haiku model has lower token rates, a different tokenizer and Messages API migration changes. Existing Haiku 4.5 users should recount requests and test response handling before switching.
By BIG CHANGE Editorial
Published: 2026-10-08T01:08:04.633Z
Updated: 2026-10-08T01:08:04.633Z
Canonical: https://bigchange.ai/blog/claude-haiku-55-prices-tokenizer-migration

AI-generated conceptual illustration by BIG CHANGE; no actual usage, invoice or Anthropic product is depicted.
Anthropic released Claude Haiku 5.5 on October 7 for the short, frequent jobs that can dominate an agent system's call volume. Its Claude API list rates start at $0.10 per million input tokens and $0.50 per million output tokens, one tenth of Haiku 4.5's rates. Yet Anthropic says the same text takes about 30% more tokens on the new model, and prompts over 100,000 tokens enter a higher price tier. A team running Haiku 4.5 needs to count its own requests again before estimating the bill or moving production traffic.
The release also changes the Messages API path. Manual thinking budgets fail on Haiku 5.5, adaptive thinking is on by default, and the first response block may be `thinking` instead of text. Those are concrete compatibility checks for the services that call a small model repeatedly to classify, extract, route or summarize work.
## The big change
- **What changed:** Anthropic has brought its new million-token context window and adaptive thinking to the lowest-priced Claude tier. Haiku 5.5 lowers the rate for short prompts, while changing how text is counted and how a response may begin.
- **Why it matters:** A team can now consider Claude for more frequent, narrowly scoped agent steps at a lower posted rate. The gain depends on the mix of prompt lengths, output and thinking tokens, cache use, and whether the new model meets that step's quality target.
- **What to watch:** Haiku 4.5 integrations need a measured migration. Requests near the 100,000-token boundary can cross into the higher tier when recounted, and code that assumes the first content block is text can mishandle an otherwise successful reply.
## The price depends on prompt length
On the Claude API, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens for prompts of up to 100,000 tokens. For prompts above that threshold, the respective rates are $0.50 and $2.50. Haiku 4.5 is listed at $1 and $5 across its 200,000-token context window. Haiku 5.5 has a one million-token context window and a 128,000-token maximum output, compared with 200,000 and 64,000 on Haiku 4.5. Anthropic lists availability through the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry. The Amazon Bedrock model ID is `anthropic.claude-haiku-5-5`; the Claude API uses `claude-haiku-5-5`. [Anthropic's model pages](https://platform.claude.com/docs/en/models/haiku-5-5/overview) supply the current platform IDs, limits and rates.
Anthropic says Haiku 5.5 costs around 75% less *on average* to run than Haiku 4.5. Its footnote combines the 90% lower list rate for prompts up to 100,000 tokens, the 50% lower rate above that threshold, its observation that 90% of Haiku 4.5 requests were in the shorter category, and the new tokenizer's higher token count for the same work. That calculation describes Anthropic's mix, not a uniform discount on every request. The approximately 30% token increase is also a vendor estimate for the same input text; the precise change depends on its contents. [Anthropic's launch announcement](https://www.anthropic.com/claude-haiku-5-5) gives both the claim and its method.
For an uncached Claude API request, take the Haiku 5.5 input and output token counts, multiply each by its rate for that *request's prompt-length tier*, and divide by one million. As arithmetic, 80,000 input tokens and 10,000 output tokens at the short-prompt rates cost $0.013. At 120,000 input and 10,000 output tokens, the long-prompt rates yield $0.085. Those are price calculations, not observed jobs or forecasts. Cache reads and writes have their own rates, and batch processing has a listed 50% discount on input and output. A useful estimate therefore groups real calls by prompt length and cache behavior before summing them. The [Haiku 5.5 model page](https://platform.claude.com/docs/en/models/haiku-5-5/overview) lists these categories.
Recount saved representative prompts with the [Message token-count endpoint](https://platform.claude.com/docs/en/api/go/messages/count_tokens), specifying `claude-haiku-5-5`. The old model's token totals cannot establish which 5.5 tier a prompt reaches. For a live sample, record the response `usage`, output length, cache fields, effort and actual charge alongside task outcomes. Include short requests, the longest repeated conversations and cases close to 100,000 tokens. BIG CHANGE has reviewed the documentation; it has not run Haiku 5.5 or measured a production workload.
## The settings that need a migration pass
Anthropic's [Haiku 5.5 migration guide](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide) gives an unusually specific checklist for clients moving from Haiku 4.5:
1. Change the model ID for the platform and recount prompts. `claude-haiku-5-5` is a fixed Claude API ID with no date suffix or separate alias. Review `max_tokens`, because thinking consumes that limit and equivalent text may take more tokens.
2. Replace `thinking: {"type":"enabled","budget_tokens":N}` with adaptive thinking or leave `thinking` unset. Set `output_config.effort` to calibrate depth; the documented default is `medium`. A low `max_tokens` can end the reply after a thinking block and before any text.
3. Read content blocks by `type`, not by position. Pass thinking blocks back unchanged with tool results. Test any conversation store that replays a block through another account or changes earlier messages, system instructions or tools.
4. Remove custom `temperature`, `top_p` and `top_k` settings. Replace a final assistant prefill with a user turn and, for constrained formats, the structured-output or tool approach Anthropic documents. Handle a `stop_reason` of `refusal` in the client.
5. If the integration uses computer use on the Claude API or Google Cloud, replace the older `computer_20250124` tool with `computer_toolset_20260801` and update the tool loop as the guide specifies. Organizations using Priority Tier capacity for Haiku 4.5 also need a separate plan: the guide says Priority Tier is unavailable on Haiku 5.5.
These changes do not prescribe one effort setting for every task. A comparison should hold the application task and acceptance criteria steady, then record answer quality, refusals, truncation, latency and billed tokens at the settings under consideration. Anthropic reports stronger results for Haiku 5.5 than 4.5 across several benchmarks, including 39.2% versus 0.0% on its Terminal-Bench 4.0 table. Those are Anthropic-reported results under its evaluation conditions, not a measurement of a reader's agent. Anthropic also describes Sonnet 5.5 and Opus 5.5 as better suited to complex agentic coding in that benchmark. The near-term decision for a Haiku 4.5 user is whether the smaller model clears the requirements of each narrow step after the API changes and new token count are included.
## Sources & further reading
- [Anthropic's Haiku 5.5 announcement](https://www.anthropic.com/claude-haiku-5-5) dates the launch, describes intended workloads and availability, and explains the company's average task-cost calculation. Its benchmark table reports Anthropic's evaluation results.
- [Claude Platform's Haiku 5.5 model page](https://platform.claude.com/docs/en/models/haiku-5-5/overview) lists IDs, context and output limits, prompt-length tiers, cache prices and platform support. The [Haiku 4.5 page](https://platform.claude.com/docs/en/models/haiku-4-5/overview) supplies predecessor rates and limits.
- [The Haiku 5.5 migration guide](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide) specifies the Messages API changes and errors to check before switching an existing integration. The [token-count API reference](https://platform.claude.com/docs/en/api/go/messages/count_tokens) documents counting a request with a chosen model.
- [Reuters' October 7 launch report](https://www.marketscreener.com/news/anthropic-launches-third-claude-5-5-model-expanding-ai-lineup-before-planned-ipo-ce785ddedd89f722) independently corroborates the release date. The technical settings and rates above come from Anthropic's current documentation.
## Sources
- [Anthropic, Claude Haiku 5.5 launch](https://www.anthropic.com/claude-haiku-5-5) — Launch date, intended tasks, vendor benchmark and average-cost claims, calculation footnote, cloud availability.
- [Claude Haiku 5.5 model overview](https://platform.claude.com/docs/en/models/haiku-5-5/overview) — Current model ID, API prices by prompt length, cache and batch rates, context, output, availability.
- [Claude Haiku 4.5 model overview](https://platform.claude.com/docs/en/models/haiku-4-5/overview) — Predecessor's current legacy status, list rates and limits.
- [Claude Haiku 5.5 migration guide](https://platform.claude.com/docs/en/models/haiku-5-5/migration-guide) — Primary Messages API change checklist and tokenizer guidance.
- [Count tokens in a Message](https://platform.claude.com/docs/en/api/go/messages/count_tokens) — Token-count endpoint and model-specific request count.
- [Reuters, Anthropic launches third Claude 5.5 model](https://www.marketscreener.com/news/anthropic-launches-third-claude-5-5-model-expanding-ai-lineup-before-planned-ipo-ce785ddedd89f722) — Independent release corroboration only.
The BIG CHANGE newsletter
The big picture. At your pace.
Recent stories on AI and robotics, the shifts worth watching and practical ideas to use. Choose a daily briefing, weekly digest or monthly perspective.
Sent at 09:00 Belgrade time: daily, Mondays or the first of the month. Your first edition arrives at the next scheduled send after you confirm.
Your privacy, your choice.
Necessary storage supports site security and remembers your choices. Optional Google Analytics stays off until you allow it. You can read every story with necessary storage only. Privacy details