Anthropic released Claude Haiku 5.5 on October 7 for the short, frequent jobs that can dominate an agent system's call volume. Its Claude API list rates start at $0.10 per million input tokens and $0.50 per million output tokens, one tenth of Haiku 4.5's rates. Yet Anthropic says the same text takes about 30% more tokens on the new model, and prompts over 100,000 tokens enter a higher price tier. A team running Haiku 4.5 needs to count its own requests again before estimating the bill or moving production traffic.

The release also changes the Messages API path. Manual thinking budgets fail on Haiku 5.5, adaptive thinking is on by default, and the first response block may be thinking instead of text. Those are concrete compatibility checks for the services that call a small model repeatedly to classify, extract, route or summarize work.

The big change

  • What changed: Anthropic has brought its new million-token context window and adaptive thinking to the lowest-priced Claude tier. Haiku 5.5 lowers the rate for short prompts, while changing how text is counted and how a response may begin.
  • Why it matters: A team can now consider Claude for more frequent, narrowly scoped agent steps at a lower posted rate. The gain depends on the mix of prompt lengths, output and thinking tokens, cache use, and whether the new model meets that step's quality target.
  • What to watch: Haiku 4.5 integrations need a measured migration. Requests near the 100,000-token boundary can cross into the higher tier when recounted, and code that assumes the first content block is text can mishandle an otherwise successful reply.

The price depends on prompt length

On the Claude API, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens for prompts of up to 100,000 tokens. For prompts above that threshold, the respective rates are $0.50 and $2.50. Haiku 4.5 is listed at $1 and $5 across its 200,000-token context window. Haiku 5.5 has a one million-token context window and a 128,000-token maximum output, compared with 200,000 and 64,000 on Haiku 4.5. Anthropic lists availability through the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry. The Amazon Bedrock model ID is anthropic.claude-haiku-5-5; the Claude API uses claude-haiku-5-5. Anthropic's model pages supply the current platform IDs, limits and rates.

Anthropic says Haiku 5.5 costs around 75% less on average to run than Haiku 4.5. Its footnote combines the 90% lower list rate for prompts up to 100,000 tokens, the 50% lower rate above that threshold, its observation that 90% of Haiku 4.5 requests were in the shorter category, and the new tokenizer's higher token count for the same work. That calculation describes Anthropic's mix, not a uniform discount on every request. The approximately 30% token increase is also a vendor estimate for the same input text; the precise change depends on its contents. Anthropic's launch announcement gives both the claim and its method.

For an uncached Claude API request, take the Haiku 5.5 input and output token counts, multiply each by its rate for that request's prompt-length tier, and divide by one million. As arithmetic, 80,000 input tokens and 10,000 output tokens at the short-prompt rates cost $0.013. At 120,000 input and 10,000 output tokens, the long-prompt rates yield $0.085. Those are price calculations, not observed jobs or forecasts. Cache reads and writes have their own rates, and batch processing has a listed 50% discount on input and output. A useful estimate therefore groups real calls by prompt length and cache behavior before summing them. The Haiku 5.5 model page lists these categories.

Recount saved representative prompts with the Message token-count endpoint, specifying claude-haiku-5-5. The old model's token totals cannot establish which 5.5 tier a prompt reaches. For a live sample, record the response usage, output length, cache fields, effort and actual charge alongside task outcomes. Include short requests, the longest repeated conversations and cases close to 100,000 tokens. BIG CHANGE has reviewed the documentation; it has not run Haiku 5.5 or measured a production workload.

The settings that need a migration pass

Anthropic's Haiku 5.5 migration guide gives an unusually specific checklist for clients moving from Haiku 4.5:

  1. Change the model ID for the platform and recount prompts. claude-haiku-5-5 is a fixed Claude API ID with no date suffix or separate alias. Review max_tokens, because thinking consumes that limit and equivalent text may take more tokens.
  2. Replace thinking: {"type":"enabled","budget_tokens":N} with adaptive thinking or leave thinking unset. Set output_config.effort to calibrate depth; the documented default is medium. A low max_tokens can end the reply after a thinking block and before any text.
  3. Read content blocks by type, not by position. Pass thinking blocks back unchanged with tool results. Test any conversation store that replays a block through another account or changes earlier messages, system instructions or tools.
  4. Remove custom temperature, top_p and top_k settings. Replace a final assistant prefill with a user turn and, for constrained formats, the structured-output or tool approach Anthropic documents. Handle a stop_reason of refusal in the client.
  5. If the integration uses computer use on the Claude API or Google Cloud, replace the older computer_20250124 tool with computer_toolset_20260801 and update the tool loop as the guide specifies. Organizations using Priority Tier capacity for Haiku 4.5 also need a separate plan: the guide says Priority Tier is unavailable on Haiku 5.5.

These changes do not prescribe one effort setting for every task. A comparison should hold the application task and acceptance criteria steady, then record answer quality, refusals, truncation, latency and billed tokens at the settings under consideration. Anthropic reports stronger results for Haiku 5.5 than 4.5 across several benchmarks, including 39.2% versus 0.0% on its Terminal-Bench 4.0 table. Those are Anthropic-reported results under its evaluation conditions, not a measurement of a reader's agent. Anthropic also describes Sonnet 5.5 and Opus 5.5 as better suited to complex agentic coding in that benchmark. The near-term decision for a Haiku 4.5 user is whether the smaller model clears the requirements of each narrow step after the API changes and new token count are included.

Sources & further reading