Anthropic released Claude Sonnet 5.5 on September 28 with the same API token prices as Sonnet 5. The company says the model generates output more than 30% faster and can cost up to 30% less per task because it uses fewer tokens. Those are separate claims: the posted price has not fallen, and a task's bill still depends on the prompt, output, tools and effort setting.
For developers, the release pairs a potentially useful efficiency gain with changes that can break Sonnet 5 integrations. Anthropic's documentation lists the model as claude-sonnet-5-5 on the Claude API, Google Cloud and Microsoft Foundry, and anthropic.claude-sonnet-5-5 on Amazon Bedrock. Its model page lists a one million token context window and a 128,000 token maximum output. Access is also listed through Claude Platform on AWS. Cloud provider billing and regional settings may differ from the Claude API's posted rates.
What the evaluations show
In Anthropic's Terminal-Bench 4.0 run, Sonnet 5.5 completed 70.6% of 66 terminal tasks, against 10.3% for Sonnet 5 in the release table. The system card describes a demanding set of science and engineering tasks in containerized command-line environments. Anthropic ran Claude Code in bare mode at maximum thinking effort, blocked internet access and cached resources needed for the tasks. Safeguards sent 1.2% of Sonnet 5.5 requests to a fallback model, affecting 1.5% of trials. Anthropic reports a standard error of 2.5 percentage points for the Sonnet 5.5 result. The score measures this setup; it does not establish the same improvement across developers' own repositories.
Cognition ran FrontierCode in Claude Code on 150 repository tasks, according to Anthropic's system card, and reported Sonnet 5.5 at 52.1% on its Main set at xhigh effort. The score fell to 46.2% at max effort, partly because the agent's additional review and edits exceeded the benchmark's scope or time limit in cases Cognition examined. Cursor ran CursorBench 4.0 in its production agent harness on tasks drawn from Cursor sessions. Its Sonnet 5.5 result of 55.5% at max effort appears in a results table sent to Anthropic; Anthropic estimated that model's per-task cost from Cursor's token counts. These external evaluations used different tasks and harnesses, and the Sonnet 5.5 figures here come through Anthropic's system card.
Artificial Analysis also ran the pre-release model on GDPval-AA, which compares work products across 220 tasks from 44 occupations, and AA-Briefcase, a set of linked knowledge-work projects. The card reports Elo scores of 1844 and 1811 at max effort, respectively, close to Opus 5.5 in those tests. Anthropic says the pre-release deployment had a structured-output bug that might have affected the results and has since been fixed. The comparison evaluates benchmark outputs under those conditions; it does not show that Sonnet 5.5 will match Opus on a reader's open-ended work. Anthropic itself says Opus remains stronger at complex work requiring sustained judgment.
Anthropic's launch quotes from early customers describe faster workflows and fewer tokens in their own tests. The quoted customers used their own tasks and methods, so their accounts cannot establish a shared productivity measure. BIG CHANGE has not run Sonnet 5.5 or checked the vendor's speed and task-cost figures in a live workflow.
Price and migration
The Claude API list rate is $2 per million input tokens and $10 per million output tokens, unchanged from Sonnet 5. A five-minute cache write costs $2.50 per million tokens and a cache read costs $0.20. Anthropic's cost-per-task charts combine list rates with the tokens consumed at specified effort levels. They cannot establish a universal 30% saving. A team comparing models should run its own representative tasks at the quality threshold it needs, recording token use and elapsed time as well as success.
Changing the model ID alone may fail for existing Messages API code. Sonnet 5's thinking: {"type": "disabled"} produces an error on 5.5; the documented replacement for avoiding up-front thinking is thinking: {"type": "between_tools"}, accepted at low, medium and high effort. Adaptive thinking is on when the field is omitted, and the Claude API defaults to high effort. Forced tool_choice values any and tool also return errors; Anthropic directs developers to auto and, where supported, strict tool schemas. For Sonnet 5.5 on Amazon Bedrock, strict tool use is unavailable: use auto without strict and validate tool inputs in application code. Code that displays text between tool calls may need to handle progress updates returned in thinking blocks. Computer-use integrations on the Claude API and Google Cloud need the newer computer_toolset_20260801 toolset, while Bedrock keeps the earlier computer_20251124 version. The migration guide details further compatibility changes.
The big change
Sonnet 5.5 gives cost-sensitive teams a new model at Sonnet 5's token rates, with higher results in several specified evaluations and an Anthropic claim of faster output. Its practical value turns on the task mix and effort setting. Existing integrations also need a settings review. The clearest next comparison is an organization's own workload measured for correct outcomes, token use and elapsed time.
Sources & further reading
- Anthropic's Sonnet 5.5 announcement gives the release date, company speed and task-cost claims, benchmark summary and early-customer accounts. The claimed performance needs workload-specific checking.
- Anthropic's Sonnet 5.5 system card documents benchmark tasks, harnesses, effort settings, fallback behavior and external evaluation provenance. It is a primary disclosure by the model vendor, including externally run results supplied to it.
- Claude Platform's model page lists API IDs, access, limits and token prices. The full pricing page confirms Sonnet 5 and 5.5 have the same base rates and explains cloud pricing differences.
- The Sonnet 5.5 migration guide records the request settings that change or return errors. Use it for the complete checklist for an existing integration.


