The easiest Sonnet 5.5 migration bug to notice is a rejected request. The harder one is an agent that keeps answering after losing the reasoning that made its previous work coherent.
Anthropic released Claude Sonnet 5.5 on September 28. Its launch claims are attractive: output generation more than 30% faster than Sonnet 5, with task costs up to 30% lower in Anthropic’s testing. But the token prices have not fallen, and the million-token context window is inherited. The opportunity is better work per run—not a blanket discount on every request.
RohitAI’s read: upgrade the session handling before chasing the savings. Sonnet 5.5 changes how applications request tools, preserve thinking, switch accounts and execute desktop actions. Some incompatibilities fail loudly. Others leave the HTTP dashboard green while continuity or user-visible progress disappears.
Our Opus 5.5 analysis covered the first wave of these changes. Sonnet adds a different question: how much capability can a workhorse deliver before extra effort and handoff costs erase its economic advantage? That calls for a configuration comparison and a migration plan, not an automatic promotion to max effort.
The familiar numbers are the baseline, not the breakthrough
The Sonnet 5.5 reference lists text and image input, text output, and availability through the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry. Those are documented routes, not a test of your region or quota.
Compare it with Sonnet 5 before calling a specification new. Prices below are standard USD per million tokens from Anthropic’s pricing table; cache reads and writes are separate billing categories.
Specification or rate | Sonnet 5 | Sonnet 5.5 |
|---|---|---|
Context window | 1M tokens | 1M tokens |
Synchronous output ceiling | 128K tokens | 128K tokens |
Fresh input / output | $2 / $10 | $2 / $10 |
Five-minute / one-hour cache write | $2.50 / $4 | $2.50 / $4 |
Cache read | $0.20 | $0.20 |
Batch input / output | $1 / $5 | $1 / $5 |
One genuine caching change: the minimum eligible prompt drops from 1,024 to 512 tokens. The tokenizer is unchanged from Sonnet 5. These are documented in what’s new.
That lower threshold may help compact, repeated agent instructions. It does not guarantee a cache hit. Likewise, faster token generation cannot shorten a build queue or a slow database call. Measure the workflow component that actually dominates the wait.
A cheaper model can still be the expensive configuration
CursorBench 4.0 gives a useful launch-day comparison. These September 28 results use Cursor’s evaluation setup and token-price-based task costs, not your production bill.
Configuration | Score | Cost per task |
|---|---|---|
Sonnet 5 max | 34.1% | $7.17 |
Sonnet 5.5 medium | 39.2% | $0.70 |
Sonnet 5.5 high | 47.8% | $1.67 |
Sonnet 5.5 xhigh | 53.1% | $3.88 |
Sonnet 5.5 max | 55.5% | $9.67 |
Opus 5.5 high | 56.0% | $3.97 |
Two calculations matter. Sonnet 5.5 medium costs about 90% less than Sonnet 5 max here, with a higher score. Yet Sonnet 5.5 max costs about 2.44 times Opus 5.5 high. The half-point score difference between those last two configurations may be noise; the cost gap is substantial.
RohitAI’s interpretation: the strongest workhorse opportunity may sit below the top effort setting. Compare model-plus-effort configurations at your acceptance threshold. A cheaper token tariff is not a routing policy.
Buy the least expensive configuration that finishes your task correctly—not the cheapest model name with every reasoning control turned up.
This also changes the trial you should run. Anthropic says Sonnet 5.5 defaults to medium effort in Claude apps and Claude Code, but high on the Claude Platform API. A pleasant interactive session is not a controlled test of your API defaults.
And the rollout can reach developers before their application changes: Claude Code v2.1.284 makes Sonnet 5.5 the default Sonnet model on the Anthropic API route. That does not mean every explicitly pinned application or cloud alias switches. Record both the model and effort in reproducible evaluations.
Start with the assumptions your integration makes
The migration guide concerns custom Messages API code. It says Claude Managed Agents requires only the model-name change; that removes protocol work, not behavioral evaluation. For a custom integration, use this failure map before editing requests.
Old assumption | Sonnet 5.5 change | Migration test |
|---|---|---|
disabled means no thinking | Use between_tools at high effort or below | Check mode and effort together |
A tool can be forced | any and tool return HTTP 400 | Require execution evidence in application state |
Thinking travels with the transcript | Model, account and prefix constraints apply | Resume, hand off and edit history deliberately |
One computer interface works everywhere | Tool versions differ by provider | Exercise the actual cloud adapter |
Advisor output is readable text | Supported advisors return encrypted results | Handle the redacted result variant |
Separate newly introduced failures from old cleanup. Manual thinking budgets and non-default sampling controls were already unsupported on Sonnet 5. A migration from Sonnet 4.x has a different starting point from a migration from 5. Don’t attribute every stale setting to this release.
between_tools is a mode choice, not a spelling fix
Sonnet 5.5 rejects thinking.type: disabled. The thinking documentation offers between_tools to suppress up-front thinking, at low, medium or high effort. Without tools, that mode returns text without preliminary thinking. Omitting the thinking field instead enables adaptive thinking.
For a well-scoped agent evaluation, this is a documentation-based starting configuration, not a complete request or a tested production recipe:
{
"model": "claude-sonnet-5-5",
"output_config": {"effort": "medium"},
"tool_choice": {"type": "auto"}
}That fragment leaves adaptive thinking enabled. If you need no up-front thinking, add "thinking": {"type": "between_tools"}. The migration reference says this mode rejects display, budget_tokens and block_binding fields, as well as mid-conversation effort changes. xhigh and max require adaptive thinking.
The design consequence is easy to miss. A router that starts cheap and later increases effort cannot assume all these controls compose. Use an explicitly tested transition or start a new work phase from a checkpoint. Otherwise, the “think harder now” recovery path may itself be invalid.
Effort changes can also affect the bill before they affect answer quality. Top-level effort changes invalidate the prompt cache; supported per-message changes preserve it, remain beta, and require adaptive thinking here. Account for cache churn when evaluating dynamic effort routing.
Required actions belong in the workflow, not the tool selector
The release-specific tool rules reject tool_choice types any and tool, including at the token-counting endpoint. auto and none remain supported.
Where available, strict tool use guarantees schema-conforming arguments when a tool is called. It does not guarantee that an optional call happens. That distinction matters more than replacing one parameter.
Consider a coding agent required to run the targeted regression test before presenting a patch as ready. Under automatic selection, a confident summary is not proof of execution. The application should require a recorded test result tied to the current patch, not accept the phrase “tests pass.”
A patch is ready only when:
the required change exists
the required check ran against that change
its result satisfies the acceptance rule
no required work remains unresolvedRohitAI’s recommendation: separate the proposed action, the executor’s result and the condition that advances the task. This makes missing calls and stale results detectable without asking another model to judge whether the first one sounded confident.
A transcript is no longer a portable session
There are three different boundaries here. Keeping them separate prevents a successful request from being mistaken for a faithful continuation.
Model compatibility. Sonnet 5.5 can read Sonnet 5 and some earlier thinking, but not Opus 5/5.5 or Fable/Mythos thinking. No other model reads Sonnet 5.5 thinking. Incompatible blocks are dropped before inference. See the compatibility rules.
Account binding. Sonnet 5.5 thinking works within its originating account or linked accounts. An unrelated account can submit the history successfully while those blocks are dropped. See account-bound thinking.
Conversation prefix. Editing the system prompt, tool definitions or earlier messages can invalidate subsequent thinking and cause HTTP 400 on synchronous requests. Prefix enforcement defaults on for accounts created from August 31, 2026, 00:00 UTC; older accounts can opt in. See prefix binding.
The third boundary creates an awkward support trap: your long-lived staging account may pass a smoke test that fails for a new customer. Test the enforced behavior explicitly, not just whichever account happens to run CI.
With adaptive thinking, prefix_mismatch_behavior: drop_block permits recovery but does not restore the lost reasoning. On Claude API and Google Cloud, the thinking-binding-controls-2026-08-01 header exposes account-related drops as organization_binding_mismatch in input_transformations. The preserved-thinking guide documents these controls.
RohitAI’s architectural read: switch models at evidence checkpoints, not arbitrary turns. Before handing off a debugging investigation, preserve the failing example, files examined, changes made, test results, constraints and unresolved questions. These are observable work products—not requests to reveal hidden reasoning.
Portable checkpoint:
objective and acceptance criteria
artifact versions and changes
tool results and verified facts
decisions, constraints and open questions
Opaque history:
preserved separately for compatible continuationKeep the canonical history separate from each route’s usable input. A server dropping a block for one request is not a reason to overwrite the original transcript. Conversely, retaining an opaque block does not make a new model capable of reading it.
This extends the state-portability issue in our Grok 4.7 analysis, without implying the vendors use identical mechanisms. My prediction is that routers will increasingly assign whole work phases. The savings from swapping every turn have to exceed both reconstruction cost and the chance of losing an important constraint.
Cloud availability does not mean interface parity
Two provider differences deserve their own fixtures. The computer-use guide changes the direct API and Google Cloud tool version, while the structured-output support matrix does not list Sonnet 5.5 support on Amazon Bedrock at this snapshot.
Route | Computer-use migration | Strict-output assumption |
|---|---|---|
Claude API / Google Cloud | Use computer_toolset_20260801; computer_20251124 is rejected | Check supported schemas and feature configuration |
Amazon Bedrock | computer_20251124 remains accepted | Do not assume Sonnet 5.5 structured outputs or strict tools are supported |
For the new computer toolset, execute member actions sequentially. Stop at the first failure, but return a result for every call, marking later actions unexecuted. Echo toolset_name: computer on results. These are execution requirements, not cosmetic field changes.
A click followed by typing and a screenshot is a dependency chain. If the click fails, typing into whichever field already has focus is not partial success. A generic parallel tool dispatcher can turn an otherwise valid model response into the wrong desktop operation.
The advisor compatibility table also needs review: Sonnet 5.5 executors reject Opus 4.8, Opus 4.7 and Sonnet 5 advisors. Supported pairings return advisor_redacted_result, not readable advice. Build explanations from visible decisions and execution evidence; do not depend on parsing private advisor text.
Scores are useful only with their acceptance rules attached
Anthropic’s Sonnet 5.5 system card reports Terminal-Bench 4.0 at 70.6% ±2.5 standard error with max effort. The run used Claude Code --bare, 66 tasks, five trials each and default fallback; fallbacks affected 1.5% of trials. This is a configured service result, not exclusively one checkpoint answering every request.
The same card reports FrontierCode Main at 52.1% with xhigh and 46.2% with max; that evaluation penalizes out-of-scope edits. For OSWorld 2.1, Sonnet 5.5 reaches 80.1% partial credit but 43.5% strict pass. Those figures are reported evaluations, not experiments reproduced for this article.
The lesson is not that more effort causes worse work everywhere. It is that more effort and more acceptable work are different variables. Anthropic’s prompting guidance warns that low effort may skip verification, while xhigh or max may add unsolicited review work. Measure both unfinished work and unnecessary work.
For a repository maintainer, an extra refactor can enlarge the review surface even when it is technically competent. For a desktop operator, completing most steps may leave the only important submission unfinished. Our Holo4 analysis makes the same distinction between progress and completion; its differently versioned OSWorld results should not be ranked directly against these.
Three failures an HTTP-success chart will miss
A quiet interface. Longer between-tool progress notes now arrive in thinking blocks. Default omitted display can hide them from a text-only renderer. Adaptive-mode updates need display: updates with thinking-display-updates-2026-08-18; between_tools returns summaries without accepting a display field. Follow the display documentation and dispatch response blocks by type.
A refusal presented as completion. A declined request can return HTTP 200 with stop_reason: refusal. On the Claude API, opt-in beta default fallback retries cyber and frontier_llm declines on Sonnet 5—not every refusal category. Inspect the served model and stop reason. Fallback documentation.
An incomplete bill. For server-side fallback, usage.iterations records attempts; top-level usage describes the returned attempt. Include billed attempts without double-counting that returned usage. With server-side-fallback-2026-07-01, the service translates between_tools to Sonnet 5’s disabled mode; a client-managed rollback must adapt settings itself. Routing and usage rules.
RohitAI’s recommendation is to track four outcomes separately: request accepted, continuity retained, useful progress shown and task accepted. Collapsing them into one “success” metric hides exactly the regressions this migration can introduce.
A rollout that can distinguish improvement from luck
Choose representative tasks with checkable outcomes: a bounded bug fix, a document grounded in supplied evidence, a mandatory-tool workflow and a resumed investigation. Add desktop tasks only if your product uses them. Keep failed attempts in the dataset.
Freeze the comparison. Record model, provider, SDK, explicit effort, thinking mode, tool schema and fallback policy. Use identical starting artifacts. Evaluate the actual user-facing defaults separately from tuned configurations.
Sweep effort before widening traffic. Try medium and high on well-scoped tasks, then test higher effort where failures justify it. Score correctness, scope compliance and completed verification—not just fluent answers.
Replay the uncomfortable sessions. Resume after interruption, change tools through supported mechanisms, test prefix enforcement and exercise model/account handoffs. Assert that constraints and evidence survive even when opaque thinking cannot.
Break the executor deliberately. Test a missing mandatory call, a failed computer action and an encrypted advisor result. The application must remain able to explain what ran and what did not.
Measure accepted work. Track total spend, cache behavior, tool steps, fallback frequency, time to useful progress, completion latency and human repair. A parseable artifact is not necessarily a complete one.
Canary with a compatible rollback. Start with a limited workload and pre-agreed acceptance thresholds. Retain the old route and a portable checkpoint. Test the rollback request shape rather than discovering its incompatibilities during an incident.
For structured tasks, Anthropic specifically recommends treating a max_tokens stop as failure even if the visible JSON parses. Hidden thinking also consumes the output allowance. Preserve enough headroom and inspect completion signals alongside schema validation.
Cost per accepted task =
total spend across successful, failed and retried runs
/ number of tasks meeting the acceptance criteria
Track human review and repair time alongside it.This is how I would judge the upgrade: can it finish the same class of work with less total expenditure and no increase in cleanup? If acceptance improves, a higher bill may still be worthwhile. If acceptance falls, cheaper attempts are a distraction.
Before you switch: quick answers
Is Sonnet 5.5 a token-price cut?
No. Standard input/output rates remain $2/$10 per million tokens, matching Sonnet 5. Anthropic attributes its lower task-cost claim to using fewer tokens. Current pricing is distinct from observed cost per finished job.
Does Sonnet 5.5 support 300K output?
The model reference lists 128K for synchronous output. Up to 300K is a Message Batches API extension using output-300k-2026-03-24, not the default ceiling for interactive requests.
Can I move every old conversation to the new model?
Do not treat transcript storage as a portability guarantee. Test the specific source model, account relationship and history edits. Preserve explicit task evidence so a handoff can continue even when hidden reasoning cannot.
The upgrade worth making
Sonnet 5.5 offers a credible reason to revisit workhorse routing. The most promising trial is a well-specified task at a measured effort setting, with a clear definition of done—not a universal max-effort default.
My expectation is that teams with explicit checkpoints and acceptance checks will capture the improvement first. They can tell whether an apparently successful run preserved its constraints, executed the required work and produced something usable. Teams relying on conversational continuity alone will have a harder time separating model gains from integration regressions.
Migrate the request shape. Test the resumed session. Then decide which configuration earned the traffic.
Evidence note: this article uses Anthropic documentation and release artifacts plus evaluator-owned Cursor results checked on September 28, 2026. API migration behavior is documentation-verified; no production migration or model benchmark was run for this article. Calculations and routing recommendations are RohitAI’s analysis.
