Anthropic announced on September 30, 2026 that Claude Sonnet 4.5 will retire from the Claude API on November 30. Teams still calling claude-sonnet-4-5-20250929 need to migrate: Anthropic recommends claude-sonnet-5-5. The deadline appears in its September 30 release notes.
Deprecation does not stop requests today, but retirement does: Anthropic says calls to retired models fail. The reviewed notice does not establish a cutoff hour or timezone, so plan an earlier internal completion date.
The immediate task is to find every remaining 4.5 dependency, adapt the requests it sends and establish a fallback that will still work after retirement. This guide focuses on that exit plan; our Sonnet 5.5 launch analysis covers the broader model comparison.
Confirm which platform owns your deadline
Anthropic says its lifecycle dates cover the Claude API, Claude Platform on AWS and Microsoft Foundry. Amazon Bedrock and Google Cloud set separate schedules. Claude Platform on AWS is not the same service as Amazon Bedrock.
As checked on October 5, AWS’s Sonnet 4.5 card labels the model Active and gives no-sooner-than retirement dates, not a fixed November 30 shutdown. Google’s model page likewise lists a retirement floor. Neither establishes that the native API deadline applies there, nor guarantees indefinite availability.
Sonnet 5.5 is listed as available across those platforms. Before choosing a replacement route, confirm the actual endpoint, model ID, region, quota and required features in your account. A documentation listing is not an access check.
Find dormant dependencies as well as live traffic
Start with the Console Usage CSV, which Anthropic documents as broken down by API key and model. Then inspect deployment configuration, gateway mappings, stored prompts, scheduled jobs, queued requests and fallback routes. Usage shows what ran; configuration review can catch a monthly job that has not run recently.
A read-only repository search can locate files containing the 4.5 identifier. This example prints filenames, not matching configuration values:
rg --files-with-matches --hidden -g '!.git/**' -g '!node_modules/**' 'claude-sonnet-4-5' .This is a starting point, not a complete inventory: ignored files, external configuration stores and dynamically constructed model names need separate inspection. Record an owner, provider, deployment location and replacement decision for each dependency. Include rollback configurations; an old model ID left there can revive the problem during an unrelated incident.
Apply the cumulative 4.5 migration changes
For custom Messages clients, Anthropic’s migration guide is cumulative: apply the every-model, 4.6-or-earlier and 4.5-or-earlier sections. Reading only the changes from Sonnet 5 misses legacy settings. The following are priority checks, not the full platform-specific checklist:
Replace manual
thinking.type: enabledandbudget_tokenswith adaptive thinking and an explicitoutput_config.effort. Remove non-defaulttemperature,top_pandtop_kvalues.Replace final-assistant prefills according to their purpose. Do not simply delete a formatting requirement and assume the application still receives the same output.
Remove obsolete context-window and
interleaved-thinking-2025-05-14headers. Replacefine-grained-tool-streaming-2025-05-14with per-tooleager_input_streamingwhere applicable. Parse tool inputs with a standard JSON parser.
For supported structured JSON responses, the current wire field is output_config.format. An SDK helper may still accept output_format and translate it internally; check the structured-output documentation before renaming helper arguments, and verify support on your platform.
Omitting thinking now enables adaptive thinking, unlike 4.5. Sonnet 5.5 rejects disabled; between_tools suppresses up-front thinking at low, medium or high effort. Read responses by block type and preserve thinking blocks in tool loops. Hidden thinking still uses max_tokens and is billed as output. These are documented thinking behaviors.
Tool-dependent workflows also need an acceptance check: tool_choice values any and tool are rejected by Sonnet 5.5. With auto, a textual answer can arrive without a tool call. Require evidence of the actual business action in application state before marking a task complete.
If the application operates a desktop, use the provider-specific computer-use migration: Sonnet 5.5 uses computer_toolset_20260801 on the Claude API and Google Cloud, but computer_20251124 on Bedrock. Review the executor and result protocol as well as the tool name.
Recalculate costs using the new tokenizer
The standard global token rates below are USD per million tokens from Anthropic’s pricing table. They exclude separate feature charges, residency premiums and negotiated discounts.
Billing category | Sonnet 4.5 | Sonnet 5.5 |
|---|---|---|
Uncached input | $3 | $2 |
Output | $15 | $10 |
Five-minute cache write | $3.75 | $2.50 |
One-hour cache write | $6 | $4 |
Cache read | $0.30 | $0.20 |
A one-third reduction in the rate is not necessarily a one-third reduction in spend. Anthropic says the newer tokenizer produces approximately 30% more tokens for the same text, with variation by content.
Illustrative calculation, not a measured saving: if identical uncached text uses 1.3 times as many tokens, its input cost becomes 1.3 × $2 ÷ $3 = 0.867 times the old cost—about 13.3% less. Changed output length, thinking, images, cache hits and retries can change the total substantially.
Use token counting with the target model ID to re-estimate representative inputs. The estimate does not simulate cache hits or determine the final bill. In a separately run evaluation, compare total spend per accepted task, including unsuccessful attempts, rather than the price of one successful response.
Set evidence-based exit criteria before the cutoff
Recommended rollout sequence:
Validate request and response handling first. Replay representative text, structured-output, tool and resumed-session cases in an isolated evaluation environment. Include desktop actions only if the application uses them. Record failures, completion signals and resulting artifacts—not just HTTP success.
Canary the accepted configuration on limited traffic. Set workload-specific limits for errors, task completion, latency and cost before the canary; increase traffic only when those conditions hold.
Close the inventory. Check deployed configurations, queues, scheduled jobs and observed traffic for remaining 4.5 references. Exercise dormant paths deliberately. Keep a separately validated active-model or manual fallback; the retiring model cannot be the post-deadline recovery plan.
Treat retirement remediation and performance tuning as separate decisions. The first is complete when the service no longer depends on 4.5 and its required behavior still works. Higher effort or a different routing policy can be evaluated after that baseline is established.
Methodology: AI-assisted reporting and analysis based on public Anthropic, AWS and Google documentation rechecked on October 5, 2026. No authenticated API, billing or workload tests were performed. The calculation and rollout criteria are analytical guidance, not reported test results.
