Article

Claude Haiku 5.5: Pricing, Token Thresholds and Migration Changes

Decide whether to migrate from Haiku 4.5: compare two-tier prices, recount prompts and check thinking, tool use, refusals and stored-session behavior.

Editorial illustration for Claude Haiku 5.5: Pricing, Token Thresholds and Migration Changes: an arrow connects an old service to its replacement. Not documentary evidence.

Anthropic released Claude Haiku 5.5 on October 7, 2026, giving teams running high-volume classification, extraction and agent subtasks a new migration option. Its release notes list availability on the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry.

The model specification lists text and image input, text output, a 1M-token context window and up to 128K output tokens for synchronous requests. Starting token rates are lower than Haiku 4.5’s, but the cheapest band stops at 100,000 prompt tokens. Our recommendation: pilot a bounded workload after recounting its prompts and adapting the client, then decide whether completed-task cost justifies switching.

Adoption is optional; compatibility work is not

Haiku 4.5 remains Active in Anthropic’s lifecycle table. The listed October 15, 2026 date is a retirement floor, not a scheduled shutdown. There is no fixed Haiku 4.5 retirement date in the reviewed notice. Do not confuse this with the separate Sonnet 4.5 migration deadline; partner-operated platforms can also have their own schedules.

A useful first pilot is extraction with known expected fields, a summary with a clear acceptance rubric, or a short tool task. Delay a full switch if the application cannot yet handle changed responses or resumed sessions. If reserved capacity is essential, note that Priority Tier does not support Haiku 5.5; an existing Haiku 4.5 commitment does not establish equivalent capacity on the new model.

Two price bands—and a tokenizer change

Recount prompts before switching models: the new tokenizer can move the same text into a higher price band. Anthropic says identical text produces approximately 30% more tokens than on Haiku 4.5, with the increase varying by content. Use the target model’s token counter for representative requests, including instructions and tool definitions; do not reuse Haiku 4.5 counts.

These are Anthropic’s standard list prices in US dollars per million tokens. Prompt length selects the Haiku 5.5 price band, including its output rate. The 1M context limit does not extend the lowest tariff across the whole window.

Billing category

Haiku 4.5

Haiku 5.5: prompt ≤100,000 tokens

Haiku 5.5: prompt >100,000 tokens

Uncached input

$1.00

$0.10

$0.50

Output

$5.00

$0.50

$2.50

Cache write: 5 minutes

$1.25

$0.125

$0.625

Cache write: 1 hour

$2.00

$0.20

$1.00

Cache read

$0.10

$0.01

$0.05

Illustrative input-only calculation—not a measured bill. Assume identical uncached text produces exactly 1.3 times as many tokens on 5.5. Apply input cost = tokens ÷ 1,000,000 × the applicable rate. Exclude output and thinking, tools, retries, taxes, residency premiums and negotiated discounts.

  • 75,000 Haiku 4.5 tokens cost $0.075. At an assumed 97,500 tokens on 5.5, the cost is $0.00975: an 87% input-cost reduction.

  • 80,000 Haiku 4.5 tokens cost $0.08. At an assumed 104,000 tokens on 5.5, the higher rate applies and the cost is $0.052: a 35% input-cost reduction.

Both examples save money under these assumptions, but the savings differ sharply across the boundary. These calculations use the published Haiku 5.5 price bands and the Haiku 4.5 baseline above; they do not predict total workflow savings. For mixed cached and uncached requests, confirm threshold accounting with current provider guidance rather than assuming cache hits remove tokens from prompt length.

There is a separate caching opportunity: on the Claude API, the minimum cacheable prompt length falls from 4,096 tokens on Haiku 4.5 to 512 on 5.5. Short reusable instructions may therefore become eligible. Exact reuse and cache lifetime still matter; confirm hits in usage data. Bedrock has separate provider instructions. For work that can run asynchronously, the model page lists a 50% Batch API input/output discount.

Audit requests, not just the model name

Use claude-haiku-5-5 on the Claude API; Bedrock lists anthropic.claude-haiku-5-5. Other provider IDs are in the model specification. Documentation availability is not confirmation of access, region or quota in your account.

For custom Messages clients, prioritize these changes from the Haiku-specific migration guide:

  • Manual thinking: requests with thinking.type: "enabled" and budget_tokens fail with HTTP 400. Omit thinking or select adaptive thinking, then set output_config.effort.

  • Sampling: remove temperature, top_p and top_k. Non-default temperature or top_p values, any top_k, and temperature plus top_p cause HTTP 400.

  • Assistant prefills: end messages with a user turn. Replace formatting prefills with structured outputs or tools. On Bedrock, use tools: the guide says structured outputs are unsupported.

  • Forced tools: tool_choice values any and a named tool still work on Haiku 5.5, but produce the tool call without a thinking block. Use auto when reasoning before tool selection matters.

  • Computer use: on the Claude API and Google Cloud, replace computer_20250124 with computer_toolset_20260801 and adapt the executor as documented; changing only the tool name is insufficient.

Anthropic says Claude Managed Agents handles this migration with only a model-name change. That exception does not establish quality or cost for your workload.

Choose effort and handle non-answer responses

Haiku 5.5 defaults to adaptive thinking and medium effort. Its effort controls range from low through medium, high, xhigh and max. Effort is a behavioral control, not a hard spending cap. Thinking can still be disabled at low, medium or high; disabling it at xhigh or max returns HTTP 400. Do not import Sonnet-only restrictions into a Haiku migration.

The behavior-change documentation says responses can begin with thinking blocks. Parse by block type, not by assuming the first block is answer text. Thinking consumes max_tokens, so a small budget can run out before a visible answer. Check truncation and empty replies explicitly instead of treating HTTP success as task completion.

The Haiku prompting guide also documents stop_reason: "refusal" and stop_details.category. Haiku 5.5 has no server-side fallback, and repeating an unchanged refused request usually repeats the refusal. Provide an explicit error or permitted-review path rather than an automatic retry loop.

Stored conversations need a separate compatibility check

For multi-tenant services, preserved-thinking rules introduce two distinct checks. Thinking replayed through an unlinked account is dropped while the request can still succeed. Separately, changing system instructions, tools or earlier messages can invalidate later thinking blocks.

Prefix enforcement defaults on for accounts created from August 31, 2026 at 00:00 UTC; older accounts opt in. Ordinary Messages requests reject invalid prefixes by default when enforcement applies, while documented drop settings and Batch behavior differ. Keep replayed history append-only and bound to its producing account.

Include save-and-resume, history trimming, tool changes and customer-account routing in compatibility evaluation. A clean response on an older development account is not evidence that every customer’s session will work. The binding diagnostics can expose dropped or mismatched blocks on supported routes.

Make the decision on accepted work

A proposed rollout gate: freeze representative inputs and acceptance rules, then compare the current route with Haiku 5.5 at explicit effort settings. Include prompts below, near and above the pricing boundary, plus resumed conversations. Record total billed usage, accepted results, retries, review rate and end-to-end median and p95 latency. Keep the current route available until the new one meets the agreed thresholds.

If the deliverable is only a fixed label or probability, reconsider whether generation is needed at all; our decision-model guide examines that narrower interface. If the task needs generated fields, explanations or tool arguments, evaluate Haiku on those outputs—not just on its lowest token price.

Methodology: AI-assisted reporting and analysis based on Anthropic’s published documentation, with prices, availability and lifecycle details rechecked on October 7, 2026. No inference, billing or latency experiment was run. The arithmetic is illustrative, and the rollout criteria are proposed checks, not reported results.