Article

Claude Haiku 5.5 in GitHub Copilot: Access, Pricing and Admin Controls

Check Claude Haiku 5.5 availability in GitHub Copilot, the 100K-token price boundary, inherited model policies and AI Credit budget controls.

Editorial illustration for Claude Haiku 5.5 in GitHub Copilot: Access, Pricing and Admin Controls: code brackets represent developer tools. Not documentary evidence.

GitHub announced Claude Haiku 5.5 in Copilot on October 7, 2026, adding a lightweight model for paid users and teams managing coding-agent costs. The launch lists Pro, Pro+, Max, Business and Enterprise and says rollout is gradual; eligibility does not mean every account already sees it.

For administrators, model access, paid usage and tool execution are three separate decisions. Our recommendation is to make Haiku available for bounded, reviewable work where existing policy permits it, while keeping a stronger-model option for difficult tasks. A lower token rate is a reason to evaluate that split, not proof that every workflow should switch.

Check billing eligibility, client support and model policy

First check the billing regime. GitHub says legacy annual Pro and Pro+ subscribers on request-based billing do not receive new models or features. The old Haiku 4.5 premium-request multiplier is not a Haiku 5.5 price.

The announced clients are VS Code, Visual Studio, Copilot CLI, cloud agent, Copilot app, github.com, GitHub Mobile, JetBrains, Xcode and Eclipse. Their capabilities are not identical: GitHub limits the 1M-token context option to VS Code and CLI, and reasoning controls to those clients plus cloud agent. Haiku-specific minimum IDE/plugin versions remain listed as TBD; keep clients current rather than assuming any older version will work.

Next inspect the effective policy. Unconfigured GA models inherit the active default-availability policy, so Haiku may already be enabled. An explicit per-model restriction can prevent that without changing unrelated models. Enterprise residency or FedRAMP restrictions still apply. The separate new-feature default policy begins on October 22; that is not a grace period for model access.

  • Enterprise owners: open AI controls → Copilot → Configure models. Check whether Haiku is enabled, disabled or delegated.

  • Organization owners: use Settings → Copilot → Models when the enterprise permits local configuration. Enforced enterprise choices cannot be overridden; the opt-in enterprise-teams access mode disables organization-level model settings.

Default enablement only makes a model available. It does not establish that Auto will route tasks to it: Haiku 5.5 is not listed in the Auto-selection table checked on October 8.

Budget for tokens, including the 100K boundary

GitHub’s Copilot rate card lists these Haiku 5.5 prices in US dollars per million tokens:

Billing category

Input ≤100,000 tokens

Input >100,000 tokens

Uncached input

$0.10

$0.50

Cached input

$0.01

$0.05

Cache write

$0.125

$0.625

Output

$0.50

$2.50

The input threshold selects the pricing tier, including the output rate. All four rates rise fivefold above 100,000 input tokens; a 1M-token capacity is not a promise of the lowest price throughout the window.

Under GitHub’s usage-based billing, one AI Credit represents $0.01 of usage. Credits measure consumption, not a fixed number of prompts; included credits and an additional cash charge are different things.

Illustrative calculation, not a measured bill: assume one call with 80,000 uncached input tokens and 4,000 billed output tokens, identical counts across models, no caching, discounts, retries, other charges or taxes. Cost = input tokens ÷ 1,000,000 × input rate + output tokens ÷ 1,000,000 × output rate.

Using the published Copilot rates, that is $0.010 on Haiku 5.5, $0.100 on Haiku 4.5 and $0.200 on Sonnet 5.5—respectively 1, 10 and 20 AI Credits. These are equal-token comparisons, not equivalent-task savings.

Actual task costs can differ. Anthropic says Haiku 5.5’s tokenizer produces approximately 30% more tokens for the same text than Haiku 4.5. Retries and reasoning also change consumption. Copilot’s default Haiku context size and mixed-cache threshold accounting were not established in this reporting. For a pilot, record actual usage rather than projecting from an old model’s token counts.

Separate shared-pool limits from paid overage

Business licenses contribute 1,900 monthly AI Credits each; Enterprise licenses contribute 3,900. These feed a shared pool at the billing entity, not automatically enforced personal buckets. Lower consumption can preserve that pool without reducing fixed seat fees. Cash savings depend on avoiding usage that would otherwise become paid overage.

For organizations and enterprises, additional paid usage is enabled by default. To allow included credits but prevent spending beyond them, explicitly disable the AI credits paid usage policy in AI Controls. Credit-consuming work then stops when the included pool is exhausted.

For finer control, GitHub’s budget documentation distinguishes these mechanisms:

  • User-level budgets cap each person’s included and metered consumption together. They always impose a hard stop, even when the shared pool still has credits.

  • Organization, cost-center and enterprise overage budgets concern metered spending. To enforce a cap, enable “Stop usage when budget limit is reached”; it is off by default. Cost-center included-usage controls are separate.

  • A $0 budget can block affected users immediately. Do not use it as shorthand for “keep included usage, forbid overage”; use the paid-usage policy for that goal.

Budget exhaustion does not trigger a fallback to a cheaper model. Paid-plan code completions and next edit suggestions remain unlimited and do not consume AI Credits, but that does not keep credit-consuming chat or agent work running.

Choose by accepted work, not the rate card alone

GitHub says its early tests matched Sonnet 5 on many coding tasks with fewer tokens and steps. The announcement provides no quantified task set or reproducible protocol. That is GitHub’s claim about Sonnet 5, not evidence of universal Sonnet 5.5 parity or a measured saving in your repository.

A practical evaluation can stay small:

  1. Start with a scoped edit, a terminal task with a known expected result, or a subagent assignment whose output is easy to check. Define acceptance criteria before running it.

  2. Record the client, chosen model, context and reasoning settings, credits across every call and retry, accepted changes, review time and latency. Compare cost per accepted result with the current workflow.

  3. Retain Haiku for tasks that meet the required quality at a useful cost. Escalate repeated failures or complex work to the stronger-model route rather than extending an unsuccessful cheap session indefinitely.

Permission to select Haiku does not expand an agent’s execution permissions. Review the separate Copilot local-sandboxing controls before allowing broader file or network access. Developers changing a custom Claude API client instead should use the Haiku 5.5 pricing and migration guide.

Methodology: AI-assisted reporting and analysis based on GitHub’s announcement, live documentation and Anthropic’s model specification, checked October 8, 2026. No hands-on Copilot test, latency measurement, billing experiment or account-specific rollout check was performed. The calculation is illustrative; the evaluation steps are a proposal, not reported results.