Article

SGLang 0.5.21: Prefill–Decode Role-Switch Acceptance Checklist

An operator checklist for SGLang 0.5.21 role switching: qualify support, drain traffic, check graph headroom, verify execution and plan cache recovery.

Editorial illustration for SGLang 0.5.21: Prefill–Decode Role-Switch Acceptance Checklist: a document represents the research briefing. Not documentary evidence.

SGLang released v0.5.21 on October 2, 2026, adding opt-in runtime switching between prefill and decode for disaggregated inference workers. For operators who separate prompt processing from token generation, the feature offers a way to change a worker’s role without reloading model weights on the successful path. Version 0.5.21 is available on PyPI; availability does not establish compatibility with a particular GPU deployment.

The acceptance question is whether the worker can drain, switch and verify readiness in your serving stack. A successful control response is not enough: it can represent a same-role no-op, and a real switch can finish without successful decode-graph capture. The version-pinned implementation also resets reusable prefix state. The checklist below translates those branches into proposed acceptance evidence, not reported test results.

1. Qualify the deployment before enabling the flag

The feature is disabled by default. Enable --enable-pd-role-switch only on a PD deployment whose model, GPU/runtime, transfer library and topology you intend to qualify. Save the version or image digest, model revision, quantization, tensor-parallel layout, cache and graph settings, and router/controller version alongside the results.

Both Mori and Mooncake explicitly implement role switching in this tag. That is a backend capability, not certification of every model, network or parallelism combination.

Check the startup exclusions before designing the experiment:

  • DP attention and configured speculative decoding are rejected with role switching enabled.

  • Expert, pipeline, data and decode-context parallel sizes above one are rejected, as is a MoE all-to-all backend other than none.

  • A positive decode-host-receive threshold is incompatible; the runtime guard also rejects an enabled staging buffer.

If your workload depends on one of these configurations, retain fixed-role pools rather than silently removing it to make the switch pass. Use the unchanged deployment as the comparison baseline.

2. Prove withdrawal and drain through your actual router

Remove the candidate from every route that can admit work, then demonstrate that new requests stop selecting it while admitted work finishes. The full-idle check must pass before a real role change. It covers more than the ordinary waiting queue: active batches, PD bootstrap/transfer work and applicable asynchronous cache operations can keep the worker busy.

Set a drain timeout from your workload and service objective. Preserve enough capacity in both roles during withdrawal. A lone one-prefill/one-decode pair has no spare worker in either role; without added capacity, plan a maintenance interval rather than assuming uninterrupted service.

The bundled MiniLB illustrates withdrawal, not-idle retries and role-pool reassignment, but is explicitly a testing/debugging load balancer. Its behavior does not prove that your production gateway performs those steps. Retain routing evidence for long requests, in-flight transfers and enabled cache operations, not just an idle health response.

3. Establish headroom for the first decode capture

For a switch to decode when no captured decode batch-size list exists, the switch handler requires a nonnegative decode_cuda_graph_memory_gb and checks it against available GPU memory before teardown. It compares an operator-supplied estimate; it does not reserve that memory or certify the estimate.

The request schema describes that value as a measured graph footprint from a matching decode peer. Use GET /server_info and its internal_states to record the peer’s decode_cuda_graph_bs and decode_cuda_graph_memory_gb. Match the model, hardware and graph configuration; the reported graph footprint is not a measurement of every transient allocation peak.

Proposed planning rule: require H ≥ M + S on each relevant worker/rank, where H is available headroom, M is the matching graph footprint and S is a justified allowance for transient allocations and variation, all in the same units. This is an operator acceptance rule, not an SGLang guarantee. Validate a finite, nonnegative budget; the sources establish neither a universal margin nor a safe GB value.

In an isolated acceptance environment, exercise missing-budget and insufficient-headroom rejection paths before a real flip. An arbitrary zero may satisfy part of the input check but is not evidence of sufficient memory. Record first-capture transitions separately from later round trips.

4. Verify a real transition and the execution mode

Use the worker’s POST /pd_role_switch endpoint with new_role set to prefill or decode. For first decode capture, supply the validated memory estimate. The optional decode_cuda_graph_bs selects capture batch sizes; omission uses the configured list. Once graphs are captured, the model runner skips recapture, so a later switch request with a new list is not a retuning guarantee.

  • Save success, message, old_role, new_role and safe_to_restore with request/response timestamps. A same-role success can test idempotency, but cannot pass the real-transition check.

  • Read live disaggregation_mode from /server_info internal states, reconcile worker results, and verify controller pool membership. Do not rely only on startup configuration or the HTTP status.

  • Run an end-to-end generation fixture through the intended route after reassignment. Hold the worker out of normal traffic if state, routing or worker outcomes disagree.

  • Inspect graph batch sizes, graph-memory reporting and capture logs. If claiming the workload actually replays graphs, retain suitable profiling or telemetry: a nonempty capture list does not show that every request is graph eligible.

The capture exception path can log a failure and still return a successful role switch, with eager execution as the fallback. Qualify that mode against your service objective explicitly, or keep the worker out until the intended execution mode is restored.

5. Measure cold-cache recovery, not just warm throughput

Avoiding a weight reload does not preserve a warm prefix cache. The teardown implementation retains the underlying token KV pool but resets radix/HiCache prefix state and allocation bookkeeping when prefix caching is enabled. External HiCache storage clearing is best-effort, so verify its outcome separately where used.

Use a fixed, representative mix of short, long and repeated-prefix requests. Compare correctness fixtures, time to first token, time per output token, tail latency, completed throughput, errors and cache behavior before switching, immediately afterward and after warmup. Keep sampling settings and load comparable; nondeterministic generation need not produce identical text.

Repeat actual prefill-to-decode and decode-to-prefill transitions with a cycle count chosen for the deployment. Track free memory, graph/KV allocations, transfer failures and transport resources over those cycles. Set tolerances from the application’s existing service objectives, not a borrowed benchmark percentage. Actual switch duration, warmup time and memory peaks remain unknown until measured on that configuration.

For background on the separate lifetimes of serving processes, weights and session caches, see SGLang 0.5.17 Gives Agent Sessions a Vote in GPU Memory. This checklist concerns changing a live worker’s PD role, not that earlier release’s recovery features.

6. Classify the outcome before restoring traffic

The following decision table combines the runtime failure branches and MiniLB restoration example with proposed operator actions. It is not a claim that external health checks automatically isolate or recover a worker.

Outcome

What it establishes

Proposed routing/recovery decision

Success; old and new roles match

No actual flip was exercised.

Record the idempotency result; still require a real transition and readback.

Rejection with safe_to_restore=true

The response explicitly permits restoring old-role routing.

Confirm the old live state and controller membership before restoration; resolve the rejected precondition.

Teardown/rebuild failure; restart required

The instance is marked unhealthy; there is no in-place rollback.

Keep it isolated, restart, then verify live role and routed generation before readmission.

Role switch succeeds; graph capture fails

The role may be usable in eager mode, without the intended graph performance.

Accept only if the independently checked eager-mode behavior meets your objective; otherwise hold it out.

Timeout, lost connection or inconsistent worker results

The completed state is unknown or inconsistent.

Keep the worker isolated and reconcile its state. Do not infer rollback or blindly restore its old route.

Retain an acceptance record containing the exact deployment manifest, input fixtures, outputs, timestamps, readbacks, routing changes, raw metrics/logs and failure-recovery outcomes. Admit the worker only when actual role, routed correctness, execution mode and cold-to-warm performance all satisfy the predeclared criteria. Otherwise keep the established fixed-role deployment.

Methodology: Prepared with AI assistance from SGLang’s published release, package metadata and version-pinned source, checked on October 5, 2026. This is source-based analysis and a proposed procedure. RohitAI did not run GPU workloads, inject failures or reproduce upstream benchmarks for this guide.