vLLM published stable v0.31.0 on October 5, 2026, at 06:44:55 UTC. For operators moving from 0.30, the release adds a preload CLI and weight-cache operational improvements, experimental initialized-engine snapshots, and a default rejection of non-empty per-request multimodal overrides. These changes affect restart procedures and client compatibility, not just package selection.
Treat runtime adoption, weight-cache adoption and snapshot experiments as separate rollout decisions. This guide proposes acceptance checks for those changes; it is not a complete breaking-change audit or a hands-on performance review. For the earlier Fast Start and state-lifetime background, see the vLLM 0.30 guide. Persistent weight caching already existed in that release.
1. Pin the distribution before testing the runtime
The GitHub release API lists eight uploaded assets. At 07:20:55 UTC on October 5, the version-specific PyPI endpoint still returned HTTP 404. GitHub publication therefore does not establish that pip install vllm==0.31.0 is available from PyPI. Recheck your chosen distribution before installation; no wheel was installed or container pulled for this guide.
Record the chosen wheel hash or image digest, installed vLLM version, model revision, GPU/driver stack and parallel layout. Verify compatibility and installation in staging before changing production. Keep the previous artifact and configuration available; this guide does not establish that any particular wheel or container will run on your hardware.
2. Match the preload daemon to the engine
The preload design retains processed weight shards on local GPUs and lets engines reuse them through interprocess communication, or IPC. Daemon and engine must run on the same node as the same user; each node needs its own launcher. This is weight reuse, not recovery of conversations or in-flight requests.
The cache identity includes the vLLM version as well as model revision, dtype, quantization, TP/DP placement and target/draft role. Operational implication: an existing 0.30 daemon is not a matching warm cache for a 0.31 engine. Budget for populating the new daemon, and retain the old engine/daemon configuration as a rollback pair.
Placement: match tensor-parallel (TP) and data-parallel (DP) ranks, GPU placement and socket paths. The daemon implementation supports tensor, expert and data parallelism, but rejects pipeline parallelism. Give its rendezvous a port distinct from the engine’s.
Draft models: verify both target and draft groups when using speculative decoding. The cacheable draft-method list covers MTP, EAGLE and EAGLE3 with a draft-model configuration; it does not make every speculative method cacheable.
Memory and lifetime: default
zero_copymode shares daemon-owned weights, so the daemon must outlive dependent engines and sleep-mode weight offloading is incompatible.copymode duplicates weights into engine-owned storage and requests cache release. Plan transient memory and restart behavior accordingly. Mode documentation.
The CLI loads weights from disk with vllm preload --model …; --load-format ipc_cache belongs on vllm serve, not on the daemon. This unexecuted syntax sketch uses a placeholder local model path. Add matching model, dtype, quantization and topology options for your deployment; it is not a snapshot recipe.
# Terminal 1: keep the daemon running
vllm preload --model /path/to/pinned-model \
--weight-cache-health-host 127.0.0.1 \
--weight-cache-health-port 8081
# Terminal 2: controlled cache-use check, same node and user
vllm serve /path/to/pinned-model \
--load-format ipc_cache \
--model-loader-extra-config '{"fallback": false}'3. Prove cache use separately from readiness
The optional preload /health endpoint returns 200 after local readiness is marked and while no child has exited; otherwise it returns 503. It is disabled unless --weight-cache-health-port is set, and its default bind address is 0.0.0.0. Choose the address deliberately. This health check does not execute inference or establish whole-cluster health.
The IPC loader enables disk fallback by default, so successful startup alone cannot prove a cache hit. With fallback=false, an unavailable daemon gets a bounded startup wait controlled by state_timeout_s (300 seconds by default). A configuration mismatch is not a readiness wait. Unsupported platforms and quantization remain errors even with fallback enabled.
Use the following proposed matrix, informed by upstream cache tests. Those tests were inspected as source, not executed for this article.
Staging case | What to establish before promotion |
|---|---|
Ordinary disk start | Baseline time to a successful request and peak GPU memory, with the cache path disabled. |
Cold daemon population | Every node and required model role becomes ready; initial loading cost is recorded separately. |
Warm engine restart, fallback disabled | Loader logs confirm tensors mapped from the daemon, followed by a successful request through the intended ingress. |
Cache unavailable, intended fallback enabled | Disk loading actually completes within your recovery and memory limits. Exercise this in an isolated test, not by killing a live shared daemon. |
Deliberate cache-identity mismatch | Disabled fallback rejects the mismatch; enabled fallback takes the documented disk path. Restore matching configuration afterward. |
Keep model revision, prompt set and serving options fixed across comparisons. Record cache-path logs, UTC timestamps, time to successful response, output checks and peak memory. Probe each node, then the real serving route. A ready daemon, a confirmed cache mapping and a usable service are three separate acceptance conditions.
4. Check snapshot eligibility before writing a restore runbook
Initialized snapshots use CRIU, a process checkpoint/restore tool, and remain experimental. The release-pinned requirements specify Linux x86-64, one NVIDIA GPU, TP1 and a single plaintext HTTP server without authentication. Other parallel sizes, TLS, middleware, Unix sockets and speculative decoding are unsupported. Only dense float16 TP1 is described as validated.
That excludes treating snapshots as an automatic add-on to a DP or speculative-decoding preload deployment. Do not remove production authentication or topology requirements to qualify. An isolated pilot needs external access controls, CRIU/CUDA tooling, root or passwordless sudo, disabled io_uring, no external established TCP peer, and a remote model ID pinned to an immutable 40-character revision. Local model directories are unsupported. Snapshot prerequisites.
Check the actual image contents. The tagged Dockerfile installs the snapshot runtime only for linux/amd64 builds with CUDA major version 13 or newer; CUDA 12.x and Arm64 builds skip it. A CUDA 12.9 release artifact therefore does not imply bundled snapshot tools.
A snapshot is not a self-contained model or a saved live KV cache. The capture/restore code releases reloadable model and KV state before capture, then reloads weights and recreates KV allocations during restore. The manifest binds compatibility to the host, GPU, driver, kernel, Python, PyTorch, vLLM and checkpoint tooling, alongside model/tokenizer revisions and engine arguments. Preserve the exact environment and external files, not merely the artifact directory.
Prepare: create before traffic, protect the artifact as process memory, and retain its model, container and generated-cache paths. The documentation makes no power-loss durability guarantee; recreate after an unclean shutdown. Artifact limitations.
Inspect and restore: inspect the manifest before executing the saved process. Restore compares a recorded single-token canary with generated output before reporting success. That controller check is narrower than application correctness; follow it with representative requests and output evaluation.
Recover: snapshot failure does not automatically start an ordinary server. Rehearse an explicit ordinary-start recovery path and verify cleanup before retrying. Do not terminate a process merely because its PID appears in an old manifest. Restore behavior.
5. Migrate multimodal request contracts
In 0.31, non-empty mm_processor_kwargs and media_io_kwargs are rejected by default, not silently ignored. The request validator permits them through this particular gate only when request-kwargs trust is enabled. Omitted fields, null and empty dictionaries do not trigger it; that does not guarantee the rest of the request is valid.
Client payload | Default server | Server with --trust-request-mm-kwargs |
|---|---|---|
Both fields omitted, null or empty objects | This gate does not reject the request. | This gate does not reject the request. |
Non-empty mm_processor_kwargs | Rejected by this gate. | Proceeds past this gate; other validation still applies. |
Non-empty media_io_kwargs | Rejected by this gate. | Proceeds past this gate; other validation still applies. |
Inventory SDK extra-body fields and preprocessing assumptions before rollout. The change description covers online chat/completion rendering and pooling/scoring paths, while offline LLM usage is unchanged. Test the routes your applications actually use, with benign non-empty values for each field separately and together. Capture the error visible through your gateway rather than assuming an exact HTTP response from source inspection. Upstream rejection and opt-in cases provide a starting point.
For fixed preprocessing policy, move settings to deployment-level --mm-processor-kwargs or --media-io-kwargs. If clients genuinely need different settings, evaluate a separately bounded trusted-client deployment. The security guidance limits --trust-request-mm-kwargs to trusted clients because overrides can change media loading and preprocessing resource use. Do not enable it globally just to clear migration errors.
Replacing varying client settings with one server default may change preprocessing and outputs even when requests stop failing. Include representative image, audio or video inputs from your own workload in the compatibility check. Keep the trust decision separate from whether the model loads successfully.
Promote only the parts that pass
Adopt the core runtime only after your client-contract and ordinary-serving checks pass. Enable cached loading after mapping, readiness and recovery checks pass. Keep snapshots an independently reversible experiment until the exact supported environment and ordinary-start recovery are demonstrated. None of these checks alone establishes better throughput, output quality or cost.
Before promotion, review the full 0.31 release notes for additional tokenizer, quantization and accelerator changes outside this guide. Preserve the old artifact, configuration and client routing until rollback has been rehearsed.
Methodology: This AI-assisted guide draws on the official release, version-pinned documentation, implementation and upstream test source. The acceptance matrix and rollout sequence are analysis derived from those sources. No package installation, GPU benchmark, snapshot restore or client request test was performed for this article.
