Posts tagged “vllm”
vLLM 0.27.0 Expands the Runtime—and the Blast Radius
vLLM 0.27.0 adds Kimi K3, MRv2 workloads, Rust control APIs, and fault recovery—but its PyTorch 2.13 jump makes this a fleet migration.
vLLM 0.26.0 Makes KV Tiering an Infrastructure Contract
vLLM 0.26.0 matures KV tiering with storage identity, cache events, model-specific execution, API controls, and a wider security boundary.