Posts tagged “inference”
vLLM 0.30 Gives Model State New Lifetimes. Budget Them.
Fast Start and HiSparse move weights and KV across new lifetimes and memory tiers. The vLLM 0.30 upgrade question is who owns warm state—and its cost.
Tencent Hy4 Preview: Why a 770B Open-Weight Agent Model Is Easier to Rent Than Run
Tencent’s Hy4 Preview offers 1M context, Apache-2.0 weights, and cheap agent APIs. Its real test is retrieval, caching, and route economics.
Nemotron 3.5 Lightning Turns Agent Speed Into a Routing Problem
NVIDIA Nemotron 3.5 Lightning is a fast open agent worker. Its real value lies in routing, validation, and deployment economics.
vLLM 0.27.0 Expands the Runtime—and the Blast Radius
vLLM 0.27.0 adds Kimi K3, MRv2 workloads, Rust control APIs, and fault recovery—but its PyTorch 2.13 jump makes this a fleet migration.
SGLang 0.5.17 Gives Agent Sessions a Vote in GPU Memory
SGLang 0.5.17 adds session-aware KV caching, faster recovery, Rust ingress, and frontier-model paths—shifting serving toward agent runtime policy.
JoyAI-Video-Edit Hits 30 FPS. The Stream Still Runs on Five Clocks.
JD’s open-weight JoyAI-Video-Edit reports 30 FPS on one B200. Here’s what latency, stream state, and single-session serving mean for builders.
LG’s K-EXAONE 2.0 Makes Open Weights a Cluster Procurement Decision
LG’s 750B K-EXAONE 2.0 is Apache 2.0, agent-ready, and built for eight H200s. Here is what its open-weight economics mean for builders.
vLLM 0.26.0 Makes KV Tiering an Infrastructure Contract
vLLM 0.26.0 matures KV tiering with storage identity, cache events, model-specific execution, API controls, and a wider security boundary.
Thinking Machines Inkling: Open Weights Built to Be Fine-Tuned
Inkling is a 975B open-weight multimodal model built for customization through Tinker—but self-hosting still demands institution-scale GPUs.
OpenAI Jalapeño: Why Its First Inference Chip Is Bigger Than an Nvidia Alternative
OpenAI Jalapeño is not just a custom AI chip. It is OpenAI’s move toward full-stack inference economics, agent latency, and compute control.