llama.cpp published prerelease b11401 on October 5, 2026, at 00:36 UTC. It includes PR #29895, which separates lifecycle messages from log output in llama-server router mode. The fix matters to operators whose router manages child processes for model downloads and serving: an unfinished log write could otherwise hide a completion message and leave a model marked as downloading.
Prioritize a staging upgrade if your router has matching stuck-download symptoms. For an apparently healthy deployment, add this change to normal lifecycle regression checks. The evidence does not establish a throughput gain or make this an emergency upgrade for every single-model server.
What changes between the child and the router
Router mode manages multiple models behind one server endpoint; the server documentation describes starting it without a model argument. In the old communication path, the router distinguished state messages from logs by their line prefix. The PR author reports that injecting output without a terminating newline before a download-completion message caused that message to be logged instead of handled. The model remained in the downloading state. This is an upstream reproduction report, not a RohitAI test.
In b11401, router child commands travel separately from ordinary logs. The child reserves its original stdout stream for commands and redirects ordinary stdout writes to stderr. The router reads the two pipes separately, handling state messages on the command pipe and forwarding the other stream as logs. This changes internal process communication, not the public REST API.
The same patch moves color resets before trailing newlines and adds Windows console handling for ANSI colors. Router children inherit the parent’s effective color setting. Cleaner progress output is useful, but it is not evidence that model state is correct.
Why a completed transfer can still look stuck: the download-child implementation sends its result and then waits for the router’s exit instruction. Losing that message can interrupt coordination after the transfer has finished. That is a reason to inspect state and process exits alongside transfer progress—not to assume every stalled download has this cause.
Which deployments should prioritize it?
The following triage is an operational recommendation based on the patch’s scope, not a measured incident-rate ranking.
Deployment or symptom | Recommended priority | What to verify |
|---|---|---|
Router keeps a finished transfer marked downloading | Prioritize an isolated candidate | Completion event, model registration and download-child exit agree. |
Router works, but progress logs contain unwanted blank lines or color escapes | Schedule a lifecycle and logging regression pass | State transitions work with the actual log collector. |
Single-model server without router children | No urgency established by this router bug alone | Review shared-logger behavior and other changes in your upgrade. |
The first affected build and a complete affected-version range are not established. Disabling colors is not a demonstrated complete workaround: the reported trigger includes arbitrary unterminated output. Network, disk and model-access failures remain separate diagnoses.
Pin the candidate and isolate the comparison
The release metadata identifies commit a7fb71fab83b474a0892b9a05aaa3a8ddca2729b and lists binaries for Windows, Linux and macOS, among other targets. Choose an asset matching your architecture and backend, or use your established source-build procedure. Record the executable revision and any container digest; do not infer inclusion from a downstream wrapper’s version number.
Keep the current binary and configuration available for rollback. Use a separate staging port and disposable model cache, while matching the production backend, model quantization, presets, autoload policy and loaded-model limit.
For fault isolation, b11400 to b11401 is a one-commit change. Identically configured builds offer a narrow comparison; your actual production revision is still the rollout baseline. If you are also adopting mixed-input integration changes, retain the separate checks in our b11400 text-and-embedding batch guide.
Check download, readiness and shutdown separately
This is a proposed acceptance sequence, not a report of tests performed. Capture timestamped HTTP responses, model listings, events, logs and child-process exits.
Subscribe before starting. Open
GET /models/sseand wait until the streaming connection is established. The tagged router test explicitly waits for its subscriber before triggering a download: a terminal event can otherwise arrive too early and be missed. A missing event alone is therefore not proof of the pipe bug.Start a download into the isolated cache. Send
POST /modelswith a valid public repository and quantization tag, replacing the placeholder below. The documented response acknowledges background work; it does not certify completion.{"model":"YOUR_PUBLIC_HF_REPO:QUANT_TAG"}Check the terminal result. Look for the matching
download_finishedevent, then useGET /modelsto confirm the model is registered and no longer downloading. Preserve anydownload_failedresult and its logs. The upstream download test checks both the event and the final listing. Include an uncached transfer, not only a fast cache hit.Load, request, unload, repeat. Use
POST /models/loadwith the exact listed model ID in themodelfield. Wait forGET /modelsto reportloaded, then send a representative inference request through the router using that ID. SendPOST /models/unloadwith the same ID; verifyunloadedand child exit before repeating. These documented residency controls distinguish a downloaded file from a model ready to answer.
Exercise the paths a clean download misses
Cancellation and failure. In staging, cancel an active transfer through
/models/unloadand exercise a controlled download failure. Check for a terminal outcome and eventual child exit rather than an indefinitely downloading entry. The state-handling code makes the result-to-exit handoff part of the behavior worth checking.Loaded-model limits. With your intended concurrency, check that downloading does not unexpectedly evict a serving model. The existing router test covers a download while another model remains loaded at
models_max=1; that assertion is a useful reference, not proof for your workload.Log collection. Repeat representative transitions with your production logging configuration. Check unwanted blank progress lines, color escapes and command-pipe warnings. Exercise colored text and JSONL separately: the logging options specify that
--log-jsonldisables colors.Preset reloads, if used. Check an unchanged staging preset for unintended unloads. Do not use
GET /models?reload=1as a passive production health probe: the documented behavior can unload running models when their source configuration changes or disappears.
Promote on lifecycle evidence, not cleaner logs
Promote only when expected downloads reach terminal outcomes, loaded models answer through the router, unloading releases the intended child processes and logs remain usable. Set workload-specific error and latency limits before comparing builds; the sources provide no universal acceptance timeout or measured speedup. If the candidate fails, hold the rollout and preserve the evidence instead of repeatedly upgrading or clearing caches.
Methodology: AI-assisted reporting and analysis based on the official release, merged PR, version-pinned source, documentation and test definitions, checked October 5, 2026. No binaries, model downloads, inference workloads or benchmarks were run for this article; upstream reproduction claims remain attributed.
