Article

llama.cpp b11401: Router-Mode Fix and Upgrade Checks

llama.cpp b11401 separates router child commands from logs. See who should upgrade and how to check downloads, model readiness, shutdown and logging.

Editorial illustration for llama.cpp b11401: Router-Mode Fix and Upgrade Checks: an arrow connects an old service to its replacement. Not documentary evidence.

llama.cpp published prerelease b11401 on October 5, 2026, at 00:36 UTC. It includes PR #29895, which separates lifecycle messages from log output in llama-server router mode. The fix matters to operators whose router manages child processes for model downloads and serving: an unfinished log write could otherwise hide a completion message and leave a model marked as downloading.

Prioritize a staging upgrade if your router has matching stuck-download symptoms. For an apparently healthy deployment, add this change to normal lifecycle regression checks. The evidence does not establish a throughput gain or make this an emergency upgrade for every single-model server.

What changes between the child and the router

Router mode manages multiple models behind one server endpoint; the server documentation describes starting it without a model argument. In the old communication path, the router distinguished state messages from logs by their line prefix. The PR author reports that injecting output without a terminating newline before a download-completion message caused that message to be logged instead of handled. The model remained in the downloading state. This is an upstream reproduction report, not a RohitAI test.

In b11401, router child commands travel separately from ordinary logs. The child reserves its original stdout stream for commands and redirects ordinary stdout writes to stderr. The router reads the two pipes separately, handling state messages on the command pipe and forwarding the other stream as logs. This changes internal process communication, not the public REST API.

The same patch moves color resets before trailing newlines and adds Windows console handling for ANSI colors. Router children inherit the parent’s effective color setting. Cleaner progress output is useful, but it is not evidence that model state is correct.

Why a completed transfer can still look stuck: the download-child implementation sends its result and then waits for the router’s exit instruction. Losing that message can interrupt coordination after the transfer has finished. That is a reason to inspect state and process exits alongside transfer progress—not to assume every stalled download has this cause.

Which deployments should prioritize it?

The following triage is an operational recommendation based on the patch’s scope, not a measured incident-rate ranking.

Deployment or symptom

Recommended priority

What to verify

Router keeps a finished transfer marked downloading

Prioritize an isolated candidate

Completion event, model registration and download-child exit agree.

Router works, but progress logs contain unwanted blank lines or color escapes

Schedule a lifecycle and logging regression pass

State transitions work with the actual log collector.

Single-model server without router children

No urgency established by this router bug alone

Review shared-logger behavior and other changes in your upgrade.

The first affected build and a complete affected-version range are not established. Disabling colors is not a demonstrated complete workaround: the reported trigger includes arbitrary unterminated output. Network, disk and model-access failures remain separate diagnoses.

Pin the candidate and isolate the comparison

The release metadata identifies commit a7fb71fab83b474a0892b9a05aaa3a8ddca2729b and lists binaries for Windows, Linux and macOS, among other targets. Choose an asset matching your architecture and backend, or use your established source-build procedure. Record the executable revision and any container digest; do not infer inclusion from a downstream wrapper’s version number.

Keep the current binary and configuration available for rollback. Use a separate staging port and disposable model cache, while matching the production backend, model quantization, presets, autoload policy and loaded-model limit.

For fault isolation, b11400 to b11401 is a one-commit change. Identically configured builds offer a narrow comparison; your actual production revision is still the rollout baseline. If you are also adopting mixed-input integration changes, retain the separate checks in our b11400 text-and-embedding batch guide.

Check download, readiness and shutdown separately

This is a proposed acceptance sequence, not a report of tests performed. Capture timestamped HTTP responses, model listings, events, logs and child-process exits.

  1. Subscribe before starting. Open GET /models/sse and wait until the streaming connection is established. The tagged router test explicitly waits for its subscriber before triggering a download: a terminal event can otherwise arrive too early and be missed. A missing event alone is therefore not proof of the pipe bug.

  2. Start a download into the isolated cache. Send POST /models with a valid public repository and quantization tag, replacing the placeholder below. The documented response acknowledges background work; it does not certify completion.

    {"model":"YOUR_PUBLIC_HF_REPO:QUANT_TAG"}
  3. Check the terminal result. Look for the matching download_finished event, then use GET /models to confirm the model is registered and no longer downloading. Preserve any download_failed result and its logs. The upstream download test checks both the event and the final listing. Include an uncached transfer, not only a fast cache hit.

  4. Load, request, unload, repeat. Use POST /models/load with the exact listed model ID in the model field. Wait for GET /models to report loaded, then send a representative inference request through the router using that ID. Send POST /models/unload with the same ID; verify unloaded and child exit before repeating. These documented residency controls distinguish a downloaded file from a model ready to answer.

Exercise the paths a clean download misses

  • Cancellation and failure. In staging, cancel an active transfer through /models/unload and exercise a controlled download failure. Check for a terminal outcome and eventual child exit rather than an indefinitely downloading entry. The state-handling code makes the result-to-exit handoff part of the behavior worth checking.

  • Loaded-model limits. With your intended concurrency, check that downloading does not unexpectedly evict a serving model. The existing router test covers a download while another model remains loaded at models_max=1; that assertion is a useful reference, not proof for your workload.

  • Log collection. Repeat representative transitions with your production logging configuration. Check unwanted blank progress lines, color escapes and command-pipe warnings. Exercise colored text and JSONL separately: the logging options specify that --log-jsonl disables colors.

  • Preset reloads, if used. Check an unchanged staging preset for unintended unloads. Do not use GET /models?reload=1 as a passive production health probe: the documented behavior can unload running models when their source configuration changes or disappears.

Promote on lifecycle evidence, not cleaner logs

Promote only when expected downloads reach terminal outcomes, loaded models answer through the router, unloading releases the intended child processes and logs remain usable. Set workload-specific error and latency limits before comparing builds; the sources provide no universal acceptance timeout or measured speedup. If the candidate fails, hold the rollout and preserve the evidence instead of repeatedly upgrading or clearing caches.

Methodology: AI-assisted reporting and analysis based on the official release, merged PR, version-pinned source, documentation and test definitions, checked October 5, 2026. No binaries, model downloads, inference workloads or benchmarks were run for this article; upstream reproduction claims remain attributed.