llama.cpp published prerelease b11412 on October 5, 2026, at 10:30 UTC. It contains a fix for unexpected graph reallocation in the qwen4exp and glm5-next model paths. Operators serving these architectures should prioritize an isolated regression run, especially when sharing prompt state between sequences or moving from prompt processing to token generation.
The repository names are architecture identifiers: the pinned converters associate qwen4exp with Qwen3.8-Flash-Next and GLM5_NEXT with GLM-5.3-Flash. This is not a new model release or evidence that every Qwen or GLM model has the same issue.
What the graph fix changes
A compute graph describes the operations the runtime schedules. According to the upstream explanation, the initial reservation could cover one graph topology, then encounter another during decoding. Sharing cached cells changed the Qwen path; switching between prompt processing and generation could change the GLM path. Re-reserving for the current state then lost the original worst-case sizing, leaving a later growth step short of space.
The fix keeps the k-pool graph topology stable as shared sequences continue decoding. Both model paths now scatter newly pooled keys before gathering them. GLM selects its gather path from fixed context settings instead of changing token counts. Stable topology does not mean tensor dimensions or memory requirements never grow. The committed patch also corrects a CUDA allocation-dependency mismatch for microbatches with no output rows.
The abort needs a qualification. In the scheduler code, GGML_SCHED_DEBUG_REALLOC=1 deliberately aborts when reallocation occurs without a change in backend assignments or scheduler graph size. It normally defaults to off, except in builds defining GGML_SCHED_NO_REALLOC. Without that diagnostic abort, this path attempts reservation and allocation again. The evidence establishes an allocation defect—not a universal production crash rate.
Compare the patch and the production upgrade separately
Use b11411 versus b11412 to isolate this patch: the b11412 commit has b11411 as its immediate parent. Also compare your real production revision with the deployment candidate; an older baseline includes other intervening changes.
Pin the source revisions: b11411 at
210791069bf5460bef3c09b982ee6f0aa9967404; b11412 atc173a53bdfca1047c710018dc934a6d67a8b010f. For a wrapper application, identify its embedded llama.cpp commit.Keep model-file hashes, quantization, backend, device, driver, offload, context capacity, cache types, batch limits and parallelism matched. Preserve the known-working binary and configuration for rollback.
The following is a proposed Bash/source-build procedure, not a test result. In each isolated checkout, enable tools and tests using the documented CMake build process. Add your backend settings—for example, -DGGML_CUDA=ON for CUDA—and use the required compiler/toolkit. Release archives may not contain the regression executables.
cmake -B build -DCMAKE_BUILD_TYPE=Release \
-DLLAMA_BUILD_TESTS=ON -DLLAMA_BUILD_TOOLS=ON
cmake --build build --config ReleaseExercise shared state, then keep decoding
Adapt the upstream diagnostic to a compatible local GGUF. Run each deployed affected model separately, in a disposable process: the diagnostic is designed to abort on failure. Use sufficient memory and supported context capacity, and add identical explicit backend/offload settings to both builds.
GGML_SCHED_DEBUG_REALLOC=1 ./build/bin/llama-batched-bench \
-m /path/to/compatible-model.gguf \
-npp 2500 -ntg 32 -npl 1,2 -c 32768 -pps -kvuHere -pps enables shared-prompt testing and -kvu selects unified KV. The batched benchmark implementation copies sequence state before continuing generation. It can also skip a workload that exceeds context capacity, so require actual results for both requested parallelism rows. A successful exit alone is insufficient.
Coverage case | What to check |
|---|---|
2,500-token prompt; one and two sequences | Confirm the two-sequence shared-state case ran and continued without the reallocation diagnostic. |
Repeat with -npp 512 | Cover a shorter prompt-to-generation transition, not just the original reproduction. |
GLM with -ub 16 and with the deployment microbatch setting | Exercise the context-selected gather path where eligible; retain the real configuration as a separate case. |
Repeat without -pps; then test representative longer contexts | Compare independent prompts and shared prompts, keeping each baseline/candidate pair otherwise matched. |
The GLM source selects gather when n_ubatch <= 16 and n_ctx > indexer_top_k + kpool - 1. The indexer threshold is a model hyperparameter, not the sampling --top-k option. Treat -ub 16 as a coverage case, not a general performance recommendation.
Calculated capacity check: with a shared prompt and unified KV, the tool budgets 2,500 + 2 × 32 = 2,564 logical KV positions for two sequences. Independent prompts need 2 × (2,500 + 32) = 5,064. These follow the tool’s capacity formula; they are not VRAM measurements. -c 32768 sets capacity—it does not make this a 32k-token prompt test.
For source-level regression coverage, run the registered test:
ctest --test-dir build -R test-recurrent-state-rollback --output-on-failureIts CTest registration enables the diagnostic and requests generated dummy models. Check the selected test count and logs for the relevant shared-sequence cases. The test implementation checks continued decoding and finite logits, with skips for some architectures; this is not an answer-quality evaluation. Windows shared-library builds lack this CTest registration at this revision.
Promotion criteria and remaining boundaries
Our acceptance recommendation is to require completed relevant diagnostic cases, then replay real application prompts at intended concurrency with both fresh and continuing shared contexts. Verify that the service actually exercises sequence sharing. Retain outputs, full logs, process status and build/model hashes; agree latency, memory and error limits before comparison. Measure normal-operation performance separately from diagnostic runs. If the baseline does not reproduce the failure, do not count that run as evidence of an abort-to-pass fix.
Vulkan needs its own checks. Follow-up PR #29988 remained open and unmerged when checked on October 5; its description targets an Nvidia DeviceLost report linked from this discussion. That proposed fix is not in b11412. The report does not establish that b11412 introduced the failure, and CUDA or Metal results cannot certify a Vulkan deployment.
Upgrades from before b11411 also cross a saved-state boundary: the b11412 header defines session-file version 11 and sequence-state-file version 4. Follow the saved KV-cache migration and restore guide: preserve the old runtime/cache pair and create candidate caches separately. This concerns persisted inference state, not the GGUF weight format.
Keep performance claims separate too. The earlier CUDA lightning-indexer guide covers a different optimization. This patch supports a targeted stability test, not a promised speedup or quantified memory saving. The first affected release and the full affected-configuration matrix remain unestablished.
Methodology: AI-assisted reporting based on upstream release metadata, the patch and source code pinned to b11412. Release status and cited links were checked on October 5, 2026. No model execution, build or benchmark was performed for this article.
