llama.cpp published prerelease b11414 on October 5, 2026, at 11:43 UTC, delivering a Vulkan backend fix for an input-reuse error that can affect long-prompt inference. The merged PR #29591 identifies GPT-J, Falcon and Command-R-style graphs that share normalized input between attention and the feed-forward network as potentially affected; it does not establish that every model in those families fails.
For operators running those workloads on Vulkan, the practical next step is an isolated correctness comparison with b11414. Confirm the backend actually used by your application or wrapper, keep the deployed build as a baseline, and test real long inputs before evaluating throughput. The specific regression test described in the PR was removed before merge, so installing the release does not supply that new test.
What the scratch-buffer fix changes
Vulkan matrix multiplication can retain a converted input in temporary storage named prealloc_y. Its reuse bookkeeping lets a later multiplication skip conversion when it recognizes the same input. The defect arises when another operation overwrites that storage without invalidating the remembered conversion, as shown in the pinned implementation. The problematic sequence is:
A first matrix multiplication converts input X into the shared scratch buffer.
An eligible Flash Attention operation or wide softmax operation writes its own temporary data into that buffer.
A later matrix multiplication reuses X, but stale bookkeeping makes it consume the overwritten buffer instead of converting X again.
b11414 corrects scratch-buffer reuse by invalidating the remembered conversion when attention or wide softmax overwrites that storage. The merged diff adds nine lines in one Vulkan source file, clearing the remembered tensor and conversion-pipeline identities in two blocks covering attention mask optimization, sparse-index storage and wide-softmax partial results. This is temporary backend storage, not a saved prompt-cache file.
Two operational conclusions follow from the branch conditions: disabling Flash Attention is not a complete workaround because wide softmax can still overwrite the buffer. Also, the softmax boundary above 16,384 input columns is a tensor-shape condition, not a universal safe prompt length. Attention uses different eligibility conditions. The reviewed sources do not establish the first affected release or a complete device, driver and model matrix.
What the upstream results establish
The contributor’s test report lists 19,007 of 19,011 backend checks passing before the fix and all 19,011 afterward on both R9700 and 7900 XT, eliminating four q4_0/q8_0 failures. The contributor characterizes the throughput differences as noise, not a speed gain. These are upstream reports, not independently reproduced measurements.
The author later confirmed removal of the targeted test, and the final commit changes no test file. Do not treat the reported counts as the expected stock b11414 suite output. General operation checks remain useful, but passing matrix multiplication, attention and softmax separately does not demonstrate that their problematic interleaving was exercised.
Build an isolated validation candidate
The release page lists Vulkan archives for Ubuntu and Windows, each in x64 and arm64 variants. For a source-built test runner, the following Linux/POSIX commands follow the pinned build documentation and CMake test option. Install the documented compiler and Vulkan prerequisites first. These commands are proposed instructions; they were not executed for this article.
git clone --branch b11414 --depth 1 https://github.com/ggml-org/llama.cpp.git llama-cpp-b11414-check
cd llama-cpp-b11414-check
git rev-parse HEADBefore building, confirm the checkout resolves to 806eee9841de5f2c20f9d43914117f157d2baacc, the release’s target commit. Keep the original runtime and configuration intact.
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_VULKAN=ON -DLLAMA_BUILD_TESTS=ON
cmake --build build --config Release -j
./build/bin/llama-cli --list-devicesUse the actual Vulkan device name printed for your system. Only if it is Vulkan0, a focused numeric check is:
./build/bin/test-backend-ops test -b Vulkan0 -o MUL_MAT,FLASH_ATTN_EXT,SOFT_MAXThe test runner supports CPU-reference comparisons and operation filters. Inspect the selected device, executed cases, failures and skips: a mismatched backend filter can skip devices while counting them as successful. A clean exit alone is insufficient. This command is a general check, not a recreation of the removed sequence-specific test.
Exercise real long prompts
The following is a proposed validation matrix, not a set of measured results. Reuse identical model files and inputs across the deployed build and candidate. Record the model hash, quantization, K/V types, GPU and driver, offload placement, prompt token count, batch limits and concurrency. The CLI reference documents the controls below.
Coverage case | How to run it | What it helps distinguish |
|---|---|---|
Short control | Evaluate a small known-answer prompt on both builds. | Basic setup problems versus long-input behavior; a pass does not establish long-context correctness. |
Real long input | Evaluate the actual long prompt, then a near-capacity case within supported limits and available memory. | Workload correctness under meaningful input length, not merely a large configured context. |
Attention modes | Where supported, repeat with explicit --flash-attn on and off. | Coverage of different execution paths; neither mode alone proves the other is correct. |
Request lifecycle | Start with fresh processes, then repeated requests and intended concurrency. | Whether results hold under the deployment’s reuse and request pattern. |
Increasing --ctx-size only sets capacity; it does not create a long test input. Verify how many prompt tokens were actually evaluated, rather than accidentally testing only a cached suffix. Keep --batch-size and --ubatch-size fixed in the main comparison. Treat batch-size variations as separate coverage cases.
For a strict CPU control, the build documentation specifies --device none; -ngl 0 alone can still allow GPU work. Hold sampling settings fixed where appropriate, but judge numeric results with suitable tolerances and application outputs against predefined answer criteria. Identical generated prose is not a universal cross-backend correctness requirement.
If you need to prove this exact defect, a dedicated regression must enforce the convert–overwrite–reuse sequence, with a real graph dependency preventing reordering. Compare the fix commit with its immediate parent to isolate the patch. Comparing a much older production build with b11414 answers the broader upgrade question and cannot attribute every difference to these nine lines.
Decide whether to promote the build
A reasonable acceptance gate is evidence that the intended Vulkan device and workloads ran, numeric comparisons meet declared tolerances, and application answers meet predefined correctness checks. Hold a candidate with unexplained failures for diagnosis. Measure prefill, generation, latency and memory only after correctness passes; the upstream report establishes no speedup.
If the old build does not reproduce the fault, record the outcome as candidate validation, not proof that your setup reproduced and fixed this bug. Retain commands, revisions, environment details, inputs, outputs and timestamps so later comparisons are interpretable.
If the upgrade also crosses b11411 and your application restores saved inference state, review the separate saved KV-cache migration guide. Preserve the old runtime/cache pair; that disk-state compatibility issue is distinct from this Vulkan scratch-buffer correction.
Methodology: AI-assisted reporting and implementation analysis based on the official release, merged code, review discussion and pinned documentation, checked on October 5, 2026. No build, inference run or benchmark was performed for this article.
