Article

llama.cpp b11413: How to Test Vulkan Sparse FlashAttention with Quantized KV

Check eligibility for llama.cpp’s Vulkan sparse-attention change, understand the 16× cache threshold, and compare performance with the follow-up fix included.

Editorial illustration for llama.cpp b11413: How to Test Vulkan Sparse FlashAttention with Quantized KV: a geometric block represents a model release. Not documentary evidence.

On October 5, 2026, llama.cpp published the b11413 prerelease, extending Vulkan sparse FlashAttention to eligible quantized key/value (KV) caches and accelerating sparse-index compaction. The main audience is local-inference operators generating tokens at long context depths with models that already use sparse attention.

The useful question is whether your cache layout and occupied depth trigger the new route—and whether it helps on your GPU. For evaluation, use a build that also contains the subsequent correctness fix, rather than treating stock b11413 as the deployment target.

Check the cache and attention path, not the model filename

For an eligible model, sparse FlashAttention gathers selected cache positions instead of processing the full attention row. The merged change also reduces the work needed to assemble those positions. It does not turn a dense-attention model into a sparse one or establish a reduction in allocated cache memory.

  • Confirm an existing sparse model graph. The PR names Qwen3.8-Flash-Next QSA and DeepSeek-style workloads as motivations; a model-family label alone does not establish eligibility.

  • Inspect K and V cache precision separately from weight precision. A Q4 GGUF filename says nothing about the active cache settings; the server documentation defaults both caches to f16.

  • Check the effective types and shader route. The tagged implementation admits quantized K/V on scalar and coopmat1 paths, while coopmat2 sparse gathering still requires effective f16 K and V.

For quantized scalar/coopmat1 attention, the depth condition is KV >= max(4096, 16 * n_kv_max). Here, n_kv_max bounds the positions retained by the sparse mask. The code also requires compatible mask, shape and attention settings; the ratio alone is insufficient.

Calculation, not a benchmark: with 2,048 retained positions, the minimum is max(4,096, 16 × 2,048) = 32,768 KV cells.

Attention KV cells

Ratio to 2,048 retained positions

Quantized depth gate

16,384

8×

Not met

24,576

12×

Not met

32,768

16×

Met; other gates still apply

65,536

32×

Met; benefit still needs measurement

Use the actual KV extent presented to the attention operation, not merely a configured context ceiling. Pooling and cache layout can change the relationship to input-token count. This boundary establishes eligibility, not a guaranteed speedup.

Read the revised benchmark claim narrowly

In an October 2 follow-up, contributor fxgsell reports Flash-Next q8_0-KV decode throughput gains of 19% at 64k and 15% at 128k on an R9700 + 7900 XT setup. These replace the PR’s earlier percentages; they are not RohitAI measurements. The update lacks a full configuration and new absolute-throughput table.

The final patch can also affect f16 workloads: another author comment reports faster f16 sparse operations after compaction, not proven f16 whole-model gains. An RTX 5090 reviewer reports mixed operation-level results. AMD-only threshold tuning and those limited NVIDIA samples do not establish general Vulkan performance or faster prompt processing.

Calculation: for equal token output, 19% more throughput means 1 − 1/1.19 ≈ 16% less decode time, not 19% less total request time. Unchanged prefill and queueing dilute that saving.

Pin a candidate containing both changes

The b11414 prerelease followed on October 5 at 11:43 UTC. Its commit, 806eee9841de5f2c20f9d43914117f157d2baacc, is one commit ahead of b11413 and adds the companion Vulkan scratch-buffer fix. It provides a reproducible evaluation target; check the embedded revision in wrappers or packaged builds.

The PR’s development benchmark baseline already carried that fix, so its percentages are not a stock-b11412-to-b11413 comparison. See the separate b11414 correctness guide for the buffer-reuse issue.

The following is a proposed procedure, not an executed test. For a Linux source build, install the documented Vulkan and compiler prerequisites. In an isolated checkout containing the candidate revision, select it and build with the tool and test targets enabled:

git switch --detach 806eee9841de5f2c20f9d43914117f157d2baacc
git rev-parse HEAD
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_VULKAN=ON -DLLAMA_BUILD_TESTS=ON -DLLAMA_BUILD_TOOLS=ON
cmake --build build --config Release
./build/bin/llama-bench --list-devices

Record the revision, build settings, model-file hashes, GPU and driver. Keep the current production binary, configuration and compatible saved state available for rollback. If the upgrade crosses the saved-cache format boundary, follow the KV-cache migration guide separately.

Compare sparse-permitted and sparse-disabled runs

The tagged llama-bench documentation supports existing context depth with -d, generation length with -n, repetitions with -r and JSON output. This template assumes a supported sparse model and a listed device named Vulkan0. Replace the model path and device; hold deployment-specific offload, split, batch and thread settings identical in both commands.

env -u GGML_VK_FA_SPARSE_DISABLE ./build/bin/llama-bench \
  -m /path/to/sparse-model.gguf -dev Vulkan0 -fa on \
  -ctk q8_0 -ctv q8_0 -p 0 -n 128 \
  -d 0,16384,32768,65536 -r 5 -o json > sparse-permitted.json

env GGML_VK_FA_SPARSE_DISABLE=1 ./build/bin/llama-bench \
  -m /path/to/sparse-model.gguf -dev Vulkan0 -fa on \
  -ctk q8_0 -ctv q8_0 -p 0 -n 128 \
  -d 0,16384,32768,65536 -r 5 -o json > sparse-disabled.json

Run each command in a fresh process. The disable switch is presence-based: setting it to 0 still disables sparse attention. Removing it permits eligible dispatch but does not prove the sparse path executed. Inspect executed cases or profile the route if that distinction matters; the switch controls sparse attention generally, not only quantized K/V.

Reduce depths to fit the model’s supported context and available memory, leaving room for the 128 generated tokens. Add a near-128k case only if supported. This matrix tests below and around the example boundary; it does not recreate the contributor’s undisclosed full setup.

Keep warm-up consistent, alternate run order and preserve every sample, error and skip. Use distinct output filenames for additional rounds. Keep power conditions and background load comparable; five repetitions are a starting point, not a statistical guarantee. A same-build toggle comparison holds the correctness fix constant. Separately compare your production build with the candidate to answer the broader upgrade question.

Accept the change on correctness and service results

  • Check numerical correctness for the attention cases your backend actually executes, then run representative application-quality tasks. The merged tests add quantized sparse cases, but coverage is not proof for every GPU, cache format or model. Record executed and skipped cases, not just an exit code.

  • Measure prompt ingestion, deep-context decode and full-request latency separately. llama-bench excludes tokenization and sampling, so follow it with real server requests at the intended concurrency. Capture first-token latency, completion latency, memory use and application failures.

  • Hold prompts, sampling and caching policy fixed. The native /completion endpoint defaults cache_prompt to true; for deliberate cold-input comparisons, set it false and check normal cache reuse separately.

  • Set acceptable quality and latency criteria before inspecting results. Keep the candidate only if it meets those criteria on the workload you serve. Do not lower the sparse threshold merely to force activation, or interpret different generated prose alone as a numerical failure.

Methodology: AI-assisted analysis of the official release, merged source, documentation and contributor reports, rechecked on October 5, 2026. No local build, inference run or benchmark was performed; commands are proposed validation steps.