On October 5, 2026 at 04:32 UTC, llama.cpp published b11402 as a prerelease, bringing a CUDA FlashAttention scheduling change for eligible DGX Spark workloads. For operators processing long, uncached prompts, it is a reason to run a controlled prefill comparison—not evidence that every CUDA server will get faster.
Prefill is the work of processing input tokens before generating a reply. On eligible DGX Spark kernels, b11402 selects whole-tile scheduling while preserving the previous selector elsewhere. The released patch is narrower than the early PR overview: do not treat its RTX 5090 benchmark figures as gains delivered by this override.
Check the execution path, not just the context length
The shipped selector applies the override only when all four conditions hold:
The CUDA backend classifies the device as DGX Spark.
The selected non-sparse MMA kernel has exactly two loading stages, enabling asynchronous key/value preloading. This is an internal kernel property, not a new command-line switch.
The attention-mask scan is active.
Calculated tile-wave efficiency is at least 75%. This describes how fully tiles fill waves of CUDA blocks, not measured GPU utilization.
The mask scan can skip attention work. The patch explains that Stream-K divides work before those skips are accounted for, potentially leaving blocks with unequal workloads. Whole-tile scheduling is the alternative for the guarded case; failing a guard leaves the existing architecture-dependent policy in control. Source: released scheduling change.
A long conversation can still involve only a short new suffix. The server's prompt-cache behavior and physical microbatch limit determine how much input is processed at once. A 64K context setting or -fa on alone does not demonstrate eligibility.
Prioritize staging on an existing Spark deployment with substantial uncached prompt work. This override establishes no direct speedup for other NVIDIA devices, CPU, Metal or Vulkan. It does not expand model context limits or memory capacity; the separate DGX Spark memory-fit guide addresses those deployment constraints.
The biggest reported uplift was not the fastest setting
An October 1 comment by the PR author reports the following Qwen3.8 27B results on DGX Spark with a 64K prompt. These are intermediate-branch measurements, not a benchmark of the final October 5 release.
Physical microbatch / reported path | Baseline tokens/s | Modified tokens/s | Reported change |
|---|---|---|---|
512 / no scan; Stream-K retained | 711.9 | 715.0 | +0.4% |
4096 / mask scan; whole tiles | 511.1 | 677.0 | +32.5% |
Calculated from this comparison: (715.0 / 677.0 - 1) × 100 = 5.61%. The modified 512 setting was about 5.6% faster than modified 4096, despite the latter's larger upgrade gain. This assumes the comment's configurations are comparable; it is not a new measurement.
The comment does not bind these results to a final-release hash pair or supply run dispersion. Use it to motivate two separate questions: does the new revision improve a fixed configuration, and which configuration is fastest afterward? It does not justify setting every server to -ub 4096.
1. Keep the baseline and candidate separate
Retain the serving binary and configuration for rollback. For a narrow comparison, use b11401 at a7fb71fab83b474a0892b9a05aaa3a8ddca2729b and b11402 at d89651a7b205c03c4a0b13cd0646d400dc929f79. The latter's two-file commit has the former as its parent. Also compare with your actual production revision before promotion.
The b11402 release lists Ubuntu ARM64 CUDA assets. Match architecture and driver/runtime requirements, or build in a new directory using the tagged CUDA build instructions. The following is a proposed procedure, not a build run for this article:
git clone --branch b11402 --depth 1 https://github.com/ggml-org/llama.cpp.git llama-cpp-b11402
cd llama-cpp-b11402
git rev-parse HEAD
# Expected: d89651a7b205c03c4a0b13cd0646d400dc929f79
cmake -B build -DGGML_CUDA=ON
cmake --build build --config ReleaseCheck the hash before continuing. Prepare b11401 in its own directory with equivalent compiler, optimization and CUDA-architecture settings. Record binary version, model-file hash and quantization, K/V types, driver/toolkit, offload layout, context, batch sizes, server slots and cache policy.
b11402 is not the latest tag: b11403 was published at 05:21 UTC on October 5. The comparison here isolates the b11402 change; it does not assess subsequent releases.
2. Sweep prompt length and microbatch size
Run the same grid against each build before tuning either side. This proposed llama-bench command covers three prompt lengths and three physical microbatch limits, with generation disabled:
./build/bin/llama-bench -m /path/to/model.gguf \
-p 8192,32768,65536 -n 0 -d 0 \
-b 4096 -ub 512,1024,4096 -ngl 99 \
-fa on -ctk f16 -ctv f16 \
-t 20 -r 5 -o jsonl --progressHere -b caps the logical batch and -ub caps the physical microbatch. Holding the former at 4096 lets the sweep compare smaller and larger physical chunks.
Replace the model path. Keep only lengths supported by the model and available memory; adapt thread and offload settings to the deployment. Include your production microbatch limit if it is absent. The grid is an evaluation plan, not an optimal configuration or a promise that every case reaches the new scheduler.
Leave warmup enabled. In the b11402 implementation, prompt warmup evaluates the full test prompt and each measured repetition clears model memory. With -d 0, the prompt test does not begin with prefilled context.
Save raw repetition samples and dispersion, not just the best rate. Alternate baseline/candidate runs to reduce order effects from temperature and background load.
Compare matching configurations across versions first. Then compare microbatch choices within each version, looking at absolute throughput as well as percentage improvement.
Run generation and combined prompt/generation cases separately. llama-bench timings exclude tokenization and sampling; this prefill-only command cannot establish decode speed or client-observed first-token latency.
3. Require a service-level result before promotion
Replay representative requests through an isolated server at intended concurrency. For the documented /completion endpoint, use cache_prompt: false to test a cold prefix, then repeat with production caching. Do not assume that field works unchanged through every API wrapper.
Track
prompt_nandcache_nso a cache hit cannot masquerade as faster full-prompt processing. Recordprompt_ms, client-observed first-token and end-to-end latency, generation rate, memory use and errors.Set acceptance limits before testing. Require a repeatable improvement on the workload you care about, with no unexplained latency, memory or task-output regressions.
Add backend correctness checks such as
test-backend-ops -o FLASH_ATTN_EXTand representative application-output checks. The PR author's reported backend results do not validate your final Spark model/server configuration.If the candidate does not meet those limits, keep the production build. Do not remove the hardware or kernel guards to chase the reported uplift.
Reporting method
This AI-assisted guide uses upstream release metadata, the released CUDA patch, version-pinned documentation and attributed PR reports, checked on October 5, 2026. No builds, inference workloads or benchmarks were run for this article. The calculation is derived from published numbers; the commands are a proposed evaluation procedure.
