On October 5, 2026, llama.cpp published b11405, a prerelease adding a tiled CUDA scoring kernel for eligible four-head lightning-indexer workloads. For operators serving Qwen4Exp-style models, the useful question is whether it improves prompt processing at the context depths their service actually uses—not whether every CUDA request becomes faster.
Prioritize a controlled comparison if that indexer runs on your GPU and long-prompt processing matters to your service. CUDA release archives are available. Treat b11405 as the introduction point for this change, not the newest deployment recommendation: b11406 was already published later on October 5.
Check the indexer path, not just the model name
The merged CUDA source selects the new path automatically when the indexer has a head size of 128, four heads and an operation batch of at least eight tokens. Smaller batches retain the vector kernel. These are indexer dimensions, not the model’s total attention-head count or your server’s number of slots.
The four-head lightning indexer reuses a shared key tile across eight tokens. The tile holds 64 keys; the released implementation stages keys in half precision but performs products and accumulation in float. The release notes record that arithmetic revision, superseding the initial PR description of half-precision products.
Confirm that
LIGHTNING_INDEXERexecutes on CUDA, with compatible tensor strides and indexer-key types. A GGUF weight-quantization label alone does not establish those properties.Check the actual operation batch. A high configured microbatch ceiling does not guarantee eight new tokens reach the indexer when a request mostly reuses a cached prefix.
Do not infer support for every Qwen model.
qwen4expis the repository architecture label used by the PR; this release is not a new Qwen model announcement.
The Qwen4Exp graph scores pooled keys and then runs top-k selection separately. Consequently, a kernel test with kv=65536 is not automatically a 65,536-token application context. This also differs from b11402’s CUDA FlashAttention scheduling change: its DGX Spark results and tuning choices should not be carried over to this indexer test.
Read the upstream gains at the right scale
The PR author’s initial RTX PRO 6000 report gives three distinct observations. They are development-branch results, not RohitAI measurements or a verified benchmark of the final b11405 binary.
Reported comparison | Before → PR build | What it establishes |
|---|---|---|
Indexer kernel; kv=65536, batch=2048 | 16.4 → 6.4 ms | A faster scoring operation, not a whole-server multiplier. |
Prompt processing at 128k, author’s configuration | 3,867 → 4,038 tokens/s | A smaller improvement in full-model prompt processing. |
Token generation in that configuration | Reported unchanged | No demonstrated generation-speed benefit in this report. |
Calculated from those throughput figures, (4038 / 3867 − 1) × 100 = 4.42%. For the same prompt-token count, that corresponds to about 4.23% less processing time. Neither figure includes a demonstrated improvement in HTTP request latency; the much larger kernel ratio measures only one component.
An October 4 author update lowered the reported kernel time to 4.76 ms. Another kernel revision followed. The available reports do not tie every number to the final release hash, so the development gains should not be compounded or relabeled as guaranteed b11405 performance.
Context depth is worth isolating. An RTX 3060 contributor’s follow-up reported pp2048 on a CPU-expert-offloaded setup rising from 122.71 ± 1.18 to 128.51 ± 1.09 tokens/s at depth 65,536, after a depth-zero comparison showed little difference. That is evidence for testing your deep-context workload, not a universal prediction for older GPUs.
Run a controlled acceptance test
The following is a proposed procedure derived from upstream code and documentation. No build, model run or benchmark was executed for this article.
1. Pin the comparison and keep the toolchain fixed
Use two isolated checkouts. The release commit and its immediate parent give a patch-isolating comparison:
Candidate b11405:
1b43d311698269073fb4827db10f9d6f195d2e7d.Baseline b11404:
a3a1c4747fdc0dcad40b3946108b89375d9a7d0e.
Also compare your chosen deployment candidate with the actual production revision before rollout. Record embedded commits for wrappers or backports; a wrapper’s version alone is insufficient. Keep the current binary and configuration available for rollback.
With CMake, a compatible compiler and CUDA toolkit installed, this source-derived build recipe follows the pinned build documentation. Run it separately in each checkout, with matching compiler, toolkit, GPU architecture flags and build type:
git rev-parse HEAD
cmake -B build -DGGML_CUDA=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build --config Release2. Separate prompt length from cached-context depth
Hold model-file hashes, weight quantization, cache types, GPU placement, CPU-expert offload, thread counts, batch settings and load mode constant. Use the same GPU and driver; warm up both builds, interleave repeated runs and retain all samples and failures.
This llama-bench template requests separate prompt-processing and generation tests at two depths. Replace the model path and add identical deployment-specific offload, batch, cache and FlashAttention flags to both builds. Reduce the depth if the model’s limits or available memory require it:
./build/bin/llama-bench -m /path/to/model.gguf -p 2048 -n 128 -d 0,65536 -r 5 -o jsonHere -d fills the cache before the timed test. Processing 2,048 new tokens after 65,536 cached tokens is different from processing a fresh 65,536-token prompt. Add a short-prompt case and a supported large-prompt case with -d 0 to represent both. Five repetitions are a starting point, not evidence that any small difference is statistically reliable.
3. If profiling the kernel, fix the baseline harness first
The patch’s test-matrix change adds four-head performance cases; the older matrix lists only 32 and 64 heads. A four-head filter on the unmodified baseline can therefore match no tests.
For a kernel-only comparison, apply just that test-matrix addition to the baseline harness, record the modification and leave its production kernel untouched. After rebuilding, this optional filter selects an existing case; change the device identifier to your intended GPU:
./build/bin/test-backend-ops perf -b CUDA0 -o LIGHTNING_INDEXER -p "nh=4,kv=65536,nb=2048,ns=1,nm=1,type_K=f16"Require matching case output from both builds. A successful exit with no matching measurement is not a baseline. Keep indexer-key count, application context depth and new prompt length separately labeled. Run correctness checks too: a performance result does not validate model outputs.
4. Accept the change on service results and output quality
The benchmark documentation excludes tokenization and sampling from its timings. Finish with an isolated server at your intended concurrency, measuring client-observed first-token and completion latency as well as generation throughput, memory use and errors.
For deliberate uncached tests on the native /completion endpoint, set cache_prompt:false; its documented default is true. Test normal production caching separately. Record prompt_n, cache_n and prompt_ms so reused tokens are not mistaken for freshly processed prompt work. Check wrapper-specific semantics before copying that request field.
The author’s WikiText-2 perplexity comparison reported differences within the stated uncertainty at 32k and 64k. That supports a narrow upstream quality check, not identical outputs or coverage of your tasks.
Set workload-specific latency, throughput, memory and quality limits before comparing results; upstream numbers do not establish a universal pass percentage.
Inspect representative task outputs and failures, not just timing averages. Promote only when repeated service measurements and correctness checks meet those limits.
Treat unchanged ordinary token-by-token generation as consistent with the reported result. If prefill gains are small or noisy in your request mix, do not promise a faster service from the kernel headline alone.
Methodology: AI-assisted reporting and analysis based on the persisted research dossier, official release metadata, pinned source, documentation and upstream discussion rechecked on October 5, 2026. Percentages labeled as calculations use reported inputs; all procedures above are unexecuted proposals, not hands-on test results.
