llama.cpp released b11404 on October 5, 2026, adding Metal kernels relevant to Apple Silicon users running speculative decoding or several generation sequences together. GitHub marks it as a prerelease, published at 06:06 UTC. It is worth a controlled upgrade test for those workloads; single-sequence generation without speculation is not the main target.
The new Metal kernels target few-row matrix multiplication during speculative verification and small-batch generation. The released implementation shares weight computation across several activation rows using small matrix tiles. This reduces repeated work inside that operation, not every source of inference latency.
Check whether your workload reaches the new kernels
Speculative decoding proposes several tokens and asks the target model to verify them together. That creates a small matrix operation instead of only the single-token work of serial generation. The runtime documentation explains why both verification cost and draft acceptance matter: faster verification alone does not establish that the complete request will finish sooner.
The final Metal dispatch selects the new path automatically when the device supports simdgroup matrix operations and the tensors qualify. It does not exclude devices with the tensor API: eligible M5 devices are included. There is no new MMA switch to enable, and no reason to disable the tensor API for this comparison.
The b11404 eligibility checks set these activation-row ranges for individual matrix operations:
Weight tensor type | Eligible row count |
|---|---|
Q4_0, Q4_1, Q8_0, Q5_K, Q6_K | 2–16 |
F16, Q4_K, Q5_0, Q5_1 | 3–16 |
F32 | 6–16 |
Other layout checks still apply: F32 activations, non-transposed inputs, aligned activation strides and compatible weight dimensions. BF16, Q2_K, Q3_K and IQ types are outside this new path’s whitelist, not necessarily unsupported by llama.cpp. These are tensor-level rules, not context-window sizes or a promise attached to a GGUF filename. Practical implication: inspect the tensors in mixed-quantization files rather than assigning one threshold to the whole model.
Also distinguish the final patch from its development notes. The merged commit history records removal of batched-copy fusion and a special Q4_0 two-row variant. Neither should be advertised as a shipped b11404 feature.
What the reported speedups establish
The PR author’s development benchmarks used an M3 Ultra with a 60-core GPU, macOS 15.7.9, Qwen3.8-27B Q4_0 and a DFlash2 Q8_0 drafter. Reported decode throughput, in tokens per second:
Workload | Baseline 836d571 | PR build |
|---|---|---|
Serial, all listed prompts/temperatures | 32.1 | 32.0 |
DFlash2, code, temperature 0 | 30.2 | 110.0 |
DFlash2, code, temperature 1 | 24.3 | 80.9 |
DFlash2, prose, temperature 0 | 16.8 | 62.6 |
DFlash2, prose, temperature 1 | 13.9 | 48.8 |
Protocol: 64 generated tokens; median of five requests after one warm-up; mean of two server runs. Calculated from that table, 110 ÷ 30.2 = 3.64× compares speculative builds, while 110 ÷ 32.0 = 3.44× compares speculation with serial on the PR build. Prose at temperature 1 gives 48.8 ÷ 32.0 = 1.53×.
Those are different comparisons, not a universal speedup. The evidence reviewed does not establish a complete benchmark rerun against the final release hash or long-context performance. Use the numbers to choose an experiment, not to promise production throughput.
Prepare an isolated candidate
Keep the deployed binary for rollback. Use the macOS ARM64 archive on the release page or a separate checkout pinned to a3a1c4747fdc0dcad40b3946108b89375d9a7d0e. Do not assume a package manager or application wrapper already contains this commit.
For a source build, the following commands adapt the version-pinned build instructions. Run them from that checkout with CMake and Apple’s command-line build tools installed. This is source-derived guidance, not an executed recipe:
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON
cmake --build build --config Release -j 8For a DFlash2 comparison, use a compatible target/drafter pair. The Qwen3.8-27B DFlash2 model card describes a target-specific drafter, not a standalone chat model. Its historical PR checkout instructions are not the build pin for this guide. With verified local Q4_0 target and Q8_0 draft files, an illustrative candidate launch is:
./build/bin/llama-server \
-m /path/to/Qwen3.8-27B-Q4_0.gguf \
-md /path/to/Qwen3.8-27B-DFlash2-Q8_0.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 \
-ngl 99 -ngld 99 -fa on -c 8192 -np 1 --jinjaThe paths are placeholders; the example is not a RAM-fit guarantee. Account for target weights, draft weights, caches and runtime allocations. For the serial control, remove -md and its path, --spec-draft-n-max 7 and -ngld 99; replace --spec-type draft-dflash with --spec-type none. Keep the target, offload, context and request sampler unchanged. Draft options and mode definitions.
Run four comparisons before deciding
Proposed evaluation: cross two runtime builds with two decoding modes. This separates the upgrade decision from the decision to enable speculation.
Old runtime, serial: the non-speculative baseline.
Old runtime, speculative: the existing speculation result.
b11404, serial: the candidate’s non-speculative control.
b11404, speculative: the proposed configuration.
Record the chip, GPU-core count, macOS version, build options, model-file hashes, quantizations, cache settings and sampler. Comparing against your deployed runtime answers the rollout question. To isolate this patch alone, its immediate parent is 9d3aba6b5e10a6dd086bd0555bd12c00366a8916; older baselines include other changes.
Start with short-context smoke checks, then use your actual prompt mix, context lengths and output lengths. Warm up every configuration and interleave repeated runs. Keep failures in the results. Do not change model quantization while trying to attribute a difference to the runtime.
The repository’s SPEED-Bench client can report decode throughput, request latency and draft acceptance, and save per-request JSON. Keep --bench, --category, --osl and --limit fixed between runs; match --concurrency to the server’s slot count and save each run with --output. Add memory-pressure and application-quality checks.
For parallel serving, sweep the active sequence counts you expect to deploy and record both total throughput and per-request latency. llama-batched-bench computes aggregate generation throughput across sequences; that number is not each user’s token rate. The kernel’s 16-row ceiling is not a recommendation to maximize slots or draft length.
Check complete outputs, not just opening characters. Investigate greedy-output disagreements and score representative application tasks. The speculative-decoding documentation notes that stochastic sampling can differ despite a fixed seed, and synthetic acceptance is not valid model output. Acceptance correctness is also distinct from kernel speed; the related b11387 n-gram upgrade guide covers a separate draft-rejection fix.
Promote the candidate only when repeated results meet your workload’s latency, throughput and correctness requirements without unacceptable memory pressure. If new speculation beats old speculation but loses to new serial, upgrade and speculation remain separate choices. Retain the configuration that passes those checks; these sources do not justify a universal minimum gain.
Methodology: AI-assisted reporting and analysis of the official release, pinned source code, documentation and attributed PR benchmarks, checked October 5, 2026. No model execution, hands-on benchmark or independently reproduced speedup is claimed here. The evaluation procedure is proposed, not performed.
