llama.cpp published b11387 on October 4, 2026, at 10:19 UTC as a prerelease. It fixes a speculative-decoding regression that caused shortened n-gram drafts to be rejected at nonzero temperature, slowing generation. Operators using n-gram speculation, including configurations that also use multi-token prediction (MTP), should check their deployed revision and stage an upgrade.
Prioritize the upgrade when an affected build uses an n-gram drafter and requests can have their drafts shortened by an output or context limit. A model name alone does not tell you whether the service is exposed.
Check the build, drafter and request limits
The regression originated in commit 1fb7ef3e3; b11387 contains fix commit 8330e9696. Check the bundled llama.cpp revision and any backports, not just a wrapper’s version. The first affected numbered release and downstream package adoption have not been established here; not every older build is affected.
Inspect the whole speculative configuration. In the b11387 argument parser the
--spec-defaultpreset enablesngram-mod, even without an explicit n-gram type in the launch command.Check request-level temperature. The server gates this path on temperature above zero. Selecting
--spec-draft-sampling greedyis not equivalent to setting the target temperature to zero and does not remove n-gram exposure.Confirm draft truncation. The server bounds draft length by remaining context space and output budget. A tight cap on a short response can therefore matter, not just a long answer. That is a code-based inference to cover in testing, not a measured workload result.
Services without speculation or with target temperature zero do not enter this failure path. Standalone draft-simple and draft-mtp handle candidate data differently and are excluded from this defect; adding an n-gram component removes that exemption. See the tagged speculative implementation and server verification dispatch.
Do not extend that exemption to every other drafter: the PR also mentions EAGLE3, DFlash and DSpark but provides no repair benchmarks for them.
Why a shortened draft lost acceptance
N-gram methods can propose tokens from repeated patterns without a separate draft model. The target still verifies those proposals. This makes repeated text and code useful candidates for n-gram speculation, though a match does not guarantee acceptance. The speculative-decoding documentation describes the implementations.
The faulty truncation code resized an empty probability-candidate container into a nonempty container of empty entries. That changed which verifier the server selected: it tried probability-based rejection sampling despite lacking usable draft probabilities. The sampler requires a positive draft probability for acceptance, so the malformed data caused the shortened draft to fail.
The fix keeps absent probability data absent when truncated n-gram drafts reach verification. It resizes candidate data only when candidates exist, preserving the sample-and-match fallback. Drafters that supply distributions still have them shortened alongside their tokens. There is no new switch to activate the repair.
Reported performance: recovery, not a general speed promise
The PR author reports RTX 5090 results at temperature 0.7 across eight seeds, using Q4_K_M models. These are contributor measurements, not RohitAI tests. Selected median decode rates:
Model and speculation | Affected build, tokens/s | Fixed build, tokens/s |
|---|---|---|
Llama-3.1-8B-Instruct · ngram-mod | 375.1 | 500.7 |
Qwen3.6-35B-A3B · ngram-mod + MTP, one slot | 301.7 | 333.9 |
Calculated from those medians, the increases are 33.5% and 10.7%, respectively: (fixed ÷ affected − 1) × 100. Llama’s pre-regression result was 509.6 tokens/s, so the fix remained about 1.7% below that baseline. This comparison describes recovered performance, not a new acceleration over the old baseline.
The post does not provide complete prompts, raw logs or uncertainty intervals. Its decode rates cannot establish whole-request latency, production capacity or cost savings. They justify a workload-specific comparison, not a promised percentage.
Stage a comparison that reaches the failing path
The following is a proposed validation procedure, not an executed experiment. Retain the current binary and configuration for rollback, then use a pinned b11387 build or a later revision verified to contain the fix. The release has uploaded platform-specific binaries; choose the asset matching your operating system and backend.
Record the baseline. Use
llama-server --versionto capture build identity. Keep the model file and hash, quantization, hardware, backend/driver, context size, batching, slot count and request settings fixed across the comparison. The server reference documents build information and slot configuration.Exercise the boundary. Include representative repetitive and nonrepetitive prompts, a tight output budget, a near-context-limit case and a case without truncation. Confirm the shortening branch is reached with diagnostic logging; match logging levels for timed runs. Compare the production nonzero temperature with a zero-temperature control and record multiple seeds. These cases follow from the server’s draft-budget calculation.
Control request history.
ngram-modshares a hash pool across server slots, so earlier requests can affect later drafting opportunities. Use a consistent warm-up and request order, and distinguish cold-process from warmed runs. Keep--spec-synth-ratesand--spec-synth-lenoff: synthetic acceptance bypasses normal verification and cannot establish valid output recovery.Measure acceptance and service behavior separately. Record accepted/generated draft tokens by implementation, output token counts and decode tokens/s. Also measure first-token and complete-request p50/p95 latency, errors, memory and task success. Start with one slot, then repeat at production concurrency. The documented completion timing fields are not a substitute for client-side latency measurements.
For a single-slot n-gram staging setup, the syntax below uses the preset’s lookup and draft-length values. It is illustrative, not a recommended optimum or a tested command; replace the model path and preserve your backend settings. An existing mixed MTP deployment should keep its original supported MTP configuration for the version comparison. Argument definitions.
llama-server -m /path/to/model.gguf \
--host 127.0.0.1 --port 8080 --parallel 1 --temp 0.7 \
--spec-type ngram-mod \
--spec-ngram-mod-n-match 24 \
--spec-ngram-mod-n-min 48 --spec-ngram-mod-n-max 64Check older scripts against that parser: --draft-max, --draft-min and --spec-ngram-size-n are rejected as removed arguments in this tag. Those removals are a compatibility check, not all new changes introduced by b11387.
Promote only after your predefined correctness, latency and throughput criteria pass. If an upgrade must wait, separately evaluate disabling the n-gram component; changing temperature to zero also changes application behavior and is not an equivalent workaround.
For Intel deployments, the related OpenVINO 2026.4.1 upgrade guide covers backend and device considerations. Its backend-specific results and this PR’s RTX figures should not be transferred between platforms.
Methodology: AI-assisted reporting and analysis based on official release metadata, the merged change and tagged documentation/source code, rechecked October 4, 2026. No model inference or benchmark commands were run for this article; performance figures are attributed to the PR author.
