Article

llama.cpp b11387: N-Gram Speculative Decoding Upgrade Guide

Check whether llama.cpp’s n-gram regression affects your deployment, what b11387 fixes, and how to validate acceptance and latency before upgrading.

Editorial illustration for llama.cpp b11387: N-Gram Speculative Decoding Upgrade Guide: code brackets represent developer tools. Not documentary evidence.

llama.cpp published b11387 on October 4, 2026, at 10:19 UTC as a prerelease. It fixes a speculative-decoding regression that caused shortened n-gram drafts to be rejected at nonzero temperature, slowing generation. Operators using n-gram speculation, including configurations that also use multi-token prediction (MTP), should check their deployed revision and stage an upgrade.

Prioritize the upgrade when an affected build uses an n-gram drafter and requests can have their drafts shortened by an output or context limit. A model name alone does not tell you whether the service is exposed.

Check the build, drafter and request limits

The regression originated in commit 1fb7ef3e3; b11387 contains fix commit 8330e9696. Check the bundled llama.cpp revision and any backports, not just a wrapper’s version. The first affected numbered release and downstream package adoption have not been established here; not every older build is affected.

Services without speculation or with target temperature zero do not enter this failure path. Standalone draft-simple and draft-mtp handle candidate data differently and are excluded from this defect; adding an n-gram component removes that exemption. See the tagged speculative implementation and server verification dispatch.

Do not extend that exemption to every other drafter: the PR also mentions EAGLE3, DFlash and DSpark but provides no repair benchmarks for them.

Why a shortened draft lost acceptance

N-gram methods can propose tokens from repeated patterns without a separate draft model. The target still verifies those proposals. This makes repeated text and code useful candidates for n-gram speculation, though a match does not guarantee acceptance. The speculative-decoding documentation describes the implementations.

The faulty truncation code resized an empty probability-candidate container into a nonempty container of empty entries. That changed which verifier the server selected: it tried probability-based rejection sampling despite lacking usable draft probabilities. The sampler requires a positive draft probability for acceptance, so the malformed data caused the shortened draft to fail.

The fix keeps absent probability data absent when truncated n-gram drafts reach verification. It resizes candidate data only when candidates exist, preserving the sample-and-match fallback. Drafters that supply distributions still have them shortened alongside their tokens. There is no new switch to activate the repair.

Reported performance: recovery, not a general speed promise

The PR author reports RTX 5090 results at temperature 0.7 across eight seeds, using Q4_K_M models. These are contributor measurements, not RohitAI tests. Selected median decode rates:

Model and speculation

Affected build, tokens/s

Fixed build, tokens/s

Llama-3.1-8B-Instruct · ngram-mod

375.1

500.7

Qwen3.6-35B-A3B · ngram-mod + MTP, one slot

301.7

333.9

Calculated from those medians, the increases are 33.5% and 10.7%, respectively: (fixed ÷ affected − 1) × 100. Llama’s pre-regression result was 509.6 tokens/s, so the fix remained about 1.7% below that baseline. This comparison describes recovered performance, not a new acceleration over the old baseline.

The post does not provide complete prompts, raw logs or uncertainty intervals. Its decode rates cannot establish whole-request latency, production capacity or cost savings. They justify a workload-specific comparison, not a promised percentage.

Stage a comparison that reaches the failing path

The following is a proposed validation procedure, not an executed experiment. Retain the current binary and configuration for rollback, then use a pinned b11387 build or a later revision verified to contain the fix. The release has uploaded platform-specific binaries; choose the asset matching your operating system and backend.

  1. Record the baseline. Use llama-server --version to capture build identity. Keep the model file and hash, quantization, hardware, backend/driver, context size, batching, slot count and request settings fixed across the comparison. The server reference documents build information and slot configuration.

  2. Exercise the boundary. Include representative repetitive and nonrepetitive prompts, a tight output budget, a near-context-limit case and a case without truncation. Confirm the shortening branch is reached with diagnostic logging; match logging levels for timed runs. Compare the production nonzero temperature with a zero-temperature control and record multiple seeds. These cases follow from the server’s draft-budget calculation.

  3. Control request history. ngram-mod shares a hash pool across server slots, so earlier requests can affect later drafting opportunities. Use a consistent warm-up and request order, and distinguish cold-process from warmed runs. Keep --spec-synth-rates and --spec-synth-len off: synthetic acceptance bypasses normal verification and cannot establish valid output recovery.

  4. Measure acceptance and service behavior separately. Record accepted/generated draft tokens by implementation, output token counts and decode tokens/s. Also measure first-token and complete-request p50/p95 latency, errors, memory and task success. Start with one slot, then repeat at production concurrency. The documented completion timing fields are not a substitute for client-side latency measurements.

For a single-slot n-gram staging setup, the syntax below uses the preset’s lookup and draft-length values. It is illustrative, not a recommended optimum or a tested command; replace the model path and preserve your backend settings. An existing mixed MTP deployment should keep its original supported MTP configuration for the version comparison. Argument definitions.

llama-server -m /path/to/model.gguf \
  --host 127.0.0.1 --port 8080 --parallel 1 --temp 0.7 \
  --spec-type ngram-mod \
  --spec-ngram-mod-n-match 24 \
  --spec-ngram-mod-n-min 48 --spec-ngram-mod-n-max 64

Check older scripts against that parser: --draft-max, --draft-min and --spec-ngram-size-n are rejected as removed arguments in this tag. Those removals are a compatibility check, not all new changes introduced by b11387.

Promote only after your predefined correctness, latency and throughput criteria pass. If an upgrade must wait, separately evaluate disabling the n-gram component; changing temperature to zero also changes application behavior and is not an equivalent workaround.

For Intel deployments, the related OpenVINO 2026.4.1 upgrade guide covers backend and device considerations. Its backend-specific results and this PR’s RTX figures should not be transferred between platforms.

Methodology: AI-assisted reporting and analysis based on official release metadata, the merged change and tagged documentation/source code, rechecked October 4, 2026. No model inference or benchmark commands were run for this article; performance figures are attributed to the PR author.