Article

llama.cpp b11390: CUDA MoE Memory-Fault Fix and Upgrade Guide

llama.cpp b11390 corrects CUDA MoE buffer padding. Identify relevant deployments, pin the fix and validate real workloads without assuming a speedup.

Editorial illustration for llama.cpp b11390: CUDA MoE Memory-Fault Fix and Upgrade Guide: a processor represents compute infrastructure. Not documentary evidence.

llama.cpp released b11390 on October 4, 2026 at 12:52 UTC as a prerelease containing a CUDA memory-access fix for quantized mixture-of-experts (MoE) inference. The merged change corrects temporary-buffer padding in the MMQ matrix-multiplication path. It matters most to operators whose CUDA MoE workloads hit illegal-memory-access errors during prompt processing.

Prioritize a staged upgrade if your deployment uses the affected path, especially if it already shows matching faults. The release is not evidence that every older build is affected or that inference will become faster.

Which deployments should move first?

Use the execution path and build identity to triage, not just the model name. The upstream report describes a 512-expert model selecting 10 experts per token, with a failure after a long generation followed by another long prompt. That is a reported configuration, not an affected-model list. The following priorities are an operator decision framework derived from that report and the patch.

Deployment evidence

Recommended next step

CUDA quantized MoE with illegal-access or MUL_MAT_ID failures

Prioritize an isolated canary containing the fix; replay the failing request sequence.

CUDA quantized MoE without a known failure

Check the deployed commit and MMQ execution path, then schedule workload-specific validation.

Dense-only CUDA, or execution wholly on another backend

This MoE padding report does not establish upgrade urgency for that workload.

A wrapper, bundled app or container with unclear upstream version

Identify its vendored llama.cpp commit or backport before assuming the fix is included.

Two configuration shortcuts can mislead. First, MMQ can be selected by default on GPUs with int8 tensor-core support; not setting GGML_CUDA_FORCE_MMQ does not rule it out. Second, an earlier issue report describes host-resident expert weights still reaching GPU MMQ under --cpu-moe. Check where computation runs, not only where weights reside.

Why the batch limit is not a safety boundary

The fix sizes MoE buffer padding from the token dimension, rather than the dimension used for dense layouts. In the tagged CUDA code, the MoE allocation now passes ne12 instead of ne11 to its tile-width helper; the dense branch still uses ne11. The tagged backend test code helps explain the distinction: the second input has a separate token axis, while the preceding axis can be one or the number of selected experts.

This makes a simple expert-count-to-batch-size comparison unreliable. The server documentation defines -ub / --ubatch-size as a physical maximum, not the size of every operation. A partial batch may have a different shape. The sources establish neither a universal safe microbatch size nor a numerical ratio that clears a deployment.

Allocation history also matters. The earlier report describes an over-read remaining inside mapped pool memory after a prior allocation, hiding the crash. The practical inference is to test both fresh processes and successive requests. One successful startup or request cannot establish that the faulty path is absent.

Pin the replacement and preserve the baseline

Before changing the service, record enough information to compare the old and candidate builds:

  • Executable version and commit, wrapper version or container digest, GPU model, driver and CUDA versions.

  • GGUF file hash, quantization, expert metadata and CPU/GPU offload placement.

  • Full launch arguments, context size, logical and physical batch limits, slot/concurrency settings and prompt-cache behavior.

Keep a known-working artifact and its configuration for rollback. A build already known to crash on the required workload is not a useful fallback.

The b11390 release page provides CUDA packages for Ubuntu and Windows. Match the asset to the operating system, architecture and supported CUDA runtime/driver combination. For a Linux source build, use a separate checkout; the following applies the tagged CUDA build recipe without modifying an existing deployment. It assumes a working CUDA toolkit and build toolchain; retain any architecture or compiler options your environment requires.

git clone --branch b11390 --depth 1 https://github.com/ggml-org/llama.cpp.git llama-cpp-b11390
cd llama-cpp-b11390
git rev-parse HEAD
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release
./build/bin/llama-server --version

Compare the checkout hash with dd266785c2595775001c1c714bd9d92b3ef34cde, the release commit. Check the candidate executable’s version and CUDA backend startup output before sending it traffic. A later build or downstream backport is also a candidate if you verify it contains the correction; adoption by particular wrappers has not been established here.

Do not substitute an arbitrary batch-size change or forced cuBLAS build for validation. Changing the batch can change the symptom, and the build documentation notes that forced FP16 cuBLAS can increase memory use and has numerical-overflow caveats. Prefer the accepted upstream change over an earlier padding snippet from an issue thread.

Replay the workload, including the easily missed prefill case

This is a proposed validation procedure, not a completed test. Start with one isolated canary using the intended model, quantization, offload placement and service settings. Define acceptable outputs, error rates and latency limits before the comparison.

  1. Replay any recorded failing sequence. As one source-based case, issue #29847 describes long generation followed by a fresh prompt longer than 512 tokens with -ub 512. Treat those numbers as reproduction context, not a universal trigger.

  2. Exercise full and partial physical batches around your deployed boundary. Include cold-process runs and repeated requests, then add the intended concurrency. Record actual evaluated token counts and request order.

  3. Control prefix reuse. On the documented /completion endpoint, cache_prompt defaults to true, so a long prompt may evaluate only its new suffix. Use distinct inputs or explicitly set cache_prompt: false for a deliberate prefill check, then also test the production cache setting. See the server request documentation.

  4. Retain inputs, outputs, sampling settings, timestamps, stderr, HTTP/process status and build identity. Check answers against the service’s task criteria as well as checking for illegal-access and MUL_MAT_ID failures. Loading successfully is not the same as completing the workload.

For a reproducing staging environment, NVIDIA Compute Sanitizer memcheck can add diagnostics for out-of-bounds CUDA accesses. A nonzero --error-exitcode lets automation detect reported sanitizer errors even when the application returns success. Keep its logs and exercised configurations; do not use instrumented timings as production performance measurements.

Promote on correctness, then measure performance separately

The issue author’s local-patch results are not a controlled b11390 speed comparison: both the padding implementation and physical batch size changed, and the tested patch differs from the merged variant. They do not establish a throughput gain for this release.

Compare baseline and candidate with the same hardware, model, quantization, offload and batch settings, documenting unavoidable differences. llama-bench supports prompt-processing and generation measurements, repetitions and JSON output. Add real-server time to first token, completion latency, peak memory and error rate at intended concurrency.

Promote only when the exercised requests meet your output and reliability criteria and remain within the predefined performance limits. A passing run supports that workload and configuration; it does not prove all GPU/model combinations are clear. The sources do not establish the first affected release or a complete list of affected configurations.

For services also using n-gram speculative decoding, keep the separate checks in the llama.cpp b11387 upgrade guide. The upstream comparison confirms b11390 descends from b11387; its intervening Vulkan RDNA4 tuning is a separate backend change, not the mechanism of this CUDA fix.

Reporting method: Source-based reporting and analysis with AI assistance, using upstream release metadata, tagged code, documentation and attributed issue reports, checked on October 4, 2026. No build, inference, sanitizer or benchmark runs were performed for this article.