Article

llama.cpp OpenVINO 2026.4.1 Update: Intel MoE Upgrade Guide

What llama.cpp b11374 changes for Intel OpenVINO users, how to interpret the Gemma MoE benchmark, and what to validate before a production upgrade.

Editorial illustration for llama.cpp OpenVINO 2026.4.1 Update: Intel MoE Upgrade Guide: a geometric block represents a model release. Not documentary evidence.

On October 3, 2026, llama.cpp published pre-release b11374, integrating OpenVINO 2026.4.1 and a substantial update to its Intel inference backend. For teams already serving local models through OpenVINO, the release brings device-selection fixes, prompt-handling corrections and mixture-of-experts (MoE) optimization work. Dedicated Ubuntu and Windows x64 OpenVINO packages are available.

The practical decision is whether these changes solve a problem in your deployment. Prioritize a staging trial if you use Intel GPU MoE models, chunked prompts or a machine with multiple compute devices. Measure prompt processing and token generation separately; the upstream performance report shows very different gains in those two phases.

What makes this update worth testing

  • Device selection and memory accounting. An unavailable requested device now produces an error listing available devices. Integrated-GPU and NPU free-memory reporting is capped by available system memory, while GPU accounting excludes host USM allocations. These are selection and reporting corrections, not extra physical memory. For capacity planning, do not add shared GPU and host headroom as though they were independent pools.

  • Mixed-platform GPU initialization. The OpenCL interop fix obtains Intel memory-function pointers from the selected device’s platform, addressing failures when another vendor’s platform was enumerated first. This concerns OpenVINO’s GPU path, not llama.cpp’s separate OpenCL backend.

  • Prompt retention. A cached-decoder fix rebinds inputs when llama.cpp supplies a different graph. The contributor describes stale bindings losing prompt content in sliding-window-attention and recurrent models when prompts are split into chunks. This gives serving teams a correctness reason to evaluate the release even without a large decode-speed gain.

  • MoE and operator coverage. The merged change includes Qwen/Gemma MoE work and additional convolution and multimodal-processing operators. That is not complete multimodal availability: the tagged backend guide still describes text-only support and unfinished accuracy validation.

Read the Gemma benchmark by phase

The fusion contributor’s report uses Gemma-4-26B-A4B on Arc B390 with GGML_OPENVINO_REQUANT_KQUANT=q4_asym64_all and llama-bench -p 512 -n 128 -r 2. It compares the fusion with an unfused GGML_OPENVINO_MOE_OP=0 baseline. A second commit adds the stateful comparison below.

Contributor-reported configuration

Prompt processing, 512 tokens (tokens/s)

Generation, 128 tokens (tokens/s)

Unfused baseline

66.16

25.94

Fusion enabled, stateless

1,608.73

26.46

Stateful; this prefill fusion does not match

66.18

29.91

Calculation from the reported values: 1,608.73 ÷ 66.16 = 24.32× prompt-processing throughput, while (26.46 ÷ 25.94 − 1) × 100 = 2.00% more generation throughput. These are configuration comparisons, not an old-release/new-release benchmark or a 24× reduction in complete-request latency.

The stateful result illustrates a separate trade-off: generation is about 13% faster than stateless-fused, but prompt processing falls back near the unfused rate. The contributor attributes this to a graph-shape mismatch that prevents the prefill fusion in stateful mode. The source comparison therefore does not justify turning on every performance switch together.

The fusion report explicitly limits the result: this optimization does not apply on CPU, without that requantization setting, or to models with separate gate/up weights. The note omits the exact checkpoint hash, full driver/OS configuration and raw repetition results. No independent reproduction was located.

Correctness remains a separate question. The stateful development note describes repetitive Gemma-4-26B-A4B output in both modes. Later prompt-binding fixes followed, so that observation is not a diagnosis of every final b11374 configuration. Equally, the reviewed sources do not establish final task-quality clearance for the benchmark model. Similar reported perplexity between fused and unfused runs is not enough to approve a service.

A controlled upgrade procedure

1. Pin the release and prove which device runs it

Keep the working binary, runtime, drivers and settings available for rollback. Obtain the dedicated b11374 OpenVINO package, or build the pinned tag with OpenVINO 2026.4.1 and -DGGML_OPENVINO=ON using the backend build instructions. Record binary and GGUF hashes, model revision, quantization, OS, drivers, device, context, slots, batch sizes and all OpenVINO environment settings.

Run llama-cli --list-devices from that build. Its device descriptions show the valid GGML_OPENVINO_DEVICE values and selected device; use that environment variable, not -dev, to choose the OpenVINO target. Confirm the same selection in startup logs.

The following Linux examples are documentation-derived templates, not commands executed for this article. They assume the documented source-build layout and an initialized OpenVINO environment. Replace the model path and use GPU.0 only if it is listed locally; package layouts may differ.

./build/ReleaseOV/bin/llama-cli --list-devices

GGML_OPENVINO_DEVICE=GPU.0 \
GGML_OPENVINO_STATEFUL_EXECUTION=0 \
./build/ReleaseOV/bin/llama-bench \
  -m /path/to/model.gguf -fa 1 -p 512 -n 128 -r 5 -o json

OpenVINO’s tool restrictions require flash attention for this benchmark. The llama-bench reference documents separate prompt/generation tests, repetitions and JSON output. Five repetitions here are a proposed evaluation choice, not a reproduction of the contributor’s two-run setup. Preserve raw results and add longer prompts and realistic generation lengths; synthetic throughput does not measure HTTP service latency.

2. Separate release, fusion and quantization changes

Start stateless with the settings your service already uses. Then run distinct comparisons, changing one intended factor at a time:

  • Release comparison: current deployment versus b11374, documenting unavoidable runtime or driver changes.

  • Fusion comparison: on b11374, hold requantization fixed and compare GGML_OPENVINO_MOE_OP=0 with GGML_OPENVINO_MOE_OP=1 for a matching MoE model.

  • Quality comparison: evaluate the existing weight-conversion setting against q4_asym64_all using the same task set and acceptance criteria. Do not infer accuracy preservation from a fusion-only comparison.

  • Execution-mode comparison: test stateful separately only if its feature limits fit the service.

The requantization documentation says q4_asym64_all also transforms Q4_K weights, trading accuracy for reduced memory traffic. That makes the runtime setting part of the deployment identity: a matching GGUF hash alone is insufficient. For the broader distinction between file format and execution path, see GGUF in Transformers; its Apple Silicon implementation is separate from this Intel backend.

3. Check correctness before raising concurrency

A proposed regression set should place known facts near the beginning, middle and end of prompts that span multiple chunks. Check whether answers retain those facts, follow the required output format and avoid repetition. Include fresh requests and multi-turn resets, retaining prompts, outputs, sampling settings and timestamps. This recommendation follows from the prompt-rebinding fix; it is not a claim that RohitAI reproduced the failure.

The b11374 mode restrictions and validation matrix make several boundaries explicit: stateful execution is experimental, defaults off and requires one CPU/GPU slot. It does not support context shift, rewind or state save/restore. In its Lunar Lake Q4_K_M checks, the matrix reports GPU-stateful accuracy issues for listed Qwen3.5 variants and failed/unsupported GPU runs for Gemma-4-E4B. NPU serving is single-slot with static graphs and needs an explicit context budget. None of the Arc benchmark numbers establishes NPU performance.

Begin with one server slot and a deliberately bounded context, then evaluate the intended supported concurrency. Before measurement, define acceptable task error rates, p50/p95 first-token and full-response latency, peak host/device memory and sustained throughput. A passing health check or a fast token counter should not substitute for accepted answers.

Treat cache-only startup as a separate rollout

If compilation dominates startup, the compiled-model cache guide offers a separate optimization. First populate GGML_OPENVINO_COMPILED_MODEL_CACHE_DIR for every intended workload and wait for exports to finish. Only then evaluate GGML_OPENVINO_COMPILED_MODEL_CACHE_ONLY=1 with matching settings and unchanged source files.

This path requires Linux/Windows mmap loading, full dynamic CPU/GPU execution and compatible OpenVINO/plugins; it does not support static/NPU execution or CPU fallback. Cache identity includes local file identity and runtime settings, so copying or replacing a GGUF invalidates entries. Missing entries fail rather than compile, and weighted blobs can approach model size per graph. Measure cold compilation, warm import and first response separately, and preserve the old deployment’s cache for rollback.

When to promote it

For an affected Intel OpenVINO service, b11374 warrants a pinned staging evaluation. Promote only when the relevant correctness checks pass and the actual workload meets its latency, memory and concurrency targets. Keep the deployed version if the model/device combination remains unsupported or the only gain appears in a synthetic phase your application rarely stresses. These OpenVINO changes do not establish a performance benefit for unrelated backends.

Methodology: AI-assisted reporting and analysis based on the persisted research dossier, upstream release metadata, tagged documentation and contributor commits, rechecked on October 3, 2026. Benchmark values are attributed upstream results; ratios are arithmetic on those values. No hardware, models or example commands were run for this article.