Article

llama.cpp b11398: When x86 CPU Vision Workloads Benefit

Check whether llama.cpp b11398’s tinyBLAS K-tail support benefits your CPU vision workload, interpret reported timings and plan a controlled upgrade.

Editorial illustration for llama.cpp b11398: When x86 CPU Vision Workloads Benefit: a geometric block represents a model release. Not documentary evidence.

On October 4, 2026, llama.cpp published prerelease b11398, adding x86 tinyBLAS support for BF16, FP16 and FP32 matrix multiplications whose K dimension leaves a partial vector. It is a targeted upgrade for CPU vision workloads, not a general promise of faster text generation.

The patch keeps eligible x86 CPU matrix multiplications on the vectorized path by handling a final partial block. Previously, a non-aligned K could send the whole multiplication to the generic CPU fallback. The merged change handles the remainder internally; it does not require changing model weights or padding input images.

Prioritize a staging comparison when image encoding is a material bottleneck and the affected operations actually run through ggml-cpu tinyBLAS. A service with GPU-only vision processing or another backend has no demonstrated benefit from this specific patch.

Check the projector, matrix shape and CPU path

A GEMM is a matrix multiplication; K is the length over which products are summed. A K tail exists when that length is not divisible by the selected vector width. The width depends on precision and the compiled CPU instruction-set path, as shown in the tagged tinyBLAS implementation.

The motivating example in PR #29806 is Qwen3.8-27B’s BF16 vision down-projection, with K=4304 and M=1152. The official model configuration also lists vision intermediate size 4304 and hidden size 1152.

The BF16 implementation selects width 32 for AVX512-BF16, 16 for its AVX512F branch without BF16 instructions, or 8 for AVX2. Calculating the remainders gives 4304 mod 32 = 16, but 4304 mod 16 = 0 and 4304 mod 8 = 0. Therefore, the same matrix has a tail on the first path and is aligned on the other two. This is source-based eligibility analysis, not a cross-CPU performance test.

An AVX2 machine should not inherit the Qwen speedup expectation simply because it uses the same model. Other non-aligned shapes may benefit. Shape arithmetic alone also does not prove dispatch: tensor layout, precision and alternative kernels matter. The entry point still rejects N below 2, so ordinary single-column token decoding is not directly accelerated by this path.

Inspect the vision component separately. llama.cpp uses a matching mmproj GGUF for image encoding and projection; a quantized language-model filename does not establish the projector’s tensor types. The multimodal documentation says projector GPU offload is enabled by default and --no-mmproj-offload disables it. A GPU language model can still use a CPU projector—and the reverse is possible.

If the relevant operations run in OpenVINO, this tinyBLAS result does not establish an OpenVINO speedup. For that separate backend decision, see the llama.cpp OpenVINO upgrade guide.

What the reported timings mean

Contributor SongXiaoXi reports these Qwen3.8-27B BF16 mmproj timings on an AMD 9950X with 16 threads in PR #29806. They concern image encoding/projector work, not a full chat response. The final column is calculated from the reported milliseconds.

Input image

Before

After

Reported speedup

Calculated time reduction

448 × 448

338.15 ms

262.81 ms

1.287×

22.28%

896 × 896

1800.33 ms

1517.03 ms

1.187×

15.74%

1120 × 1120

3410.17 ms

2999.66 ms

1.137×

12.04%

Time reduction is (before − after) / before. A 1.287× speed ratio is therefore about 22.28% less time, not 28.7%. The largest image saves more milliseconds—410.51 ms versus 75.34 ms for the smallest—but has a smaller relative improvement. Three points do not establish a scaling rule.

For an illustrative request budget, assume mmproj accounts for half of baseline serial request time and improves from 338.15 to 262.81 ms. Total time would fall by 0.5 × (1 − 262.81 / 338.15) = 11.14%. This calculation assumes all other work is unchanged, equal outputs, no overlap and no queueing change. It is not a measured service result.

The published record lacks raw repetitions, confidence intervals and exact baseline/candidate binary identities; it does not establish that every table was rerun on the final merge. FP16/FP32 figures in the discussion are operation benchmarks, not equivalent full-mmproj results. Independent reproduction is not established by the sources reviewed.

Stage an upgrade without changing the workload

The following Linux x86 commands are documentation-derived examples, not an executed test or a reproduction of the contributor’s harness. Keep the current working binary or container and configuration for rollback. Release packages are available, but a wrapper’s version alone does not establish which upstream commit or CPU path it includes.

1. Pin the candidate and record the build

Use a separate b11398 source checkout and confirm git rev-parse HEAD returns a7b94df2c616bc1f62a73b964b4a71cb0dcc488e. Record the baseline revision, binary hashes, CPU model, selected instruction set, OS, compiler and build options. The release record identifies the candidate; the build guide documents the CMake workflow.

git rev-parse HEAD
cmake -S . -B build-b11398 -DCMAKE_BUILD_TYPE=Release -DGGML_LLAMAFILE=ON -DGGML_NATIVE=ON -DLLAMA_BUILD_TESTS=ON
cmake --build build-b11398 --config Release

This assumes a working CMake/C++ toolchain on the intended x86 CPU. GGML_LLAMAFILE defaults to enabled in standalone llama.cpp, but embedded builds and existing caches can differ. Check the actual cache. GGML_NATIVE targets the build machine; do not treat this as a portable binary recipe for a mixed-CPU fleet. Match backend and compiler settings between baseline and candidate.

2. Inspect the actual mmproj file

With the checkout’s gguf-py dependencies available, the GGUF dump tool can list tensor names, types and shapes without running inference:

python3 gguf-py/gguf/scripts/gguf_dump.py /path/to/mmproj.gguf --json

For the motivating projector, the PR identifies BF16 v.blk.*.ffn_down.weight tensors with GGUF dimensions [4304,1152]. Relate your actual multiplication’s K to the selected kernel width. Metadata identifies candidates; confirming execution routing requires runtime evidence. Do not infer K from input image width or a Q4_K_M filename. Record model and projector revisions/hashes, offload settings, thread counts and image-token limits before comparing versions.

3. Smoke-test placement, then measure the real service

For a supported matching model/projector pair, this multimodal CLI example places both components on CPU. Replace the file paths and set context and image limits appropriate to the model:

./build-b11398/bin/llama-mtmd-cli \
  -m /path/to/model.gguf --mmproj /path/to/mmproj.gguf \
  --no-mmproj-offload -ngl 0 --image /path/to/image.png \
  -t 16 -tb 16 -n 64 -p "Describe this image."

For a GPU-language-model/CPU-projector deployment, retain the existing language-model offload setting instead of introducing -ngl 0. Sixteen threads are an example, not a universal optimum. The argument definitions distinguish generation and batch thread controls; mtmd-cli uses -t for projector threads. Its projector log timings and generated answer are not HTTP latency measurements.

  • Hold inputs constant: model/projector hashes, image bytes, preprocessing limits, context, batching, cache policy, CPU placement, threads and sampling settings.

  • Measure phases separately: image encoding, text prefill, decoding, client-side first-token latency and full-response latency. Include memory, errors and task accuracy.

  • Use representative small, medium and large images, cold and warmed runs, and real service concurrency. Set repetitions, warm-up policy and acceptance thresholds in advance; retain every timing, not just the best run. Three images are not a reliable p95 sample.

  • Include aligned-shape controls. The contributor reports a small aligned-shape slowdown, without an uncertainty interval establishing its significance.

A production-baseline comparison answers whether to deploy. If you need to isolate this patch, separately compare parent 0eb6d9a8137ff13deb1b3a755b01b1e5b7f89229 with the merge commit under the same build configuration.

4. Check numerical results before promotion

The merged backend tests add cases around vector boundaries. The CPU dispatcher change also bypasses tinyBLAS in reference mode, avoiding a comparison that routes both sides through the optimized path. An operation-level starting point is:

./build-b11398/bin/test-backend-ops test -b CPU -o MUL_MAT

Confirm applicable cases actually execute rather than being skipped. Then evaluate your real vision tasks against accepted answers; passing matrix-operation checks alone does not establish OCR accuracy, image understanding or service reliability.

Promote when the matched comparison meets your correctness and latency targets, including aligned workloads. Otherwise, keep investigating placement and bottlenecks rather than assuming the upstream ratio applies. Lack of a performance benefit is not a reason to defer unrelated correctness or security updates.

Reporting and analysis: Based on upstream release metadata, merged code, documentation and contributor benchmarks checked on October 4, 2026, with AI assistance. No builds, model inference or performance tests were run for this article. Calculations and proposed validation steps are distinct from the contributor’s measurements.