Article

llama.cpp b11415: Media Truncation Fix and Upgrade Checks

llama.cpp b11415 fixes media-chunk boundary checks. What multimodal server operators should validate in prompt caches, client errors and staged upgrades.

Editorial illustration for llama.cpp b11415: Media Truncation Fix and Upgrade Checks: a geometric block represents a model release. Not documentary evidence.

llama.cpp published b11415 on October 5, 2026 at 12:02 UTC as a prerelease with downloadable builds. It includes a server fix for trimming cached media tokens, relevant to operators serving multimodal conversations with reused prompt prefixes. The change rejects a cut inside a media chunk and permits a valid cut between adjacent chunks.

Upgrade recommendation: prioritize a staged evaluation if your multimodal service reuses prompt caches. For a wrapper or downstream package, verify that its embedded llama.cpp includes commit 8e1642198dcd4e408f8776222d6ae31b74d01187. The changed guard is specific to media-bearing token buffers; this line alone establishes no benefit for a text-only deployment.

What the boundary check changes

The one-line patch changes find_chunk(n - 1) to find_chunk(n) inside server_tokens::keep_first(n). Keeping the first n tokens means n is the first position removed, not the last position kept. When both sides contain media placeholders, that position must start a new chunk. The lookup throws if no chunk starts there.

The upstream issue illustrates the difference with text at positions 0–4, media A at 5–7 and media B at 8–9. These are explanatory internal indices, not measured image-token counts or portable request settings. The reporter explicitly describes code inspection, not a runtime reproduction.

Internal operation

Before the fix

With b11415

keep_first(6): retain text and only the first token of A

Allowed the partial chunk

Rejects the cut

keep_first(8): retain text and all of A; remove B

Rejected the valid boundary

Allows the cut

The guard must reject cuts inside media chunks while allowing a cut between complete chunks.

Checking only that the new build rejects malformed input misses half of this fix. In the released implementation, the guard runs before the media map is trimmed and the token vector resized.

Do not confuse cache trimming with context capacity

At this revision, loading a multimodal projector disables context shifting and the separate shifted cache-reuse option. Ordinary prompt-prefix reuse remains: the server calculates a common prefix and later calls keep_first(n_past). The server documentation lists cache_prompt as enabled by default and multimodal serving as experimental; check what your wrapper actually sends.

An over-capacity prompt has a separate admission check. A request can fit the context budget yet encounter an invalid media boundary when cached tokens are trimmed. Increasing --ctx-size addresses capacity, not this indexing error. An oversized-image smoke test therefore does not establish coverage of the corrected branch.

A staged acceptance plan

The following checks are proposed from the source analysis, not reported test results. The merged commit dropped its proposed vision test because the available fixtures did not exercise the required adjacent-media boundary with a reused cache. A generic HTTP success or failure is not enough to prove that your test reached it.

1. Pin the candidate and preserve a rollback pair. Use b11415 or verify that a later candidate contains the fix. Keep the production binary, configuration, model and projector identities. Record backend/build options, chat template, preprocessing, concurrency, effective per-slot context and cache_prompt settings. Compare the candidate with your actual production revision, since an older baseline crosses more than this patch.

If isolating the one-line change in a development build, its immediate parent is 806eee9841de5f2c20f9d43914117f157d2baacc. Match build options and inputs across that comparison.

2. Separate ordinary requests from boundary coverage. Start with text-only and cold multimodal controls. Then repeat and alter multimodal prompts in a known slot. Record whether a nonzero prefix was reused, the resulting cut index and the media-chunk map using development instrumentation where needed.

  • Valid boundary: arrange complete adjacent media chunks and confirm trimming between them remains legal.

  • Invalid interior cut: in an isolated development harness, exercise a cut inside a media chunk and check rejection before buffer mutation.

  • Capacity control: send a separate over-budget request to distinguish context admission from media-structure validation.

If the model architecture or template forces full reprocessing or inserts text between chunks, record the intended boundary case as unexercised. Do not label it passed. The illustrative keep_first(6) call is not a universal n_keep=6 HTTP reproducer.

3. Check the client’s failure path. Record the status, structured error, any stream error payload, process health and outcome of the next clean request. Ensure a rejected request is not displayed as a completed answer.

The slot exception handler classifies this exception as a server error; the error formatter maps it to server_error with code 500. The separate capacity error maps to exceed_context_size_error with code 400. This is source-derived behavior, not a captured response through your proxy or SDK.

Operational implication: bound retries instead of assuming every 5xx is transient. The streaming response code also handles an initial error differently from an error after a stream has begun. Verify both error-handling paths in your client integration; do not assume this particular cache-trim failure occurs midstream.

Promotion, rollback and limits

Promote only after your relevant workload checks preserve complete media, expose failures clearly and keep retries bounded. Check follow-up requests and latency/memory under intended concurrency. Budget prompts at complete message or attachment boundaries; slicing encoded media tokens is not a repair.

If the upgrade crosses b11411, account separately for its saved-state compatibility change using the KV-cache migration and restore guide. Keep candidate cache files separate from the preserved old-runtime/cache pair. That migration predates this media-boundary patch.

The reviewed evidence does not establish a complete affected-version or model/backend matrix, a CVE designation, or a throughput or accuracy gain. Treat b11415 as a targeted correctness change whose deployment acceptance depends on your serving path.

Methodology: AI-assisted reporting and analysis of the official release, merged diff, issue discussion and commit-pinned server sources, checked on October 5, 2026. No inference workload or benchmark was run for this article; the acceptance plan describes proposed validation, not observed outcomes.