llama.cpp published prerelease b11411 on October 5, 2026, at 10:05 UTC. It includes a fix that saves KV-cache rotation metadata and rejects mismatched restores. Operators using disk prompt caches, saved server slots or serialized application state should treat the upgrade as a saved-state migration.
Plan to regenerate old-format session and sequence files from authoritative prompts or conversation history, while preserving the old runtime/cache pair for rollback. The file-version checks also reject older unrotated files—not just caches exposed to the reported rotation bug.
What changes at restore time
A key-value (KV) cache holds intermediate attention state for previously processed tokens. The PR author reports that restoring rotated values into a context expecting unrotated values, or the reverse, could silently misinterpret the data and produce bad output. The contributor describes this as an edge case; its prevalence is not established.
The b11411 implementation saves the K and V rotation dimensions, n_rot_k and n_rot_v, with the cache payload. It compares them during restore and reports incompatible key rotation or incompatible value rotation when they differ. A visible rejection lets an application stop or rebuild instead of continuing with misinterpreted cached values.
The merged patch also changes LLAMA_SESSION_VERSION from 10 to 11 and LLAMA_STATE_SEQ_VERSION from 3 to 4. Because file loaders require exact equality, migration extends beyond the bug’s triggering condition. This changes saved inference state, not GGUF model weights. Editing an old file’s version number is not a migration: the serialized payload has changed too.
Preserve a rollback path, then rebuild
Keep the old runtime and its cache directory together; rebuild saved state in a separate directory for the candidate build.
Inventory persistence. Locate
llama-completion --prompt-cachefiles, server files under--slot-save-path, and application-managed serialized buffers. If every restart creates a fresh context and no older state is restored, there are no persisted files to migrate for this change.Pin the candidate. Stage b11411, whose release points to commit
210791069bf5460bef3c09b982ee6f0aa9967404. Record the model/GGUF identity, adapters and tokenizer/template where applicable, context settings, K/V cache types, backend, Flash Attention setting and startup rotation dimensions. Keep these fixed for the matching-restore check. Release metadata identifies this build as a prerelease.Recreate state from known inputs. For
llama-completion, use a new--prompt-cachefilename and replay the required prompt/history. Its load path exits with status 1 if an existing non-empty session file cannot load; it does not promise automatic regeneration. Preserve authoritative conversation history separately from the cache.
For custom buffer stores, version your cache envelope and retain the actual serialized byte count. The C API header warns that llama_state_get_size is for saving, not determining a restore blob’s length from a destination context. Check restore return values: llama_state_load_file reports failure as false, while sequence restoration reports failure as zero. File-header validation does not make raw buffers portable across builds.
Exercise the server’s save-and-restore path
On a disposable staging slot, process a representative prompt with the candidate build. The server documentation specifies save and restore requests with a filename inside the configured --slot-save-path. This request-body example uses slot 0; substitute the actual disposable slot ID and a unique filename:
POST /slots/0?action=save
Content-Type: application/json
{"filename":"b11411-restore-check.bin"}After a controlled restart with the same build, model and settings, send the same body to POST /slots/0?action=restore. Inspect the response, n_restored, n_read and lower-level logs, then route a representative continuation to the restored slot using the expected saved prompt prefix.
The server restore implementation can return a broad invalid-save-file or insufficient-KV-space error. Use lower-level diagnostics to separate a version failure, rotation mismatch and capacity problem. A failed restore is not a resumed session; route recovery through a clean context and prompt replay where the application supports it.
For router-mode deployments, also retain the separate model-startup and lifecycle checks in the b11401 router-mode upgrade guide. Those checks do not replace persisted-state validation.
Use separate checks for format and rotation
This proposed acceptance matrix follows the file loaders and rotation checks. It has not been executed for this article.
Staging case | Expected result or acceptance target | What it establishes |
|---|---|---|
Old session-v10 or sequence-v3 file loaded by b11411 | Rejected at the file-version check | The format boundary; not the rotation check |
New b11411 state restored with the same model and configuration | Restore succeeds | The matching save/restore path works for this case |
New-format state restored with genuinely different K or V rotation dimensions | Rejected with the relevant rotation diagnostic | The rotation guard is exercised |
Clean context rebuilt from authoritative inputs | Representative continuation meets task criteria | The application can recover without the old cache |
For the rotation-mismatch case, use a supported configuration that actually applies rotation and verify that n_rot_k or n_rot_v differs between the saving and restoring contexts. Merely toggling LLAMA_ATTN_ROT_DISABLE is insufficient: some model paths have no effective rotation, while some indexer paths force K rotation. Keep other settings unchanged to isolate the cause.
Check loading, prefix reuse and output quality separately. A successful restore response does not prove how much prompt processing the next request avoids. The completion documentation also warns that a restored prompt need not reproduce the same token sequence, even with a fixed seed. Use reuse diagnostics and task-specific output criteria; identical prose is not a universal pass condition.
Decide when to roll out
Prioritize staging if your application restores saved state, especially when rotation settings can differ between producer and consumer. Retain the old binary, configuration and cache directory until both candidate restore and clean-rebuild paths pass. Budget for cold prompt processing while caches are regenerated, and stagger the rollout if rebuilding them together would exceed local capacity.
Matching rotation dimensions are not complete cache provenance. As an application-level safeguard, include build, model, adapter and relevant configuration identities in your cache key or manifest. This specific metadata check does not establish interchangeability across arbitrary models, hardware or future builds.
No supported automatic converter was identified in the reviewed patch or pinned documentation. The first affected release and a complete affected-model/backend matrix remain unknown; this guide does not imply every earlier deployment produced incorrect output.
Methodology: AI-assisted reporting and source-code analysis of the official release, merged PR and documentation pinned to b11411, checked on October 5, 2026. No binary, inference run, cache restore or benchmark was executed. The steps above are proposed validation work, not hands-on results.
