Article

llama.cpp b11393 Fixes Tool-Call Parser Use-After-Free

llama.cpp b11393 clears a stale tool-call pointer. See which parser configurations warrant an upgrade and why the earlier b11390 CUDA fix is separate.

Editorial illustration for llama.cpp b11393 Fixes Tool-Call Parser Use-After-Free: a processor represents compute infrastructure. Not documentary evidence.

llama.cpp released b11393 on October 4, 2026, fixing a use-after-free in its shared tool-call parser. The build is marked prerelease. Operators whose llama-server deployments let less-trusted callers supply chat-parser definitions should prioritize a staged upgrade.

The PR author reports a remotely triggered worker abort through the native /completion endpoint using a caller-supplied chat_parser. The report establishes a reason to address memory safety, but does not demonstrate reliable remote code execution or exploitation in the wild.

What the patch fixes

The parser mapper kept a pointer to a temporary tool-call object after a tool-close event destroyed that object. A later tool-ID event could then write through the stale pointer. The author reports both a double-free abort and an AddressSanitizer use-after-free finding in the upstream report; those results have not been independently reproduced for this article.

The fix clears the pointer when the pending tool call is destroyed, preventing a later tool-ID assignment from using it. The merged patch adds only current_tool = nullptr; in one source file. It does not include the regression test mentioned in the initial PR description.

Exposure depends on who controls the parser

The upstream discussion says built-in parsers do not emit the problematic ordering. Ordinary tool calling or structured JSON output alone therefore does not establish exposure. However, the request schema accepts a caller-supplied parser definition.

Operator triage, based on those source findings:

Deployment configuration

Recommended action

Callers can submit custom chat parsers

Prioritize deploying the fix. Restrict unnecessary parser overrides while staging the update.

Only fixed, trusted parsers are intended

Verify that wrappers and proxies actually reject overrides; intended use is not proof of the enforced boundary.

Check route aliases as well as request fields: the server registers both /completion and /completions to the same native handler. Filtering only the singular route can leave that review incomplete. These are exposure-reduction recommendations, not tested complete mitigations; behavior through every OpenAI-compatible route and wrapper has not been established.

Deploy the parser fix, then check real workflows

Use b11393 or a verified descendant or backport containing fix commit dbe4c3e. Check the actual running binary or vendored llama.cpp revision, not just a container or wrapper version. The first affected release and full historical version range remain unknown.

If you already upgraded for the b11390 CUDA MoE memory-fault fix, that build lacks this pointer reset. b11393 descends from b11390, but this correction concerns the shared parser, not CUDA settings or model weights.

Suggested staging checks, not tests performed for this report:

  • Keep the model, template and request settings constant; verify single and parallel tool calls, IDs, arguments and no-tool responses.

  • Check streaming, nonstreaming and legitimate continuation/prefill behavior. Confirm expected output and worker stability before rollout.

Methodology: AI-assisted reporting and code-based analysis of the release, merged patch and version-pinned sources. No inference, exploit reproduction or performance benchmark was run.