Article

llama.cpp b11400: How to Test Mixed Text and Embedding Batches

Evaluate llama.cpp b11400 mixed inputs: architecture gates, non-causal batch sizing, API steps and the limits of current multimodal server integration.

Editorial illustration for llama.cpp b11400: How to Test Mixed Text and Embedding Batches: documents enter a shared index with two query paths. Not documentary evidence.

llama.cpp published prerelease b11400 on October 5, 2026, at 00:04 UTC, bringing mixed-input support to builders maintaining custom multimodal integrations. The core can now accept text token IDs and input embedding rows in one batch through llama_batch_ext, following the merge of PR #29622.

Prioritize a staging evaluation if your model needs image-derived inputs and text evaluated together, especially with non-causal prompt attention. This is an integration capability, not evidence of faster generation for every llama.cpp deployment.

What ships—and what still needs integration

Here, embeddings are vectors supplied as model inputs, such as projected image representations. This is not a facility for combining a vector-search embedding job with an unrelated text-generation request. The tagged batch allocator supports a mixture of token-only entries and embedding-only entries; it does not turn those entries into independent jobs.

The application boundary matters. In b11400, the standard mtmd helper still loops over separate text and media chunks. The linked Clef vision follow-up, PR #29969, remained open, draft and unmerged at the October 5 source check. Installing b11400 therefore does not establish automatic mixed batching for every llama-server multimodal request.

Those boundaries suggest the following priorities; these are integration recommendations, not a compatibility certification for downstream packages.

Your integration

Recommended action

Reason

Custom C/C++ multimodal input path

Evaluate in staging now if joint input is needed.

You control batch construction and attention handling.

Python, mobile or other wrapper

Check its embedded core revision and bindings first.

A wrapper version alone does not prove the extended-batch path is exposed.

Stock multimodal server

Track the relevant helper/server integration.

The tagged helper still processes chunks separately.

Text-only chat or retrieval embedding output

Do not prioritize this feature alone.

The evidence establishes mixed model inputs, not a benefit for these workloads.

Why attention changes the correctness check

The motivation is clearest in a model such as PaliGemma. Its architecture paper describes image and text-prefix tokens attending across the whole input, while the generated suffix uses autoregressive masking. Earlier image positions can therefore depend on later prompt text.

Our interpretation: processing an image chunk first and merely appending the question afterward cannot reproduce that joint prefix computation. For a causal model, a matched chunked path can be a useful reference; for a non-causal prefix, use a trusted implementation that evaluates the prefix jointly. Upstream’s mixed-batch comparison test explicitly restricts its chunked-equivalence check to causal attention with memory.

PaliGemma illustrates the requirement; it is not proof that this release alone provides its complete converter, projector, attention-mask and server integration.

Construct a candidate path in three steps

1. Check architecture and context eligibility

The architecture guard excludes six internal identifiers: COGVLM, DEEPSEEK4, GRANITE_SWITCH, EAGLE3, DFLASH and GEMMA4_ASSISTANT. These are exact architecture names, not exclusions of every model in a similarly named family. Other architectures pass this guard, but still need end-to-end validation.

The context implementation additionally requires LLAMA_CONTEXT_TYPE_DEFAULT. An unsupported mixed batch is rejected, not automatically split into a working fallback. Keep the guards intact. Retain a known-correct chunked fallback only where the model’s attention semantics permit it.

2. Preserve the model’s input layout

The b11400 public header defines the following API sequence. This is documentation-derived guidance, not an executed example; the declarations already existed before this change.

  1. Initialize the batch with llama_batch_ext_init for your context.

  2. Append entries in model-defined order using llama_batch_ext_add_token or llama_batch_ext_add_embd, supplying the appropriate sequence ID. Check every returned index. Embedding data must have the model’s expected input width and row layout, not arbitrary retrieval-vector dimensions.

  3. Set positions with llama_batch_ext_set_pos and request the needed outputs, for example through llama_batch_ext_set_output_logits. Check setter results. Media positions can be multidimensional for M-RoPE models; do not copy synthetic test positions into a real vision model without checking its convention.

  4. Call llama_process with the appropriate LLAMA_PROCESS_TYPE_DECODE or LLAMA_PROCESS_TYPE_ENCODE, handle its documented return codes, and release the batch with llama_batch_ext_free when finished.

Do not confuse mixed entry types with an entry carrying both a token ID and an embedding. The allocator treats the latter separately, including for multi-token prediction paths, and forbids mixing those both-fields entries with the other entry types.

3. Size for the whole non-causal submission

For non-causal decode, the tagged capacity checks require the complete submitted batch to fit both n_batch and the physical microbatch capacity, n_ubatch. Validate the effective capacities before submission: these conditions are enforced with assertions, not just recoverable capability errors.

Illustrative calculation—not a tested model configuration: assume 256 input embedding rows and 128 text tokens, including all prompt markers.

256 embedding rows + 128 text tokens = 384 total entries
Required for this non-causal submission:
effective n_batch >= 384
effective n_ubatch >= 384

A microbatch capacity of 256 is insufficient even though each modality separately fits. The example also assumes adequate context space and memory; it does not establish a new context-window limit or a universal PaliGemma setting.

Validate correctness before judging performance

Use the tagged build guidance to create a separate candidate at b11400’s commit 0bb496dbd3af0add77ff82c406a915b41e839d56. Keep your production baseline available for rollback. To isolate this patch, compare with its parent under matching compiler and backend settings; a broader version jump includes unrelated changes.

  • Record the model and projector hashes, hardware, offload, thread count, cache policy and context settings. Use the attention-appropriate reference described above, with numerical tolerances and task-quality criteria chosen before running.

  • Exercise text-only, embedding-only and mixed inputs; short and long prefixes; modality transitions; output/logit mapping; and expected rejection cases. Inspect which upstream architecture-test cases execute or skip. A compiled target is not a passing workload test.

  • After correctness passes, measure media encoding, prefix evaluation, time to first output, full-request latency, peak host/device memory and errors. Keep inputs fixed, include cold and warm runs, and test representative concurrency.

Memory deserves a separate check. The mixed graph implementation adds dedicated token, slot and embedding inputs and copies the embedding tensor before inserting token-derived rows. This is a resource tradeoff, not a measured total-VRAM multiplier. The reviewed evidence does not establish a final-release latency or throughput gain.

For CPU vision compute, see the separate b11398 tinyBLAS eligibility guide. Its performance evidence should not be transferred to this mixed-input change.

Adopt the mixed path when the exact model/backend combination meets your correctness criteria and resource budget. If the missing piece is server integration or a valid non-causal reference, keep the evaluation in staging.

Methodology: AI-assisted reporting and source analysis using the persisted research dossier, official release/PR metadata, tagged code and the PaliGemma paper. Availability and key links were rechecked on October 5, 2026. No builds, inference tests or benchmarks were run for this article.