freenode
AI & ML

llama.cpp mixes embeddings and tokens in one batch

Multimodal prefill can pack text and image chunks together, after a fix for scheduler aborts on graph-shape switches.

llama.cpp now accepts both embeddings and raw tokens in a single inference batch, so multimodal runs can pack text and image chunks into one prefill pass instead of alternating separate batches. The goal is faster prefill, especially for video.

On hardware tests with reallocation disabled in the scheduler, sihanyu03 found that vision server paths aborted on a simple one-image chat with tinygemma3. Switching between image and text graphs produced slightly different shapes, so the scheduler replanned and then hit an unexpected reallocation even when the reported graph size had not grown. Results otherwise matched the prior behavior.

ngxson adjusted embedding construction so mixed and ordinary batches keep a stable graph shape across those switches. Text token embeddings are scattered into the correct slots; the trade-off is higher graph memory use.