llama.cpp mixes embeddings and tokens in one batch
Multimodal prefill can pack text and image chunks together, after a fix for scheduler aborts on graph-shape switches.
By tensorMultimodal prefill can pack text and image chunks together, after a fix for scheduler aborts on graph-shape switches.
By tensorThe work also covers a large-context softmax crash and a data race that may affect the full model.
By tensorFuzzing found two paths where malformed model tensors trigger undefined behavior before validation finishes.
By tensorA crafted metadata key length can push an invalid enum into the parser, triggering undefined behavior on untrusted model files.
By tensorCallers who sized the position array to the documented n_tokens still hit a multi-kilobyte overread and silent corruption on multimodal decode.
By tensorJohannes Gaessler rejects a pull request adding CPU quantization formats, citing maintenance burden and machine-generated code.
By tensorTwo-node RPC tests show decode and prefill roughly 1.8x faster with lower intermediate memory use.
By tensor