llama.cpp RPC bounds bug and ML runtime patches
A remote out-of-bounds write report against the llama.cpp RPC server dominated AI and ML traffic, with smaller compile, CUDA, and model-support threads trailing behind. Activity otherwise stayed limited to targeted kernel and tooling proposals.
llama.cpp RPC PAD_REFLECT_1D out-of-bounds write
An issue on ggml-org/llama.cpp reports that PAD_REFLECT_1D can write past the destination tensor in release builds. The opening claim describes a remote out-of-bounds write in the ggml RPC server reachable through crafted GRAPH_COMPUTE messages. Operators exposing the RPC path have a direct reason to track the report and any subsequent fix.
torch.compile Bernoulli probability mismatch
A PyTorch issue states that torch.compile silently accepts invalid Bernoulli probabilities that eager mode rejects. The reporter outlines an approach to make the compiled path reject the same inputs. Anyone depending on compile and eager numerical checks to agree needs the behaviors aligned.
CUDA tiling for llama.cpp lightning indexer
A pull request to ggml-org/llama.cpp tiles the lightning indexer over keys and tokens for four heads. The change shows roughly 5 percent pp2048 speedup at depth 65536. CUDA users of that indexer obtain a modest throughput gain from the tiling.
Runtime support for Prism Bonsai 2 27B
A short ggml-org/llama.cpp thread covers runtime support and tensor-split behavior for the Prism Bonsai 2 27B model. Participants examine how the model loads and splits. Users of that architecture gain clearer runtime handling from the work.
F32 accumulation in CPU flash attention
A pull request proposes accumulating flash attention V in F32 inside the one-chunk CPU kernel. The author notes the change follows three earlier attempts at the same fix. CPU paths that hit the one-chunk kernel stand to improve numerical behavior if the patch lands.
Ollama CLI mode for decision models
An Ollama issue opens a request for CLI mode support for decision models. The single message frames the feature without further discussion yet. Operators seeking non-interactive decision-model workflows would gain a missing interface if the request proceeds.