llama.cpp model support and PyTorch proposals
llama.cpp saw the most concrete progress with a multi-participant pull request adding GLM model support, while PyTorch received early-stage proposals on optimizations and tensor APIs. Overall activity was light and largely proposal-driven.
GLM-5.3-Flash support in llama.cpp
A pull request in ggml-org/llama.cpp adds support for GLM-5.3-Flash, also called GLM5-Next. Five participants discussed a softmax crash fix, KV pooling changes, and a data race in a dummy model. Inference developers tracking new model families have concrete integration work to follow.
Polyhedral optimizations proposed as default in PyTorch
A work-in-progress pull request in pytorch/pytorch seeks to turn polyhedral optimizations on by default. The thread contains only a bot comment so far. Compiler and performance-minded PyTorch users may want to watch whether the change advances beyond the WIP stage.
Deferred tool loading requested for llama.cpp
A feature request in ggml-org/llama.cpp asks for deferred tool loading support. No discussion has appeared on the single-message thread. The capability would matter to anyone building agent-style or tool-using workflows on the library.
NumPy-free Tensor.tobytes path in PyTorch
A pull request in pytorch/pytorch proposes an owned CPU Tensor.tobytes conversion that avoids NumPy. The change is blocked by a missing CLA on the first commit. The API would simplify serialization paths that need to stay free of the NumPy dependency.