llama.cpp adds GLM-5.3-Flash support
The work also covers a large-context softmax crash and a data race that may affect the full model.
llama.cpp is adding support for GLM-5.3-Flash (also called GLM5-Next), widening the set of models the local inference engine can run without a separate stack.
Alongside the model wiring, contributors fixed a softmax crash that hit once the key-value cache reached 262144 entries. The failure came from a grid-dimension overflow on the k-pool gate path, in the same family of bugs the project has been chasing on related softmax and norm code.
Review feedback also pushed on KV pooling behavior: full re-pools on every memory-index edit are heavier than necessary when the cache and pool entries are retained, and partial re-pooling from the change point would keep multi-token prediction and similar paths working cleanly.
Georgi Gerganov reported a data race against the dummy GLM5 test model under a CPU-only build with the thread sanitizer enabled. He asked for it to be investigated, noting it could signal a problem in the full model as well. Save/load state testing for the new model was called out as a required check.