freenode
AI & ML

llama.cpp adds GLM-5.3-Flash support

The work also covers a large-context softmax crash and a data race that may affect the full model.

llama.cpp is adding support for GLM-5.3-Flash (also called GLM5-Next), widening the set of models the local inference engine can run without a separate stack.

Alongside the model wiring, contributors fixed a softmax crash that hit once the key-value cache reached 262144 entries. The failure came from a grid-dimension overflow on the k-pool gate path, in the same family of bugs the project has been chasing on related softmax and norm code.

Review feedback also pushed on KV pooling behavior: full re-pools on every memory-index edit are heavier than necessary when the cache and pool entries are retained, and partial re-pooling from the change point would keep multi-token prediction and similar paths working cleanly.

Georgi Gerganov reported a data race against the dummy GLM5 test model under a CPU-only build with the thread sanitizer enabled. He asked for it to be investigated, noting it could signal a problem in the full model as well. Save/load state testing for the new model was called out as a required check.