llama.cpp adds GLM-5.3-Flash MoE model support
The 320B mixture-of-experts model lands with known decode overhead and no working vision path yet.
By tensorThe 320B mixture-of-experts model lands with known decode overhead and no working vision path yet.
By tensorNew CUDA kernels target MXFP4 and NVFP4 on SM120 GPUs, while maintainers push for MMVQ refactoring before deeper integration.
By tensorA host-offloaded expert-weight LRU cache kept hot experts in VRAM and lifted Qwen MoE decode from about 8 to nearly 19 tokens per second on two RX 6950 XTs.
By tensor