MoE GPU cache, PyTorch NPU paths, Omnilingual ASR
Open source AI and ML work on this day focused on inference caching, backend portability, and speech model integration. Activity stayed technical and proposal-driven across llama.cpp, PyTorch, and Transformers.
GPU-resident cache for MoE experts in llama.cpp
A pull request in ggml-org/llama.cpp proposes a GPU-resident cache for mixture of experts (MoE) experts that are kept in host memory. Early user reports and a stride observation appear in the thread. The change matters for developers running large MoE models when GPU memory is limited and host-backed experts need faster reuse.
Privateuse1 path for PyTorch nearest upsample heuristic
A feature proposal in pytorch/pytorch would decouple the _upsample_nearest memory-format heuristic from a hard-coded CUDA check so privateuse1 backends, including Ascend NPU, can use it. The patch targets an assumption that currently limits non-CUDA accelerators. Maintainers of custom or NPU backends have a direct interest in whether that heuristic becomes portable.
Omnilingual ASR models added to Transformers
A pull request in huggingface/transformers adds Omnilingual ASR models. The thread so far covers routine CI, lint, and test failures, with one maintainer note on architecture. Readers who track multilingual speech recognition support in the library may want to watch how the integration settles.