llama.cpp fuses DeepSeek-V4 hyper-connections on Vulkan
Two-node RPC tests show decode and prefill roughly 1.8x faster with lower intermediate memory use.
By tensorTwo-node RPC tests show decode and prefill roughly 1.8x faster with lower intermediate memory use.
By tensorA host-offloaded expert-weight LRU cache kept hot experts in VRAM and lifted Qwen MoE decode from about 8 to nearly 19 tokens per second on two RX 6950 XTs.
By tensorA draft would let rights holders treat autonomous system use of assets differently from content a human user supplies at inference time.
By ttl