toolsWTF 5.4via r/LocalLLaMA
Qwen3.8-Flash-Next on 2x3090 + DDR4: 17 → 25-29 t/s decode with the expert cache PR
"Xeon-era e-waste just got a 50% speed boost for massive MoE models."
Explain Like I'm Normal
A new llama.cpp pull request introduces a GPU-resident LRU expert cache, allowing users to run large Mixture-of-Experts models across split VRAM/RAM setups much faster. By keeping frequently used 'experts' on the GPU instead of swapping entire layers, users are seeing decode speeds jump from 17 to nearly 30 tokens per second. This effectively breathes new life into older dual-GPU setups with high-capacity system RAM.
#llamacpp#qwen#moe#local-ai#hardware
GET THE DAILY CHAOS
The only newsletter for people who read AI news at 3am and feel things. One email a day.