SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
toolsWTF 5.4via r/LocalLLaMA

Qwen3.8-Flash-Next on 2x3090 + DDR4: 17 → 25-29 t/s decode with the expert cache PR

"Xeon-era e-waste just got a 50% speed boost for massive MoE models."

Explain Like I'm Normal

A new llama.cpp pull request introduces a GPU-resident LRU expert cache, allowing users to run large Mixture-of-Experts models across split VRAM/RAM setups much faster. By keeping frequently used 'experts' on the GPU instead of swapping entire layers, users are seeing decode speeds jump from 17 to nearly 30 tokens per second. This effectively breathes new life into older dual-GPU setups with high-capacity system RAM.

Read original ↗
#llamacpp#qwen#moe#local-ai#hardware

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.