SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
toolsWTF 6.0via r/LocalLLaMA

UPDATE: Qwen3.8-Flash-Next on 2x3090 + DDR4 (Part 2): 25-29 -> 37-41 t/s decode (UD-Q4_K_XL + expert cache + MTP), plus a branch you can build

"Local maxxing at its finest: turning 2016 Xeons into an inference powerhouse."

Explain Like I'm Normal

A developer achieved a 2x speedup on local Qwen inference by combining expert caching with Multi-Token Prediction (MTP) on consumer hardware. The setup manages to offload massive expert layers to system RAM while maintaining high decode speeds, proving that optimization beats raw compute for local enthusiasts.

Read original ↗
#llm-inference#local-llama#optimization#mtp#hardware-hacking

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.