SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
toolsWTF 5.9via r/LocalLLaMA

Qwen3.8-Flash-Next MTP merged in ik_llama.cpp (integrated head or separate -md file)... 45 → 90 tok/s on a 5090 + 128GB, works down to a 12GB 4070

"Local LLM speeds just doubled because we finally figured out how to use the 'extra' brain."

Explain Like I'm Normal

Multi-Token Prediction (MTP) support for Qwen3-Flash has been officially merged into llama.cpp, allowing models to use a secondary 'head' to predict multiple tokens at once. This speculative decoding technique effectively doubles generation speeds for coding tasks on consumer GPUs like the RTX 4070 and 5090. While a massive win for logic-heavy tasks, performance varies on creative prose where prediction acceptance rates are lower.

Read original ↗
#llm#inference#qwen#quantization#optimization

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.