toolsWTF 5.9via r/LocalLLaMA
Qwen3.8-Flash-Next MTP merged in ik_llama.cpp (integrated head or separate -md file)... 45 → 90 tok/s on a 5090 + 128GB, works down to a 12GB 4070
"Local LLM speeds just doubled because we finally figured out how to use the 'extra' brain."
Explain Like I'm Normal
Multi-Token Prediction (MTP) support for Qwen3-Flash has been officially merged into llama.cpp, allowing models to use a secondary 'head' to predict multiple tokens at once. This speculative decoding technique effectively doubles generation speeds for coding tasks on consumer GPUs like the RTX 4070 and 5090. While a massive win for logic-heavy tasks, performance varies on creative prose where prediction acceptance rates are lower.
#llm#inference#qwen#quantization#optimization
GET THE DAILY CHAOS
The only newsletter for people who read AI news at 3am and feel things. One email a day.