SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
toolsWTF 5.4via r/LocalLLaMA

I benchmarked Unsloth's Qwen3.6-27B NVFP4 on 1x/2x 5090s. MTP is great until it really isn't.

"MTP is a speed demon until your context window decides to wake up and violence."

Explain Like I'm Normal

Independent testing of Unsloth's Qwen3.6-27B quantized to NVFP4 reveals that speculative decoding (nspec=3) provides massive gains for single users on 5090 GPUs. However, these performance wins evaporate under heavy batching or large context loads, where hardware constraints eventually cause Multi-Token Prediction to actually slow down generation.

Read original ↗
#llm-inference#5090#quantization#vllm

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.