SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
breakthroughsWTF 7.3via r/LocalLLaMA

Increasing active parameters per token in MOE (Qwen 35B A4B+) reduce reasoning token by 8.5% - and you don't need to train or finetune!

"Why train more when you can just let the late-layer experts cook?"

Explain Like I'm Normal

Researchers discovered that increasing the number of active experts specifically in the final layers of an MoE model reduces 'rambling' without needing any fine-tuning. By expanding the expert budget at the end of the transformer stack, the model reaches correct conclusions 8.5% faster, effectively shortening reasoning chains while maintaining accuracy.

Read original ↗
#moe#inference-optimization#qwen#efficiency

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.