SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
toolsWTF 5.4via r/LocalLLaMA

Updated my benchmark with a new vLLM based recipe for Qwen 3.8 Flash Next : now up to 98/100 (instead of 91 previously)

"LocalLLaMA wizard turns 91 into 98 by squeezing weights through a new quantization funnel."

Explain Like I'm Normal

A developer achieved a significant jump in benchmark performance for the Qwen 3.8 Flash model by switching to a specific vLLM recipe and 4-bit PLE quantization. While the inference speed is slower for low-concurrency tasks, the accuracy improvement suggests that 'lost' performance in smaller models can often be recovered through better optimization stacks. This highlights how crucial the serving architecture is compared to raw model weights alone.

Read original ↗
#quantization#qwen#vllm#local-llm#benchmarking

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.