SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
toolsWTF 5.3via r/LocalLLaMA

Those in the 1000+ prefill and 100+ decode range on Qwen3.6 35B at Q4, what hardware are you running?

"The 'runs on a potato' arc is over; we're now in the 'ROCm or death' era."

Explain Like I'm Normal

Local LLM enthusiasts are crowdsourcing hardware specs to hit high-speed inference targets on the Qwen 2.5 32B/35B models. The discussion highlights the shift from mere feasibility to optimizing tokens-per-second, focusing on VRAM bandwidth and ROCm optimizations for AMD cards. It serves as a real-world benchmark for builders trying to run heavy reasoning models without enterprise-grade hardware.

Read original ↗
#qwen2.5#local-llm#hardware#inference#benchmarks

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.