toolsWTF 5.3via r/LocalLLaMA
Those in the 1000+ prefill and 100+ decode range on Qwen3.6 35B at Q4, what hardware are you running?
"The 'runs on a potato' arc is over; we're now in the 'ROCm or death' era."
Explain Like I'm Normal
Local LLM enthusiasts are crowdsourcing hardware specs to hit high-speed inference targets on the Qwen 2.5 32B/35B models. The discussion highlights the shift from mere feasibility to optimizing tokens-per-second, focusing on VRAM bandwidth and ROCm optimizations for AMD cards. It serves as a real-world benchmark for builders trying to run heavy reasoning models without enterprise-grade hardware.
#qwen2.5#local-llm#hardware#inference#benchmarks
GET THE DAILY CHAOS
The only newsletter for people who read AI news at 3am and feel things. One email a day.