toolsWTF 6.7via r/LocalLLaMA
Qwen3.8-Flash-Next: 256k context, 16tok/s on DDR4 and a Tesla T4
"Running a 180B model on a vintage GPU while you make a sandwich."
Explain Like I'm Normal
A developer successfully ran the Qwen 180B Mixture-of-Experts model at usable speeds on a $150 legacy Tesla T4 GPU and DDR4 RAM. By utilizing specific PRs for llama.cpp and aggressive quantization, they achieved a massive 256k context window on consumer-accessible server hardware. This proves that high-end frontier capabilities are increasingly decoupling from the latest H100 clusters.
#qwen#local-llm#inference#unsloth#hardware
GET THE DAILY CHAOS
The only newsletter for people who read AI news at 3am and feel things. One email a day.