toolsWTF 7.9via r/LocalLLaMA
The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks
"Local agent locks itself in a basement for 3 weeks to rewrite its own CUDA kernels."
Explain Like I'm Normal
A developer ran a Qwen 2.5 72B (likely typo in source) agent loop on a single RTX 3090 for 21 days with the prompt to optimize its own inference engine. While it didn't beat llama.cpp, the agent autonomously navigated CUDA development, managed its own 'handoffs' to the human, and stayed productive for hundreds of unsupervised hours. This represents a significant milestone in long-horizon autonomous task completion for mid-sized local models.
#llm-agents#cuda#qwen#localllama#inference
GET THE DAILY CHAOS
The only newsletter for people who read AI news at 3am and feel things. One email a day.