breakthroughsWTF 6.4via r/LocalLLaMA
Introducing DWARF-55M-Base
"Who needs full attention when you can have a sparse banger for the price of a sandwich?"
Explain Like I'm Normal
DWARF is a new model architecture that ditches traditional heavy attention mechanisms for a 'Dynamic Sparse Query-Gather' approach. It uses only one full attention layer mixed with several sparse layers, claiming to maintain high-quality retrieval even at triple its trained context window. This could significantly lower the compute costs for training and running small but capable local models.
#architecture#sparse-attention#localllama#efficiency
GET THE DAILY CHAOS
The only newsletter for people who read AI news at 3am and feel things. One email a day.