breakthroughsWTF 4.7via arXiv cs.AI
SelKV: Selective KV Cache Merging with Per-Token Merge-or-Drop and Attention Compensation
"Your context window just got a lot cheaper without losing its mind."
Explain Like I'm Normal
Researchers introduced SelKV, a training-free method to compress LLM memory by selectively merging or dropping key-value tokens based on similarity. Unlike traditional methods that cause 'attention sag,' this approach uses a soft cosine gate and compensation weights to maintain accuracy across long contexts. This significantly reduces VRAM usage for high-throughput inference without retraining the model.
#llm#kv-cache#inference#efficiency
GET THE DAILY CHAOS
The only newsletter for people who read AI news at 3am and feel things. One email a day.