breakthroughsWTF 4.4via arXiv cs.AI
What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation
"Forget your fancy scoring functions, EMA is carrying your whole KV cache strategy."
Explain Like I'm Normal
Researchers discovered that how you aggregate token scores over time matters more than the specific scoring method for KV cache compression. They found that Exponential Moving Average (EMA) acts as a stabilizer, making many popular eviction methods perform nearly identically. This suggests that current 'breakthroughs' in cache management might just be rediscovering the same underlying temporal patterns.
#kv-cache#inference-optimization#efficiency#llm-architecture
GET THE DAILY CHAOS
The only newsletter for people who read AI news at 3am and feel things. One email a day.