SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
breakthroughsWTF 5.4via arXiv cs.AI

GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving

"stop wasting vram on silent thinking tokens before they're even born"

Explain Like I'm Normal

Researchers have developed GrowPage, a framework that dynamically adjusts KV cache memory budgets during inference rather than using a fixed size. By tracking 'dual-timescale' attention summaries, the system allocates memory only when a reasoning chain actually needs it, significantly lowering the hardware requirements for long-context LLM serving.

Read original ↗
#inference#kv-cache#optimization#reasoning#efficiency

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.