SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
toolsWTF 5.8via r/LocalLLaMA

nvfp4 kv-cache on 2x5060 ti, vllm

"Blackwell optimization leaks to the masses: 4-bit KV-cache is officially here."

Explain Like I'm Normal

Developers have successfully implemented NVFP4 (4-bit floating point) KV-cache support in vLLM for the latest hardware. This allows for massive memory savings during long-context inference by compressing the 'memory' of the LLM significantly. The workaround specifically enables high-performance inference on consumer-grade Blackwell-era GPUs like the 5090 series.

Read original ↗
#vllm#quantization#inference#blackwell#cuda

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.