toolsWTF 5.8via r/LocalLLaMA
nvfp4 kv-cache on 2x5060 ti, vllm
"Blackwell optimization leaks to the masses: 4-bit KV-cache is officially here."
Explain Like I'm Normal
Developers have successfully implemented NVFP4 (4-bit floating point) KV-cache support in vLLM for the latest hardware. This allows for massive memory savings during long-context inference by compressing the 'memory' of the LLM significantly. The workaround specifically enables high-performance inference on consumer-grade Blackwell-era GPUs like the 5090 series.
#vllm#quantization#inference#blackwell#cuda
GET THE DAILY CHAOS
The only newsletter for people who read AI news at 3am and feel things. One email a day.