toolsWTF 6.0via r/LocalLLaMA
~22% less weight VRAM, lossless: base-3 packing for ternary GGUFs
"LocalLLaMA wizard compresses models with the power of base-3 arithmetic."
Explain Like I'm Normal
A new packing method called B3S allows ternary models (which only use -1, 0, and 1 values) to run with 22% less VRAM than standard 2-bit formats. By using base-3 logic instead of binary powers, a 9B model can shrink down to just 2GB of weights without losing any mathematical precision. This is specifically for BitNet-style models rather than general FP16 conversions.
#llm#quantization#ternary#gguf#local-ai
GET THE DAILY CHAOS
The only newsletter for people who read AI news at 3am and feel things. One email a day.