SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
toolsWTF 4.1via r/LocalLLaMA

How to estimate tokens/sec for your hardware

"POV: You realize your $2,000 GPU is just a glorified memory bandwidth calculator."

Explain Like I'm Normal

A new community guide simplifies the math behind LLM inference speeds, noting that memory bandwidth—not compute—is the primary bottleneck for token generation. By dividing your VRAM bandwidth by the model's weight size in GB, you can calculate the theoretical maximum tokens per second your hardware can handle. This provides a reality check for local LLM enthusiasts trying to optimize dense models versus MoE architectures.

Read original ↗
#llm#hardware#inference#vram#optimization

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.