SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT • SAM ALTMAN SAYS AGENTS ARE COMING • CHATGPT GAINED SENTIENCE FOR 4 SECONDS • GOOGLE RELEASES 40th LLM THIS WEEK • NVIDIA MARKET CAP EXCEEDS REALITY • ANTHROPIC ENGINEER DISCOVERS NEW FORM OF GRIEF • MISTRAL RAISES AT VALUATION OF GROSS DOMESTIC PRODUCT •
toolsWTF 5.1via r/LocalLLaMA

DS V4 on single b300. only 770 tok/s batched in vLLM

"POV: you bought a B300 and realize bottlenecks don't care about your budget."

Explain Like I'm Normal

A developer testing DeepSeek V4 performance on high-end B300 hardware discovered significant throughput bottlenecks with standard vLLM configurations. The investigation highlights that advanced MoE kernels often require multi-GPU setups to hit peak efficiency, and that speculative decoding can actually hinder performance on saturated batch jobs. This real-world benchmarking serves as a guide for builders trying to scale high-throughput reasoning models on restricted hardware footprints.

Read original ↗
#deepseek#vllm#gpu-optimization#moe

GET THE DAILY CHAOS

The only newsletter for people who read AI news at 3am and feel things. One email a day.