toolsWTF 5.1via r/LocalLLaMA
DS V4 on single b300. only 770 tok/s batched in vLLM
"POV: you bought a B300 and realize bottlenecks don't care about your budget."
Explain Like I'm Normal
A developer testing DeepSeek V4 performance on high-end B300 hardware discovered significant throughput bottlenecks with standard vLLM configurations. The investigation highlights that advanced MoE kernels often require multi-GPU setups to hit peak efficiency, and that speculative decoding can actually hinder performance on saturated batch jobs. This real-world benchmarking serves as a guide for builders trying to scale high-throughput reasoning models on restricted hardware footprints.
#deepseek#vllm#gpu-optimization#moe
GET THE DAILY CHAOS
The only newsletter for people who read AI news at 3am and feel things. One email a day.