ImportantHardware

Vera Rubin NVL72 inference tests show up to 7x better token throughput per MW vs. Blackwell on a 1.6T DeepSeek model, above Huang's 3x claim for 1T-3T LLMs (Bryan Shan/SemiAnalysis)

SemiAnalysis reports that Nvidia’s Vera Rubin NVL72 inference tests delivered up to 7x better token throughput per megawatt than Blackwell on a 1.6T DeepSeek model, exceeding Nvidia’s earlier 3x-perf claim for large models. If these results hold broadly, they could materially lower cost per token, improve data center power efficiency, and shift the economics of AI deployment toward larger-scale inference platforms—making power, cooling, and GPU roadmap decisions a strategic priority for CIOs and technology leaders.

Bryan ShanTechMeme2 min read
Read full article
Vera Rubin NVL72 inference tests show up to 7x better token throughput per MW vs. Blackwell on a 1.6T DeepSeek model, above Huang's 3x claim for 1T-3T LLMs (Bryan Shan/SemiAnalysis)

Read the full story at TechMeme →