Nvidia: performance-per-watt is now the metric that decides AI infra profitability
Nvidia's argument: since power is now AI infrastructure's hard constraint, token-generation capacity per watt — not raw throughput — is what determines revenue and profitability under a fixed power budget. The concrete numbers: GB300 NVL72 claims up to 25x performance-per-watt over Hopper on DeepSeek V4 Pro, 20x on GLM5.1, and 10x on Kimi K2.6 (an agentic-task model); Nvidia also says software optimization alone improved DeepSeek V4 performance up to 5x in one month, and its DSX MaxLPS power-management feature packs "40% more GPUs within the same power budget." The more interesting framing than the numbers themselves: Nvidia pushes Pareto curves over single-point benchmarks (different workloads want different optimization targets) and cites production validation from Anthropic, OpenAI, Perplexity, and CoreWeave rather than lab-only figures. Read this as a vendor argument, not independent verification — the benchmarks are Nvidia's own.
Sources & depth
- blogs.nvidia.comWhy Performance per Watt Is the Ultimate Metric for AI Infrastructure Efficiency2026-07-14 20:51 IST