Up to 30x More Work per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents
NVIDIA, Monday, August 24th, 2026
NVIDIA measures Vera Rubin NVL72 delivering 30x higher throughput per megawatt and 35x lower token costs than GB300 NVL72.
According to OpenRouter data cited by NVIDIA, agentic AI workloads consume 15 times more tokens than a simple chat request.
The reason is visible in what an agent does: researching a company for an investment decision means querying financial databases, searching news and filings, invoking a sub-agent to run peer comparisons and model valuations, then synthesizing everything.
New on-silicon performance data measured by NVIDIA using real-world agentic coding trajectories shows Vera Rubin NVL72 systems deliver 30 times higher throughput per megawatt and 35 times lower token costs than NVIDIA GB300 NVL72.