NVIDIA Vera Rubin NVL72 Delivers 30x Higher Agentic AI Throughput Per Watt
NVIDIA has unveiled preliminary performance data demonstrating that its upcoming Vera Rubin NVL72 architecture sets a new efficiency benchmark for agentic artificial intelligence workloads. According to measurements conducted using the SemiAnalysis AgentX benchmark, Vera Rubin NVL72 delivers up to 30 times higher inference throughput per megawatt compared to the current-generation Blackwell GB300 NVL72 platform when processing real-world agentic coding sessions. This efficiency leap directly translates to a 35 percent reduction in cost per million tokens, positioning the system as a critical solution for power-constrained AI data centers. The performance gains address a fundamental shift in AI usage patterns. Agentic workflows, which involve multi-step reasoning, tool invocation, and sub-agent coordination, generate exponentially more tokens than traditional chat or summarization tasks. Industry analytics indicate that single agentic requests consume up to 15 times more tokens than standard chat interactions, with context growing continuously across turns. This accumulating context places heavy demands on long-sequence processing, KV-cache reuse, and distributed expert parallelism, requiring infrastructure that can maintain interactive latency while scaling throughput. Vera Rubin NVL72 meets these demands through extreme hardware and software codesign. The platform leverages fifth-generation Tensor Cores, third-generation Transformer Engines, and NVFP4 quantization to optimize both prefill and decode stages without sacrificing output quality. Communication bottlenecks are mitigated by sixth-generation NVLink switches, which offer ten times higher packet rates and three times lower latency than conventional Ethernet. These components operate within a 72-GPU scale-up domain enabled by NVLink technology, allowing large mixture-of-experts models to distribute workloads across the rack with minimal latency. On the software layer, NVIDIA Dynamo session-aware routing and TensorRT-LLM Wide Expert Parallelism dynamically manage concurrency and KV-cache overlap, ensuring responsive multi-turn serving as agent sessions progress. The Blackwell generation already established a significant efficiency baseline, with GB300 NVL72 delivering up to 15 times greater throughput per megawatt than Hopper architectures on agentic workloads. Vera Rubin extends this advantage across the performance frontier, enabling AI factories to deploy up to 40 percent more GPUs within identical power budgets while sustaining target interactivity metrics. Early results, currently pending independent review, reflect optimized hardware configurations before full Vera CPU tool-execution capabilities are integrated. Nevertheless, the data underscores a strategic pivot toward architecture designs optimized for stateful, long-context AI workflows. As agentic applications expand into software development, enterprise research, and automated operations, infrastructure efficiency will dictate deployment scalability. Vera Rubin NVL72 demonstrated token-per-watt improvements directly impact AI factory revenue margins by maximizing utility conversion and minimizing operational expenditure. The complete platform architecture, incorporating specialized processors, accelerators, DPUs, and interconnects, is engineered to orchestrate complex agent trajectories across heterogeneous resources. With ongoing software refinement, the system is positioned to redefine throughput economics for next-generation AI infrastructure.
