NVIDIA has released the first real-chip performance benchmarks for Vera Rubin, showing up to 30x higher throughput per megawatt and up to 35x lower token costs for agent workloads compared to the previous generation GB300 NVL72. The tests used SemiAnalysis's AgentX workload, which replays real production-style agent sessions to evaluate a system's ability to handle real agent requests.
Unlike traditional chat tasks, agent sessions feature long chains with dynamically changing context lengths. In the real agent trajectory NVIDIA demonstrated, the main agent's context can grow from about 60,000 tokens to nearly 400,000 tokens, with multiple sub-agent requests interspersed. NVIDIA defines the metric as 'throughput per megawatt' at the AI factory level, while also focusing on end-to-end latency and time-to-first-token.
Using the DeepSeek V4 Pro model, the Vera Rubin NVL72 achieved 30x higher throughput per megawatt compared to GB300 NVL72 under a 160 tokens/s per user interaction target. This reflects a shift in agent infrastructure evaluation from fixed-prompt tests to workloads that closely resemble real production environments.