Back to feed
News Story
机器之心
1 sources

NVIDIA Vera Rubin Debuts: 30x Agent Throughput Improvement

NVIDIA has released the first real-chip benchmark results for Vera Rubin, showing up to 30x higher throughput per megawatt and 35x lower token costs compared to the previous GB300 NVL72 in agent workloads. The tests used SemiAnalysis' AgentX benchmark with the DeepSeek V4 Pro model. Meanwhile, SpaceXAI announced plans to deploy NVIDIA Vera CPUs and expand its AI infrastructure for Grok based on the Vera Rubin architecture.

SynthePulse Insight · AI deep readingMembers

NVIDIA Vera Rubin Debut: The Infrastructure Race in the Agent Era

Version 1 · 1 source

NVIDIA has released the first real-chip performance benchmarks for Vera Rubin, showing up to 30x higher throughput per megawatt and 35x lower token costs for agent workloads, and announced SpaceXAI's deployment of Vera CPUs. This marks a shift in AI competition from model capability to the compute infrastructure that powers continuous agent actions.

  • NVIDIA has released the first real-chip performance benchmarks for Vera Rubin, showing up to 30x higher throughput per megawatt and up to 35x lower token costs for agent workloads compared to GB300 NVL72.
  • The tests used SemiAnalysis's AgentX workload with the DeepSeek V4 Pro model, targeting 160 tokens/s per user interaction.
  • Vera Rubin NVL72 results are rack-scale, co-designing Rubin GPU, HBM4, NVLink 6, and a new inference software stack, with key mechanisms including Prefill/Decode separation and Distributed KV Cache.
Open section navigationA New Benchmark for Agent Workloads

A New Benchmark for Agent Workloads

NVIDIA has released the first real-chip performance benchmarks for Vera Rubin, showing up to 30x higher throughput per megawatt and up to 35x lower token costs for agent workloads compared to the previous generation GB300 NVL72. The tests used SemiAnalysis's AgentX workload, which replays real production-style agent sessions to evaluate a system's ability to handle real agent requests.

Unlike traditional chat tasks, agent sessions feature long chains with dynamically changing context lengths. In the real agent trajectory NVIDIA demonstrated, the main agent's context can grow from about 60,000 tokens to nearly 400,000 tokens, with multiple sub-agent requests interspersed. NVIDIA defines the metric as 'throughput per megawatt' at the AI factory level, while also focusing on end-to-end latency and time-to-first-token.

Using the DeepSeek V4 Pro model, the Vera Rubin NVL72 achieved 30x higher throughput per megawatt compared to GB300 NVL72 under a 160 tokens/s per user interaction target. This reflects a shift in agent infrastructure evaluation from fixed-prompt tests to workloads that closely resemble real production environments.

Free for now

Read the full analysis

4 more sections of analysis, plus the full takeaway

Loading

Credibility boundary

This article is primarily based on a report by Machine Intelligence (机器之心) covering NVIDIA's official announcements, and is a secondhand account. Performance figures (30x, 35x, etc.) are NVIDIA's official claims and have not been independently verified. Third-party test data from Artificial Analysis is as reported by the source, without the original report. The orbital computing plans are based on statements from Musk and are forward-looking, subject to uncertainty.

Primary report

机器之心

Primary source