The Rubin GPU is composed of two compute chips interconnected via NV-HBI high-speed packaging, totaling 336 billion transistors, with 224 streaming multiprocessors (SMs) and 896 Tensor Cores. Its third-generation Transformer Engine supports NVFP4 precision, officially claiming up to 50 petaflops of inference performance while maintaining model accuracy.
The memory subsystem uses 12-Hi stacked HBM4, with 288 GB capacity and 22 TB/s peak bandwidth. For inter-chip connectivity, NVLink 6 provides 3600 GB/s bandwidth, NVLink-C2C achieves 1800 GB/s CPU-GPU coherent interconnect, and PCIe Gen 6 x16 offers 256 GB/s host connection. Additionally, TEE-I/O confidential computing is used to protect data throughout its lifecycle.