Back to feed
News Story
BStandard53
NVIDIA Developer Blog
1 sources

NVIDIA Vera Storage Benchmarks: Faster Encryption, Compression, Integrity Checking, and Recovery for AI-Native Storage

NVIDIA released Vera storage benchmarks for AI-native storage, showing significant performance improvements in encryption, compression, integrity checking, and recovery. These benchmarks highlight the critical role of storage in AI workflows, especially for agentic AI applications.

SynthePulse Insight · AI deep readingMembers

NVIDIA Vera Storage Benchmarks: CPU Performance Leap for AI-Native Storage

Version 1 · 1 source

NVIDIA releases benchmarks for the Vera BlueField-4 STX storage processor, showing significant advantages over x86 CPUs in storage primitives such as encryption, compression, integrity checks, and recovery, delivering higher throughput and efficiency for AI-native storage platforms.

  • Vera integrates 88 Olympus Armv9.2 cores, supporting 176 threads, with SCF and SOCAMM2 memory, providing up to 3.4 TB/s bisection bandwidth and 1.2 TB/s memory bandwidth.
  • Compared to x86 CPUs, Vera is 1.43x and 1.29x faster in encryption and decryption, 3.26x faster in Reed-Solomon recovery, and 3.67x faster in CRC32C integrity checks.
  • Compression and decompression performance improve by 3.29x and 1.72x, respectively, with multi-stage pipeline operations improving by 3.21x.
Open section navigationStorage Processing Bottlenecks and Vera's Positioning

Storage Processing Bottlenecks and Vera's Positioning

In agentic AI workflows, storage operations are frequent and concurrent, with each agent step potentially triggering multiple storage operations. Traditional CPU scaling requires more cores, power, and cooling, with performance limited by the slowest component. NVIDIA introduces the Vera BlueField-4 STX storage processor, bringing the Vera CPU architecture into the storage data path to address this bottleneck.

Vera employs 88 Olympus Armv9.2 cores, supporting 176 Spatial Multithreading threads, combined with Scalable Coherency Fabric (SCF) and SOCAMM2 LPDDR5X memory, providing high single-thread performance and bandwidth-intensive processing capabilities. SCF offers up to 3.4 TB/s bisection bandwidth and 164 MB unified L3 cache, with memory bandwidth of 1.2 TB/s (14 GB/s per core).

These features enable Vera to sustain more CPU-side storage processing in concurrent data streams without proportionally increasing CPU resources, power, and cooling, thereby supporting higher service density and concurrent data streams.

Free for now

Read the full analysis

3 more sections of analysis, plus the full takeaway

Loading

Credibility boundary

This report is based on benchmarks published on NVIDIA's official blog, which are vendor-released information and have not been independently verified. The test methodology excludes I/O and network factors, so results may not represent real workloads. All performance data are vendor claims and should be interpreted with caution.

Primary report

NVIDIA Developer Blog

Primary source