NVIDIA Blackwell Ultra MLPerf Inference Benchmark Results
- New GB300 NVL72 system powered by Blackwell Ultra sets MLPerf Inference v5.1 records, boosting AI factory efficiency.
- NVIDIA's GB300 NVL72 rack-scale system, utilizing the newly unveiled NVIDIA Blackwell Ultra architecture, achieved record-breaking performance on the MLPerf Inference v5.1 benchmark suite. The system demonstrated up to...
- These results, announced less than six months after the initial unveiling at NVIDIA GTC 2024, highlight important advancements in AI inference capabilities. The GB300 NVL72 excelled across a...
NVIDIA Blackwell Ultra Achieves Record-Breaking AI Inference Performance
Table of Contents
New GB300 NVL72 system powered by Blackwell Ultra sets MLPerf Inference v5.1 records, boosting AI factory efficiency.
Published: November 19, 2024
What Happened?
NVIDIA’s GB300 NVL72 rack-scale system, utilizing the newly unveiled NVIDIA Blackwell Ultra architecture, achieved record-breaking performance on the MLPerf Inference v5.1 benchmark suite. The system demonstrated up to 45% higher DeepSeek-R1 inference throughput compared to systems based on the previous-generation NVIDIA Blackwell GB200 NVL72.
These results, announced less than six months after the initial unveiling at NVIDIA GTC 2024, highlight important advancements in AI inference capabilities. The GB300 NVL72 excelled across a range of benchmarks, including deepseek-R1, Llama 3.1 405B Interactive, Llama 3.1 8B, and Whisper.
Why Inference Performance Matters
Inference, the process of using a trained AI model to make predictions, is a crucial component of the AI factory. Higher inference throughput-the amount of data processed per unit of time-directly impacts the economics of AI.
Increased throughput translates to faster processing of tokens, leading to higher revenue generation, reduced total cost of ownership (TCO), and improved overall system productivity. Efficient inference is notably critical for large language models (LLMs) and other demanding AI workloads.
Blackwell Ultra: Key Architectural Improvements
The Blackwell Ultra architecture builds upon the foundation of the Blackwell architecture, delivering substantial performance gains. key improvements include:
- Increased AI Compute: Blackwell Ultra features 1.5x more NVFP4 AI compute compared to Blackwell.
- Enhanced attention Layer Acceleration: It provides 2x more acceleration for attention layers, vital for processing sequential data like text.
- Expanded Memory Capacity: Each GPU in the Blackwell Ultra architecture boasts up to 288GB of HBM3e memory.
These enhancements collectively contribute to the significant performance improvements observed in the mlperf Inference v5.1 benchmarks.
MLPerf Inference v5.1 Results: A Detailed Look
The GB300 NVL72 system demonstrated leading performance across multiple
