Product Information

NVIDIA Blackwell MLPerf Training v4.1 LLM Performance Results NEWS DETAIL

Current Position:Home > News and Insights
Category: News and Insights Author: Zhongke Xinyuan Content Reviewer: Zhongke Xinyuan Review Published: 2025-01-22 Updated: 2026-07-22 Source: Existing page; verify sources
NVIDIA Blackwell MLPerf Training v4.1 LLM Performance Results

NVIDIA Blackwell delivered higher normalized per-GPU performance than NVIDIA Hopper in NVIDIA's MLPerf Training v4.1 submissions. The source reports a 2.0x improvement for GPT-3 pretraining and a 2.2x improvement for Llama 2 70B LoRA fine-tuning, based on NVIDIA's Blackwell submissions. These results are relevant to teams assessing infrastructure for large-language-model training or task-specific model adaptation, but they should be treated as benchmark-specific evidence rather than an estimate of every production workload.

What the MLPerf Training v4.1 results show

MLPerf Training includes an LLM pretraining benchmark based on GPT-3 and an LLM fine-tuning benchmark using LoRA, a parameter-efficient fine-tuning method, on Llama 2 70B. NVIDIA reported Blackwell results across every benchmark in this MLPerf Training round.

Benchmark workloadReported Blackwell per-GPU improvement versus latest H100 performance
LLM LoRA fine-tuning2.2x
LLM pretraining2.0x
Graph neural network2.0x
Text-to-image1.7x
Recommendation system1.6x
Object detection1.6x
Natural language processing1.4x

For GPT-3 175B, the source states that an HGX B200 configuration could run the benchmark using 64 GPUs without reducing per-GPU performance. It contrasts this with a 256-GPU HGX H100 submission size used to achieve optimal per-GPU performance. This indicates that GPU memory capacity, memory bandwidth, and parallel mapping can materially affect the efficient scale selected for a training run.

Platform elements described in the submission

The reported systems contained eight Blackwell GPUs per system, with a runtime thermal design power of 1,000W per GPU. Fifth-generation NVLink and NVLink Switch were used inside the nodes, while NVIDIA ConnectX-7 SuperNICs and NVIDIA Quantum-2 InfiniBand switches connected nodes.

The source attributes the results to both platform and software work. It identifies optimized GEMM, convolution, and multi-head attention kernels; improved overlap of computation and communication in multi-GPU execution; higher memory-bandwidth utilization; and parallel mappings enabled by larger HBM capacity. It also references cuBLAS, cuDNN Runtime Fusion Engines, cuDNN kernels, and NVIDIA Transformer Engine. A procurement or engineering evaluation should therefore assess the full hardware, network, library, framework, and model configuration rather than evaluating the GPU in isolation.

Suitable evaluation scenarios

These benchmark results are most applicable when an organization needs to compare platforms for large-scale LLM pretraining or LoRA-based customization of an existing model. Faster fine-tuning may matter where model variants must be prepared for specific tasks, while stronger per-GPU pretraining performance may influence cluster sizing for foundation-model development.

  • Evaluate GPT-style pretraining where model scale, HBM capacity, bandwidth, and multi-GPU communication are material constraints.
  • Evaluate LoRA fine-tuning where the objective is to adapt a pretrained LLM for a defined task.
  • Assess multi-node designs that use NVLink within a node and InfiniBand between nodes.
  • Compare the cluster footprint, power design, networking, and software environment required for the intended training workflow.

How to validate a deployment decision

  1. Define the target model, parameter count, dataset, sequence lengths, precision settings, and acceptance metric.
  2. Obtain dated official NVIDIA documentation and a complete system SKU/BOM to confirm GPU, memory, interconnect, server, and network details.
  3. Run a representative training or fine-tuning test using the intended framework and software versions.
  4. Measure time to the required training quality, scaling behavior, memory headroom, network utilization, and operational power requirements.
  5. Compare results with an existing Hopper environment using the same workload definition and evaluation method.

FAQ

Do the reported 2.0x and 2.2x gains apply to every LLM workload?

No. The figures are per-GPU results reported for specific MLPerf Training v4.1 workloads: GPT-3 pretraining and Llama 2 70B LoRA fine-tuning. Actual results depend on the model, data, parallelization approach, software stack, cluster topology, and operating configuration. Validate performance with a project-specific test.

Does the source establish the configuration for a production AI cluster?

It describes NVIDIA's submitted systems and their networking components, but it does not provide a complete deployment design, bill of materials, workload sizing, pricing, availability, or operating requirements. Those details require dated official product documentation, a complete SKU/BOM, and engineering validation for the intended environment.

Conclusion

NVIDIA's MLPerf Training v4.1 submissions position Blackwell as a platform with materially higher reported per-GPU performance than Hopper across the listed benchmarks, including LLM pretraining and LoRA fine-tuning. Use the results as a focused comparison input, then verify the complete system design and reproduce relevant measurements with the models and operating conditions that matter to the project.

After reviewing NVIDIA Blackwell MLPerf Training v4.1 LLM Performance Results, continue with NVIDIA products and networking solutions for related evaluation paths.