Nvidia Blackwell Trains Llama 3.1 405B in 7 Minutes on Azure
Synopsis
Key Takeaways
Chip giant Nvidia on Wednesday, June 17, 2026, publicly praised Microsoft chief executive Satya Nadella and the Azure cloud platform after the two companies completed one of the largest MLPerf Training submissions to date, deploying 8,192 NVIDIA GB200 NVL72 Blackwell-generation GPUs to train Meta's Llama 3.1 405B model in just 7.07 minutes.
Context
The post, shared from Nvidia's official corporate account, reads: 'Great work with Azure on one of the largest MLPerf Training submission to-date on NVIDIA Blackwell: 8,192 GPUs on NVIDIA GB200 NVL72 systems, Llama 3.1 405B training target met in only 7.07 minutes.' The benchmark result was submitted to MLCommons, the industry body that runs the MLPerf suite — widely regarded as the most credible public scorecard for AI training hardware.
The GB200 NVL72 is a rack-scale system built around Nvidia's Blackwell architecture, announced in March 2024 as the successor to the Hopper generation. Each NVL72 rack integrates 72 Blackwell GPUs with high-bandwidth NVLink interconnects, designed specifically for the extreme memory and compute demands of large language model training at scale.
Policy Backdrop
Nvidia and Microsoft deepened their multi-year Azure partnership across 2023 and 2024, committing to deploy large Nvidia GPU clusters for OpenAI and other frontier model developers. That relationship now extends to Blackwell-generation infrastructure, with Azure acting as one of the first hyperscalers to field GB200 NVL72 systems at benchmark-ready scale.
Successive MLPerf Training rounds have evolved into a competitive arena where cloud operators and hardware vendors jointly publish results to signal readiness for enterprise AI workloads. Each new submission round effectively sets a new performance baseline that competitors must match or exceed to remain credible in the hyperscale GPU market.
Stakeholders and Impact
For hyperscale cloud providers — including Google Cloud, Amazon Web Services, and Microsoft Azure — MLPerf results directly influence enterprise procurement decisions, particularly for organisations training or fine-tuning large foundation models. A training time of 7.07 minutes for a 405-billion-parameter model at this cluster size sets a new reference point for time-to-train economics.
For AI model developers and research labs, faster training cycles translate into shorter iteration loops, lower compute costs per experiment, and the practical ability to scale model complexity further. The result also reinforces Nvidia's continued dominance in the AI accelerator market at a moment when rivals are investing heavily in alternative chip architectures.
What's Next
Nvidia's post closes with the phrase 'More to come,' signalling further benchmark submissions or capability announcements tied to the Blackwell platform. Analysts and cloud customers will watch for the next MLPerf Training submission round and the commercial availability timeline for GB200 NVL72 capacity on Azure through 2026.
The broader race among hyperscalers to claim leadership in large-scale LLM training infrastructure shows no sign of slowing, with each MLPerf cycle raising the floor for what is considered production-grade AI compute at scale.