Nvidia Blackwell Trains Llama 3.1 405B in 7 Minutes on Azure

Share:
Audio Loading voice…
Nvidia Blackwell Trains Llama 3.1 405B in 7 Minutes on Azure

Synopsis

Nvidia's corporate account highlighted a record-scale MLPerf Training submission by Microsoft Azure using 8,192 Blackwell GB200 NVL72 GPUs, achieving the Llama 3.1 405B training target in 7.07 minutes — one of the largest such benchmarks to date.

Key Takeaways

8,192 NVIDIA GB200 NVL72 GPUs were deployed on Microsoft Azure for the MLPerf Training submission.
The Llama 3.1 405B training target was completed in just 7.07 minutes , one of the fastest results at this scale.
The submission is described as one of the largest MLPerf Training submissions to date on Nvidia Blackwell hardware.
Nvidia's Blackwell GB200 NVL72 is a rack-scale system with 72 GPUs per rack, succeeding the Hopper architecture announced in March 2024 .
Nvidia and Microsoft have a multi-year Azure partnership covering large-scale GPU deployment for frontier AI model training.
Nvidia's post signals further announcements are forthcoming, with the phrase 'More to come.'

Chip giant Nvidia on Wednesday, June 17, 2026, publicly praised Microsoft chief executive Satya Nadella and the Azure cloud platform after the two companies completed one of the largest MLPerf Training submissions to date, deploying 8,192 NVIDIA GB200 NVL72 Blackwell-generation GPUs to train Meta's Llama 3.1 405B model in just 7.07 minutes.

Context

The post, shared from Nvidia's official corporate account, reads: 'Great work with Azure on one of the largest MLPerf Training submission to-date on NVIDIA Blackwell: 8,192 GPUs on NVIDIA GB200 NVL72 systems, Llama 3.1 405B training target met in only 7.07 minutes.' The benchmark result was submitted to MLCommons, the industry body that runs the MLPerf suite — widely regarded as the most credible public scorecard for AI training hardware.

The GB200 NVL72 is a rack-scale system built around Nvidia's Blackwell architecture, announced in March 2024 as the successor to the Hopper generation. Each NVL72 rack integrates 72 Blackwell GPUs with high-bandwidth NVLink interconnects, designed specifically for the extreme memory and compute demands of large language model training at scale.

Policy Backdrop

Nvidia and Microsoft deepened their multi-year Azure partnership across 2023 and 2024, committing to deploy large Nvidia GPU clusters for OpenAI and other frontier model developers. That relationship now extends to Blackwell-generation infrastructure, with Azure acting as one of the first hyperscalers to field GB200 NVL72 systems at benchmark-ready scale.

Successive MLPerf Training rounds have evolved into a competitive arena where cloud operators and hardware vendors jointly publish results to signal readiness for enterprise AI workloads. Each new submission round effectively sets a new performance baseline that competitors must match or exceed to remain credible in the hyperscale GPU market.

Stakeholders and Impact

For hyperscale cloud providers — including Google Cloud, Amazon Web Services, and Microsoft Azure — MLPerf results directly influence enterprise procurement decisions, particularly for organisations training or fine-tuning large foundation models. A training time of 7.07 minutes for a 405-billion-parameter model at this cluster size sets a new reference point for time-to-train economics.

For AI model developers and research labs, faster training cycles translate into shorter iteration loops, lower compute costs per experiment, and the practical ability to scale model complexity further. The result also reinforces Nvidia's continued dominance in the AI accelerator market at a moment when rivals are investing heavily in alternative chip architectures.

What's Next

Nvidia's post closes with the phrase 'More to come,' signalling further benchmark submissions or capability announcements tied to the Blackwell platform. Analysts and cloud customers will watch for the next MLPerf Training submission round and the commercial availability timeline for GB200 NVL72 capacity on Azure through 2026.

The broader race among hyperscalers to claim leadership in large-scale LLM training infrastructure shows no sign of slowing, with each MLPerf cycle raising the floor for what is considered production-grade AI compute at scale.

Point of View

Memorable data point that will circulate in enterprise procurement conversations. For India's rapidly expanding cloud and AI ecosystem — where hyperscalers are racing to build GPU capacity — results like this set the performance benchmarks that domestic cloud buyers and AI startups will use to evaluate infrastructure choices. The broader arc here is Nvidia cementing its position as the indispensable layer of the AI stack, with each successive benchmark making the case that alternatives remain a generation behind.
NationPress
1 Aug 2026

Frequently Asked Questions

What is MLPerf Training and why does it matter?
MLPerf Training is an industry-standard benchmark suite run by MLCommons that measures how quickly different hardware and software combinations can train AI models to a set accuracy target. It matters because it provides a publicly verifiable, apples-to-apples comparison across chip vendors and cloud platforms, making it a key reference for enterprise AI infrastructure decisions.
What is the NVIDIA GB200 NVL72 system?
The NVIDIA GB200 NVL72 is a rack-scale AI computing system built on Nvidia's Blackwell GPU architecture, housing 72 Blackwell GPUs per rack connected by high-bandwidth NVLink interconnects. It was designed for the extreme memory and compute requirements of training very large language models.
How fast did Azure train Llama 3.1 405B on Nvidia Blackwell GPUs?
Microsoft Azure trained Meta's Llama 3.1 405B model to its MLPerf target in just 7.07 minutes using 8,192 NVIDIA GB200 NVL72 Blackwell GPUs, according to Nvidia's official post on June 17, 2026.
What is the relationship between Nvidia and Microsoft Azure?
Nvidia and Microsoft have a multi-year partnership under which Azure deploys large clusters of Nvidia GPUs — including the latest Blackwell-generation hardware — to power AI training workloads for clients including OpenAI and other frontier model developers.
What does this MLPerf result mean for AI infrastructure in India?
For India's growing cloud and AI sector, results like this set the performance benchmarks that domestic enterprises, AI startups, and research institutions use when evaluating GPU cloud infrastructure. As hyperscalers expand their India data centre footprint, Blackwell-class hardware availability on Azure will increasingly influence where Indian AI workloads are run.
Nation Press
The Trail

Connected Dots

Tracing the thread behind this story — newest first.

8 Dots
  1. Latest 2 weeks ago
  2. 2 weeks ago
  3. 2 weeks ago
  4. 2 weeks ago
  5. 3 weeks ago
  6. 1 month ago
  7. 1 month ago
  8. 1 month ago
Google Prefer NP
On Google