NVIDIA Developer Blog→ original

NVIDIA Exemplar Cloud: почему одинаковые AI-кластеры теряют до 12% производительности

NVIDIA Exemplar Cloud: два AI-кластера на идентичных H100 или GB200 NVL72 регулярно показывают разрыв в 8–12% по пропускной способности обучения при одинаковой задаче и модели. Причина — не в железе, а в стеке конфигурационных решений на уровне ядра ОС. Программа Exemplar Cloud помогает облачным провайдерам закрыть этот разрыв, сравнивая их развёртывания с эталонной архитектурой NVIDIA.

AI-processed from NVIDIA Developer Blog; edited by Hamidun News
NVIDIA Exemplar Cloud: почему одинаковые AI-кластеры теряют до 12% производительности
Source: NVIDIA Developer Blog. Collage: Hamidun News.
◐ Listen to article

NVIDIA published a breakdown of the Exemplar Cloud program on Developer Blog: two AI clusters on identical H100, GB200 NVL72, or GB300 NVL72 accelerators regularly show a training throughput gap of 8% to 12% — with the same task, the same model, and the same global batch size.

Why identical clusters produce different results

The cause of the gap is not in the hardware, but in configuration choices. According to NVIDIA Developer Blog, the 8% to 12% discrepancy between partner deployments and the company's Reference Architecture (RA) arises from a stack of configuration decisions at the operating system kernel level. Two clouds with identical H100 or GB200 NVL72 GPUs can show fundamentally different training throughput if their software layers are configured differently from what NVIDIA recommends.

To understand the scale: in a cluster of a thousand accelerators, a 10% gap equals dozens of hours of compute time per week. At market GPU rental prices, this means losses of hundreds of thousands of dollars in productivity per year.

  • Systems compared: NVIDIA H100, GB200 NVL72, GB300 NVL72
  • Recorded gap: 8% to 12% in training throughput
  • Condition: same task, same model, same global batch size
  • Source of the gap: "a stack of configuration decisions at the kernel level," according to NVIDIA Developer Blog

What is NVIDIA Exemplar Cloud

Exemplar Cloud is an NVIDIA partner program created to help cloud providers achieve Reference Architecture performance. The company analyzes specific partner deployments, compares them to the RA, and identifies configuration "bottlenecks" that quietly reduce cluster efficiency. Typical optimization areas include network stack parameters, NVLink and InfiniBand settings, Linux kernel configuration, and NUMA topology.

The program is particularly relevant right now: GB200 NVL72 and GB300 NVL72 are NVIDIA's latest racks based on the Blackwell architecture, which are being massively deployed by the world's largest providers in 2025–2026. Competition for AI clients is conducted at the level of real throughput, and a 10% gap between two identically equipped clouds can determine the outcome of a major tender.

"We consistently observe a gap of 8% to 12% between partner installations and our

Reference Architecture on the same task, the same model, and the same global batch size," the NVIDIA Developer Blog publication states.

What this means

For ML engineers and DevOps teams managing GPU clusters, the key takeaway is simple: having the right hardware is a necessary but not sufficient condition. The 8–12% of performance lost to configuration represents real money and real model training time. Exemplar Cloud offers a systematic approach: compare your deployment against NVIDIA's reference and eliminate specific bottlenecks, rather than guessing why throughput is below the theoretical maximum. For cloud providers, achieving RA metrics becomes both a competitive advantage and an operational savings.

Frequently Asked Questions

What is NVIDIA Reference Architecture?

NVIDIA Reference Architecture (RA) is the reference hardware and software stack configuration that NVIDIA publishes for H100, GB200 NVL72, and GB300 NVL72 systems. It serves as a baseline: partner deployment performance is compared against the RA to identify specific discrepancies.

Why is the 8–12% gap critical for cloud AI?

In a cluster of hundreds or thousands of GB200 NVL72 accelerators, a 10% gap in training throughput means an equivalent 10% in extra compute costs — or the same amount of lost GPU rental revenue compared to a competitor on identical systems but with a better-configured stack.

ZK
Hamidun News
AI news without noise. Daily editorial selection from 50+ sources. A product by Zhemal Khamidun, Head of AI at Alpina Digital.

Want to stop reading about AI and start using it?

AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.

What do you think?
Loading comments…