Rack-scale

GB200 NVL72

The GB200 NVL72 is a rack, not a node: 72 Blackwell GPUs in one coherent NVLink 5 domain, 13.4 TB of HBM3e, direct-to-chip liquid cooling at 130 kW.

Specification

GPU memory
186 GB HBM3e/GPU · 13.4 TB per rack
Intra-node
72-GPU coherent NVLink 5 domain
Inter-node
Quantum-X800 XDR between racks
Node config
Whole or half rack · liquid-cooled
Best for
400B+ dense / MoE, trillion-param

Inside the rack: a 72-GPU coherent NVLink 5 domain at 1.8 TB/s per GPU. Between racks: Quantum-X800 InfiniBand XDR at 800G per link, 1:1, dedicated per tenant.

Why this machine

The point of NVL72 is that 72 GPUs behave like one very large accelerator. Model-parallel traffic that would cross an InfiniBand hop on an 8-GPU node stays inside a coherent NVLink 5 domain at 1.8 TB/s per GPU instead — which is what makes trillion-parameter dense and MoE models practical rather than merely possible.

Rack-scale is also where operations stop being incidental. These racks are liquid-cooled at 130 kW in halls we designed and run ourselves, instrumented down to per-GPU power draw, at a PUE of 1.12 annualized. Capacity is quoted per rack — whole or half — because supply is the constraint, not pricing policy.

What it is for

  • 400B+ dense and MoE training

    Model-parallel traffic that stays inside the NVLink domain rather than crossing the network on every step.

  • Trillion-parameter runs

    One coherent 72-GPU domain per rack, with Quantum-X800 XDR carrying traffic between racks.

  • Reasoning-heavy inference

    Very large models served where interconnect, not arithmetic, sets the achievable batch and latency.

GB200 NVL72 questions

  • What does a GB200 NVL72 rack cost?

    Rack-scale systems are quoted per rack rather than listed, because the constraint is supply rather than secrecy. Send GPU count, interconnect and timeline and a firm quote comes back within 48 hours with a delivery date.

  • Can I reserve less than a full rack?

    Whole or half rack. The NVL72 is a single coherent domain, so it is allocated in rack units rather than 8-GPU nodes.

  • GB200 or GB300?

    GB300 carries 288 GB of HBM3e per GPU and 20.7 TB per rack, against the GB200’s 186 GB and 13.4 TB. Both are quoted per rack. GB300 is the answer for frontier MoE and reasoning-heavy inference; GB200 for 400B+ dense and MoE training.

  • How is a rack-scale cluster accepted?

    The same way as every other cluster: a 72-hour burn-in with NCCL all-reduce sweeps, HBM bandwidth checks and thermals under sustained load, and a report you sign off before you pay.