A rack that behaves like one giant GPU. The GB200 NVL72 links 72 Blackwell GPUs and 36 Grace CPUs into a single NVLink domain with 130 TB/s of all-to-all bandwidth — an exascale-class inference and training system delivered as an integrated, liquid-cooled ~120kW rack.
Key specifications for the NVIDIA GB200 NVL72. Hard specs mirror the live catalog; confirm exact revision and configuration at quote.
| Specification | GB200 NVL72 |
|---|---|
| Architecture | Blackwell — GB200 superchip (2× B200 + Grace) |
| GPUs per unit | 72 Blackwell GPUs + 36 Grace CPUs (per rack) |
| Memory | Up to ~13.5 TB HBM3e + ~17 TB LPDDR5X (rack) |
| Memory bandwidth | 576 TB/s aggregate HBM (rack) |
| FP8 tensor | 720 PFLOPS training (rack) |
| FP16 / BF16 tensor | 360 PFLOPS (rack) |
| NVLink | 130 TB/s all-to-all (72-GPU NVLink domain) |
| Board power (TDP) | ~120 kW per rack |
| Form factor | NVL72 integrated rack (liquid-cooled) |
| Cooling | Integrated direct-to-chip liquid |
| Typical lead time | Configured build — confirmed at quote |
| Indicative price | Project quote (rack-scale system) |
The NVL72 is for workloads that outgrow a single server: trillion-parameter training and real-time inference on the largest models. By turning a whole rack into one 72-GPU NVLink domain, it lets a model be served or trained as if on one enormous GPU — with an order-of-magnitude jump in inference throughput on giant models. It arrives as an integrated, factory-tested, liquid-cooled rack at roughly 120kW, so the facility conversation is power and heat rejection, not individual boards.
A single 72-GPU NVLink domain removes cross-node bottlenecks for the largest training runs, scaling out across multiple NVL72 racks.
Delivers roughly an order-of-magnitude inference speedup on very large models versus an equivalent Hopper deployment.
Ships as an integrated, liquid-cooled ~120kW rack — plan facility power, CDU/heat-rejection capacity and floor loading rather than per-server builds.
Density decides the facility. Size the electrical capacity, heat rejection and cooling method before you rack a single GB200 NVL72 — our free tools turn the specs above into facility numbers.
Add it to your quote list — free, no commitment — and we confirm a firm, itemized price, availability and lead time after review, usually within a business day.
The GB200 NVL72 is a rack-scale system that connects 72 Blackwell GPUs and 36 Grace CPUs into a single NVLink domain with 130 TB/s of all-to-all GPU bandwidth. It behaves like one very large GPU for training and inference on the biggest models.
A GB200 NVL72 rack draws on the order of 120kW with integrated direct-to-chip liquid cooling. At a 1.4 PUE that is roughly 168kW of facility power per rack before redundancy — an order of magnitude beyond a traditional enterprise rack.
Yes. At ~120kW per rack the NVL72 uses integrated direct-to-chip liquid cooling; air cooling is not an option at this density. Sizing CDUs, coolant loops and heat rejection to the rack is central to any deployment.
It is sourced as a rack-scale system, not a single board — we scope the racks, coolant distribution, power train and network fabric together as one quote. Start from the AI-compute catalog or talk to an engineer to size the deployment.
Our solutions engineers size compute, power and cooling together and confirm real lead times — no payment, no commitment, quotes back in about one business day.
Our engineers scope power, cooling and compute together. No payment, and quotes come back in about a business day.