Two different units of deployment. The GB200 NVL72 is an integrated 72-GPU liquid-cooled rack that acts like one giant GPU; the H200 is a discrete 700W accelerator you rack eight-to-a-server anywhere. This compares a rack-scale system against a flexible building block.
Every spec side by side. Hard specs mirror the live catalog; confirm exact configuration at quote.
| Specification | GB200 NVL72 | H200 |
|---|---|---|
| Architecture | Blackwell — GB200 superchip (2× B200 + Grace) | Hopper (GH100, TSMC 4N) |
| GPUs per unit | 72 Blackwell GPUs + 36 Grace CPUs (per rack) | 1 (8 per HGX/DGX server) |
| Memory | Up to ~13.5 TB HBM3e + ~17 TB LPDDR5X (rack) | 141 GB HBM3e |
| Memory bandwidth | 576 TB/s aggregate HBM (rack) | 4.8 TB/s |
| FP8 tensor | 720 PFLOPS training (rack) | 3,958 TFLOPS (with sparsity) |
| FP16 / BF16 tensor | 360 PFLOPS (rack) | 1,979 TFLOPS (with sparsity) |
| NVLink | 130 TB/s all-to-all (72-GPU NVLink domain) | 900 GB/s (4th-gen NVLink) |
| Board power (TDP) | ~120 kW per rack | 700 W |
| Form factor | NVL72 integrated rack (liquid-cooled) | SXM5 |
| Cooling | Integrated direct-to-chip liquid | Air or direct-to-chip liquid |
| Typical lead time | Configured build — confirmed at quote | 10–14 wk |
| Indicative price | Project quote (rack-scale system) | From $31,000 / unit |
The decision comes down to workload, facility and timeline. Here is the call by scenario.
One 72-GPU NVLink domain removes cross-node bottlenecks for the largest training and real-time inference — roughly an order-of-magnitude inference speedup on giant models versus discrete Hopper.
Deploy from a single server upward, air- or liquid-cooled, into existing halls. The right pick for most training and inference that fits within a handful of GPUs.
A 700W GPU slots into standard racks and cooling; the ~120kW NVL72 rack needs integrated liquid cooling and facility-scale power. Match the unit to your facility.
Open either product, add it to your quote list — free, no commitment — or size the facility with our calculators first.
The GB200 NVL72 is a rack-scale system — 72 Blackwell GPUs and 36 Grace CPUs in one NVLink domain, liquid-cooled at ~120kW. The H200 is a discrete 700W Hopper GPU you deploy eight-to-a-server. One is a system, the other a building block.
There is no clean one-to-one: an NVL72 packs 72 Blackwell GPUs into a single high-bandwidth NVLink domain, so beyond raw GPU count it removes the cross-node communication limits that cap large-model training and inference on discrete GPUs. For the largest models the NVL72 delivers throughput a same-count H200 deployment cannot match.
Yes — the NVL72 uses integrated direct-to-chip liquid cooling at ~120kW per rack. The H200, at 700W, can be air-cooled, which is why it fits facilities that are not yet liquid-ready.
Our solutions engineers size compute, power and cooling together and confirm real lead times — no payment, no commitment, quotes back in about one business day.
Our engineers scope power, cooling and compute together. No payment, and quotes come back in about a business day.