System comparison · Rack-scale vs discrete

NVIDIA GB200 NVL72 vs H200

Two different units of deployment. The GB200 NVL72 is an integrated 72-GPU liquid-cooled rack that acts like one giant GPU; the H200 is a discrete 700W accelerator you rack eight-to-a-server anywhere. This compares a rack-scale system against a flexible building block.

  • Unit72-GPU NVLink rackSingle discrete GPUsystem vs board
  • GPU interconnect130 TB/s NVLink domain900 GB/s NVLinkrack-wide vs node
  • Power~120 kW / rack700W / GPUfacility vs rack scale
  • CoolingIntegrated liquidAir or liquidDLC required for NVL72

GB200 NVL72 vs H200 — specifications

Every spec side by side. Hard specs mirror the live catalog; confirm exact configuration at quote.

NVIDIA GB200 NVL72 vs NVIDIA H200 Tensor Core GPU
SpecificationGB200 NVL72H200
ArchitectureBlackwell — GB200 superchip (2× B200 + Grace)Hopper (GH100, TSMC 4N)
GPUs per unit72 Blackwell GPUs + 36 Grace CPUs (per rack)1 (8 per HGX/DGX server)
MemoryUp to ~13.5 TB HBM3e + ~17 TB LPDDR5X (rack)141 GB HBM3e
Memory bandwidth576 TB/s aggregate HBM (rack)4.8 TB/s
FP8 tensor720 PFLOPS training (rack)3,958 TFLOPS (with sparsity)
FP16 / BF16 tensor360 PFLOPS (rack)1,979 TFLOPS (with sparsity)
NVLink130 TB/s all-to-all (72-GPU NVLink domain)900 GB/s (4th-gen NVLink)
Board power (TDP)~120 kW per rack700 W
Form factorNVL72 integrated rack (liquid-cooled)SXM5
CoolingIntegrated direct-to-chip liquidAir or direct-to-chip liquid
Typical lead timeConfigured build — confirmed at quote10–14 wk
Indicative priceProject quote (rack-scale system)From $31,000 / unit

Which should you choose?

The decision comes down to workload, facility and timeline. Here is the call by scenario.

GB200 NVL72Trillion-parameter scale

One 72-GPU NVLink domain removes cross-node bottlenecks for the largest training and real-time inference — roughly an order-of-magnitude inference speedup on giant models versus discrete Hopper.

H200Node-level & flexible deployment

Deploy from a single server upward, air- or liquid-cooled, into existing halls. The right pick for most training and inference that fits within a handful of GPUs.

H200Brownfield & lower power density

A 700W GPU slots into standard racks and cooling; the ~120kW NVL72 rack needs integrated liquid cooling and facility-scale power. Match the unit to your facility.

Source the GB200 NVL72 or H200

Open either product, add it to your quote list — free, no commitment — or size the facility with our calculators first.

NVIDIANVIDIA GB200 NVL72Project quote
Details
NVIDIANVIDIA H200 Tensor Core GPUFrom $31,000 / unit
Details

GB200 NVL72 vs H200, answered

What is the difference between the GB200 NVL72 and the H200?

The GB200 NVL72 is a rack-scale system — 72 Blackwell GPUs and 36 Grace CPUs in one NVLink domain, liquid-cooled at ~120kW. The H200 is a discrete 700W Hopper GPU you deploy eight-to-a-server. One is a system, the other a building block.

How many H200 GPUs equal a GB200 NVL72?

There is no clean one-to-one: an NVL72 packs 72 Blackwell GPUs into a single high-bandwidth NVLink domain, so beyond raw GPU count it removes the cross-node communication limits that cap large-model training and inference on discrete GPUs. For the largest models the NVL72 delivers throughput a same-count H200 deployment cannot match.

Do I need liquid cooling for the GB200 NVL72?

Yes — the NVL72 uses integrated direct-to-chip liquid cooling at ~120kW per rack. The H200, at 700W, can be air-cooled, which is why it fits facilities that are not yet liquid-ready.

Not sure which fits your build?

Our solutions engineers size compute, power and cooling together and confirm real lead times — no payment, no commitment, quotes back in about one business day.

Want a second opinion on a build?

Our engineers scope power, cooling and compute together. No payment, and quotes come back in about a business day.

Talk to an engineer