NVIDIA Hopper · Tensor Core GPU

NVIDIA H200 Tensor Core GPU

The memory-upgraded Hopper. The H200 keeps the H100’s compute and 700W envelope but nearly doubles memory to 141GB of HBM3e at 4.8 TB/s — the most-quoted accelerator for memory-bound inference, long-context serving and larger models that would spill off an H100.

  • Hopper
  • 141GB HBM3e
  • 4.8 TB/s
  • Most quoted

H200 specifications

Key specifications for the NVIDIA H200 Tensor Core GPU. Hard specs mirror the live catalog; confirm exact revision and configuration at quote.

NVIDIA H200 Tensor Core GPU — reference specifications
SpecificationH200
ArchitectureHopper (GH100, TSMC 4N)
GPUs per unit1 (8 per HGX/DGX server)
Memory141 GB HBM3e
Memory bandwidth4.8 TB/s
FP8 tensor3,958 TFLOPS (with sparsity)
FP16 / BF16 tensor1,979 TFLOPS (with sparsity)
NVLink900 GB/s (4th-gen NVLink)
Board power (TDP)700 W
Form factorSXM5
CoolingAir or direct-to-chip liquid
Typical lead time10–14 wk
Indicative priceFrom $31,000 / unit

What the H200 is for

The H200 exists for one reason: memory. Same Hopper compute, same 700W, but 141GB of HBM3e at 4.8 TB/s. That extra capacity and bandwidth is decisive for inference throughput and for fitting larger models or longer contexts onto fewer GPUs — often reducing the GPU count (and the fabric and power around it) for a given serving target.

Large-model & long-context inference

The 141GB / 4.8 TB/s memory holds bigger models and KV caches per GPU, lifting tokens-per-second on memory-bound serving without a fabric change.

Drop-in Hopper upgrade

Same SXM5 socket, NVLink and 700W as the H100 — the same servers, cooling and power designs carry over, so it is the low-risk capacity bump.

Mixed training + inference

Identical compute to the H100 for training, with headroom to consolidate inference onto fewer, larger-memory GPUs.

Power & cooling for the H200

Density decides the facility. Size the electrical capacity, heat rejection and cooling method before you rack a single H200 — our free tools turn the specs above into facility numbers.

Source the H200

Add it to your quote list — free, no commitment — and we confirm a firm, itemized price, availability and lead time after review, usually within a business day.

NVIDIANVIDIA H200 Tensor Core GPU141GB HBM3e · 4.8TB/s · 700W · SXM510–14 wkFrom $31,000 / unit
View full product

H200, answered

How much memory does the NVIDIA H200 have?

The H200 has 141GB of HBM3e with 4.8 TB/s of memory bandwidth — about 76% more capacity and 40% more bandwidth than the H100’s 80GB HBM3 at 3.4 TB/s, on the same Hopper silicon and 700W board.

Is the H200 faster than the H100?

For compute-bound work they are the same — identical GH100 dies deliver the same FP8/FP16 TFLOPS. The H200 wins on memory-bound workloads (inference, long context, larger models), where its extra HBM3e capacity and bandwidth raise real-world throughput substantially.

How much power does an H200 use?

The H200 SXM5 keeps the 700W TDP of the H100, so an 8-GPU HGX H200 server still lands near 10kW. Existing H100-class power and cooling designs carry over directly.

Should I buy the H200 or wait for Blackwell (B200)?

The H200 is air-coolable, available now and drops into H100 infrastructure. The B200 is a generational leap in compute and memory bandwidth but needs direct liquid cooling and higher power. Choose H200 for near-term, brownfield or air-cooled capacity; B200 for frontier training in a liquid-ready hall.

Turn the spec into a scoped quote

Our solutions engineers size compute, power and cooling together and confirm real lead times — no payment, no commitment, quotes back in about one business day.

Want a second opinion on a build?

Our engineers scope power, cooling and compute together. No payment, and quotes come back in about a business day.

Talk to an engineer