Network · AI fabric

AI Data Center Networking for GPU Clusters

AI training lives or dies on the fabric between GPUs. A modern cluster needs a lossless, high-radix Ethernet (or InfiniBand) network built on 800G switches and structured MPO fiber — sized so the interconnect never starves the accelerators it exists to feed.

What it is

AI data center networking is the high-bandwidth, low-latency fabric that connects GPUs across servers and racks into one training cluster. Distributed training exchanges enormous gradient traffic every step, so the network must be non-blocking and lossless — dropped packets stall the whole job. Modern designs use high-radix 800G Ethernet switches (with RoCE v2 for lossless transport) or InfiniBand, arranged in a leaf/spine (Clos) topology so any GPU can reach any other at full bandwidth.

The physical layer is as important as the switches. A 51.2 Tb/s switch with 64× 800G ports terminates hundreds of links, all carried on structured fiber: pre-terminated MPO trunks and OM4 multimode (or single-mode for longer reach) with matched OSFP transceivers. Getting the optics, connector polarity and trunk lengths right against the switch radix is what turns a rack diagram into a working, low-loss fabric.

When you need it

Any multi-node GPU cluster needs a purpose-built fabric. The trigger is scale, bandwidth per GPU and the oversubscription you can tolerate.

  • Multi-node GPU training clusters where GPUs must exchange gradients at full bandwidth every step.
  • Per-GPU bandwidth needs met by 400G/800G NICs and a matching 800G leaf/spine fabric.
  • Lossless transport requirements — RoCE v2 Ethernet or InfiniBand, engineered non-blocking.
  • High port counts: 51.2 Tb/s switches and structured MPO/OM4 fiber to terminate hundreds of links.
  • Oversubscription targets (typically 1:1 for training) that set leaf/spine switch and trunk counts.

Quick facts & specs

Switch radix64× 800G
Throughput51.2 Tb/s
TransportEthernet (RoCE v2)
FiberMPO / OM4 multimode
Key specifications — Spectrum-4 800G Switch
Ports64× 800G OSFP
Throughput51.2 Tb/s
TypeEthernet (RoCE v2)
Latency~1 µs
Form factor2U
CoolingFront-to-back air
Size it

AI Cluster Network Planner

Size a leaf/spine fabric — switch count, transceivers and MPO fiber trunks — from GPU count and oversubscription.

Open the calculator

Source it

Add the equipment to your quote list — pricing, availability and lead times are confirmed after you ask, usually within a business day.

Compare
Corning · Network & Fiber

MPO Fiber Trunk OM4

MPO trunk · OM4 · 24-fiber

6–8 wkNew · Used · Refurb

From $240 / unit

Min order 24 units

Browse all network & fiber equipment

Frequently asked

What makes AI data center networking different?

Distributed GPU training exchanges huge gradient traffic on every step, so the fabric must be non-blocking and lossless — a single dropped packet stalls the whole job. That drives high-radix 800G switches, RoCE v2 (or InfiniBand) for lossless transport, and a leaf/spine topology where any GPU reaches any other at full bandwidth. It is engineered very differently from a general enterprise LAN.

What switches and fiber do I need for a GPU cluster?

A modern AI fabric uses high-radix 800G Ethernet switches — for example a 51.2 Tb/s unit with 64× 800G ports — in a leaf/spine design, cabled with structured MPO fiber trunks and OM4 multimode (single-mode for longer runs) plus matched OSFP transceivers. Switch radix, oversubscription and GPU count set how many leaf and spine switches and trunks you need.

How do I size the network for a GPU cluster?

Work from GPU count, NICs (and bandwidth) per GPU and your oversubscription target (usually 1:1 for training) to the number of leaf and spine switches, transceivers and MPO fiber trunks. Our AI cluster network planner computes the full leaf/spine bill of materials so the fabric matches the compute.

Ethernet (RoCE) or InfiniBand for AI training?

Both deliver lossless, low-latency fabrics. InfiniBand is long-established for HPC/AI; 800G Ethernet with RoCE v2 has become a leading open alternative at the same bandwidths, with a broad multi-vendor ecosystem. The choice hinges on your software stack, operational familiarity and scale — we scope switches, optics and fiber for either.

Want a second opinion on a build?

Our engineers scope power, cooling and compute together. No payment, and quotes come back in about a business day.

Talk to an engineer