Free tool · AI data center

AI Cluster Network Planner

Plan the network fabric for a GPU cluster in seconds. Enter your GPU count, NICs per GPU and oversubscription target and get a leaf/spine switch count, transceiver and MPO fiber-trunk totals — with the 800G switch and fiber that fit.

Cluster fabric

= 1,024 network endpoints
Spectrum-4: 64 × 800G

Leaf/spine (Clos) fabric. Ports split between GPU downlinks and spine uplinks by the oversubscription ratio.

Total switches

48

32 leaf + 16 spine

Endpoints

1,024

1,024 GPUs × 1 NIC

Leaf switches

32

32 GPU ports each

Spine switches

16

≈ leaf ÷ 2 (non-blocking)

Transceivers

4,096

(endpoints + links) × 2

MPO fiber trunks

86

≈ endpoints ÷ 12

A 64-port 800G leaf/spine fabric with MPO fiber trunks — the standard building block for a non-blocking GPU cluster network.

Recommended network equipment

Browse all network & fiber →

Assumptions: leaf/spine (Clos) fabric; endpoints = GPUs × NICs/GPU; leaf downlinks = ports × oversub ÷ (oversub + 1) (rest are spine uplinks); leaf switches = ceil(endpoints ÷ downlinks); spine switches ≈ ceil(leaf ÷ 2); transceivers = (endpoints + inter-switch links) × 2; MPO fiber trunks ≈ ceil(endpoints ÷ 12). Excludes redundancy and management/storage networks. Planning estimate only — confirm the design with a solutions engineer before purchase.

How the AI cluster network planner works

The network fabric is what turns a pile of GPUs into a cluster. Modern AI training needs a non-blocking, low-latency fabric so every GPU can talk to every other at full bandwidth. This planner sizes that leaf/spine fabric — switches, optics and fiber — from a GPU count and a few design choices.

Each step is a simple, labelled calculation so you can see exactly where the numbers come from and adjust them to your design.

  1. EndpointsGPU count × NICs per GPU
  2. Leaf downlinksports × oversub ÷ (oversub + 1) — rest are spine uplinks
  3. Leaf switchesceil(endpoints ÷ leaf downlinks)
  4. Spine switches≈ ceil(leaf switches ÷ 2) — rough non-blocking
  5. Transceivers(endpoints + inter-switch links) × 2
  6. MPO fiber trunks≈ ceil(endpoints ÷ 12)

Leaf/spine design for GPU clusters

A Clos (leaf/spine) fabric is the standard for AI clusters because it scales bandwidth predictably and keeps every GPU the same small number of hops from every other. Each GPU NIC lands on a leaf (top-of-rack) switch; leaf switches connect up to a layer of spine switches. The whole design turns on one choice: how you split each leaf switch\'s ports between downlinks to GPUs and uplinks to the spine.

That split is oversubscription. At 1:1 — non-blocking — half the leaf ports face the GPUs and half go up to the spine, so there is never contention; this is what large-scale training needs. At 2:1 or 4:1 you dedicate more ports to GPUs and fewer to the spine, which cuts switch and cable count but introduces contention that can slow collective operations. With 800G NICs and 64-port 800G switches, a non-blocking fabric puts 32 GPUs on each leaf and ties the leaves together through the spine.

Optics and fiber follow the switch count. Every endpoint link and every inter-switch link needs a transceiver at each end, and the runs are carried on MPO fiber trunks. These are long-lead, high-value items, so sizing them early — alongside compute, power and cooling — is what keeps the network from becoming the thing that gates the cluster. The recommended switch and fiber link straight into our catalog.

Oversubscription reference

How the oversubscription ratio splits a 64-port leaf switch between GPU downlinks and spine uplinks, and what it means for the fabric. Planning bands; confirm against your switch and traffic pattern.

Port split by oversubscription (64-port leaf)
RatioGPU downlinksSpine uplinksUse
1:13232Non-blocking — large-scale training
2:14222Cost-optimised — mixed workloads
4:15113Dense — inference / tolerant traffic

AI cluster networking, answered

How do you design a GPU cluster network?

AI clusters use a leaf/spine (Clos) fabric. Each GPU NIC is an endpoint that connects to a leaf (top-of-rack) switch; leaf switches connect up to spine switches. You size it by splitting each leaf switch's ports between downlinks to GPUs and uplinks to the spine according to your oversubscription ratio, then count how many leaf and spine switches the endpoint total needs.

How many switches does a GPU cluster need?

It scales with GPUs and ports per switch. With a 64-port 800G leaf at 1:1 (non-blocking), half the ports face the GPUs, so each leaf serves 32 endpoints: leaf switches = ceil(endpoints ÷ 32). Spine switches are roughly ceil(leaf switches ÷ 2). A 1,024-GPU single-NIC cluster needs about 32 leaf and 16 spine switches — 48 total — before redundancy.

What is oversubscription in a leaf/spine fabric?

Oversubscription is the ratio of downlink (to GPUs) to uplink (to spine) bandwidth on a leaf switch. 1:1 is non-blocking — every GPU can talk to every other at full rate, which AI training needs. 2:1 or 4:1 dedicate more ports to GPUs and fewer to the spine, cutting switch and cable count at the cost of contention. This planner lets you compare all three.

How do you size an 800G AI fabric?

Start from endpoints = GPUs × NICs per GPU, using 400/800G NICs on a switch like the 64-port Spectrum-4 800G. Split ports by oversubscription, count leaf and spine switches, then size the optics and fiber: transceivers ≈ (endpoints + inter-switch links) × 2, and MPO fiber trunks ≈ ceil(endpoints ÷ 12). This tool computes all of them and links the switch and fiber.

Turn your fabric plan into a scoped quote

Our solutions engineers size the network alongside compute, power and cooling and confirm real lead times — no payment, no commitment, quotes back in about one business day.

Want a second opinion on a build?

Our engineers scope power, cooling and compute together. No payment, and quotes come back in about a business day.

Talk to an engineer