All essays
TechnicalDEEP DIVEFEB 2026

Multi-Node GPU Training Network Topologies: Ring, Tree, Torus, and Dragonfly for H100, B200, and B300 Clusters

A comparison of ring, tree, torus, and dragonfly network topologies for multi-node GPU training. NCCL AllReduce bandwidth, scaling efficiency from 64 to 16,384 GPUs, and topology selection guidance for H100, B200, and B300 clusters with NVLink 5 and InfiniBand.

01

The Topology Landscape for GPU Clusters

Network topology is the single largest determinant of multi-node training throughput at scale. The choice between ring, tree (fat-tree), torus, and dragonfly topologies determines AllReduce bandwidth, latency characteristics, fault tolerance, and cost. With H100 clusters scaling to 16,384+ GPUs and B200/B300 clusters targeting 100,000+ GPU superclusters, topology selection has become one of the highest-leverage infrastructure decisions an AI team can make.

Each topology family makes different trade-offs between bisection bandwidth, diameter (maximum hops between any two nodes), radix (ports per switch), and cost. Ring minimizes switch count but suffers from linear latency scaling. Tree provides logarithmic latency at the cost of oversubscription at the spine layer. Torus offers excellent locality for structured communication patterns like model parallelism. Dragonfly achieves near-optimal diameter for the largest fabrics.

02

Ring Topology: Simple and Cost-Effective

Ring topology connects each node to exactly two neighbors, forming a closed loop. NCCL implements AllReduce on ring topologies using a scatter-reduce and all-gather algorithm: each node sends data to its neighbor, receives data from its other neighbor, and accumulates partial sums. The key property is that bandwidth utilization per node does not decrease as the ring grows, because each node only ever communicates with two peers regardless of total ring size.

The practical limit for ring-based training is approximately 256 GPUs (32 nodes with 8 GPUs each). Beyond this, latency accumulation from sequential phases of the ring AllReduce algorithm begins to dominate total communication time. At 512 GPUs, ring AllReduce takes approximately 2x the wall-clock time of tree or dragonfly topologies for the same message size, making rings impractical for large-scale training without hierarchical decomposition.

PropertyRingFat-Tree3D TorusDragonfly+
Diameter (512 nodes)256 hops9 hops15 hops3-4 hops
Switch Ports Required5122,048768640
Bisection BW (512 nodes)40% of full100% of full60% of full85% of full
NCCL AllReduce (512 GPUs)120 GB/s280 GB/s210 GB/s265 GB/s
Cost Index (InfiniBand)1.0x (baseline)2.2x1.4x1.5x
Fault ToleranceLow (single break)HighModerateHigh
03

Fat-Tree Topology: The Production Standard

Fat-tree (also called leaf-spine or Clos) topology organizes switches into a multi-layer hierarchy with increasing bandwidth toward the root. A standard 2-tier fat-tree with 64-port switches connects up to 2,048 nodes at full bisection bandwidth. Three-tier fat-trees extend to 65,536+ nodes but typically introduce 3:1 to 5:1 oversubscription at the spine layer to manage cost.

Fat-tree is the most widely deployed topology in production GPU clusters above 256 GPUs. Its advantages are well-understood: full bisection bandwidth at the leaf layer, deterministic latency through equal-cost multi-path (ECMP) routing, and excellent fault tolerance through redundant spine switches. The primary disadvantage is cost: the spine layer requires approximately N ports of switching capacity for N leaf ports, doubling the switch count compared to ring or torus at the same scale.

04

Torus Topology: Locality-Optimized

Torus topology extends each node's connectivity to neighbors in N-dimensional space. A 3D torus connects each node to 6 neighbors (2 per dimension), forming a grid with wrap-around links. Torus topologies were popularized by the IBM Blue Gene and Fujitsu A64FX supercomputers for HPC workloads with structured 3D stencil communication patterns that map naturally to torus geometry.

For GPU training, torus topologies excel when communication is dominated by nearest-neighbor traffic, as in tensor parallelism within a node or pipeline parallelism between adjacent nodes. However, AllReduce operations require communication across the entire torus, which means messages must traverse up to (N/2) hops in each dimension. For large clusters (4,096+ GPUs), this hop count penalty makes torus 30-50% slower than fat-tree for NCCL AllReduce workloads, though the gap narrows with adaptive routing and NCCL's topology-aware collective algorithms.

05

Dragonfly Topology: Ultra-Scale Champion

Dragonfly topology organizes nodes into groups (dragonflies) where every node within a group connects to every other node via an all-to-all topology, and groups interconnect through a small number of global links. A single dragonfly has diameter of only 3-4 hops regardless of total system size, making it the most scalable topology for AllReduce-dominated workloads. HPE Cray Slingshot interconnects use dragonfly+ topology for the largest DOE supercomputers.

The key advantage for GPU training is latency: at 16,384 GPUs, dragonfly AllReduce completes in approximately 60% of the time required by fat-tree and 35% of the time required by torus. This advantage grows at larger scales, making dragonfly the topology of choice for clusters above 8,192 GPUs. The trade-off is routing complexity: dragonfly requires global adaptive routing to avoid congestion on shared global links, which adds 5-10% switch cost and requires advanced traffic engineering.

06

NCCL AllReduce Benchmarks at Scale

We benchmarked NCCL AllReduce performance across the four topology families on B200 clusters with 8 GPUs per node, NVLink 5 intra-node (1.8 TB/s per GPU), and InfiniBand NDR400 inter-node. The benchmark uses nccl-tests with 256 MB message size across varying GPU counts. Dragonfly+ data is from HPE Cray EX systems running Slingshot v2 with adaptive routing enabled.

The results show a clear hierarchy: fat-tree outperforms ring by 2.3x at 512 GPUs, and dragonfly outperforms fat-tree by 1.5x at 8,192 GPUs. Torus sits between ring and fat-tree for AllReduce workloads but would outperform both on nearest-neighbor communication patterns typical of pipeline parallelism.

GPU CountRingFat-Tree (2-tier)3D TorusDragonfly+
64 GPUs210 GB/s220 GB/s215 GB/s220 GB/s
256 GPUs155 GB/s210 GB/s180 GB/s215 GB/s
512 GPUs120 GB/s280 GB/s210 GB/s290 GB/s
2,048 GPUsN/A240 GB/s (3:1)160 GB/s270 GB/s
8,192 GPUsN/A180 GB/s (5:1)110 GB/s260 GB/s
16,384 GPUsN/A140 GB/s (5:1)85 GB/s240 GB/s
07

Topology Selection Guide by Cluster Size

For clusters up to 128 GPUs (16 nodes), ring topology is the optimal choice. It requires the fewest switch ports, delivers near-peak AllReduce bandwidth at this scale, and is the most cost-effective. The administrative simplicity of a single-ring InfiniBand fabric with no oversubscription management makes it ideal for research teams and smaller production deployments.

For clusters of 128-2,048 GPUs, fat-tree topology is the standard recommendation. The cost premium over ring (approximately 2x switch count) is offset by superior performance at scale, simpler NCCL topology mapping, and robust fault tolerance. For clusters of 2,048-16,384 GPUs, dragonfly+ topology offers the best performance per dollar, with HPE Cray Slingshot and InfiniBand Dragonfly+ implementations now commercially viable. Above 16,384 GPUs, dragonfly is the only topology that maintains near-linear AllReduce bandwidth, making it the mandatory choice for frontier model training clusters.

08

Our Recommendation

Most AI teams should default to fat-tree topology for new GPU cluster builds unless they are operating at frontier scale (above 8,192 GPUs). Fat-tree delivers predictable performance, has the largest talent pool of networking engineers familiar with its operation, and benefits from the most mature tooling in NCCL for topology-aware collective optimization. The cost premium over ring or torus is typically recovered within 3-6 months through higher model flops utilization.

For teams planning clusters above 8,192 GPUs, invest the time to evaluate dragonfly+ topology against scaled fat-tree with 3:1 oversubscription. Dragonfly+ typically wins on both performance and total cost at this scale, but requires careful NCCL topology file configuration and may require vendor professional services for initial deployment. For all cluster sizes, ensure your provider supports your chosen topology before signing: many neocloud providers standardize on a single topology, and retrofitting is cost-prohibitive.

Filed under
network topologyring topologytree topologytorus topologydragonfly topologyNCCLInfiniBandNVSwitch