All essays
TechnicalDEEP DIVEFEB 2026

Multi-Node GPU Training: Network Topologies and Performance Scaling

Compare ring, tree, and fully-connected network topologies for multi-node GPU training. NCCL performance benchmarks, scaling efficiency from 8 to 1024 GPUs, and NVLink vs InfiniBand tradeoffs.

01

The Network is the Bottleneck

Network topology is the single largest determinant of multi-node training throughput. A poorly configured topology can reduce model flops utilization (MFU) from 55% to under 20% at 64 GPUs. The choice between ring, tree, and fully-connected topologies determines AllReduce bandwidth, latency variance, and fault tolerance characteristics.

Three trends make topology more critical in 2026: model parallelism increases inter-GPU communication volume, longer sequence lengths in long-context models increase AllReduce payload sizes, and the gap between compute throughput and network bandwidth per GPU widens with each generation. Hopper GPUs compute at 989 FP16 TFLOPS, but inter-node bandwidth per GPU has only grown from 200 to 400 Gbps over the same period.

02

Ring Topology: The Default Choice

Ring topology connects each GPU to two neighbors in a circular chain. NCCL implements ring AllReduce by partitioning data across the ring and passing partial sums around the circuit. Ring achieves near-linear bandwidth scaling with GPU count under ideal conditions because each GPU only sends and receives from two peers regardless of total ring size.

The practical limit for ring-based AllReduce on InfiniBand is around 256 GPUs. Beyond that, latency accumulation from sequential data transfer phases begins to dominate throughput. Ring also suffers from single-link failure: one broken connection splits the ring and forces NCCL tree fallback, which reduces bandwidth by 25-40% until the fabric is repaired.

03

Tree (Leaf-Spine) Topology

Tree topology reduces communication steps from N-1 to log2(N) for AllReduce using recursive doubling. NCCL automatically switches to tree for GPU counts above 512 or when ring bandwidth degrades. Tree topology is also the default for hierarchical reduce-scatter and all-gather operations when the InfiniBand fabric supports SHARP in-network computing.

Most production InfiniBand clusters use a hybrid approach: NVLink within the node and fat-tree across nodes. The intra-node NVLink domain provides 900 GB/s for H100, while inter-node InfiniBand runs at 200-400 Gbps per NIC. This 18:1 or 36:1 bandwidth ratio between intra-node and inter-node communication makes topology optimization at the rack level essential for scaling efficiency.

04

Fully-Connected Topology

Fully-connected topology theoretically provides optimal AllReduce bandwidth but does not scale past 8-16 nodes due to physical port limits. NVIDIA NVSwitch achieves full bisection bandwidth for 8 GPUs on a single HGX baseboard, and NVSwitch 3rd gen extends this to 256 GPUs within a single NVLink domain.

Beyond one NVLink domain, fully-connected topologies become physically impossible. A 64-node cluster with 8 GPUs each would need 512 ports per node. The practical maximum for any-to-any connectivity is one NVLink domain (256 GPUs with NVSwitch 3rd gen). Beyond that, hierarchical topologies combining NVLink and InfiniBand are mandatory.

05

NCCL Performance by Topology

NCCL benchmarks show that hybrid topologies combining NVLink and tree outperform pure-ring configurations by 30-100% at scale. The advantage grows with GPU count as ring latency accumulates. At 256 GPUs, hybrid NVLink+Tree delivers 2.3x the AllReduce bandwidth of pure ring on InfiniBand.

The table below shows NCCL AllReduce bandwidth measurements across common configurations. Data collected using nccl-tests with 256 MB message size on H100 clusters with InfiniBand NDR400 fabric.

Topology8 GPUs64 GPUs256 GPUs
Ring (InfiniBand NDR400)220 GB/s185 GB/s120 GB/s
Tree (Leaf-Spine NDR400)210 GB/s175 GB/s155 GB/s
NVLink Domain (H100)880 GB/s780 GB/s520 GB/s
Hybrid NVLink+Ring880 GB/s210 GB/s140 GB/s
Hybrid NVLink+Tree880 GB/s350 GB/s280 GB/s
06

Network Bandwidth Requirements

The bandwidth a model requires depends on the parallelism strategy. Data parallelism requires AllReduce of gradients equal to model size per step. Tensor parallelism needs all-gather and reduce-scatter per transformer layer. Pipeline parallelism uses point-to-point activations per microbatch. The total requirement is the maximum of these patterns, not the sum.

For large-scale training using 3D parallelism (DP+TP+PP), tensor-parallel communication is the most bandwidth-intensive and typically determines the minimum network configuration. The table below shows minimum network bandwidth requirements for common model sizes using recommended parallelism configurations.

Model SizeDP Only3D ParallelMin Network
7B params2 Gbps8 Gbps200 Gbps
70B params20 Gbps80 Gbps400 Gbps
405B params120 Gbps400 Gbps800 Gbps
1T params (MoE)-800+ Gbps1.6 Tbps
Filed under
network topologyNCCLInfiniBandNVLinkmulti-node trainingscaling efficiency