The Network is the Bottleneck
Network topology is the single largest determinant of multi-node training throughput. A poorly configured topology can reduce model flops utilization (MFU) from 55% to under 20% at 64 GPUs. The choice between ring, tree, and fully-connected topologies determines AllReduce bandwidth, latency variance, and fault tolerance characteristics.
Three trends make topology more critical in 2026: model parallelism increases inter-GPU communication volume, longer sequence lengths in long-context models increase AllReduce payload sizes, and the gap between compute throughput and network bandwidth per GPU widens with each generation. Hopper GPUs compute at 989 FP16 TFLOPS, but inter-node bandwidth per GPU has only grown from 200 to 400 Gbps over the same period.
Ring Topology: The Default Choice
Ring topology connects each GPU to two neighbors in a circular chain. NCCL implements ring AllReduce by partitioning data across the ring and passing partial sums around the circuit. Ring achieves near-linear bandwidth scaling with GPU count under ideal conditions because each GPU only sends and receives from two peers regardless of total ring size.
The practical limit for ring-based AllReduce on InfiniBand is around 256 GPUs. Beyond that, latency accumulation from sequential data transfer phases begins to dominate throughput. Ring also suffers from single-link failure: one broken connection splits the ring and forces NCCL tree fallback, which reduces bandwidth by 25-40% until the fabric is repaired.
Tree (Leaf-Spine) Topology
Tree topology reduces communication steps from N-1 to log2(N) for AllReduce using recursive doubling. NCCL automatically switches to tree for GPU counts above 512 or when ring bandwidth degrades. Tree topology is also the default for hierarchical reduce-scatter and all-gather operations when the InfiniBand fabric supports SHARP in-network computing.
Most production InfiniBand clusters use a hybrid approach: NVLink within the node and fat-tree across nodes. The intra-node NVLink domain provides 900 GB/s for H100, while inter-node InfiniBand runs at 200-400 Gbps per NIC. This 18:1 or 36:1 bandwidth ratio between intra-node and inter-node communication makes topology optimization at the rack level essential for scaling efficiency.
Fully-Connected Topology
Fully-connected topology theoretically provides optimal AllReduce bandwidth but does not scale past 8-16 nodes due to physical port limits. NVIDIA NVSwitch achieves full bisection bandwidth for 8 GPUs on a single HGX baseboard, and NVSwitch 3rd gen extends this to 256 GPUs within a single NVLink domain.
Beyond one NVLink domain, fully-connected topologies become physically impossible. A 64-node cluster with 8 GPUs each would need 512 ports per node. The practical maximum for any-to-any connectivity is one NVLink domain (256 GPUs with NVSwitch 3rd gen). Beyond that, hierarchical topologies combining NVLink and InfiniBand are mandatory.
NCCL Performance by Topology
NCCL benchmarks show that hybrid topologies combining NVLink and tree outperform pure-ring configurations by 30-100% at scale. The advantage grows with GPU count as ring latency accumulates. At 256 GPUs, hybrid NVLink+Tree delivers 2.3x the AllReduce bandwidth of pure ring on InfiniBand.
The table below shows NCCL AllReduce bandwidth measurements across common configurations. Data collected using nccl-tests with 256 MB message size on H100 clusters with InfiniBand NDR400 fabric.
| Topology | 8 GPUs | 64 GPUs | 256 GPUs |
|---|---|---|---|
| Ring (InfiniBand NDR400) | 220 GB/s | 185 GB/s | 120 GB/s |
| Tree (Leaf-Spine NDR400) | 210 GB/s | 175 GB/s | 155 GB/s |
| NVLink Domain (H100) | 880 GB/s | 780 GB/s | 520 GB/s |
| Hybrid NVLink+Ring | 880 GB/s | 210 GB/s | 140 GB/s |
| Hybrid NVLink+Tree | 880 GB/s | 350 GB/s | 280 GB/s |
Network Bandwidth Requirements
The bandwidth a model requires depends on the parallelism strategy. Data parallelism requires AllReduce of gradients equal to model size per step. Tensor parallelism needs all-gather and reduce-scatter per transformer layer. Pipeline parallelism uses point-to-point activations per microbatch. The total requirement is the maximum of these patterns, not the sum.
For large-scale training using 3D parallelism (DP+TP+PP), tensor-parallel communication is the most bandwidth-intensive and typically determines the minimum network configuration. The table below shows minimum network bandwidth requirements for common model sizes using recommended parallelism configurations.
| Model Size | DP Only | 3D Parallel | Min Network |
|---|---|---|---|
| 7B params | 2 Gbps | 8 Gbps | 200 Gbps |
| 70B params | 20 Gbps | 80 Gbps | 400 Gbps |
| 405B params | 120 Gbps | 400 Gbps | 800 Gbps |
| 1T params (MoE) | - | 800+ Gbps | 1.6 Tbps |
NVLink vs InfiniBand for Multi-Node
NVLink provides 900 GB/s bidirectional bandwidth per GPU (H100) for intra-node communication. InfiniBand NDR400 provides 50 GB/s per port. The 18:1 ratio drives the hierarchical communication strategy used by all production training frameworks. NVLink 5 on B200/B300 increases per-GPU bandwidth to 1.8 TB/s, widening the gap further.
InfiniBand BWR (800 Gbps) arriving in late 2026 will improve the per-port ratio to 9:1, but NVLink continues to scale faster each generation. Ethernet with RoCE v2 remains an option for cost-sensitive clusters but adds 15-25% latency overhead for collective operations compared to InfiniBand.
The decision framework: use NVLink for intra-node and InfiniBand for inter-node in clusters up to 1024 GPUs. Switch to InfiniBand-only or HPE Slingshot for clusters above 4096 GPUs, where NVLink domain boundaries and hierarchical topology coordination become the dominant complexity factor.
