All essays
TechnicalDEEP DIVEFEB 2026

Distributed Training at 10K+ GPU Scale: Network Topologies, Collective Communication, and the AllReduce Wall

Technical analysis of distributed training at 10,000+ GPU scale: fat tree vs dragonfly vs torus topologies, NCCL ring vs tree vs all-to-all collectives, and how the AllReduce wall determines cluster efficiency.

01

The AllReduce Wall

Scaling distributed training beyond 10,000 GPUs is a communication problem, not a compute problem. At this scale, gradient synchronization via AllReduce becomes the dominant term in per-step time. For a 1 trillion-parameter model trained with full sharding, each step requires synchronizing approximately 2.5 GB of gradient data across all 10,000+ GPUs.

The AllReduce wall describes the point where adding more GPUs does not reduce training time because communication overhead negates the compute gains. For a typical 10,000 GPU cluster using InfiniBand NDR (400 Gbps per link), the all-reduce time for a 70B model gradient is approximately 180 ms in a ring topology, compared to roughly 35 ms of compute per step. This 5:1 communication-to-compute ratio means the cluster achieves approximately 17% theoretical peak utilization on all-reduce bound steps.

02

Network Topology Comparison

The three dominant topologies for 10K+ GPU clusters are fat tree, dragonfly, and 3D torus. Fat tree, used in the majority of H100 clusters, provides full bisection bandwidth between any two nodes at the cost of higher cable count. A 10,000 GPU fat tree using 512-port InfiniBand switches requires approximately 40 switches and 5,000+ optical cables in the spine layer alone.

Dragonfly topology, pioneered by Cray and adopted in the B200 generation, groups switches into groups with all-to-all connectivity between groups. This reduces the cable count by approximately 60% compared to fat tree for the same scale, at the cost of non-uniform latency. The maximum hop count between any two GPUs in a dragonfly network is 3, compared to 5-6 in a fat tree of equivalent size.

3D torus, used in some custom Google TPU v5p and AMD MI300X deployments, connects each node to 6 neighbors in a three-dimensional grid. This minimizes cable count but suffers from congestion on multi-hop paths. Torus performance degrades non-linearly as cluster size increases beyond 4,096 GPUs due to the increased average path length.

PropertyFat Tree (3-level)Dragonfly3D Torus
Cables for 10K GPUs~5,500~2,200~1,500
Max Hop Count5-6312-16
Bisection BWFullFullPartial
Latency VariationLowMediumHigh
Routing ComplexitySimpleModerateComplex
Power (Switches)~65 kW~35 kW~20 kW
Adoption (2026)Most H100/H200B200/B300 clustersTPU/MI300X
03

Collective Communication Algorithms

NCCL (NVIDIA Collective Communications Library) implements three primary algorithms for AllReduce: ring, tree, and all-to-all. Ring AllReduce splits the gradient into N chunks and passes them through a logical ring of GPUs, reducing bandwidth usage from O(N) to O(N) total communication, but increasing latency proportional to the ring size.

Tree AllReduce uses a binary-tree reduction where each GPU reduces data from two children before passing results upward. This reduces latency O(log N) but requires more aggregate bandwidth. On clusters with 10,000+ GPUs, the NCCL tree algorithm completes all-reduce approximately 2.5x faster than ring, though at the cost of higher peak bandwidth consumption on root nodes.

All-to-all AllReduce, introduced in NCCL 2.22, excels for small message sizes common in MoE model training where expert-parallel all-to-all communication dominates. For a 10,000 GPU cluster with Mixture-of-Experts routing, all-to-all AllReduce reduces token dispatch time by up to 4x compared to ring AllReduce.

04

Interconnect Technology: InfiniBand vs Spectrum-X

InfiniBand NDR (400 Gbps per link) dominates 10K+ GPU clusters in 2026, with approximately 85% market share. NVIDIA Quantum-2 and Quantum-X switches provide 64 ports at NDR speed with adaptive routing and congestion control. The key limitation is cost: an NDR InfiniBand switch costs approximately $250K, making the network fabric roughly 12-15% of total cluster cost.

Spectrum-X Ethernet, NVIDIA's answer to InfiniBand's cost problem, achieves 400 Gbps with RoCEv2 and adaptive congestion control via the BlueField-3 DPU. At 10,000 GPU scale, Spectrum-X reduces network cost by approximately 25% compared to InfiniBand, with latency within 5-10% of InfiniBand for most collective patterns. The trade-off is higher CPU overhead for protocol processing, even with DPU offload.

05

Scaling Efficiency Measurements

Real-world scaling efficiency at 10,000+ GPU varies dramatically by model architecture and parallelization strategy. For dense transformer models using FSDP with ring AllReduce over InfiniBand, scaling efficiency at 10,240 H100 GPUs is approximately 38% (relative to single-GPU throughput). This drops to 28% at 20,480 GPUs, confirming the AllReduce wall in practice.

MoE models achieve significantly better scaling efficiency. DeepSeek's MoE training at 10,000+ GPU scale reports approximately 55% scaling efficiency, benefiting from the sparse activation pattern that reduces the all-reduce payload size. Expert parallelism combined with all-to-all communication allows MoE training to scale beyond the AllReduce wall that limits dense models.

06

Cluster Design Choices for Cross-10K Scale

The optimal cluster design for 10,000+ GPUs depends on the training workload. For dense transformer training, a fat tree topology with InfiniBand NDR and tree AllReduce minimizes the communication penalty. NVLink domains (256 or 576 GPUs) within the cluster reduce the effective number of nodes for all-reduce, as intra-domain communication uses the much faster NVLink fabric.

For MoE training, dragonfly topology with all-to-all NCCL and Spectrum-X Ethernet offers the best cost-performance ratio. The dragonfly's lower cable count and Ethernet's lower per-port cost make this combination approximately 20-30% cheaper than a fat tree InfiniBand cluster of equivalent scale, with comparable performance on MoE communication patterns.

For multi-tenant clusters serving diverse workloads, a hybrid design with fat tree for dense training partitions and dragonfly for MoE partitions, linked through a shared InfiniBand spine, maximizes flexibility. ClusterBid works with multiple providers that offer both topology options and can route workloads to the optimal fabric based on model architecture.

Filed under
Distributed training10K GPU clusterNetwork topologyNCCL performanceAllReduce wallDragonfly topologyFat tree networkInfiniBand vs Ethernet