NCCL COLLECTIVE ALGORITHMS: RING, TREE, AND ALL-TO-ALL
NCCL implements three fundamental collective communication algorithms, each optimized for different GPU topology and message size regimes. The ring algorithm partitions data across N GPUs in a logical ring topology, where each GPU sends a chunk to the next GPU and receives from the previous GPU in N-1 steps. Ring all-reduce achieves bandwidth-optimal performance for large messages (>1 MB) because each GPU participates with full duplex bandwidth throughout the operation. On 8x H100 with NVSwitch, ring all-reduce for 256 MB reaches 450 GB/s (95% of the 475 GB/s theoretical NVSwitch bandwidth).
The tree algorithm uses a butterfly or binary-tree topology where data is exchanged in log(N) steps. Tree all-reduce reduces startup latency because fewer steps are required (log2(8)=3 steps for 8 GPUs versus N-1=7 steps for ring). For small messages (<128 KB), tree all-reduce achieves 2.3-3.0x lower latency than ring: 8 microseconds versus 24 microseconds for 8 KB messages. NCCL 2.22+ automatically selects between ring and tree based on message size, using a heuristic threshold that defaults to 256 KB on H100.
| Message Size | Ring Algorithm | Tree Algorithm | NCCL Auto-Select | Best Algorithm |
|---|---|---|---|---|
| 8 KB | 24 us, 0.3 GB/s | 8 us, 0.9 GB/s | Tree | Tree |
| 128 KB | 28 us, 4.6 GB/s | 14 us, 9.1 GB/s | Tree | Tree |
| 1 MB | 42 us, 23.8 GB/s | 32 us, 31.3 GB/s | Tree (tie) | Tree or Ring |
| 16 MB | 90 us, 177.8 GB/s | 120 us, 133.3 GB/s | Ring | Ring |
| 256 MB | 580 us, 441.4 GB/s | 820 us, 312.2 GB/s | Ring | Ring |
| 1 GB | 2,300 us, 443.5 GB/s | 3,400 us, 300.0 GB/s | Ring | Ring |
ALL-TO-ALL AND SPECIALIZED COLLECTIVES
All-to-all collectives serve workloads where each GPU needs distinct data from every other GPU, notably for Mixture of Experts (MoE) model training. NCCL's ncclAllToAll primitive implements a direct peer-to-peer exchange: each of N GPUs sends N distinct buffers and receives N distinct buffers in N-1 steps. On 8x H100 with NVSwitch, all-to-all for MoE token routing achieves 320 GB/s bandwidth, or 72% of the theoretical NVSwitch bandwidth.
NCCL 2.22+ introduces the ncclGroupEndAllToAll API which batches multiple all-to-all operations into a single kernel launch, reducing launch overhead from 4-6 microseconds per all-to-all to a single 2-microsecond launch. NCCL also supports ncclReduceScatter and ncclAllGather primitives used in FSDP gradient accumulation. FSDP uses reduce-scatter for gradient reduction followed by all-gather for parameter updates. NCCL's reduce-scatter on 8x H100 achieves 92% of all-reduce bandwidth for the same message size.
NETWORK-AWARE ALGORITHM SELECTION AND TOPOLOGY DETECTION
NCCL 2.22+ includes a topology-aware graph optimizer (NCCL_TOPO_FILE) that profiles the GPU interconnect topology at startup and selects the optimal collective algorithm for each message size and GPU placement. The topology detection uses CUDA Runtime API to query NVLink connectivity, PCIe hierarchy, and network interface location. The topology is encoded as an XML file at /etc/nccl-topology.xml and can be overridden with NCCL_TOPO_FILE=/path/custom.xml. Each link is assigned a bandwidth weight.
For multi-node all-reduce, NCCL uses a hierarchical decomposition: intra-node NVLink all-reduce followed by inter-node all-reduce across the remaining aggregated data. The algorithm uses a 2D-Torus topology for homogeneous InfiniBand fabrics: 4 nodes x 4 GPUs per node achieves 82-88% of theoretical bandwidth. The CollNetDirect collective uses NVIDIA's CollNet accelerator (available on H100 ConnectX-8 SmartNIC) to offload all-reduce accumulation to the network adapter, reducing GPU utilization from 15-20% to 2-4% during gradient synchronization.
| Algorithm / Feature | Symbol | Intra-Node (8xH100 NVSwitch) | Inter-Node (InfiniBand NDR400) | Use Case |
|---|---|---|---|---|
| Ring AllReduce | AR_RING | 441 GB/s (95% peak) | 180 GB/s (80% peak) | Large gradients > 1 MB |
| Tree AllReduce | AR_TREE | 312 GB/s (67% peak) | 145 GB/s (64% peak) | Small gradients < 128 KB |
| CollNetDirect | AR_COLLNET | N/A (NIC-only) | 200 GB/s (89% peak) | GPU-offloaded all-reduce |
| ReduceScatter | RS_RING | 405 GB/s (87% peak) | 165 GB/s (73% peak) | FSDP gradient reduction |
| AllGather | AG_RING | 410 GB/s (88% peak) | 170 GB/s (76% peak) | FSDP parameter gather |
| AllToAll (MoE) | A2A_DIRECT | 320 GB/s (72% peak) | 80 GB/s (36% peak) | MoE token routing |
| Hierarchical 2D-Torus | N/A | NVSwitch (full BW) | 82-88% aggregate | Multi-node training |
NCCL ENVIRONMENT VARIABLE TUNING FOR PRODUCTION
NCCL exposes 40+ environment variables for fine-tuning communication performance. The most impactful for H100/B200 clusters: NCCL_IB_TIMEOUT=22 (InfiniBand timeout, increase from default 18 for congested fabrics), NCCL_IB_RETRY_CNT=7 (retry count), NCCL_SOCKET_IFNAME=ib0,ib1,ib2,ib3 (bind to InfiniBand interfaces), NCCL_NET_GDR_LEVEL=5 (enable GPU Direct RDMA), NCCL_ALGO=Ring,Tree,CollNetDirect (algorithm priority), NCCL_PROTO=Simple,LL,LL128 (protocol selection). The NCCL_DEBUG=INFO flag shows algorithm selection per collective.
For multi-node training at scale (64+ GPUs): NCCL_IB_HCA=mlx5_0:1,mlx5_1:1 (specify InfiniBand HCAs), NCCL_NVLS_ENABLE=1 (enable NVLink SHARP in-network reduction on H100 ConnectX-8), NCCL_TUNER_PLUGIN=1 (adaptive algorithm selection). For RoCE fabrics, set NCCL_IB_DISABLE=0 and NCCL_NET_GDR_LEVEL=5 plus NCCL_IB_QPS_PER_CONNECTION=8. The nccl-tuner tool validates performance: nccl-tuner --bench all_reduce --msg_size 256M --niter 100.
NVLink VS INFINIBAND VS ROCE: IMPACT ON ALGORITHM CHOICE
The optimal NCCL algorithm is determined by the interconnect hierarchy. NVLink 4.0 (H100) provides 900 GB/s bidirectional bandwidth between GPUs within a node. InfiniBand NDR400 provides 50 GB/s per link between nodes. The 18:1 bandwidth ratio between intra-node and inter-node links means the algorithm must minimize inter-node communication. NCCL's hierarchical algorithm achieves this by performing NVLink-speed intra-node all-reduce, then sending a single aggregated gradient per GPU across the inter-node link.
RoCE (RDMA over Converged Ethernet) with 200 Gbps links achieves approximately 80% of InfiniBand NDR400 performance but introduces 2x higher tail latency due to packet loss. NCCL detects RoCE fabric via NCCL_IB_DISABLE=0 and automatically reduces the message chunk size from 256 KB to 64 KB to minimize the impact of packet retransmission. For clusters mixing NVLink and RoCE, NCCL_ALGO=Ring handles asymmetric bandwidth environments better than tree.
| Interconnect | Per-Link Bandwidth | Latency | NCCL Intra-Node | NCCL Inter-Node | Algorithm Preference |
|---|---|---|---|---|---|
| NVLink 4.0 (H100) | 900 GB/s bidirectional | 0.5 us | 441 GB/s (ring, 256 MB) | N/A | Ring for > 1 MB |
| NVSwitch (H100) | 900 GB/s per GPU pair | 0.8 us | 450 GB/s (ring+all-to-all) | N/A | Ring or CollNet |
| InfiniBand NDR400 | 50 GB/s per link | 1.2 us | N/A | 180 GB/s (8 links) | Hierarchical 2D |
| InfiniBand HDR200 | 25 GB/s per link | 1.5 us | N/A | 90 GB/s (8 links) | Hierarchical |
| RoCE 200 Gbps | 40 GB/s per link (effective) | 3.0 us | N/A | 130 GB/s (8 links) | Ring, smaller chunks |
| Ethernet 100 Gbps | 11.5 GB/s per link | 5.0 us | N/A | 45 GB/s (8 links) | Tree (lower BW efficiency) |
NCCL-OPTIMIZED GPU CLUSTERS ON CLUSTERBID
ClusterBid's provider network includes NCCL-optimized GPU cluster configurations with validated interconnect topologies. The --interconnect filter identifies NVLink, InfiniBand, and RoCE connectivity options. NVSwitch-connected H100 instances (8x GPU per node, full NVLink bandwidth) are available from 8 providers starting at $24/hour per node. The --nccl-version filter selects instances with specific NCCL versions pre-installed.
The recommended NCCL validation workflow on ClusterBid instances: (1) Provision an 8-node GPU cluster with matching interconnects. (2) Run nccl-tuner --bench all_reduce --msg_size 256M --niter 100 for baseline bandwidth. (3) Compare against expected H100 benchmarks: 441 GB/s intra-node, 180 GB/s inter-node. (4) Tune NCCL_NVLS_ENABLE=1 for CollNetDirect. (5) Set NCCL_ALGO=Ring,Tree,CollNetDirect for optimal auto-selection.
