All essays
TechnicalDEEP DIVEFEB 2026

NCCL Performance Tuning: Ring vs Tree vs All-to-All, Network-Aware Algorithm Selection

NCCL performance tuning deep dive: ring, tree, and all-to-all collective algorithm comparison, network-aware selection for NVLink/InfiniBand/RoCE, NCCL 2.22+ features, and optimization for H100/B200 multi-node training.

01

NCCL COLLECTIVE ALGORITHMS: RING, TREE, AND ALL-TO-ALL

NCCL implements three fundamental collective communication algorithms, each optimized for different GPU topology and message size regimes. The ring algorithm partitions data across N GPUs in a logical ring topology, where each GPU sends a chunk to the next GPU and receives from the previous GPU in N-1 steps. Ring all-reduce achieves bandwidth-optimal performance for large messages (>1 MB) because each GPU participates with full duplex bandwidth throughout the operation. On 8x H100 with NVSwitch, ring all-reduce for 256 MB reaches 450 GB/s (95% of the 475 GB/s theoretical NVSwitch bandwidth).

The tree algorithm uses a butterfly or binary-tree topology where data is exchanged in log(N) steps. Tree all-reduce reduces startup latency because fewer steps are required (log2(8)=3 steps for 8 GPUs versus N-1=7 steps for ring). For small messages (<128 KB), tree all-reduce achieves 2.3-3.0x lower latency than ring: 8 microseconds versus 24 microseconds for 8 KB messages. NCCL 2.22+ automatically selects between ring and tree based on message size, using a heuristic threshold that defaults to 256 KB on H100.

Message SizeRing AlgorithmTree AlgorithmNCCL Auto-SelectBest Algorithm
8 KB24 us, 0.3 GB/s8 us, 0.9 GB/sTreeTree
128 KB28 us, 4.6 GB/s14 us, 9.1 GB/sTreeTree
1 MB42 us, 23.8 GB/s32 us, 31.3 GB/sTree (tie)Tree or Ring
16 MB90 us, 177.8 GB/s120 us, 133.3 GB/sRingRing
256 MB580 us, 441.4 GB/s820 us, 312.2 GB/sRingRing
1 GB2,300 us, 443.5 GB/s3,400 us, 300.0 GB/sRingRing
02

ALL-TO-ALL AND SPECIALIZED COLLECTIVES

All-to-all collectives serve workloads where each GPU needs distinct data from every other GPU, notably for Mixture of Experts (MoE) model training. NCCL's ncclAllToAll primitive implements a direct peer-to-peer exchange: each of N GPUs sends N distinct buffers and receives N distinct buffers in N-1 steps. On 8x H100 with NVSwitch, all-to-all for MoE token routing achieves 320 GB/s bandwidth, or 72% of the theoretical NVSwitch bandwidth.

NCCL 2.22+ introduces the ncclGroupEndAllToAll API which batches multiple all-to-all operations into a single kernel launch, reducing launch overhead from 4-6 microseconds per all-to-all to a single 2-microsecond launch. NCCL also supports ncclReduceScatter and ncclAllGather primitives used in FSDP gradient accumulation. FSDP uses reduce-scatter for gradient reduction followed by all-gather for parameter updates. NCCL's reduce-scatter on 8x H100 achieves 92% of all-reduce bandwidth for the same message size.

03

NETWORK-AWARE ALGORITHM SELECTION AND TOPOLOGY DETECTION

NCCL 2.22+ includes a topology-aware graph optimizer (NCCL_TOPO_FILE) that profiles the GPU interconnect topology at startup and selects the optimal collective algorithm for each message size and GPU placement. The topology detection uses CUDA Runtime API to query NVLink connectivity, PCIe hierarchy, and network interface location. The topology is encoded as an XML file at /etc/nccl-topology.xml and can be overridden with NCCL_TOPO_FILE=/path/custom.xml. Each link is assigned a bandwidth weight.

For multi-node all-reduce, NCCL uses a hierarchical decomposition: intra-node NVLink all-reduce followed by inter-node all-reduce across the remaining aggregated data. The algorithm uses a 2D-Torus topology for homogeneous InfiniBand fabrics: 4 nodes x 4 GPUs per node achieves 82-88% of theoretical bandwidth. The CollNetDirect collective uses NVIDIA's CollNet accelerator (available on H100 ConnectX-8 SmartNIC) to offload all-reduce accumulation to the network adapter, reducing GPU utilization from 15-20% to 2-4% during gradient synchronization.

Algorithm / FeatureSymbolIntra-Node (8xH100 NVSwitch)Inter-Node (InfiniBand NDR400)Use Case
Ring AllReduceAR_RING441 GB/s (95% peak)180 GB/s (80% peak)Large gradients > 1 MB
Tree AllReduceAR_TREE312 GB/s (67% peak)145 GB/s (64% peak)Small gradients < 128 KB
CollNetDirectAR_COLLNETN/A (NIC-only)200 GB/s (89% peak)GPU-offloaded all-reduce
ReduceScatterRS_RING405 GB/s (87% peak)165 GB/s (73% peak)FSDP gradient reduction
AllGatherAG_RING410 GB/s (88% peak)170 GB/s (76% peak)FSDP parameter gather
AllToAll (MoE)A2A_DIRECT320 GB/s (72% peak)80 GB/s (36% peak)MoE token routing
Hierarchical 2D-TorusN/ANVSwitch (full BW)82-88% aggregateMulti-node training
04

NCCL ENVIRONMENT VARIABLE TUNING FOR PRODUCTION

NCCL exposes 40+ environment variables for fine-tuning communication performance. The most impactful for H100/B200 clusters: NCCL_IB_TIMEOUT=22 (InfiniBand timeout, increase from default 18 for congested fabrics), NCCL_IB_RETRY_CNT=7 (retry count), NCCL_SOCKET_IFNAME=ib0,ib1,ib2,ib3 (bind to InfiniBand interfaces), NCCL_NET_GDR_LEVEL=5 (enable GPU Direct RDMA), NCCL_ALGO=Ring,Tree,CollNetDirect (algorithm priority), NCCL_PROTO=Simple,LL,LL128 (protocol selection). The NCCL_DEBUG=INFO flag shows algorithm selection per collective.

For multi-node training at scale (64+ GPUs): NCCL_IB_HCA=mlx5_0:1,mlx5_1:1 (specify InfiniBand HCAs), NCCL_NVLS_ENABLE=1 (enable NVLink SHARP in-network reduction on H100 ConnectX-8), NCCL_TUNER_PLUGIN=1 (adaptive algorithm selection). For RoCE fabrics, set NCCL_IB_DISABLE=0 and NCCL_NET_GDR_LEVEL=5 plus NCCL_IB_QPS_PER_CONNECTION=8. The nccl-tuner tool validates performance: nccl-tuner --bench all_reduce --msg_size 256M --niter 100.

06

NCCL-OPTIMIZED GPU CLUSTERS ON CLUSTERBID

ClusterBid's provider network includes NCCL-optimized GPU cluster configurations with validated interconnect topologies. The --interconnect filter identifies NVLink, InfiniBand, and RoCE connectivity options. NVSwitch-connected H100 instances (8x GPU per node, full NVLink bandwidth) are available from 8 providers starting at $24/hour per node. The --nccl-version filter selects instances with specific NCCL versions pre-installed.

The recommended NCCL validation workflow on ClusterBid instances: (1) Provision an 8-node GPU cluster with matching interconnects. (2) Run nccl-tuner --bench all_reduce --msg_size 256M --niter 100 for baseline bandwidth. (3) Compare against expected H100 benchmarks: 441 GB/s intra-node, 180 GB/s inter-node. (4) Tune NCCL_NVLS_ENABLE=1 for CollNetDirect. (5) Set NCCL_ALGO=Ring,Tree,CollNetDirect for optimal auto-selection.

Filed under
NCCL PerformanceNCCL Ring AlgorithmNCCL Tree AlgorithmNCCL All-to-AllNetwork-Aware NCCLNCCL 2.22Multi-Node GPU Training