Why Multi-GPU Communication Matters
Multi-GPU training performance is determined as much by inter-GPU communication as by GPU compute capability. The NCCL (NVIDIA Collective Communications Library) all-reduce operation, which averages gradients across all GPUs, must complete synchronously before the next training step can begin. Every millisecond of communication overhead is multiplied by the number of training steps -- a 5ms all-reduce delay over 100,000 training steps adds 500 seconds (8.3 minutes) to a training run.
At scale (256+ GPUs), communication overhead can dominate training time. A poorly configured cluster may spend 40-60% of training time on gradient synchronisation, compared to 10-20% for a well-optimised cluster. The difference between a 5-day training run and a 7-day run is purely network configuration.
This post covers NCCL and ROCm communication optimisation, NVLink topology, InfiniBand fabric tuning, and the tools for diagnosing multi-GPU communication bottlenecks.
NCCL Performance Fundamentals
NCCL performance depends on three factors: intra-node bandwidth (NVLink between GPUs on the same HGX baseboard), inter-node bandwidth (InfiniBand or RoCEv2 between GPU nodes), and NCCL algorithm selection (ring, tree, or NVLink ring). The optimal algorithm depends on message size and network topology.
For intra-node communication, NVLink 4.0 (H100) provides 900 GB/s bidirectional bandwidth per GPU in an 8-GPU fully connected topology via the NVSwitch. This means any GPU can communicate with any other GPU at 900 GB/s with no oversubscription. In practice, NCCL achieves 800-870 GB/s for large message all-reduce on 8-GPU H100, approximately 90-95% of theoretical peak.
For inter-node communication, the bandwidth drops to InfiniBand rates. With 8x 400 Gbps InfiniBand per node (one per GPU), the inter-node all-reduce bandwidth is approximately 320-370 GB/s per node (80-90% of the 400 Gbps theoretical). The ratio of intra-node to inter-node bandwidth (approximately 2.5:1) means that crossing node boundaries significantly increases communication time.
NVLink Topology and NCCL Ring Configuration
The NVLink topology within an HGX baseboard determines NCCL communication patterns. The H100 HGX has a full NVSwitch mesh: all 8 GPUs connect to a single NVSwitch, which can route between any GPU pair at full bandwidth. The B200 HGX uses a similar architecture with NVLink 5.0, maintaining full bisection bandwidth across the 8-GPU baseboard.
When training across multiple HGX baseboards (within the same node or across nodes), the NVLink topology becomes hierarchical. GPUs on the same baseboard communicate via NVSwitch at 900 GB/s. GPUs on different baseboards communicate via the CPU PCIe root complex or NVLink bridge (if available). The NCCL topology discovery (`ncclTopoFromXml`) maps these connections and selects the optimal communication rings.
The NCCL ring order should match the physical NVLink topology. A misconfigured ring order (e.g., mixing GPUs from different baseboards in the same ring) causes some communication to traverse slower PCIe paths instead of NVSwitch, reducing all-reduce bandwidth by 30-50%. Use `nvidia-smi topo -m` to verify the topology and `NCCL_DEBUG=INFO` to check the ring construction during NCCL initialisation.
Inter-Node InfiniBand Tuning
InfiniBand fabric performance for NCCL depends on: subnet configuration (single subnet vs multi-subnet, fabric partitioning), adaptive routing (enabled vs disabled, affects load balancing across the fabric), congestion control (enabled vs disabled, prevents incast during all-reduce), and GPU-to-NIC affinity (NUMA alignment, each GPU should use the NIC connected to its NUMA domain).
The most impactful tuning parameter is adaptive routing. On NVIDIA Quantum-2 InfiniBand, enabling adaptive routing distributes all-reduce traffic across multiple paths, improving bandwidth utilisation from 60-70% to 85-95% for large messages. The trade-off is slightly higher latency for small messages (< 1 MB), which is negligible for training workloads.
GPU-to-NIC NUMA affinity is the second most important factor. Each GPU should use the NIC that shares its NUMA node. Misaligned GPU-NIC pairing forces data through the CPU interconnect (UPI or PCIe bridge), adding 3-8 microseconds of latency per transfer. For a multi-GPU all-reduce with 100+ message exchanges, this adds 0.3-0.8 milliseconds to each training step.
NCCL Algorithm Selection: Ring, Tree, and NVLink
NCCL supports three primary algorithms for all-reduce. The ring algorithm splits the message into N chunks and passes them around the ring. It achieves high bandwidth for large messages but has latency proportional to the number of GPUs. The tree algorithm uses a binary tree for reduce-scatter and all-gather. It has lower latency for small messages but can be bandwidth-limited for large ones. The NVLink algorithm uses the NVSwitch for direct GPU-to-GPU copies, achieving the highest bandwidth on compatible hardware.
NCCL automatically selects the algorithm based on message size and topology. The thresholds are: messages < 128 KB: tree algorithm for lower latency. Messages 128 KB - 32 MB: ring algorithm for balanced latency and bandwidth. Messages > 32 MB: NVLink algorithm (on H100/B200) or ring algorithm (on older hardware).
You can override algorithm selection with NCCL_ALGO for benchmarking: `NCCL_ALGO=Ring`, `NCCL_ALGO=Tree`, `NCCL_ALGO=NVLink`. Run the nccl-tests benchmark (`all_reduce_perf -b 1M -e 1G -f 2`) with each algorithm to determine the best configuration for your workload's message size distribution.
ROCm Communication: RCCL and AMD Multi-GPU
AMD's ROCm platform uses RCCL (ROCm Collective Communications Library) as the equivalent of NCCL. RCCL supports the same collective operations (all-reduce, all-gather, reduce-scatter) and API interface as NCCL, enabling PyTorch and TensorFlow to use it transparently. However, performance characteristics differ from NCCL.
AMD MI350X GPUs connect via Infinity Fabric, providing 128 GB/s perGPU bidirectional bandwidth (compared to NVLink's 900 GB/s). The lower interconnect bandwidth means multi-GPU communication is more expensive on AMD hardware, and the optimal parallelism strategy differs: tensor parallelism should minimise the number of GPUs (ideally 2-4 per model replica rather than 8), and pipeline parallelism may be preferred over tensor parallelism for large models.
RCCL supports the same algorithm options as NCCL (ring, tree) but does not have an equivalent of NCCL's NVLink algorithm. The Infinity Fabric does not provide the full-mesh connectivity that NVSwitch provides, so inter-GPU communication latency varies by GPU pair. RCCL topology detection maps these distances and optimises communication patterns accordingly.
Diagnosing Multi-GPU Communication Bottlenecks
The diagnostic toolkit for multi-GPU communication includes: `nccl-tests` (the standard NCCL benchmark, run all_reduce_perf at various message sizes), Nsight Systems (timeline profiling showing communication vs computation overlap), `nvidia-smi topo -m` (topology matrix showing GPU-NVLink-GPU and GPU-NIC connectivity), and fabric manager logs (InfiniBand fabric health, link errors, congestion).
The diagnostic process: run nccl-tests and compare measured bandwidth to theoretical peak (NVLink: 900 GB/s, InfiniBand HDR: 50 GB/s per port). If bandwidth is below 80% of theoretical, check: topology (are GPUs on the same NVSwitch?), NUMA affinity (matched GPU-NIC?), InfiniBand link errors (any ports showing errors?), and adaptive routing (enabled on the fabric?).
The most common multi-GPU communication issues at mid-2026: InfiniBand cable faults (loose or damaged cables cause intermittent errors, detected via fabric manager), NUMA misalignment (GPU uses wrong NIC, detected via `nvidia-smi topo -m`), oversubscribed fabric (more GPUs than expected, detected via nccl-tests bandwidth degradation at scale), and NCCL version mismatch (mixed NCCL versions in distributed training, detected via initialisation log messages).
