WHY COMMUNICATION BENCHMARKS MATTER
Multi-GPU training performance is dominated by communication efficiency. A 70B model training on 128 GPUs spends 15-30% of time in all-reduce synchronization. With suboptimal NCCL configuration, this can reach 50-60%. Understanding the all-reduce bandwidth curve across message sizes is essential for training optimization.
The choice of NCCL algorithm (Ring, Tree, or NVLink Tree) determines throughput at different scales. Ring is optimal for small messages (<1 MB) on large clusters, Tree for medium messages (1-64 MB), and NVLink Tree for large messages (>64 MB) on single-node runs.
| NCCL Algorithm | Best Message Size | Scalability | GPU Memory Usage | Typical Bandwidth (8x H100), 512 MB |
|---|---|---|---|---|
| Ring | <1 MB | Excellent (256+ GPUs) | Low | 395 GB/s |
| Tree | 1-64 MB | Good (64+ GPUs) | Medium | 428 GB/s |
| NVLink Tree | >64 MB | Single node only | High | 450 GB/s |
| Rabbit | Custom | Multi-node NVLink | High | 440 GB/s |
NVIDIA NCCL: ALGORITHMS AND TOPOLOGY AWARENESS
NCCL 2.21+ automatically detects topology and selects algorithms. NVLink-connected GPUs within a node achieve 450 GB/s all-reduce. Cross-node via InfiniBand NDR400 achieves 50 GB/s per link. With 8 nodes and 4 IB links each: 185 GB/s inter-node all-reduce.
NCCL_NET_GDR_LEVEL enables GPUDirect RDMA for bypassing CPU memory, improving bandwidth by 15-25%. NCCL_ALGO selects algorithm family. In production training at 512 H100 scale, optimal configuration achieves 92% of theoretical peak bisection bandwidth.
| GPU Type | Intra-Node All-Reduce (512 MB) | Inter-Node 8x NDR400 (512 MB) | NCCL Version | CUDA Version |
|---|---|---|---|---|
| H100 SXM 80GB | 450 GB/s | 185 GB/s | 2.21+ | 12.4+ |
| A100 SXM 80GB | 300 GB/s | 105 GB/s | 2.18+ | 12.0+ |
| A100 PCIe 80GB | 150 GB/s | 85 GB/s | 2.18+ | 12.0+ |
| B100 (expected) | 900 GB/s | 350 GB/s | 3.0+ | 13.0+ |
AMD RCCL: PERFORMANCE AND MI300X STATUS
AMD RCCL (based on NCCL) supports MI300X with Infinity Fabric providing 896 GB/s intra-node bandwidth. MI300X all-reduce: 420 GB/s intra-node (8 GPUs), 95 GB/s inter-node via 4x InfiniBand NDR200. Approximately 7% lower than equivalent H100 configuration on inter-node.
RCCL configuration requires specific environment variables: RCCL_MSCCL_ENABLE=1 for MSCCL-based all-reduce (introduced in ROCm 6.1). Multi-rack MI300X training (256 GPUs) achieves 92% of H100-equivalent throughput for FP8 training, narrowing to 5% gap with ROCm 6.2+ and optimized XGMI topology.
| Metric | H100 (NVLink 4.0 + NDR400) | MI300X (Infinity Fabric + NDR200) | Ratio H100/MI300X |
|---|---|---|---|
| Intra-node BW (512 MB) | 450 GB/s | 420 GB/s | 1.07x |
| Inter-node BW 8 nodes (512 MB) | 185 GB/s | 95 GB/s | 1.95x |
| All-reduce latency (4 GPUs, 1 MB) | 24 us | 31 us | 1.29x |
| Training throughput (70B, FP8, 256 GPUs) | 875 tok/s/GPU | 805 tok/s/GPU | 1.09x |
BENCHMARKING METHODOLOGY
Use the NCCL 'sendrecv' and 'allreduce' benchmarks from nccl-tests. Key configurations: (1) test at message sizes 1B to 512 MB (log scale), (2) run 5 warmup iterations + 50 measured, (3) report busbw (bus bandwidth, not algorithm bandwidth), (4) set CUDA_VISIBLE_DEVICES to test specific GPU subsets.
For cluster-wide profiling: run nccl-tests/all_reduce_perf -b 1 -e 512M -f 2 -g 8 -t 8 -w 5 -n 50 on each unique topology configuration. The busbw at 512 MB is the single-number summary. A cluster scoring >400 GB/s at 8 GPUs for H100 is well-configured. Below 320 GB/s indicates topology or driver issues.
AMD users run rocprof with RCCL to trace communication events. RCCL performance tuning typically requires 2-3 weeks of iteration for new cluster deployments, compared to 3-5 days for NCCL due to younger tooling ecosystem.
