GPU FABRIC ARCHITECTURE AND BASELINE EXPECTATIONS
GPU cluster networks operate at performance levels that leave zero margin for misconfiguration. A single H100 GPU on NDR400 InfiniBand can saturate 400 Gbps bidirectional. A node with 8x H100 GPUs connected to a fabric of 64-256 nodes creates aggregate traffic loads exceeding 10 Tbps for NCCL all-reduce operations. The two dominant fabric choices are HDR200/NDR400 InfiniBand (60 percent of large GPU clusters) and RoCE v2 (Remote Direct Memory Access over Converged Ethernet, 35 percent of clusters, primarily in latency-tolerant training workloads). The remaining 5 percent use NVIDIA Spectrum-X Ethernet with adaptive routing and congestion control for GPU-enhanced Ethernet.
Establishing baseline fabric behavior is the first troubleshooting step. For InfiniBand, the `ibstatus` and `ibstat` commands report link state, rate (400 Gbps = 4X NDR, 200 Gbps = 4X HDR), and physical width (4X, 8X, 12X). Expected output: `rate: 400 Gbps (4X NDR), link_layer: InfiniBand, state: 4 (ACTIVE)`. For RoCE v2, `mlxstat` and `ibdev2netdev -v` map RDMA devices to ethernet interfaces and report link speed. The NCCL `topology.xml` file defines the cluster's hierarchical topology (NVSwitch within DGX, InfiniBand switch between DGX, core spine layer). Any deviation from the physical wiring topology in this file will cause NCCL to use suboptimal all-reduce algorithms.
| Parameter | NDR400 InfiniBand (Healthy) | HDR200 InfiniBand (Healthy) | RoCE v2 200 GbE (Healthy) |
|---|---|---|---|
| Link Rate (ibstatus) | 400 Gbps (4X NDR) | 200 Gbps (4X HDR) | 200 Gbps (Ethernet, PFC on) |
| Link Layer | InfiniBand | InfiniBand | Ethernet |
| Link Active state | state 4 (ACTIVE) | state 4 (ACTIVE) | Link Up, LLDP enabled |
| NCCL AllReduce 8xH100 (1 node) | 25.2 GB/s | 18.5 GB/s | 21.8 GB/s |
| NCCL AllReduce 64xH100 (8 nodes) | 160-180 GB/s | 110-130 GB/s | 120-145 GB/s |
| NCCL AllReduce 512xH100 (64 nodes) | 800-1100 GB/s | 500-700 GB/s | 600-850 GB/s |
| Max Pause Frames (rdma_pause) | N/A (lossless) | N/A (lossless) | < 0.01% of total frames |
| Switch Buffer | 256 MB per port (Mellaness QM9790) | 96 MB per port (QM9700) | 42 MB per port (SN4600C) |
NCCL TROUBLESHOOTING: TIMEOUTS, TOPOLOGY, AND ALGORITHM SELECTION
NCCL (NVIDIA Collective Communications Library) is the most frequent source of training failures in multi-node GPU clusters. The most common error is NCCL timeout: `ncclInternalError: NCCL operation timed out` after a default 30-second watchdog, which means an all-reduce or all-gather operation did not complete within the timeout window. The first diagnostic step sets `NCCL_DEBUG=INFO` and `NCCL_DEBUG_SUBSYS=ALL` in the training container environment variables, which produces detailed timing logs for each collective operation. The key line to look for: `nvidia[rank][n]: NCCL INFO comm 0x... rank n nranks m busId ... time 12.345` shows the time taken for the debug NCCL communicator initialization.
NCCL timeout causes break into three categories: topology mismatch (NCCL thinks GPUs are on different nodes, causing cross-node collectives that should be intra-node), switch congestion (buffer exhaustion on core spine links, causing backpressure that delays completion beyond timeout), and silent link failures (NVLink cable or backplane failure that causes retransmission at the link layer, increasing completion time from 50 microseconds to 5+ seconds). For topology mismatch, `nvidia-smi topo -m` prints the GPU-to-GPU connectivity matrix; if NVLink1/NVLink2 connections show as PIX (via PCIe) instead of NV# (direct NVLink), the NVLink bridge or backplane needs inspection. For switch congestion, `perfquery` on InfiniBand switches reports port counters including `SymbolErrorCounter`, `LinkErrorRecoveryCounter`, and `PortRcvConstraintErrors`.
| NCCL Error | Root Cause | Diagnostic Command | Fix / Workaround |
|---|---|---|---|
| NCCL Timed Out (30s) | Congested spine link | perfquery on IB switch, look for XmitWait > 1e12 | Re-route: set NCCL_IB_TIMEOUT=22, increase agg buffer |
| NCCL Internal Error | Topology mismatch in NCCL_TOPO_FILE | nvidia-smi topo -m, compare pinning | Correct NCCL topology.xml or use NCCL_TOPO_DUMP_FILE |
| NCCL Remote Error | Node crash during all-reduce | dmesg on peer node, check for GPU fallen off bus | Replace or reboot failed node, NCCL_IB_RETRY_CNT=7 |
| NCCL Invalid Usage | CUDA stream conflict | NCCL_DEBUG=INFO, check stream synchronization | Use NCCL_GROUP_CUDART_VERSION=12040 |
| NCCL Connection Failure | IB link down on one node | ibstatus on all nodes, compare link state | Reseat IB cable, restart opensm or subn manager |
| NCCL Unhandled Cuda Error | OOM during collective buffer allocation | nvidia-smi --query-gpu=memory.free --format=csv | Reduce --gradient-accumulation-steps or batch size |
RDMA DIAGNOSTICS: INFINIBAND AND ROCE V2 TROUBLESHOOTING
InfiniBand fabric issues manifest as CRC errors, link flaps, and symbol errors. The primary diagnostic tool is `perfquery` (part of infiniband-diags), which queries switch and HCA port counters. The critical counters are `SymbolErrorCounter` (should be zero; non-zero indicates physical layer issues with optical transceivers or cables), `LinkErrorRecoveryCounter` (should be zero; non-zero indicates link training failures), `PortRcvErrors` (should be less than 0.001 percent of total received packets; higher indicates signal integrity problems), and `ExcessiveBufferOverrunErrors` (non-zero indicates buffer allocation mismatch between sender and receiver). The command `perfquery -x -P 0 $SWITCH_LID $PORT` dumps extended counters with per-VL (Virtual Lane) breakdown.
RoCE v2 troubleshooting requires a different toolkit because the fabric uses standard Ethernet with Priority Flow Control (PFC, IEEE 802.1Qbb). The most common RoCE v2 issue is PFC storm: when one flow causes buffer exhaustion on an egress port, the switch sends PFC pause frames that halt all traffic on that priority class, including unrelated RDMA traffic. `mlnx_perfquery` or `ethtool -S $IFACE | grep -E "pfc|pause|deadlock"` reports PFC pause frame counts. If `tx_pause` or `rx_pause` exceed 0.1 percent of total packets, the fabric has a head-of-line blocking problem. The fix involves enabling Explicit Congestion Notification (ECN) with DCQCN (Data Center Quantized Congestion Notification), `echo 1 > /sys/class/net/$IFACE/ecn/roce_enable`, and configuring RED/ECN thresholds on switches to mark congested packets before PFC is triggered.
PACKET LOSS DETECTION AND CONGESTION MANAGEMENT
Packet loss in GPU fabrics has a disproportionate impact: a single lost packet in an NCCL all-reduce required 8 GB buffer transfer triggers retransmission of the entire buffer, increasing completion time from 50 milliseconds to 500+ milliseconds. For InfiniBand, packet loss is invisible at the application level because the fabric is nominally lossless (credit-based flow control), but buffer overruns at congested switch ports cause the same effective retransmission penalty. The indicator is `XmitWait` counter in `perfquery` output: this measures the number of symbol periods a port waited to send because no credits were available. A `XmitWait` value exceeding 1e12 between two consecutive queries (5-second interval) indicates persistent congestion on that port.
RoCE v2 packet loss is directly measurable from `ethtool -S` counters. The command `ethtool -S $IFACE | grep -E "rx_dropped|tx_dropped|errors"` reports physical packet drops. Additionally, `rdma statistic show` reveals RDMA-level retransmission: `hir` (hardware inbound retransmission) and `hor` (hardware outbound retransmission) counts above 0 for a stable fabric warrant investigation. The remediation for RoCE v2 clusters experiencing >0.01 percent packet loss is to implement adaptive routing: NVIDIA Spectrum switches support Dynamic Load Balancing (DLB) that distributes RDMA flows across available paths, reducing congestion on hot spots. The configuration is `mlnx_qos -i $IFACE --trust dscp && echo 1 > /sys/class/net/$IFACE/adaptive_routing". For InfiniBand, adaptive routing is enabled through the OpenSM subnet manager with `--advanced-routing 1` flag.
FABRIC TOPOLOGY OPTIMIZATION FOR NCCL ALL-REDUCE
NCCL all-reduce performance depends directly on fabric topology. The optimal topology for large GPU clusters is a multi-layer fat tree (Clos topology) with full bisection bandwidth: each leaf switch connects to exactly the same number of spine switches, and each spine switch connects to every leaf switch. For a 512-GPU H100 cluster on NDR400 InfiniBand, the topology uses 8 leaf switches (16 ports each, connecting 8 DGX H100 nodes per leaf) connected to 8 spine switches in a full-mesh. Any oversubscription ratio above 1:1 at the spine layer directly reduces all-reduce bandwidth proportionally: a 2:1 oversubscription reduces bisection bandwidth to 50 percent, increasing all-reduce time from 160 GB/s to 80 GB/s for a 512-GPU cluster.
The NCCL topology file (`topology.xml`) must reflect the physical wiring exactly. NVIDIA provides a `nccl-topology-dump` tool that generates the topology file from `nvidia-smi topo -m` output combined with switch fabric discovery. The file specifies GPU-to-NIC affinity (which GPU shares a PCIe switch with which InfiniBand HCA), NIC-to-switch port mapping, and inter-switch spine connections. Common mistakes include assuming NVSwitch topology from a single node applies identically to all nodes (H100 DGX requires per-node topology validation), incorrect `guid` values for InfiniBand switch ports that cause NCCL to compute incorrect routing, and missing `link type=NVSWITCH` entries that force NCCL to use PCIe P2P instead of NVLink for intra-node communication. The `nccl-tests` benchmark (allreduce_perf -b 8 -e 128M -f 2) validates the topology by measuring bandwidth across all 512 GPUs.
| Topology | 512-GPU Bisection Bandwidth | NCCL AllReduce (512M) | Cost Comparison | Recommendation |
|---|---|---|---|---|
| Fat Tree, 1:1 oversubscription | 100% (full bisection) | 160-180 GB/s | Baseline | Training clusters, any size |
| Fat Tree, 2:1 oversubscription | 50% | 80-100 GB/s | 25-30% lower switch cost | Inference clusters only |
| DGX SuperPOD (NVSwitch spine) | 100% | 180-210 GB/s | 30-40% premium vs separate IB | Optimal for H100 DGX clusters |
| 3-tier Clos (spine + aggregation) | 100% (correct sizing) | 140-160 GB/s | Same as 1:1 fat tree | Clusters > 1024 GPUs |
| Dragonfly+ topology | Near 100% (Slingshot) | 130-150 GB/s | 15-20% less cabling | HPE Cray EX clusters |
