All essays
TechnicalDEEP DIVEFEB 2026

GPU Network Troubleshooting: RDMA, NCCL Issues, Packet Loss, and Congestion

Diagnose and fix GPU cluster network problems. RDMA (InfiniBand vs RoCE v2) debugging, NCCL timeout analysis, packet loss detection, fabric congestion, and topology optimization for H100/B200 clusters.

01

GPU FABRIC ARCHITECTURE AND BASELINE EXPECTATIONS

GPU cluster networks operate at performance levels that leave zero margin for misconfiguration. A single H100 GPU on NDR400 InfiniBand can saturate 400 Gbps bidirectional. A node with 8x H100 GPUs connected to a fabric of 64-256 nodes creates aggregate traffic loads exceeding 10 Tbps for NCCL all-reduce operations. The two dominant fabric choices are HDR200/NDR400 InfiniBand (60 percent of large GPU clusters) and RoCE v2 (Remote Direct Memory Access over Converged Ethernet, 35 percent of clusters, primarily in latency-tolerant training workloads). The remaining 5 percent use NVIDIA Spectrum-X Ethernet with adaptive routing and congestion control for GPU-enhanced Ethernet.

Establishing baseline fabric behavior is the first troubleshooting step. For InfiniBand, the `ibstatus` and `ibstat` commands report link state, rate (400 Gbps = 4X NDR, 200 Gbps = 4X HDR), and physical width (4X, 8X, 12X). Expected output: `rate: 400 Gbps (4X NDR), link_layer: InfiniBand, state: 4 (ACTIVE)`. For RoCE v2, `mlxstat` and `ibdev2netdev -v` map RDMA devices to ethernet interfaces and report link speed. The NCCL `topology.xml` file defines the cluster's hierarchical topology (NVSwitch within DGX, InfiniBand switch between DGX, core spine layer). Any deviation from the physical wiring topology in this file will cause NCCL to use suboptimal all-reduce algorithms.

ParameterNDR400 InfiniBand (Healthy)HDR200 InfiniBand (Healthy)RoCE v2 200 GbE (Healthy)
Link Rate (ibstatus)400 Gbps (4X NDR)200 Gbps (4X HDR)200 Gbps (Ethernet, PFC on)
Link LayerInfiniBandInfiniBandEthernet
Link Active statestate 4 (ACTIVE)state 4 (ACTIVE)Link Up, LLDP enabled
NCCL AllReduce 8xH100 (1 node)25.2 GB/s18.5 GB/s21.8 GB/s
NCCL AllReduce 64xH100 (8 nodes)160-180 GB/s110-130 GB/s120-145 GB/s
NCCL AllReduce 512xH100 (64 nodes)800-1100 GB/s500-700 GB/s600-850 GB/s
Max Pause Frames (rdma_pause)N/A (lossless)N/A (lossless)< 0.01% of total frames
Switch Buffer256 MB per port (Mellaness QM9790)96 MB per port (QM9700)42 MB per port (SN4600C)
02

NCCL TROUBLESHOOTING: TIMEOUTS, TOPOLOGY, AND ALGORITHM SELECTION

NCCL (NVIDIA Collective Communications Library) is the most frequent source of training failures in multi-node GPU clusters. The most common error is NCCL timeout: `ncclInternalError: NCCL operation timed out` after a default 30-second watchdog, which means an all-reduce or all-gather operation did not complete within the timeout window. The first diagnostic step sets `NCCL_DEBUG=INFO` and `NCCL_DEBUG_SUBSYS=ALL` in the training container environment variables, which produces detailed timing logs for each collective operation. The key line to look for: `nvidia[rank][n]: NCCL INFO comm 0x... rank n nranks m busId ... time 12.345` shows the time taken for the debug NCCL communicator initialization.

NCCL timeout causes break into three categories: topology mismatch (NCCL thinks GPUs are on different nodes, causing cross-node collectives that should be intra-node), switch congestion (buffer exhaustion on core spine links, causing backpressure that delays completion beyond timeout), and silent link failures (NVLink cable or backplane failure that causes retransmission at the link layer, increasing completion time from 50 microseconds to 5+ seconds). For topology mismatch, `nvidia-smi topo -m` prints the GPU-to-GPU connectivity matrix; if NVLink1/NVLink2 connections show as PIX (via PCIe) instead of NV# (direct NVLink), the NVLink bridge or backplane needs inspection. For switch congestion, `perfquery` on InfiniBand switches reports port counters including `SymbolErrorCounter`, `LinkErrorRecoveryCounter`, and `PortRcvConstraintErrors`.

NCCL ErrorRoot CauseDiagnostic CommandFix / Workaround
NCCL Timed Out (30s)Congested spine linkperfquery on IB switch, look for XmitWait > 1e12Re-route: set NCCL_IB_TIMEOUT=22, increase agg buffer
NCCL Internal ErrorTopology mismatch in NCCL_TOPO_FILEnvidia-smi topo -m, compare pinningCorrect NCCL topology.xml or use NCCL_TOPO_DUMP_FILE
NCCL Remote ErrorNode crash during all-reducedmesg on peer node, check for GPU fallen off busReplace or reboot failed node, NCCL_IB_RETRY_CNT=7
NCCL Invalid UsageCUDA stream conflictNCCL_DEBUG=INFO, check stream synchronizationUse NCCL_GROUP_CUDART_VERSION=12040
NCCL Connection FailureIB link down on one nodeibstatus on all nodes, compare link stateReseat IB cable, restart opensm or subn manager
NCCL Unhandled Cuda ErrorOOM during collective buffer allocationnvidia-smi --query-gpu=memory.free --format=csvReduce --gradient-accumulation-steps or batch size
03

RDMA DIAGNOSTICS: INFINIBAND AND ROCE V2 TROUBLESHOOTING

InfiniBand fabric issues manifest as CRC errors, link flaps, and symbol errors. The primary diagnostic tool is `perfquery` (part of infiniband-diags), which queries switch and HCA port counters. The critical counters are `SymbolErrorCounter` (should be zero; non-zero indicates physical layer issues with optical transceivers or cables), `LinkErrorRecoveryCounter` (should be zero; non-zero indicates link training failures), `PortRcvErrors` (should be less than 0.001 percent of total received packets; higher indicates signal integrity problems), and `ExcessiveBufferOverrunErrors` (non-zero indicates buffer allocation mismatch between sender and receiver). The command `perfquery -x -P 0 $SWITCH_LID $PORT` dumps extended counters with per-VL (Virtual Lane) breakdown.

RoCE v2 troubleshooting requires a different toolkit because the fabric uses standard Ethernet with Priority Flow Control (PFC, IEEE 802.1Qbb). The most common RoCE v2 issue is PFC storm: when one flow causes buffer exhaustion on an egress port, the switch sends PFC pause frames that halt all traffic on that priority class, including unrelated RDMA traffic. `mlnx_perfquery` or `ethtool -S $IFACE | grep -E "pfc|pause|deadlock"` reports PFC pause frame counts. If `tx_pause` or `rx_pause` exceed 0.1 percent of total packets, the fabric has a head-of-line blocking problem. The fix involves enabling Explicit Congestion Notification (ECN) with DCQCN (Data Center Quantized Congestion Notification), `echo 1 > /sys/class/net/$IFACE/ecn/roce_enable`, and configuring RED/ECN thresholds on switches to mark congested packets before PFC is triggered.

04

PACKET LOSS DETECTION AND CONGESTION MANAGEMENT

Packet loss in GPU fabrics has a disproportionate impact: a single lost packet in an NCCL all-reduce required 8 GB buffer transfer triggers retransmission of the entire buffer, increasing completion time from 50 milliseconds to 500+ milliseconds. For InfiniBand, packet loss is invisible at the application level because the fabric is nominally lossless (credit-based flow control), but buffer overruns at congested switch ports cause the same effective retransmission penalty. The indicator is `XmitWait` counter in `perfquery` output: this measures the number of symbol periods a port waited to send because no credits were available. A `XmitWait` value exceeding 1e12 between two consecutive queries (5-second interval) indicates persistent congestion on that port.

RoCE v2 packet loss is directly measurable from `ethtool -S` counters. The command `ethtool -S $IFACE | grep -E "rx_dropped|tx_dropped|errors"` reports physical packet drops. Additionally, `rdma statistic show` reveals RDMA-level retransmission: `hir` (hardware inbound retransmission) and `hor` (hardware outbound retransmission) counts above 0 for a stable fabric warrant investigation. The remediation for RoCE v2 clusters experiencing >0.01 percent packet loss is to implement adaptive routing: NVIDIA Spectrum switches support Dynamic Load Balancing (DLB) that distributes RDMA flows across available paths, reducing congestion on hot spots. The configuration is `mlnx_qos -i $IFACE --trust dscp && echo 1 > /sys/class/net/$IFACE/adaptive_routing". For InfiniBand, adaptive routing is enabled through the OpenSM subnet manager with `--advanced-routing 1` flag.

05

FABRIC TOPOLOGY OPTIMIZATION FOR NCCL ALL-REDUCE

NCCL all-reduce performance depends directly on fabric topology. The optimal topology for large GPU clusters is a multi-layer fat tree (Clos topology) with full bisection bandwidth: each leaf switch connects to exactly the same number of spine switches, and each spine switch connects to every leaf switch. For a 512-GPU H100 cluster on NDR400 InfiniBand, the topology uses 8 leaf switches (16 ports each, connecting 8 DGX H100 nodes per leaf) connected to 8 spine switches in a full-mesh. Any oversubscription ratio above 1:1 at the spine layer directly reduces all-reduce bandwidth proportionally: a 2:1 oversubscription reduces bisection bandwidth to 50 percent, increasing all-reduce time from 160 GB/s to 80 GB/s for a 512-GPU cluster.

The NCCL topology file (`topology.xml`) must reflect the physical wiring exactly. NVIDIA provides a `nccl-topology-dump` tool that generates the topology file from `nvidia-smi topo -m` output combined with switch fabric discovery. The file specifies GPU-to-NIC affinity (which GPU shares a PCIe switch with which InfiniBand HCA), NIC-to-switch port mapping, and inter-switch spine connections. Common mistakes include assuming NVSwitch topology from a single node applies identically to all nodes (H100 DGX requires per-node topology validation), incorrect `guid` values for InfiniBand switch ports that cause NCCL to compute incorrect routing, and missing `link type=NVSWITCH` entries that force NCCL to use PCIe P2P instead of NVLink for intra-node communication. The `nccl-tests` benchmark (allreduce_perf -b 8 -e 128M -f 2) validates the topology by measuring bandwidth across all 512 GPUs.

Topology512-GPU Bisection BandwidthNCCL AllReduce (512M)Cost ComparisonRecommendation
Fat Tree, 1:1 oversubscription100% (full bisection)160-180 GB/sBaselineTraining clusters, any size
Fat Tree, 2:1 oversubscription50%80-100 GB/s25-30% lower switch costInference clusters only
DGX SuperPOD (NVSwitch spine)100%180-210 GB/s30-40% premium vs separate IBOptimal for H100 DGX clusters
3-tier Clos (spine + aggregation)100% (correct sizing)140-160 GB/sSame as 1:1 fat treeClusters > 1024 GPUs
Dragonfly+ topologyNear 100% (Slingshot)130-150 GB/s15-20% less cablingHPE Cray EX clusters
Filed under
NCCL DebuggingRDMA InfiniBand TroubleshootingRoCE v2 IssuesGPU Fabric CongestionNVIDIA Mellanox Switch DebugNCCL Topology OptimizationGPU Network Packet Loss