Why InfiniBand Routing Matters for 10,000+ GPU Training
For training runs spanning 1,000 to 10,000+ GPUs, the network fabric is frequently the bottleneck - not compute, not memory bandwidth, but collective communication across nodes. Each all-reduce operation in a training step synchronizes gradients across all participating GPUs. As cluster sizes grow, the communication-to-computation ratio shifts unfavorably: the gradient synchronization volume scales with model size (fixed per step), but the number of participants multiplies the communication complexity. A 70B model using mixed-precision training produces roughly 140GB of gradients per step. On a 1,000-GPU cluster, that is 140MB of data per GPU to reduce across all participants. On a 10,000-GPU cluster, the same 140MB must traverse more hops through a deeper fabric topology.
InfiniBand adaptive routing, introduced with InfiniBand NDR (400 Gbps per lane) and refined in ConnectX-7/8 switch firmware, addresses the fundamental inefficiency of static routing: uneven link utilization. In a static-routed fabric, flows are pinned to specific paths based on a hash of their source/destination pair. When multiple large NCCL all-reduce operations land on the same path, congestion builds on those links while parallel links remain underutilized. Adaptive routing breaks this by making per-packet routing decisions based on real-time fabric congestion, distributing traffic across all available paths.
The throughput impact is substantial. NVIDIA's internal benchmarks on 1,024-GPU clusters show adaptive routing delivering 15-25% higher effective all-reduce bandwidth versus static routing at scale, with the benefit increasing as cluster size grows. At 10,000+ GPUs, the differential widens to 30-40% because the probability of hash collisions in static routing scales super-linearly with the number of concurrent communication streams. The net effect on training throughput (model tokens per second) varies by model architecture but frequently lands in the 10-20% improvement range for large-scale dense transformer training.
Packet Spraying vs Dynamic Path Selection: How NDR Adaptive Routing Works
InfiniBand adaptive routing in NDR400 fabrics operates at two granularities: per-flowlet and per-packet. Per-flowlet adaptive routing, the more commonly deployed mode, groups consecutive packets from the same flow into flowlets separated by a gap of at least the configured inter-flowlet time (typically 1-5 microseconds). Each flowlet can take a different path through the fabric. Flowlet-based routing preserves packet ordering within a flowlet (avoiding TCP-like reordering penalties for NCCL) while still enabling dynamic load balancing across available paths.
Per-packet adaptive routing (often called packet spraying) is the more aggressive mode: every individual packet independently selects its path based on the switch's real-time congestion state. This maximizes load balancing granularity but can cause packet reordering within a single all-reduce operation. NCCL 2.20+ handles packet reordering natively via its out-of-order receive buffers, which reorder packets in GPU memory before passing them to the reduction engine. The performance tradeoff: packet spraying delivers 5-8% higher throughput than flowlet routing in congestion-heavy scenarios but adds roughly 1-2% GPU memory overhead for reorder buffers.
The switch-level mechanism uses a congestion-aware hash function. Each NDR800 switch (the current generation) maintains per-port credit counters and queue depths. The routing decision for each packet evaluates the set of candidate output ports (typically 4-8 ports for a given destination), ranks them by available credit, and selects probabilistically - not deterministically - to avoid thundering-herd effects where all flows converge on the same newly-available port simultaneously. The congestion state is updated every 10-50 nanoseconds, allowing the switch to react to microbursts at per-packet granularity.
Fabric Topology Design for Adaptive Routing: Rail-Optimized vs Dragonfly+
Adaptive routing's effectiveness depends on path diversity in the fabric topology. Rail-optimized topologies (where GPU N of each node connects to the same switch) provide up to 8 paths between any two nodes on an 8-GPU H100 node. Dragonfly+ topologies (used in NVIDIA's DGX SuperPOD reference architecture) provide approximately 4-6 paths for inter-group traffic. Adaptive routing extracts meaningful gains from both, but the topology choice interacts with routing performance in surprising ways.
In rail-optimized fabrics, adaptive routing thrives because the path count is high and the path lengths are equal. A packet from GPU 3 on node A to GPU 3 on node B can traverse any of the 8 leaf switches in the rail group. Flowlet adaptive routing in this configuration achieves near-optimal load balancing with minimal overhead because the equal-cost multi-path (ECMP) set is large and symmetrical. The real-world throughput benefit over static routing in rail-optimized fabrics is approximately 12-15% for NCCL all-reduce at scale.
Dragonfly+ topologies, which use hierarchical groups with high-radix spine connections, benefit more dramatically from adaptive routing because the path lengths and costs vary significantly. Static routing in Dragonfly+ tends to overload the minimal-path (direct group-to-group) links while leaving non-minimal (two-hop through a spine) links underutilized. Adaptive routing detects congestion on the direct path and shifts some traffic to the non-minimal path, potentially increasing per-packet latency by 1-2 microseconds but reducing overall congestion by 30-50% and improving effective all-reduce bandwidth by 20-25% in NCCL benchmarks on 2,048-GPU clusters.
Tuning Adaptive Routing: NWN, NCC, and NCCL Environment Variables
ConnectX-7/8 firmware exposes the adaptive routing mode through the Progressive Congestion Management (PCM) profile. Three profiles ship with firmware 8.2+ for InfiniBand NDR400: profile 0 (static routing, no adaptation), profile 1 (flowlet-based adaptive, 3us flowlet gap), and profile 2 (packet-spraying adaptive, no flowlet gap). Profile selection is per-port, not per-fabric, meaning you can configure different profiles for different traffic classes. Most production deployments use profile 1 for NCCL all-reduce traffic (trading some load-balancing granularity for ordering guarantees) and profile 2 for all-to-all traffic where out-of-order delivery is naturally handled.
The NCCL environment variable NCCL_IB_ADAPTIVE_ROUTING controls how NCCL interacts with the fabric's adaptive routing. Setting NCCL_IB_ADAPTIVE_ROUTING=2 enables out-of-order receive buffers required for packet-spraying mode. Without this setting, NCCL will discard packets arriving out of order, causing catastrophic performance degradation. NCCL_NET_BUFFER_SIZE should be increased to 64MB (from the default 8MB) when using adaptive routing to accommodate larger reorder windows. NCCL_IB_TIMEOUT should be relaxed from 22 to 27 to account for the variable latency of non-minimal paths in Dragonfly+ topologies.
The most misconfigured parameter on adaptive-routed clusters is NCCL_NET_SHARED_BUFFERS. When set to 0 (default), each NCCL communicator allocates its own receive buffers. With adaptive routing's out-of-order delivery, shared buffer pools (NCCL_NET_SHARED_BUFFERS=1) improve memory utilization by permitting receivers to borrow buffer space from idle communicators, reducing OOM risk during large all-reduce operations. The tradeoff: slightly higher buffer management overhead (1-3% CPU utilization) and the requirement that all communicators use the same buffer size, which limits per-communicator tuning flexibility.
Firmware and HCA Compatibility: What to Verify Before Deploying
Adaptive routing in NDR400 fabrics requires both switch firmware and HCA firmware support. The minimum switch firmware for adaptive routing on NVIDIA QM9700 (NDR400) switches is v8.0; v8.2 is recommended for production use. The minimum HCA firmware for ConnectX-7 is v28.38 or later; for ConnectX-8, v42.20 or later. Clusters running older firmware silently fall back to static routing when adaptive routing is configured - the fabric will not error, but you will not get the performance benefit. On ClusterBid-provisioned H100 clusters, all inventory runs firmware v8.2+ on QM9700 switches and v28.40+ on ConnectX-7 HCAs.
Subnet manager configuration is the second critical dependency. OpenSM version 3.5+ (or UFM Enterprise) must enable adaptive routing at the subnet level. The relevant OpenSM option is adaptive_routing 2 (enables per-packet adaptive) or adaptive_routing 1 (flowlet-based). OpenSM must also set routing_algorithm to 6 (adaptive) rather than the default 0 (min-hop). Restarting the subnet manager is a cluster-wide event that temporarily halts all InfiniBand traffic, so this configuration change should be made during maintenance windows. On UFM Enterprise, the equivalent configuration is through the Adaptive Routing Profile under Fabric Management > QoS & Optimization Settings.
Diagnostics tools have evolved to support adaptive routing verification. The ibdiagnet 3.0+ tool includes an adaptive routing plugin (--plugin adaptive_routing) that measures flow distribution across available paths and reports the Jain's fairness index. A fairness index below 0.85 indicates that adaptive routing is not functioning as expected - the plugin will suggest corrective actions such as adjusting the OpenSM adaptive_routing_policy parameters or checking for firmware version mismatches between switches. Running ibdiagnet --plugin adaptive_routing is the first diagnostic step when suspecting suboptimal adaptive routing performance.
When Adaptive Routing Hurts: Small Clusters, High-Precision All-Reduce, and NVL
Adaptive routing is not universally beneficial. On clusters under 256 GPUs, the path diversity is typically too limited for adaptive routing to outperform a well-configured static routing topology. The overhead of maintaining per-packet congestion state (roughly 3-5% switch ASID utilization for adaptive routing logic) can actually reduce performance on small fabrics because the limited number of alternative paths means the routing decisions rarely find a better path than the static hash. For single-node (8-GPU) training using NVLink-only communication, adaptive routing has zero effect - NCCL all-reduce never traverses the fabric.
High-precision (FP32) gradient all-reduce operations benefit less from adaptive routing than FP16 or FP8 operations. The reason: FP32 all-reduce requires 2x the data transfer of FP16, making the operations bandwidth-bound rather than latency-bound. In bandwidth-bound regimes, static routing's worst-case path is closer to the average path performance because the link saturation masks the distribution inefficiency. Adaptive routing's benefit shrinks from 15-25% to 5-10% when operating in bandwidth-saturated mode. Most large-scale training uses FP16 master weights with FP32 accumulation, keeping the all-reduce in FP16 and preserving the adaptive routing advantage.
NVIDIA's NVLink Switch System (for DGX GB200 and larger NVL72 configurations) changes the calculus completely. NVLink Switch provides 1,800 GB/s GPU-to-GPU bandwidth within a 72-GPU domain - roughly 180x faster than a single NDR400 InfiniBand link. For training workloads that fit within a single NVLink domain, adaptive routing on InfiniBand is irrelevant. The InfiniBand fabric only carries inter-domain communication, which is typically 10-20% of total gradient data in model-parallel training configurations. Adaptive routing benefits remain real for the inter-domain traffic but apply to a much smaller fraction of the total communication volume. Teams migrating to NVL72 should reassess whether the complexity and per-switch ASID overhead of adaptive routing is worthwhile given the reduced inter-domain bandwidth fraction.
