Why GPU Networking Is Different
GPU cluster networking must support communication patterns that are fundamentally different from traditional data centre traffic. All-reduce operations during training require every GPU to exchange gradients with every other GPU simultaneously, creating incast congestion that standard TCP/IP networking handles poorly. A single all-reduce on 1,024 GPUs can generate 8 TB of internal traffic in milliseconds.
The networking requirements scale with cluster size. An 8-GPU node uses NVLink for intra-node communication (900 GB/s on B200). For multi-node training, InfiniBand or high-speed Ethernet handles inter-node traffic (50-100 GB/s per GPU). The ratio of intra-node to inter-node bandwidth determines training efficiency, and this ratio must be carefully engineered.
At mid-2026, the dominant GPU networking architectures are NVLink + InfiniBand (60% of AI clusters), NVLink + RoCEv2 (25%), and pure Ethernet with Spectrum-X (15%). The choice depends on cluster size, budget, and latency sensitivity. This post covers each topology with real deployment configurations and cost data.
Top-of-Rack Design: GPU Node Cabling
The Top-of-Rack (ToR) design for GPU clusters centres on the HGX baseboard. An 8-GPU HGX H100 or B200 baseboard integrates NVLink switches for intra-node GPU communication. Each GPU connects to the NVSwitch via NVLink 4.0 (H100) or NVLink 5.0 (B200), providing 900 GB/s bidirectional bandwidth per GPU. The 8 GPUs form a fully connected mesh through the NVSwitch with no oversubscription.
For inter-node networking, each GPU HBA (Host Fabric Interface) connects to the ToR switch. The standard configuration is 8x 400 Gbps InfiniBand or Ethernet ports per HGX (one per GPU), connecting to 2x ToR switches for redundancy. Each ToR switch serves 16-32 GPU nodes (128-256 GPUs) at 400-800 Gbps per port, providing non-blocking bandwidth at the ToR level.
The cabling complexity is significant. A 1,024-GPU cluster requires approximately 1,024 GPU-to-ToR cables (400 Gbps each), 16-32 ToR switches, 8-16 spine switches, and 512-2,048 inter-switch cables depending on oversubscription ratio. Cable management and documentation are not optional -- a single mislabelled or loose cable can cause intermittent NCCL timeouts that are extremely difficult to diagnose.
Spine-Leaf Architecture for GPU Traffic
The spine-leaf architecture is standard for GPU cluster networking because it provides predictable, low-latency connectivity between any two nodes. Each ToR (leaf) switch connects to every spine switch, creating a Clos topology where any GPU can reach any other GPU with at most one hop through the spine.
The key design parameter is the oversubscription ratio: the ratio of ToR-to-GPU bandwidth to ToR-to-spine bandwidth. At 1:1 oversubscription (full bisection bandwidth), each ToR switch has equal uplink bandwidth to the spine as downlink bandwidth to GPUs. This is the recommended ratio for training clusters. At 2:1 or 4:1 oversubscription (cost-reduced), some communication patterns experience congestion.
The table below shows standard spine-leaf configurations for different cluster sizes. The cost difference between 1:1 and 4:1 oversubscription is approximately 2-3x, but the training throughput impact depends on workload communication patterns.
| Cluster Size | ToR Switches | Spine Switches | Oversubscription | Network Cost/GPU/Month |
|---|---|---|---|---|
| 128 GPUs | 2 (64 ports) | 2 (32 ports) | 1:1 | $0.28 |
| 256 GPUs | 4 (64 ports) | 4 (32 ports) | 1:1 | $0.25 |
| 512 GPUs | 8 (64 ports) | 8 (64 ports) | 1:1 | $0.22 |
| 1,024 GPUs | 16 (64 ports) | 16 (64 ports) | 1:1 | $0.20 |
| 1,024 GPUs | 8 (64 ports) | 8 (64 ports) | 4:1 | $0.08 |
| 2,048 GPUs | 32 (64 ports) | 32 (64 ports) | 2:1 | $0.18 |
InfiniBand vs RoCEv2 vs Spectrum-X
The inter-node networking choice is the most consequential decision in GPU cluster design. At mid-2026, the three main options are InfiniBand (NVIDIA Quantum-2 and Quantum-X800), RoCEv2 (standard Ethernet with lossless configuration), and Spectrum-X (NVIDIA's Ethernet solution with adaptive routing and congestion control).
InfiniBand remains the gold standard for GPU clusters, offering 800 Gbps HDR (Quantum-2) with sub-microsecond latency, hardware-based congestion control, and native support for NCCL RDMA. Market share for new GPU clusters in 2026: InfiniBand 55%, RoCEv2 25%, Spectrum-X 20%. RoCEv2 has gained ground due to lower cost and familiarity with Ethernet operations. Spectrum-X positions between the two, offering InfiniBand-like performance on Ethernet infrastructure.
The cost differential is significant: InfiniBand adds $0.15-0.30/GPU-hour for a 1,024-GPU cluster. RoCEv2 adds $0.08-0.15/GPU-hour. Spectrum-X adds $0.12-0.22/GPU-hour. For a cluster running 24/7, this translates to $100,000-500,000/year in networking costs depending on the choice. The decision depends on whether the 5-15% training throughput difference justifies the cost.
NVLink Fabric Integration with Inter-Node Networking
The interface between NVLink (intra-node) and InfiniBand/Ethernet (inter-node) is the critical performance boundary. Data leaving the NVLink domain must traverse PCIe to the NIC, then the network fabric to the destination node, then PCIe back to GPU memory. This path has significantly higher latency and lower bandwidth than NVLink direct.
Optimising this boundary involves: NIC placement in the same NUMA domain as the GPU to minimise PCIe traversal, GPUDirect RDMA (GDR) for NIC-to-GPU direct memory access without CPU involvement, and balanced NIC-to-GPU ratio (typically 1 NIC per GPU for training clusters).
The GDR performance on B200 with Quantum-2 InfiniBand achieves 380-420 Gbps per GPU for NCCL all-reduce, compared to the NVLink intra-node bandwidth of 900 GB/s (7,200 Gbps). The 17:1 ratio between intra-node and inter-node bandwidth means that training efficiency drops rapidly as the number of nodes increases. At 16 nodes (128 GPUs), inter-node communication dominates training time for models with high communication-to-computation ratio.
Network Cost Optimisation Strategies
Network costs can represent 10-20% of total GPU cluster expenditure. Optimisation strategies include: right-sizing oversubscription ratio based on workload communication patterns, using RoCEv2 for latency-tolerant workloads and reserving InfiniBand for communication-heavy training, leveraging multi-network topologies (e.g., separate networks for training and storage traffic), and negotiating multi-year contracts with InfiniBand vendors for 20-30% discounts.
A common cost-optimisation approach is a tiered network fabric: full 1:1 non-blocking InfiniBand for the core training cluster (60% of GPUs), 2:1 oversubscribed InfiniBand for development and experimentation (30%), and RoCEv2 for inference, storage, and management traffic (10%). This reduces network costs by 25-35% compared to a uniform fabric.
The total network cost for a 1,024-GPU B200 cluster over 3 years is approximately $1.5-3M (InfiniBand), $0.8-1.5M (RoCEv2), or $1.2-2.2M (Spectrum-X). At 1:1 oversubscription, network hardware costs approximately $250-350 per GPU port, which translates to $0.20-0.30/GPU-hour amortised over 3 years.
Future Trends: 1.6T Ethernet and NVLink Switch Systems
Two major networking developments are shaping GPU cluster design. 1.6 Tbps Ethernet (800G per lane) is standardising through IEEE 802.3dj and will reach production GPU clusters in late 2026 to early 2027. This provides a 2x bandwidth improvement over current 800 Gbps InfiniBand, potentially making Ethernet the preferred choice for next-generation clusters.
NVIDIA's NVLink Switch System (introduced with DGX SuperPOD and expanded with B200-based systems) creates a fully non-blocking NVLink fabric across up to 256 GPUs. This eliminates the inter-node bandwidth bottleneck by treating the entire cluster as a single GPU memory domain. Early adopters report 1.5-2x training throughput improvement for communication-heavy workloads compared to InfiniBand-connected clusters.
The roadmap for GPU networking: 2026 clusters use 800 Gbps InfiniBand or RoCEv2 with NVLink 5.0 within nodes. 2027 clusters will use 1.6T Ethernet with NVLink Switch for inter-node GPU fabric, achieving 400+ GB/s per GPU for both intra-node and inter-node communication. This convergence of intra-node and inter-node bandwidth is the single most important architectural trend for GPU networking.
