All essays
InfrastructureINFRASTRUCTUREFEB 2026

GPU Networking Deep Dive: InfiniBand vs RoCE vs Proprietary Interconnects

In-depth comparison of GPU cluster networking technologies covering InfiniBand NDR400, RoCE v2, NVIDIA NVLink, and AWS EFA with real benchmark data for distributed training.

01

THE GPU CLUSTER NETWORKING LANDSCAPE

GPU cluster networking has become the primary bottleneck for distributed training at scale. Large language model training on clusters exceeding 256 GPUs spends 15-40 percent of total time on communication rather than computation. The choice of interconnect technology directly determines training efficiency, with InfiniBand NDR400 delivering 400 Gbps per port versus RoCE v2 at 200-400 Gbps and NVLink reaching 900 GB/s within an HGX baseboard.

NVIDIA NVLink/NVSwitch dominates intra-node communication with 900 GB/s bidirectional bandwidth across 8 H100 GPUs. For inter-node, InfiniBand holds 76 percent market share in GPU clusters above 1,000 GPUs. RoCE v2 adoption grew from 12 percent in 2023 to 24 percent in 2025, driven by lower cost and Broadcom/Mellanox competition.

TechnologyBandwidthLatencyTopology$/PortEfficiency @ 1K GPUs
NVLink 4.0900 GB/s0.1-0.3 usHybrid cube$0 (included)N/A (intra-node)
InfiniBand NDR400400 Gbps0.6-1.0 usFat tree$2,80092-96%
InfiniBand HDR200200 Gbps0.8-1.2 usFat tree$1,50085-90%
RoCE v2 400G400 Gbps1.5-3.0 usCLOS$1,60070-78%
AWS EFA200 Gbps2.0-5.0 usTree$0.65/hr65-72%
02

SCALING EFFICIENCY BENCHMARKS

Scaling efficiency measures how effectively a distributed training job utilizes added GPUs. All-Reduce benchmark tests on 512 H100 GPUs show InfiniBand NDR400 achieves 96 percent scaling efficiency versus 82 percent for RoCE v2 at 400 Gbps. The difference widens at 1,024 GPUs: InfiniBand maintains 92 percent while RoCE drops to 68 percent due to PFC pauses and hash collision packet drops.

Ring All-Reduce bandwidth on InfiniBand NDR400 reaches 380 Gbps per direction, utilizing 95 percent of link capacity. RoCE v2 with DCQCN congestion control achieves 270-310 Gbps (68-78 percent utilization). SHARP in-network computing on InfiniBand switches reduces All-Reduce completion time by 35-50 percent compared to receiver-side reduction.

03

NETWORKING COST AND TCO ANALYSIS

For a 256-GPU cluster with 32 nodes, InfiniBand NDR400 adds approximately $89,600 in switch and cable costs versus $51,200 for RoCE v2 at equivalent port count. However, the 14 percent training efficiency advantage of InfiniBand means jobs complete faster, offsetting the $38,400 networking premium within 3-5 months for clusters running at 70 percent utilization.

Total networking cost as a percentage of cluster spend varies: InfiniBand NDR400 represents 8-12 percent of total cluster cost, RoCE v2 represents 5-8 percent, and NVLink is included in HGX baseboard cost. For clusters exceeding $10 million total cost, the networking delta of 3-4 percent translates to $300,000-$400,000 in infrastructure budget.

04

NETWORK TOPOLOGY DESIGN FOR TRAINING

Fat-tree topology dominates GPU cluster deployments, with 2:1 blocking ratio as the standard design for training clusters. Non-blocking fat tree with 1:1 ratio adds 40-60 percent networking cost but delivers 8-12 percent better scaling efficiency at 1,024 GPUs. Dragonfly+ topology reduces cable count by 35 percent versus fat tree while maintaining 90 percent of InfiniBand performance.

Over-subscription ratio is the critical design parameter. A 4:1 oversubscribed network with RoCE v2 on 128 GPUs loses 35 percent scaling efficiency compared to 2:1 ratio. For clusters dedicated to inference, 8:1 oversubscription is acceptable since inference traffic is primarily point-to-point rather than collective.

Filed under
InfiniBandRoCENVLinkGPU NetworkingDistributed TrainingAWS EFAInterconnect