All essays
BenchmarkCOMPARISONFEB 2026

NVLink vs InfiniBand vs Ethernet: GPU Cluster Networking Decoded for AI Buyers

Lambda deployed InfiniBand CPO and CoreWeave deployed 102.4 Tb/s Spectrum-X. Accessible buyer

01

THE GPU NETWORKING LANDSCAPE

GPU cluster networking has bifurcated into three tiers. NVLink handles GPU-to-GPU within racks at 900 GB/s bidirectional per GPU. InfiniBand NDR400 connects racks at 400 Gb/s per link with sub-1 microsecond latency. Ethernet with RoCEv2 and Spectrum-X reaches 400-800 Gb/s with 3-5 microsecond latency. The choice determines training throughput, scaling efficiency, and per-port cost.

03

INFINIBAND: THE TRAINING STANDARD

InfiniBand NDR400 delivers 400 Gb/s per link with sub-1 microsecond latency and 0.00001% packet loss. Lambda's CPO integration reduces power by 30% and latency by 15%. Scaling efficiency reaches 90-95% for models up to 1T parameters across 1,024 GPUs. Cost is $1,500-2,500 per port versus $500-800 for high-end Ethernet.

FeatureNVLink 5InfiniBandSpectrum-X
Bandwidth/port1.8 TB/s400 Gb/s800 Gb/s
Latency<100 ns<1 us3-5 us
Max GPUs/domain57665,53632,000
Cost/portN/A$1,500-2,500$600-1,200
Scaling eff.95-98%90-95%80-90%
04

ETHERNET: THE DARK HORSE

CoreWeave deployed 102.4 Tb/s Spectrum-X Ethernet with RoCEv2 achieving 95% of InfiniBand training performance at 50-70% cost. Ethernet advantages include existing operational tooling, broader vendor ecosystem, and lower training costs. For clusters under 512 GPUs, modern Ethernet delivers 85-95% of InfiniBand scaling efficiency at 40-60% lower networking cost.

05

COST COMPARISON

For a 256-GPU cluster: InfiniBand networking adds $400,000-600,000 ($1,500-2,500/port) versus $150,000-250,000 for Spectrum-X Ethernet. The networking cost premium is 8-15% of total cluster cost for InfiniBand versus 3-6% for Ethernet. For clusters under 1,000 GPUs, Ethernet savings of $200,000-500,000 can fund 20-50 additional GPU-hours per day.

06

SELECTION FRAMEWORK

Choose NVLink-only for single-rack deployments (8-72 GPUs) without inter-rack training. Choose InfiniBand for clusters above 512 GPUs or models above 100B parameters requiring 95%+ scaling efficiency. Choose Ethernet for clusters under 512 GPUs, budget-constrained deployments, or teams with existing Ethernet operational expertise. Most production deployments combine all three: NVLink within racks, InfiniBand or Ethernet between racks.

Filed under
NVLinkInfiniBandEthernetGPU NetworkCluster FabricAI TrainingInterconnect