The 2026 Fabric Cost Landscape
GPU compute cost gets all the attention. At $2.02/GPU/hr for H200 SXM5 or $3.80/GPU/hr for B200, every dollar of GPU-hour spend is visible in your cloud bill. The fabric connecting those GPUs is usually invisible - bundled into the node rate, quoted as a single line item, or simply not disclosed. For a 256-GPU cluster running 90 days, fabric alone can add $150,000 to $500,000 to the total cost, depending on the interconnect you choose and the oversubscription ratio the provider actually builds.
Three interconnects dominate the mid-2026 landscape. InfiniBand NDR400 is the incumbent for high-performance AI training, with Quantum-X800 XDR at 800 Gb/s per port starting deployments in H2 2026. Spectrum-X Ethernet has closed the latency gap to within striking distance for pipeline-parallel workloads. RoCEv2 on commodity Ethernet remains the budget option. Per-port pricing ranges from roughly $500 for a basic RoCEv2 NIC to over $5,000 for a fully loaded InfiniBand HCA with SHARP licensing. The total fabric cost depends as much on your topology as on your interconnect choice, and many providers oversubscribe the spine layer to hit a target node price.
This analysis covers three cluster scales - 64 GPUs, 256 GPUs, and 1,024 GPUs - using realistic mid-2026 pricing for networking hardware, optics, cabling, and the power and cooling overhead the fabric draws. We exclude GPU cost and compute node cost to isolate the fabric decision. Every cluster size is modeled with both 1:1 non-blocking and 4:1 oversubscribed topologies, because the biggest fabric pricing mistake buyers make is paying for non-blocking InfiniBand when their parallelism strategy does not benefit from it.
Per-Port Pricing: What Each Technology Costs at the Component Level
InfiniBand NDR400 carries the highest per-port cost. A single ConnectX-7 NDR400 HCA (dual-port) runs $2,200-2,800 list per card. Quantum QM9700 switch ports cost approximately $1,200-1,600 per port in a fully configured chassis. Active optical cables at 400G add $300-500 per link for 3-meter lengths, climbing to $1,200+ for 30-meter runs between racks. Total per-GPU port cost in a leaf-spine design: roughly $2,800-3,800 when you average HCA and switch port costs across the cluster. SHARP in-network compute licensing adds $400-700 per port if enabled.
Spectrum-X SN6600 with ConnectX-8 NICs comes in lower. ConnectX-8 dual-port 400G NICs run $1,600-2,100 list, roughly 25% below equivalent InfiniBand HCAs. SN6600 switch ports are approximately $800-1,200 per port in a fully loaded chassis. Optics are identical to InfiniBand - same 400G FR4/LR4 optics or AOCs - so no savings there. Total per-GPU port cost with Spectrum-X: $1,800-2,600. The gap versus InfiniBand narrows when you include SN6600's adaptive routing and congestion control licensing, but the hardware base remains cheaper.
RoCEv2 on commodity Broadcom or Marvell switches is the cheapest path. ConnectX-7 or ConnectX-8 NICs are the same hardware as InfiniBand - the difference is firmware. A ConnectX-7 running RoCEv2 firmware costs the same $2,200-2,800 as the InfiniBand version, but the switch port cost drops dramatically. A Broadcom Tomahawk 5 51.2T switch costs roughly $300-500 per 400G port. Total per-GPU port cost with commodity RoCEv2: $1,200-1,800. The savings come entirely from the switch layer and from avoiding the InfiniBand subnet manager operational overhead.
| Component | InfiniBand NDR400 | Spectrum-X Ethernet | Commodity RoCEv2 |
|---|---|---|---|
| NIC/HCA per GPU (dual-port 400G) | $2,200-2,800 | $1,600-2,100 | $2,200-2,800 |
| Switch port cost (per 400G port) | $1,200-1,600 | $800-1,200 | $300-500 |
| AOC 3m cable | $300-500 | $300-500 | $300-500 |
| SHARP / congestion control license | $400-700/port | $200-400/port | N/A |
| Total per-GPU port (leaf-spine) | $2,800-3,800 | $1,800-2,600 | $1,200-1,800 |
| Fabric power overhead per port | 18-25W | 14-20W | 10-15W |
64-GPU Fabric Cost: $48K to $152K Total
A 64-GPU cluster running modern training workloads needs 8 leaf switches and 2 spine switches in a standard leaf-spine topology, assuming 1:1 non-blocking. Eight GPUs per node means 8 nodes, each with 8 NIC ports (one per GPU). That is 64 leaf-side ports and 64 spine-side ports. The math is the same regardless of interconnect because every GPU needs a fabric connection - the cost per port is where the technologies diverge.
At 64 GPUs, the total fabric cost difference between InfiniBand NDR400 and commodity RoCEv2 is approximately $104,000. That delta matters for a single cluster but is not the primary decision driver at this scale - you are already spending $300,000-600,000 on GPU compute over a six-month run. The decision at 64 GPUs is whether your training recipe benefits from InfiniBand's 1-2 microsecond latency and SHARP in-network reduction. For small-to-medium training runs typical of single-node tensor parallelism with 8-way data parallelism, Spectrum-X at $76,000 for the fabric is the pragmatic choice.
RoCEv2 on commodity switches at $48,000 is viable only if your workload does not saturate the fabric with all-reduce operations. RoCEv2's sensitivity to packet loss under incast patterns means a fully utilized 64-GPU fabric can see all-reduce throughput drop 30-50% versus InfiniBand during peak gradient synchronization. If your training job averages below 70% fabric utilization, RoCEv2 works. Above that, the savings disappear in idle GPU time waiting for gradients.
| Fabric Type | 64 GPUs (1:1) | 64 GPUs (4:1 oversubscribed) |
|---|---|---|
| InfiniBand NDR400 | $152,000 | $89,000 |
| Spectrum-X | $76,000 | $52,000 |
| Commodity RoCEv2 | $48,000 | $34,000 |
256-GPU Fabric Cost: $136K to $608K, Where Most Decisions Are Made
256 GPUs is the most common cluster size for serious model training in mid-2026. This is the scale where a startup has raised its Series A, signed a 6-12 month reserved contract, and is training models in the 7B-70B parameter range. At 32 nodes of 8-GPU machines, the fabric requires 32 leaf switches and 8 spine switches for 1:1 non-blocking - 256 leaf ports, 256 spine ports. The cost range widens considerably because the switch count is high enough that the per-switch price difference between InfiniBand and commodity Ethernet is magnified.
The gap between InfiniBand at $608,000 and RoCEv2 at $136,000 for a non-blocking fabric is $472,000. That is real money. Over a 12-month reserved contract, $472,000 spread across 256 GPUs adds $1.54/GPU/hr to your effective cost above the base compute rate. If your base compute rate for H200 SXM5 is $2.02/GPU/hr, InfiniBand at non-blocking 1:1 would add roughly 76% to your effective hourly rate. Most teams do not need non-blocking InfiniBand at 256 GPUs, and the providers who sell it at non-blocking ratios often do not disclose the fabric cost breakdown in their quotes.
Spectrum-X at $304,000 for non-blocking hits a sweet spot at 256 GPUs. The fabric cost adds roughly $0.99/GPU/hr to an effective rate of $3.01/GPU/hr including compute. For pipeline-parallel training at this scale - which is what most 70B model training looks like - Spectrum-X delivers the same effective throughput as InfiniBand because pipeline parallelism's latency tolerance eliminates the InfiniBand latency advantage. At 4:1 oversubscription, Spectrum-X drops to $124,000 fabric cost, adding only $0.40/GPU/hr to the base compute rate.
| Fabric Type | 256 GPUs (1:1) | 256 GPUs (4:1) |
|---|---|---|
| InfiniBand NDR400 | $608,000 | $232,000 |
| Spectrum-X SN6600 | $304,000 | $124,000 |
| Commodity RoCEv2 | $136,000 | $66,000 |
1,024-GPU Fabric Cost: $544K to $2.43M - Hyperscaler Math
At 1,024 GPUs (128 nodes), you are building a fabric equivalent to a small supercomputer. Non-blocking InfiniBand requires 128 leaf switches and 32 spine switches, totaling 1,024 leaf-side ports and 1,024 spine-side ports. The fabric cost reaches $2.43 million at InfiniBand pricing - before power and cooling, before cabling labor, before the dedicated InfiniBand subnet manager staffing. This is why most hyperscaler AI clusters above 5,000 GPUs run Ethernet variants rather than InfiniBand: the fabric cost at InfiniBand pricing for a 10,000-GPU cluster exceeds $20 million and becomes a material fraction of total cluster cost.
Spectrum-X at $1.05 million for non-blocking at 1,024 GPUs is the choice that most large-scale AI training teams are making in mid-2026. The fabric cost premium versus RoCEv2 ($544,000) is $506,000, but the throughput reliability under all-reduce patterns at scale justifies it. At 4:1 oversubscription - which many providers use for inference-heavy or fine-tuning clusters where continuous all-reduce is less frequent - Spectrum-X drops to $280,000, and RoCEv2 to $178,000. For training clusters running 50B+ parameter models with sustained all-reduce across all 128 nodes, 4:1 oversubscription hurts throughput regardless of interconnect choice.
The 1,024-GPU scale is also where co-packaged optics (CPO) economics begin to matter. Lambda's June 2026 CPO deployment at GTC showed 30-40% lower power per port at 400G, translating to roughly $12,000-18,000 annual power savings per switch. Over the 32-spine switch layer of a 1,024-GPU cluster, CPO saves $384,000-576,000 in power costs over a 3-year equipment lifecycle. CPO is currently available only at InfiniBand pricing with NVIDIA's Quantum-X800 line, but Broadcom and Marvell have announced Spectrum-X and Ethernet CPO variants for late 2026 delivery.
| Fabric Type | 1,024 GPUs (1:1) | 1,024 GPUs (4:1) | 3-year power cost |
|---|---|---|---|
| InfiniBand NDR400 | $2,430,000 | $890,000 | $1,120,000 |
| Spectrum-X SN6600 | $1,050,000 | $280,000 | $790,000 |
| Commodity RoCEv2 | $544,000 | $178,000 | $580,000 |
Bandwidth, Latency, and Model FLOP Utilization: What Actually Matters
Model FLOP utilization (MFU) is the only metric that connects fabric cost to training cost. A cluster that achieves 75% MFU completes a 30-day training run in 40 wall-clock days at 55% MFU. That 15-day schedule slip costs approximately $495,000 in additional GPU-hours at 512-GPU scale. The fabric cost delta between InfiniBand and Spectrum-X at 512 GPUs is roughly $300,000 for non-blocking. The fabric premium pays for itself if it improves MFU by more than approximately 60% of the delta - which it does for tensor-parallel-heavy workloads but not for pipeline-parallel-heavy workloads.
NVLink 5's 1.8 TB/s per GPU (bidirectional) sets the internal bandwidth baseline differently. NVLink 5 is roughly 36x faster per GPU than a single InfiniBand NDR400 port. But NVLink is intra-rack only - it does not replace fabric. The correct comparison for fabric technologies is between InfiniBand's approximately 1-2 microsecond node-to-node latency, Spectrum-X's 2-4 microseconds, and commodity RoCEv2's 5-10 microseconds under load. For a single all-reduce across 32 nodes, the latency difference between 2 microseconds (InfiniBand) and 8 microseconds (RoCEv2) is invisible in wall-clock training time if bandwidth is sufficient - the all-reduce is bandwidth-bound at scale, not latency-bound.
GPU Direct RDMA support differentiates the technologies in practice. InfiniBand has full GPU Direct RDMA (GDR) support with peer-to-peer access to GPU memory via PCIe BAR mapping, giving zero-copy data transfer between NIC and GPU memory. Spectrum-X also supports GPU Direct RDMA but with higher PCIe overhead due to the Ethernet protocol stack - approximately 5-10% lower GDR throughput in NCCL benchmarks. Commodity RoCEv2 with GPU Direct RDMA works but requires careful PCIe topology planning: NIC must be on the same PCIe root complex as the GPU for optimal performance, which not all node designs accommodate.
| Metric | InfiniBand NDR400 | Spectrum-X | Commodity RoCEv2 |
|---|---|---|---|
| Node-to-node latency (idle) | 1-2 us | 2-3 us | 3-5 us |
| Node-to-node latency (under all-reduce load) | 2-3 us | 3-5 us | 5-10 us |
| NVLS all-reduce throughput (64 GPUs) | 48 GB/s per GPU | 42 GB/s per GPU | 28 GB/s per GPU |
| GPU Direct RDMA support | Full GDR | Full GDR (5-10% PCIe overhead) | GDR (topology-dependent) |
| NCCL all-reduce (256 GPUs, 4M msg) | 12.8 GB/s | 11.2 GB/s | 7.4 GB/s |
| Typical MFU (70B TP=8, PP=4) | 74-78% | 70-76% | 58-65% |
| Typical MFU (70B TP=8, PP=16) | 62-68% | 60-66% | 52-58% |
Decision Framework: Cluster Size, Workload Type, and Budget Sensitivity
The decision matrix at each cluster size is straightforward. At 8-64 GPUs, use NVLink wherever possible (NVL72 racks) and skip inter-node fabric entirely. If multi-node at 64 GPUs, Spectrum-X is the correct choice unless you run TP=8 or higher across nodes. At 64-256 GPUs, the decision hinges on your parallelism strategy: pipeline-parallel workloads (most 70B+ training runs) should use Spectrum-X and save the fabric premium; tensor-parallel-heavy workloads (MoE models with large expert hidden dimensions) should use InfiniBand NDR with SHARP enabled. At 256-1,024 GPUs, Spectrum-X is the default with InfiniBand reserved for circumstances where SHARP's in-network reduction measurably lifts MFU - which requires profiling, not guessing. Above 1,024 GPUs, Ethernet variants dominate because InfiniBand's operational complexity at scale exceeds its throughput advantage.
NCCL and RCCL support matters for AMD or Intel GPU clusters. InfiniBand has the most mature NCCL integration and is the reference platform for NVIDIA's own testing. Spectrum-X is fully supported in NCCL 2.22+ and IBVERBS-compatible for direct InfiniBand-to-Ethernet translation. AMD's RCCL supports RoCEv2 natively but InfiniBand requires the same IBVERBS abstraction - RCCL all-reduce throughput is approximately 10-15% lower on InfiniBand than equivalent NCCL, due to less optimized collective algorithms. For multi-vendor clusters mixing NVIDIA and AMD GPUs, Spectrum-X provides the most consistent cross-platform performance.
Budget sensitivity changes the calculus for reserved versus on-demand clusters. If you are signing a 12-month reserved contract, paying the fabric premium for InfiniBand non-blocking adds approximately $0.80-1.50/GPU/hr to your effective rate. At 256 GPUs over 12 months, that is $420,000-788,000 in additional cost. The question is not whether InfiniBand is better than Ethernet on technical specs - it is. The question is whether that technical advantage translates to measurable training throughput improvement at your specific parallelism strategy, batch size, and model architecture. Most teams we work with should choose Spectrum-X for new contracts in 2026 and reserve InfiniBand for clusters where they have already profiled their training recipe and confirmed the MFU gap is 5 points or more.
| Cluster Scale | Recommended Fabric | Key Decision Factor |
|---|---|---|
| 8-64 GPUs (1 node) | NVLink only | Stay in NVL72; no fabric decision |
| 64 GPUs (multi-node) | Spectrum-X | Skip InfiniBand unless TP across nodes |
| 64-256 GPUs | Spectrum-X (pipeline); InfiniBand (tensor) | Match fabric to parallelism strategy |
| 256-1,024 GPUs | Spectrum-X default | Profile before paying InfiniBand premium |
| 1,024+ GPUs | Spectrum-X or custom Ethernet fabric | InfiniBand ops cost exceeds benefit |
