All essays
BenchmarkCOMPARISONFEB 2026

RDMA, GPU Memory Semantic Fabrics, and CXL: The Interconnect Stack That Determines Distributed AI Performance

Deep technical analysis of GPU interconnects: RDMA (InfiniBand/RoCEv2), GPU memory semantic fabrics (NVLink/NVSwitch), and CXL disaggregation for distributed AI.

01

The GPU Interconnect Stack: Three Layers That Define Performance

Distributed AI performance is determined by the interconnect hierarchy connecting GPUs within a node, across nodes in a cluster, and across clusters in a data center. The three layers of this stack are: intra-node GPU-to-GPU interconnects (NVLink, AMD Infinity Fabric), node-to-node RDMA networks (InfiniBand, RoCEv2, Ultra Ethernet), and emerging memory semantic fabrics (CXL 3, NVLink over fabric) that blur the boundary between local and remote memory.

Each layer presents different bandwidth, latency, and programmability characteristics. NVLink 5 delivers 1.8 TB/s per GPU within an 8-GPU node with microsecond latency. InfiniBand NDR400 delivers 400 Gbps per port with 1-2 microsecond latency between nodes. CXL 3 over PCIe 6 delivers 128 GB/s per link with sub-microsecond latency for cache-coherent memory access. Understanding which layer bottlenecks a specific workload determines whether cluster performance is compute-bound, memory-bound, or communication-bound.

The trend in 2026 is toward flattened memory hierarchies. NVIDIA's Dynamo disaggregated inference architecture and CXL-based memory pooling are erasing the distinction between local GPU memory and remote fabric-attached memory. For buyers, this means the interconnect fabric is becoming as important a specification as GPU FLOPs. A B200 cluster connected via 400 Gbps InfiniBand performs differently than one connected via 200 Gbps RoCEv2, even when the GPUs are identical.

02

RDMA Deep Dive: InfiniBand NDR400 vs RoCEv2 vs Ultra Ethernet

InfiniBand NDR400 is the current gold standard for GPU cluster networking, delivering 400 Gbps per port with sub-2-microsecond latency and native GPU Direct RDMA (GDR) support. NVIDIA's ConnectX-8 adapters integrate NDR400 with the BlueField DPU architecture, offloading NCCL communication from the GPU and achieving line-rate all-reduce performance. The cost premium over Ethernet-based solutions is approximately 30-50% per port, including the HDR/NDR switch infrastructure.

RoCEv2 (RDMA over Converged Ethernet) provides RDMA semantics over standard Ethernet infrastructure at 200-400 Gbps. The key limitation is that RoCEv2 requires lossless Ethernet fabric configuration (Priority Flow Control, ECN marking) to maintain RDMA performance. Without proper tuning, RoCEv2 throughput collapses at less than 0.1% packet loss. Well-tuned RoCEv2 achieves 85-95% of InfiniBand performance at 50-70% of the cost, with the caveat that tuning effort is significant (2-4 weeks for a large cluster).

Ultra Ethernet Consortium (UEC) is the newest entrant, designing an Ethernet physical layer and transport protocol specifically optimized for AI/ML workloads. Early UEC implementations target 800 Gbps with sub-microsecond latency and native RDMA semantics without the lossless fabric requirements of RoCEv2. Production silicon availability is expected in late 2026 to early 2027. For greenfield cluster builds, UEC represents the most promising path to Ethernet-based AI networking at InfiniBand-comparable performance.

CharacteristicInfiniBand NDR400RoCEv2 (400G)Ultra Ethernet
Bandwidth per port400 Gbps200-400 Gbps800 Gbps (target)
Latency (node-to-node)1.0-1.5 us2.0-4.0 us<1.0 us (target)
Lossless required?NativeYes (PFC/ECN)No (designed for AI)
GPU Direct RDMANativeSupportedSupported
Cost premium vs. Ethernet30-50%0-10%0-20% (est.)
Market maturityMature (15+ years)MaturePre-production
03

GPU Memory Semantic Fabrics: NVLink 5, NVSwitch, and Beyond

NVLink 5 is the third generation of NVIDIA's GPU-to-GPU interconnect, delivering 1.8 TB/s of bidirectional bandwidth per GPU across 18 links. This enables fully connected 8-GPU nodes where every GPU can communicate with every other GPU simultaneously at full bandwidth. NVLink 5 also introduces remote atomic operations and improved cache coherence, reducing the software overhead for distributed tensor operations.

The NVSwitch 5 fabric extends NVLink across multiple nodes, connecting up to 576 GPUs in a single NVLink domain. Within this domain, any GPU can access any other GPU's memory at NVLink bandwidth (not the reduced bandwidth of a network hop). This capability is transformative for model parallelism strategies that require frequent cross-GPU communication: tensor parallelism, sequence parallelism, and expert parallelism in MoE models all benefit directly.

The architectural limitation is domain size. An NVLink domain of 576 GPUs is a single failure domain: if the NVSwitch fabric experiences a fault, all 576 GPUs are affected. Scale-out beyond 576 GPUs requires crossing the NVLink-to-network boundary, introducing a 5-10x bandwidth reduction (1.8 TB/s NVLink to 50 GB/s 400 Gbps network). Multi-domain training strategies must partition models so that cross-domain communication is minimized, typically by placing each model replica within a single NVLink domain.

04

CXL 3: Memory Disaggregation and the Shared Memory Pool

Compute Express Link (CXL) 3.0 introduces cache-coherent memory sharing across CPU sockets and, crucially, between GPUs and CPUs. CXL 3 runs over PCIe 6 electrical layers, delivering 128 GB/s per x16 link with sub-microsecond latency for cache-coherent loads. For GPU workloads, this means host memory can serve as a transparent extension of GPU memory, with the CXL controller handling page migration between local GPU HBM and remote CXL-attached memory.

The practical application for inference is expanded memory capacity without the NVLink domain constraints. A B200 GPU with 288 GB of HBM3e can access an additional 512 GB or more of CXL-attached memory from a shared pool. This is slower than HBM (100-200 GB/s CXL bandwidth vs 8 TB/s HBM) but significantly faster than NVMe-based swap (8-32 GB/s). For KV cache storage in long-context inference, CXL-attached memory provides a cost-effective middle tier: faster than disk-based offloading, cheaper than expanding GPU HBM capacity.

CXL 3 adoption in GPU clusters is accelerating in 2026. AMD's MI400 architecture will debut with native CXL 3 support. NVIDIA is integrating CXL 3 into Grace Blackwell Superchip configurations, using CXL for GPU-to-GPU communication across Grace CPUs. Intel's upcoming Falcon Shores GPU architecture is designed around CXL as the primary scale-out interconnect. The trend is clear: the memory hierarchy is flattening, and CXL is the interconnect that enables it.

Memory TierMediumBandwidthLatencyCapacity per Node
GPU HBM3eDirect (GPU die)8 TB/s (B200)~50 ns288 GB (B200)
CXL 3 PoolPCIe 6 x16128 GB/s~200-500 ns512 GB - 2 TB
Host DDR5CPU memory bus~500 GB/s~100 ns1-4 TB
NVMe StoragePCIe 5 x48-32 GB/s~3-10 us8-30 TB
05

Workload-to-Interconnect Mapping: Which Fabric for Which Pattern

Training large dense models with full data parallelism: the critical interconnect is node-to-node RDMA for gradient all-reduce. Every training step requires synchronizing gradients across all data-parallel ranks. With Megatron-LM or FSDP, gradient sizes range from 1-10 GB per step for a 70B model. At 400 Gbps InfiniBand, all-reduce completes in 20-80ms per step. NDR400 InfiniBand with SHARP in-network computing reduces this by 30-50% by offloading reduction operations to the switch fabric.

Tensor parallelism and sequence parallelism require intra-node NVLink bandwidth. Splitting individual layers or sequence dimensions across GPUs creates communication at every transformer layer boundary. With 8 GPUs per node and model parallelism across all 8, each layer requires a full all-reduce across the NVLink domain. NVLink 5's 1.8 TB/s bandwidth makes this communication effectively free relative to compute time for most configurations below 8-way tensor parallelism.

Inference serving with disaggregated prefill and decode: the emerging pattern in 2026 uses separate GPU pools for prefill (compute-bound, needs high FLOPs) and decode (memory-bound, needs high bandwidth). The interconnect between prefill and decode pools carries KV cache transfers from prefill GPUs to decode GPUs. This transfer uses RDMA (InfiniBand or RoCEv2) and the bandwidth determines how many prefill GPUs can feed the decode pool. NVIDIA's Dynamo architecture implements this pattern with NCCL-based KV cache transfer over RDMA.

06

Cost-Performance of Interconnect Choices for Cluster Buyers

The interconnect fabric represents 15-30% of total GPU cluster cost, depending on the choice of networking technology. For a 256-GPU H100 cluster, InfiniBand NDR400 fabric adds approximately $1.5-2.5M in networking cost (switches, cables, NICs). RoCEv2 at 400 Gbps reduces this to $0.8-1.5M. The savings of $0.7-1.0M must be weighed against the communication performance impact (typically 5-15% slower training throughput on communication-bound workloads).

The break-even analysis depends on cluster utilization and workload characteristics. A cluster running at 80% utilization on communication-sensitive workloads (large-scale data-parallel training) loses more value from the 10% throughput reduction than it saves on the networking cost. A cluster running mixed workloads with significant inference components (which are less sensitive to RDMA bandwidth) may achieve better overall TCO with RoCEv2.

For clusters under 128 GPUs, the interconnect choice matters less because communication overhead is a smaller fraction of total step time. Most 64-GPU training runs on medium-sized models (7-30B parameters) see under a 5% throughput difference between InfiniBand and well-tuned RoCEv2. For clusters above 512 GPUs training large models (70B+), InfiniBand with SHARP and in-network computing delivers measurable throughput advantages that typically justify the cost premium.

Cluster SizeRec. InterconnectNetworking Cost/GPUComm. Overhead
8-64 GPUsRoCEv2 (200G)$2,000 - $3,000<5%
64-256 GPUsRoCEv2 (400G) or IB NDR200$3,000 - $6,0005-10%
256-1,024 GPUsInfiniBand NDR400$5,000 - $10,00010-15%
1,024+ GPUsInfiniBand NDR400 + SHARP$8,000 - $15,00015-25%
07

Future Directions: What the Interconnect Stack Looks Like in 2027-2028

NVLink 6, expected with the NVIDIA Rubin architecture in 2027, will further increase per-GPU interconnect bandwidth to approximately 3.6 TB/s and expand the NVLink domain to 1,152 GPUs. This effectively eliminates the multi-domain boundary for most training workloads, enabling single-domain training at the full cluster scale. The increased domain size reduces software complexity: no need for hierarchical all-reduce strategies that partition across NVLink domains.

Ultra Ethernet at 800 Gbps with native RDMA will challenge InfiniBand's dominance in GPU networking. If UEC delivers on its latency and lossless-fabric-free design goals, the cost advantage over InfiniBand (30-50%) will make it the default choice for new GPU cluster builds. InfiniBand will retain its position in NVIDIA's reference architectures, but the broader market may shift toward Ethernet-based solutions that offer more vendor diversity.

CXL-based memory pooling will become standard in GPU clusters, not just for KV cache offloading but for active training data caching. The ability to maintain training datasets in CXL-attached memory (500 GB/s+ bandwidth) rather than reading from NVMe storage will reduce data loading bottlenecks that currently idle GPUs for 5-15% of total training time. The result is effectively free throughput improvement for I/O-bound training pipelines.

Filed under
RDMAInfiniBand NDR400RoCEv2NVLink 5CXL 3GPU memory semantic fabricDistributed AI networkingNVIDIA Dynamo