All essays
TechnicalDEEP DIVEFEB 2026

GPU Interconnect Topology: NVLink, NVSwitch, and Fabric Design

GPU interconnect topology design covering NVLink and NVSwitch architecture, GPU fabric topologies for training at scale, and optimization for collective communication patterns.

01

NVLINK AND NVSWITCH ARCHITECTURE

NVLink 4.0 connects H100 GPUs within an HGX baseboard with 900 GB/s bidirectional bandwidth per GPU, 18x PCIe Gen5 bandwidth per lane. Eight H100 GPUs form a fully connected NVSwitch fabric with 7.2 TB/s aggregate bisection bandwidth. Each NVSwitch chip provides 64 NVLink ports at 450 GB/s switching capacity for 4 NVSwitches totaling 1,800 GB/s.

NVSwitch hierarchy: single NVSwitch connects 8 GPUs in a full bisection bandwidth topology. For 256 GPUs across 32 nodes, 8 NVSwitches in a 2-tier topology provide 3.2 TB/s inter-node bandwidth via InfiniBand or NVLink-over-Ethernet. NVLink domain spans a single HGX baseboard; inter-domain communication requires InfiniBand or converged Ethernet.

InterconnectBandwidth/GPUTopologyLatencyGPU Count RangeScaling Efficiency (512 GPUs)
NVLink 4.0 (intra-node)900 GB/sFully connected0.1-0.3 us8N/A
NVSwitch 4.0 (intra-node)1,800 GB/s aggregateNon-blocking0.2-0.5 us8N/A
InfiniBand NDR400 (inter-node)50 GB/s per linkFat tree / Dragonfly0.6-1.0 us16-32,00092-96%
NVLink-over-Ethernet25 GB/s per linkCLOS1.5-3.0 us16-8,00080-85%
PCIe Gen5 x1632 GB/s per linkTree0.5-1.0 us2-4N/A (host connect)
02

FABRIC TOPOLOGY FOR DISTRIBUTED TRAINING

Fabric topology choice determines scaling efficiency. Fat-tree topology with 2:1 blocking ratio supports 92-96 percent scaling efficiency up to 1,024 GPUs. Dragonfly+ topology reduces cable count by 35-50 percent versus fat-tree while maintaining 88-92 percent scaling efficiency at 8,192 GPUs. Torus topology popularized by TPU pods achieves 90-95 percent efficiency for specifically optimized collective patterns.

All-Reduce bandwidth is the key metric. On 512 H100 GPUs with NDR400 InfiniBand fat-tree, All-Reduce achieves 380 Gbps per GPU. Ring All-Reduce algorithm with 512 GPUs completes in approximately 200 microseconds. Tree-based All-Reduce with SHARP in-network computing completes in 120 microseconds (40 percent faster).

03

TOPOLOGY-AWARE JOB SCHEDULING

Topology-aware scheduling places jobs to minimize inter-node communication. A 16-GPU training job should fit within a single HGX baseboard when possible, eliminating inter-node communication entirely. 32-GPU jobs span 2 nodes within the same InfiniBand leaf switch. 128-GPU jobs occupy a full leaf-spine quadrant. Non-topology-aware scheduling reduces scaling efficiency by 15-30 percent.

Implementation via SLURM topology plugin or Kubernetes NRI topology-aware device plugin. Hwloc library provides node-level topology information (NVLink domains, PCIe hierarchy, NUMA nodes). Job allocation algorithm: attempt intra-node first, then same-rack, then same-leaf, expanding outward. Allocation success rate: 85 percent intra-node for --gpus-per-node=8, 70 percent same-rack for 16-32 GPUs.

04

FUTURE INTERCONNECT TECHNOLOGIES

NVIDIA B200 introduces NVLink 5.0 with 1,800 GB/s per GPU (2x H100). Ultra Ethernet Consortium targets 800 Gbps-1.6 Tbps for AI networking, challenging InfiniBand dominance with lower cost at equivalent performance. CXL 3.0 with fabric capabilities enables GPU memory pooling across nodes at 256 GB/s, reducing data movement for out-of-core training patterns.

Optical interconnect emerging as long-term solution. Co-packaged optics with 1.6 Tbps per port at 5-10 pJ/bit versus 15-25 pJ/bit for electrical. Coherent optical engines at 800G per lambda enable 200-meter reach at data center scale. Production deployment expected 2027-2028 for large-scale AI fabrics.

Filed under
NVLinkNVSwitchGPU TopologyInterconnect FabricDragonflyFat TreeCollective Communication