NVLINK AND NVSWITCH ARCHITECTURE
NVLink 4.0 connects H100 GPUs within an HGX baseboard with 900 GB/s bidirectional bandwidth per GPU, 18x PCIe Gen5 bandwidth per lane. Eight H100 GPUs form a fully connected NVSwitch fabric with 7.2 TB/s aggregate bisection bandwidth. Each NVSwitch chip provides 64 NVLink ports at 450 GB/s switching capacity for 4 NVSwitches totaling 1,800 GB/s.
NVSwitch hierarchy: single NVSwitch connects 8 GPUs in a full bisection bandwidth topology. For 256 GPUs across 32 nodes, 8 NVSwitches in a 2-tier topology provide 3.2 TB/s inter-node bandwidth via InfiniBand or NVLink-over-Ethernet. NVLink domain spans a single HGX baseboard; inter-domain communication requires InfiniBand or converged Ethernet.
| Interconnect | Bandwidth/GPU | Topology | Latency | GPU Count Range | Scaling Efficiency (512 GPUs) |
|---|---|---|---|---|---|
| NVLink 4.0 (intra-node) | 900 GB/s | Fully connected | 0.1-0.3 us | 8 | N/A |
| NVSwitch 4.0 (intra-node) | 1,800 GB/s aggregate | Non-blocking | 0.2-0.5 us | 8 | N/A |
| InfiniBand NDR400 (inter-node) | 50 GB/s per link | Fat tree / Dragonfly | 0.6-1.0 us | 16-32,000 | 92-96% |
| NVLink-over-Ethernet | 25 GB/s per link | CLOS | 1.5-3.0 us | 16-8,000 | 80-85% |
| PCIe Gen5 x16 | 32 GB/s per link | Tree | 0.5-1.0 us | 2-4 | N/A (host connect) |
FABRIC TOPOLOGY FOR DISTRIBUTED TRAINING
Fabric topology choice determines scaling efficiency. Fat-tree topology with 2:1 blocking ratio supports 92-96 percent scaling efficiency up to 1,024 GPUs. Dragonfly+ topology reduces cable count by 35-50 percent versus fat-tree while maintaining 88-92 percent scaling efficiency at 8,192 GPUs. Torus topology popularized by TPU pods achieves 90-95 percent efficiency for specifically optimized collective patterns.
All-Reduce bandwidth is the key metric. On 512 H100 GPUs with NDR400 InfiniBand fat-tree, All-Reduce achieves 380 Gbps per GPU. Ring All-Reduce algorithm with 512 GPUs completes in approximately 200 microseconds. Tree-based All-Reduce with SHARP in-network computing completes in 120 microseconds (40 percent faster).
TOPOLOGY-AWARE JOB SCHEDULING
Topology-aware scheduling places jobs to minimize inter-node communication. A 16-GPU training job should fit within a single HGX baseboard when possible, eliminating inter-node communication entirely. 32-GPU jobs span 2 nodes within the same InfiniBand leaf switch. 128-GPU jobs occupy a full leaf-spine quadrant. Non-topology-aware scheduling reduces scaling efficiency by 15-30 percent.
Implementation via SLURM topology plugin or Kubernetes NRI topology-aware device plugin. Hwloc library provides node-level topology information (NVLink domains, PCIe hierarchy, NUMA nodes). Job allocation algorithm: attempt intra-node first, then same-rack, then same-leaf, expanding outward. Allocation success rate: 85 percent intra-node for --gpus-per-node=8, 70 percent same-rack for 16-32 GPUs.
FUTURE INTERCONNECT TECHNOLOGIES
NVIDIA B200 introduces NVLink 5.0 with 1,800 GB/s per GPU (2x H100). Ultra Ethernet Consortium targets 800 Gbps-1.6 Tbps for AI networking, challenging InfiniBand dominance with lower cost at equivalent performance. CXL 3.0 with fabric capabilities enables GPU memory pooling across nodes at 256 GB/s, reducing data movement for out-of-core training patterns.
Optical interconnect emerging as long-term solution. Co-packaged optics with 1.6 Tbps per port at 5-10 pJ/bit versus 15-25 pJ/bit for electrical. Coherent optical engines at 800G per lambda enable 200-meter reach at data center scale. Production deployment expected 2027-2028 for large-scale AI fabrics.
