Switch Topology and Domain Design
NVLink Switch is a dedicated ASIC that connects up to 576 GPUs in a single fully-connected NVLink domain. Each GPU exposes 72 NVLink 5 lanes running at 100 GB/s per direction, yielding a total per-GPU bandwidth of 1.8 TB/s unidirectional or 3.6 TB/s bidirectional. The switch fabric provides full bisection bandwidth: any GPU can communicate with any other GPU in the domain at the full 1.8 TB/s bandwidth without congestion.
The topology uses a two-level switch hierarchy. Four NVLink Switch ASICs connect 16 B200 GPUs in a single DGX B200 NVL rack. Eight such racks (128 GPUs) connect through a second tier of NVLink Switch ASICs. At the maximum configuration, 36 switch ASICs in 18 1U chassis create a 576-GPU domain. Each ASIC dissipates 80W and provides 7.2 Tbps of aggregate switching capacity.
| Config Level | GPUs | Switch ASICs | Aggregate BW | Bisection BW |
|---|---|---|---|---|
| Single DGX B200 NVL | 16 | 4 | 28.8 TB/s | 14.4 TB/s |
| Full rack domain | 128 | 16 | 230.4 TB/s | 115.2 TB/s |
| Max NVLink domain | 576 | 36 | 1,036.8 TB/s | 518.4 TB/s |
| Multi-domain (fabric) | 4,608+ | 288+ | N/A | Limited by IB/ETH |
Eliminating the All-Reduce Bottleneck
Before NVLink Switch, distributed training across 256+ GPUs relied on InfiniBand for inter-node all-reduce. A standard IB NDR400 link provides 50 GB/s per direction, 36x slower than NVLink 5's 1.8 TB/s. The all-reduce step in a 70B parameter model with sharded gradients requires 1.2 GB of data movement per GPU per step. On InfiniBand, this takes 24ms; on NVLink 5 within the domain, it takes 0.67ms, a 36x reduction.
This latency gap is the dominant scaling inefficiency for models above 100B parameters. At 4,096 GPUs with InfiniBand, communication overhead consumes 42% of each training step at 70B scale. Within the 576-GPU NVLink domain, the communication overhead drops to 4% for the same model size. The remaining 96% of training step time is pure compute, enabling near-linear scaling across the domain.
NVLink 5 Protocol Details
NVLink 5 doubles per-lane throughput from NVLink 4's 50 GB/s to 100 GB/s by switching from NRZ (non-return-to-zero) to PAM4 (pulse-amplitude modulation 4-level) signaling. The physical layer operates at 112 Gbps per differential pair over 36 lanes per direction, aggregated across 7 HSIO (high-speed I/O) tiles per GPU. The total GPU-to-switch connection uses 1,792 differential pairs running at 112 Gbps each.
The NVLink 5 protocol implements adaptive routing with per-packet load balancing across all available lanes. When one lane experiences bit errors or signal degradation, traffic is redistributed across remaining lanes within 1 microsecond with zero packet loss. This fault tolerance is critical for 576-GPU domains where a single PCB trace failure should not impact training throughput. The chipkill-level lane redundancy guarantees 99.999% fabric uptime.
All-to-All Communication Patterns
Foundation model training relies on four primary communication patterns: all-reduce (gradient synchronization), all-gather (parameter unsharding), reduce-scatter (gradient partitioning), and all-to-all (sequence parallelism in attention layers). NVLink Switch optimizes all four through its flat topology. All-to-all for sequence parallelism, which is the most bandwidth-intensive, benefits the most from the switch's full bisection bandwidth.
A 576-GPU domain executing an all-to-all operation for a 128k context window in Ring Attention completes the exchange in 1.2ms. The same operation across 576 GPUs connected through four InfiniBand leaf switches takes 38ms due to the three-tier Clos topology that forces traffic through oversubscribed spine links at deeper layers. The 31.7x reduction in all-to-all latency enables sequence parallelism to scale to 1M+ token context windows within the domain.
Multi-Domain Scaling Beyond 576 GPUs
Training the largest models (1T+ parameters) requires aggregating multiple NVLink domains through a higher-level fabric. NVIDIA recommends Spectrum-X Ethernet at 800 Gbps per link for inter-domain connections, with 128 domains supporting 73,728 GPUs total. The inter-domain bandwidth is 400 GB/s per domain (800G x 4 links), a 4.5x reduction from intra-domain bandwidth. This asymmetric bandwidth works because only 0.5% of total data movement crosses domain boundaries in optimized training pipelines.
The NCCL topology-aware collective algorithm selects the optimal communication strategy based on whether data stays within a domain or crosses domains. Intra-domain all-reduce uses the NVLink fast-path. Inter-domain all-reduce uses hierarchical ring reduction across domain gateways, trading bandwidth for reduced domain controller overhead. On a 4,608-GPU cluster (8 domains), NCCL achieves 92% of peak inter-domain bandwidth utilization at 1T parameter scale.
NVLink Switch vs InfiniBand vs Ethernet
NVLink Switch operates at the GPU memory semantic level, exposing load/store semantics to peer GPU memory. InfiniBand and Ethernet operate at the message passing level with RDMA verbs. NVLink's memory semantics reduce overhead by 60% for small messages under 256 KB, which dominate the all-gather operations in sharded training. For large messages above 1 MB, the bandwidth advantage of NVLink 5 over NDR400 (1.8 TB/s vs 400 GB/s per GPU) is the dominant factor.
The TCO comparison favors NVLink Switch for domains of 64+ GPUs. The incremental cost of NVLink Switch ASICs and PCB routing adds roughly $6,000 per GPU for a 576-GPU domain, compared to $2,500 per GPU for InfiniBand HDR fabric. However, the 36x faster all-reduce enables 22% higher model FLOPs utilization (MFU) for large models, which translates to 22% more training throughput over the hardware's lifetime. The hardware payback period is 7 months for continuously utilized training clusters.
Domain Configuration Recommendations
Teams training models at 100B parameters or above should configure a full 576-GPU domain for each training workload. For models under 100B, smaller domains of 128 or 256 GPUs suffice and reduce the NVLink hardware premium. The sweet spot for B200 and B300 training on ClusterBid is the 256-GPU domain (16 DGX B200 racks), which supports 405B parameter training with 96% MFU and 97% linear scaling efficiency.
For multi-domain clusters, allocate one 576-GPU domain per training job and use inter-domain links for checkpoint replication and data loading only. Running a single training job across 2+ domains introduces a 4% MFU penalty from the asymmetric bandwidth. ClusterBid's rack booking interface allows selecting contiguous NVLink domains with guaranteed fabric allocation, eliminating the risk of fragmented GPU assignments that degrade interconnect performance.
