What NVSwitch Does That Ethernet Cannot
NVSwitch is a crossbar fabric purpose-built for GPU-to-GPU communication. Unlike Ethernet switches that route packets through multi-hop paths with variable latency, NVSwitch provides a fully connected, any-to-any topology within a single domain. Every GPU in the domain talks to every other GPU at the same bandwidth without contention. The H100 NVSwitch backplane delivers 900 GB/s of bisection bandwidth per direction per GPU in a DGX H100 system, which is NVLink 4 at full speed. There are no store-and-forward delays, no packet loss recovery, and no TCP overhead cutting into throughput. For your training cluster, this means gradients and activations move between GPUs at rates that Ethernet cannot practically achieve at any cost.
The architectural difference matters most during all-reduce operations in distributed training. In a standard Ethernet-based multi-node setup, ring all-reduce passes data in a sequential loop where each node sends, receives, and sums gradients in stages. Bandwidth utilization peaks around 60-70% of the physical link speed. With NVSwitch and NVLink, the same operation completes in near-constant time regardless of the number of GPUs in the domain, because the fabric can route all GPU pairs simultaneously. This is the difference between a gradient sync taking 500 microseconds versus 5 milliseconds in a 64-GPU configuration.
NVSwitch does not replace your cluster network. It sits alongside it as a computational interconnect. Your cluster still needs InfiniBand or RoCE for inter-node communication between NVSwitch domains. The NVSwitch fabric is local to a single DGX node or a single NVSwitch system rack configuration. Thinking of it as your intra-node or intra-rack ultra-fast lane while InfiniBand handles inter-rack traffic is the correct mental model for cluster architecture.
NVSwitch Generations: NVLink 4 vs NVLink 5 on B200
H100 DGX systems use NVLink 4 with NVSwitch 4. Each GPU gets 18 NVLink 4 lanes running at 50 GB/s per direction per lane, totaling 900 GB/s of bi-directional bandwidth. The NVSwitch 4 ASIC inside a DGX H100 provides 64 ports of NVLink 4, supporting up to 8 GPUs in a fully connected topology within a single node. The switch itself has a switching capacity of 7.2 TB/s, non-blocking across all ports. That is enough that every GPU-to-GPU communication within the node happens at full NVLink bandwidth regardless of the communication pattern.
B200 and the HGX B200 baseboard introduce NVLink 5, doubling the per-GPU bandwidth to 1.8 TB/s per direction. The NVSwitch 5 ASIC scales to 128 ports with a switching capacity approaching 14.4 TB/s. This is not merely a speed bump. NVLink 5 supports Grace Hopper superchip interconnects and enables memory pooling across the fabric - a GPU in a B200 domain can access the full memory capacity of every other GPU in the domain through NVLink, not just its local HBM. This changes the memory model from discrete per-GPU pools to a unified fabric-attached memory space.
The practical implication: B200 clusters can deploy models that require 256GB or more of memory footprint without pipeline parallelism splitting the model across nodes connected by slower interconnects. For model sizes between 100B and 400B parameters, this eliminates one of the most expensive sources of inefficiency in H100 deployments - the latency overhead of model shard communication over InfiniBand or Ethernet when pipeline parallelism is a bottleneck.
| Feature | NVLink 4 (H100) | NVLink 5 (B200) |
|---|---|---|
| Per-GPU Bandwidth | 900 GB/s | 1.8 TB/s |
| NVSwitch Capacity | 7.2 TB/s | 14.4 TB/s |
| Max Ports per Switch | 64 | 128 |
| Memory Pooling | No | Yes (fabric-attached) |
| Topology per Node | 8-GPU fully connected | 8-GPU fully connected |
NVLink Domain Topology: How GPUs Connect Within a Node
An NVLink domain is the set of GPUs that share a fully connected NVLink fabric through one or more NVSwitch ASICs. In a DGX H100, the eight H100 SXM5 GPUs form a single NVLink domain. Each GPU connects to all four NVSwitch ASICs on the DGX baseboard, and the switches collectively provide any-to-any connectivity. This is a full mesh topology, meaning any single GPU fault does not partition the domain - NVSwitch reroutes around it automatically at the fabric layer.
The physical layout matters for thermal and power planning. Each NVSwitch 4 ASIC dissipates roughly 40W, and the four-switch configuration adds 160W to the system power budget beyond the GPUs themselves. In a DGX H100 rated at 10.2 kW peak, the NVSwitch subsystem accounts for roughly 1.6% of total power but enables the communication bandwidth that makes the GPU utilization efficient. Skimping on the interconnect to save power is false economy in a training cluster - GPU idle time waiting for gradients over slow interconnects wastes far more power than the NVSwitch fabric draws.
B200 systems double switch density. An HGX B200 baseboard carries two NVSwitch 5 ASICs, each with 128 ports, supporting eight B200 GPUs in a fully connected domain. The topology is simpler with fewer switches because each NVSwitch 5 handles double the port count of NVSwitch 4. The power per switch is higher at roughly 55W, but total switch power drops to 110W for the domain despite providing 2x the per-GPU bandwidth. This is a meaningful efficiency improvement for dense cluster deployments where rack power density is often the limiting factor.
Extending the Fabric Across Nodes: NVLink Domain Scaling Challenges
NVLink domains are bounded by physical switch capacity and cabling constraints. An NVLink domain in current NVIDIA architectures is limited to 8 GPUs per node within a single NVSwitch fabric. To scale beyond one domain, you must bridge domains through the host network - InfiniBand NDR400 or RoCE v2 at 400 Gbps per port in typical cluster designs. The bandwidth cliff between intra-domain (900-1800 GB/s per GPU) and inter-domain (50-100 GB/s per GPU via InfiniBand) is steep and directly impacts training throughput for models that require model parallelism across domain boundaries.
The solution NVIDIA introduced with the DGX SuperPOD architecture is NVLink domain aggregation through the second-generation NVSwitch in the compute tray. The DGX H100 SuperPOD connects 32 nodes (256 GPUs) in a single pod with NVLink domain connectivity. Each SuperPOD uses 16 NVSwitch systems external to the DGX nodes, each providing 128 ports of NVLink 4 interconnected through a multi-stage Clos topology. The aggregate bisection bandwidth across the pod reaches 12.8 TB/s, which maintains acceptable inter-domain communication for tensor parallelism up to 64 GPUs.
For custom cluster designs - which many teams build to reduce cost versus buying DGX systems - scaling NVLink domains beyond a single node requires NVLink-connected switches that can aggregate at the rack level. This is where the market bifurcates. DGX customers get turnkey fabric scaling through validated SuperPOD designs. Custom cluster builders using H100 HGX baseboards must design their own compute-to-switch topologies, often settling for InfiniBand-based inter-node communication because sourcing NVSwitch systems independently is costly and complex.
Fabric Cost Analysis: NVLink Premium vs InfiniBand Alternative
The cost difference between NVLink-based fabric and an InfiniBand alternative is significant at cluster scale. A DGX H100 system includes the NVSwitch fabric in the baseboard cost, which runs roughly $35,000-40,000 per node premium over an equivalent H100 HGX baseboard without NVSwitch. For a 32-node cluster, that is over $1 million in interconnect cost baked into the DGX pricing. If you are sizing a 256-GPU cluster, the NVSwitch fabric represents about 12-15% of the total cluster hardware cost.
The alternative is building a cluster with H100 PCIe GPUs and connecting them exclusively through InfiniBand NDR400. An 8-GPU node using H100 PCIe costs roughly $220,000-250,000 depending on the host server configuration, versus $300,000+ for a DGX H100. The per-node savings of $50,000-80,000 are significant, but the tradeoff is inter-GPU bandwidth. H100 PCIe has no NVLink - GPUs communicate through the PCIe 5.0 x16 host bus at 64 GB/s per direction, which is 14x less than NVLink 4 bandwidth. For tensor parallelism across GPUs, this gap materially impacts training throughput.
The breakeven analysis depends on your model parallelism strategy. If you train models under 30B parameters and use data parallelism exclusively, the PCIe InfiniBand cluster may achieve 85-90% of the DGX training throughput at 60-70% of the cost. The lower interconnect utilization makes NVSwitch's premium unnecessary. For models above 70B requiring tensor parallelism across 4-8 GPUs, the DGX route with NVSwitch fabric delivers 35-50% higher throughput, making the per-node premium cost-effective against the alternative of longer training runs eating into your compute budget.
Monitoring NVSwitch Fabric Health: What to Watch
NVSwitch fabric issues manifest differently than host network problems. Link flapping on NVLink connections typically causes silent data corruption or hangs in CUDA kernels, not packet loss or retransmission as with Ethernet. The indicators to monitor are NVLink CRC error counters, NVSwitch temperature and power telemetry, and GPU link status reported through nvidia-smi nvlink mode. Any CRC error count that increments over multiple monitoring checks indicates a fabric issue that will eventually degrade or crash training runs.
DCGM (Data Center GPU Manager) exposes NVLink error counters through its telemetry pipeline. The key metrics to track: nvlink_flit_crc_error_total for flit-level CRC errors, nvlink_data_crc_error_total for data-integrity CRC errors, and nvlink_replay_error_total for link replay events. A training cluster operating normally should show zero CRC errors across all counters. Even a single recurring error indicates a failing cable or NVSwitch port that should be addressed before it triggers a collective communication hang.
Thermal management of NVSwitch ASICs is critical. NVSwitch 4 operates within a case temperature range of 0-85C, with performance degradation above 70C. In dense DGX deployments where switches share airflow with GPU exhaust, NVSwitch inlet temperatures can run 5-10C higher than GPU inlet temperatures. If your cluster monitoring shows NVSwitch temperatures consistently above 65C, investigate airflow patterns and consider adjusting the data center cooling strategy. A thermal shutdown of the NVSwitch fabric takes down all eight GPUs in the domain simultaneously.
Fabric Roadmap: NVLink Domains Beyond Blackwell
The NVLink domain concept is evolving toward full-cluster fabric connectivity. NVIDIA's roadmap shows NVLink domains scaling from 8-GPU domains to 72-GPU domains with the GB200 NVL72 rack-scale design. This system connects 72 B200 GPUs across multiple NVSwitch 5 ASICs in a single NVLink domain, eliminating the InfiniBand bridge entirely for up to 72 GPUs. The bisection bandwidth within this domain reaches 130 TB/s, which fundamentally changes cluster architecture for large model training.
The implication for cluster planners: the DGX node as the atomic unit of GPU deployment may be replaced by the rack-scale domain. Instead of designing clusters around 8-GPU nodes with InfiniBand interconnects between them, the 72-GPU domain becomes the smallest deployable unit. This changes how you budget for GPU capacity, how you design your data center power distribution, and how you plan for failure domain isolation. A single NVSwitch 5 ASIC failure in a NVL72 domain affects all 72 GPUs, making fabric fault tolerance a first-order operational concern rather than a theoretical discussion.
For teams planning GPU acquisitions in mid-2026, the safe bet is to design cluster networking that can accommodate both 8-GPU NVLink domains and future rack-scale domains. Invest in InfiniBand or Ethernet fabrics that scale to 512+ GPUs rather than hard-limiting designs to 256 GPUs. The interconnect cost as a percentage of total cluster budget will increase as domain sizes grow, but the training throughput gains for large model workloads make it impossible to ignore. A well-designed fabric today is the foundation for a cluster that remains performant through the next two GPU generations.
