The Physical Difference
The H100 ships in two physically incompatible form factors. The SXM module is a mezzanine card that plugs into a proprietary socket on an NVIDIA HGX baseboard. The PCIe card uses the standard PCIe Gen5 slot found in commodity servers. You cannot put an SXM module into a PCIe slot, and you cannot wire a PCIe card into an HGX baseboard. The choice of form factor is a choice of server platform.
The SXM module measures roughly 80x100mm and connects through a 4,000-pin socket. The PCIe card is a full-height, full-length dual-slot card measuring 266x111mm. The size difference reflects fundamentally different approaches to power delivery and thermal management.
Power and Thermal Limits
The H100 SXM is rated at 700W TDP, with thermal design that assumes liquid-cooled or high-flow air-cooled chassis. The PCIe version is capped at 350W, constrained by the PCIe slot power specification and the thermal capacity of standard server chassis. This 2x power envelope difference is the single largest performance differentiator.
At 700W, the SXM module can sustain higher clock speeds under load, delivering the full 989 TFLOPS FP16 and 3.35 TB/s memory bandwidth. The PCIe card must downclock under sustained load to stay within 350W, typically achieving 85-90% of the SXM's compute throughput on memory-bandwidth-bound workloads and 75-85% on compute-bound workloads.
| Spec | H100 SXM | H100 PCIe |
|---|---|---|
| TDP | 700W | 350W |
| FP16 TFLOPS (peak) | 989 | 989 |
| Sustained FP16 | ~950 | ~750 |
| Memory bandwidth | 3.35 TB/s | 2.0 TB/s |
| NVLink bandwidth | 900 GB/s | None |
| GPU-to-GPU topology | Full NVLink mesh | PCIe switch |
NVLink: The Decisive Difference
The SXM H100 connects to six other SXM modules through 18 NVLink4 links, providing 900 GB/s of bi-directional bandwidth per GPU. This allows all 8 GPUs in an HGX baseboard to communicate as a unified memory pool. The PCIe H100 has no NVLink support at all. Inter-GPU communication uses the PCIe Gen5 x16 slot, capped at 128 GB/s per direction.
This bandwidth gap is the decisive factor for training. In a multi-GPU training loop, each gradient synchronization step requires every GPU to exchange tens of megabytes of gradient data. With NVLink, this completes in microseconds. With PCIe, the same exchange takes 5-10x longer, and the gap grows with model size and parallelism strategy.
Cluster Density Differences
An HGX baseboard packs 8x H100 SXM modules into a 4-slot footprint, drawing 5.6 kW per baseboard. A standard DGX H100 system holds one baseboard. High-density clusters like the H100 NVL (NVIDIA's reference architecture) pair two baseboards in a single 8U chassis for 16 GPUs at 11.2 kW. The result is 64 GPUs per 42U rack before networking.
PCIe H100 clusters typically fit 4-8 GPUs per 4U server. An 8x H100 PCIe server draws roughly 3.2 kW (8 GPUs at 350W plus CPUs, memory, and fans). A 42U rack fits 7 such servers for 56 GPUs, but the networking switch footprint and cabling complexity reduce usable density to roughly 48 GPUs per rack. The SXM cluster delivers higher density and simpler cabling.
Performance Benchmarks
On well-parallelized training workloads like LLM pre-training with tensor and pipeline parallelism, 8x SXM H100 delivers roughly 1.8x the throughput of 8x PCIe H100. The gap comes from two sources: NVLink bandwidth reducing gradient sync time, and the higher sustained clock speed of the SXM module under full load.
On inference workloads, the gap narrows. Single-GPU inference is primarily memory-bandwidth-bound, and the PCIe card's restricted bandwidth (2.0 TB/s vs 3.35 TB/s) costs roughly 25-30% throughput on large-batch inference. On latency-optimized single-request inference with batch size 1, the difference shrinks to 10-15%.
| Workload | 8x SXM (tokens/s) | 8x PCIe (tokens/s) | SXM advantage |
|---|---|---|---|
| LLaMA 70B training | ~380k | ~210k | ~1.8x |
| LLaMA 70B inference (batch 64) | ~14,500 | ~10,800 | ~1.34x |
| LLaMA 70B inference (batch 1) | ~185 | ~162 | ~1.14x |
When PCIe Makes Sense
PCIe H100 is the right choice when your workload does not need inter-GPU communication at NVLink speeds. This includes single-GPU inference, batch inference with model parallelism that fits within PCIe bandwidth, and fine-tuning single models on one or two GPUs. The PCIe card also works in existing server infrastructure without chassis upgrades.
Cost is the other factor. PCIe H100 clusters typically rent for 25-35% less per GPU-hour than SXM clusters. For inference teams where the GPU is already memory-bandwidth-bound and the workload fits on 1-4 GPUs, paying for SXM's NVLink and higher sustained clocks is wasted spend. The PCIe form factor is also easier to procure: lead times for PCIe H100 are 4-8 weeks shorter than SXM in 2026.
When SXM Is Non-Negotiable
SXM H100 is non-negotiable for any workload that spans more than 4 GPUs with significant inter-GPU communication. Pre-training runs using tensor parallelism, pipeline parallelism, or fully sharded data parallelism (FSDP) depend on NVLink to keep GPU utilization above 80%. Without it, the gradient sync overhead drops utilization to 50-60% and extends training time by 40-60%.
If you are running multi-node training with model sizes above 30B parameters, SXM is not a luxury. It is a requirement for efficient scaling. We see teams waste $200k-500k on extended training runs trying to use PCIe GPUs for distributed training before migrating to SXM. The savings from cheaper PCIe hardware are erased by the extra wall-clock time and lower utilization.
