The Parallelism Landscape
Deploying large language models across multiple GPUs requires distributing the model across devices. Three primary strategies exist: tensor parallelism (TP), pipeline parallelism (PP), and sequence parallelism (SP). Each makes different tradeoffs between inter-device communication, memory utilization, and end-to-end latency.
The choice of parallelism strategy is the single largest factor in inference performance for models that do not fit on a single GPU. A 405B parameter Llama 4 model requires 810 GB of FP16 weights, exceeding the 288 GB HBM3e capacity of a single B300 GPU. Even at FP4 precision, models above 200B parameters require multi-GPU distribution for both memory and compute throughput.
Tensor Parallelism Deep Dive
Tensor parallelism splits individual weight matrices across GPUs. In a 2-way TP configuration, a single matrix multiplication is divided: each GPU holds half the rows or columns and performs a partial matmul, followed by an all-reduce to combine results. TP is the most communication-intensive strategy, requiring an all-reduce for every transformer layer.
TP works best within a single node where NVLink provides 900 GB/s to 1.8 TB/s of inter-GPU bandwidth. Across nodes, TP degrades rapidly because NCCL all-reduce over InfiniBand or Ethernet (50 GB/s per link) becomes the bottleneck. A B200 with 8-way TP over NVLink achieves 2.3x throughput versus 8-way TP split across two nodes over 400 Gbps Ethernet, due entirely to the all-reduce bottleneck.
| TP Degree | Interconnect | All-Reduce Time | Throughput (tok/s) | Scaling Efficiency |
|---|---|---|---|---|
| TP=2 | NVLink 5 (1.8 TB/s) | 12 us | 2,450 | 98% |
| TP=4 | NVLink 5 (1.8 TB/s) | 24 us | 4,680 | 94% |
| TP=8 (1 node) | NVLink 5 (1.8 TB/s) | 48 us | 8,100 | 82% |
| TP=8 (2 nodes) | 400 Gbps Ethernet | 810 us | 3,200 | 32% |
| TP=4 (4 nodes) | 200 Gbps InfiniBand | 940 us | 1,850 | 19% |
Pipeline Parallelism Deep Dive
Pipeline parallelism places different transformer layers on different GPUs. Input tokens flow through GPU 0 layers 1-10, then GPU 1 layers 11-20, and so on. PP requires far less inter-GPU communication: only the activations passed between pipeline stages. Communication is proportional to hidden dimension times sequence length per microbatch, not to the full model parameter count.
The primary cost of PP is the pipeline bubble: idle GPU time during the warm-up and cool-down phases when not all stages are processing. With 4 pipeline stages and 8 microbatches, pipeline utilization is approximately 75%. Increasing microbatches to 64 raises utilization to 95%. PP also requires higher total batch sizes, which increases end-to-end latency for interactive applications.
Sequence Parallelism and Hybrid Strategies
Sequence parallelism splits the sequence dimension across GPUs, allowing longer context windows than any single GPU can process. SP is often combined with TP (SP + TP = SPT), where the sequence dimension is split for attention computations while weights are split for feed-forward layers. This hybrid approach is the standard deployment configuration for models processing 128K+ token contexts.
The most common production configuration for 100B+ parameter models is TP + PP + SP combined. Typical allocation: TP=8 within a node (using NVLink), PP across 2 or more nodes, and SP for long-context sequences. This three-dimensional parallelism achieves 75-80% scaling efficiency for 1T+ parameter models across 16 B300 GPUs, compared to 40-50% for TP-only or PP-only approaches.
Network Requirements for Each Strategy
TP demands the highest inter-device bandwidth and benefits most from NVLink domains. B300 clusters with NVLink 5 at 1.8 TB/s per GPU can sustain TP=8 within a node with minimal overhead. For clusters using Ethernet or InfiniBand only, keeping TP below 4 is essential to avoid all-reduce becoming the bottleneck. Each 1 us of all-reduce latency at TP=8 adds approximately 0.5 ms to per-token latency on a 70B model.
PP and SP are far more network-tolerant. A pipeline parallel deployment across 4 nodes requires only 10-20 GB/s per link for activation transfer, well within the capability of a single 200 Gbps InfiniBand connection. SP requires gradient-quality communication for attention reduction but can use the same relatively modest bandwidth. This makes PP+SP the recommended approach for multi-rack deployments without NVLink.
| Strategy | Bandwidth Required | Latency Sensitivity | Best Interconnect | Max Practical Scale |
|---|---|---|---|---|
| Tensor (TP) | 500 GB/s+ | Very high | NVLink 5 | 8 GPUs/node |
| Pipeline (PP) | 10-20 GB/s | Low | Ethernet/IB | O(100) GPUs |
| Sequence (SP) | 20-50 GB/s | Medium | Ethernet/IB | O(100) GPUs |
| TP+PP+SP | 500 GB/s (TP) + 20 GB/s (PP) | Mixed | NVLink + IB | O(1,000) GPUs |
Practical Deployment Guidance
For teams deploying models below 70B parameters, single-GPU inference with quantization is almost always optimal. The B200 with FP4 can run a 70B model entirely on one GPU, eliminating parallelism overhead entirely. For models between 70B and 200B, TP=4 with NVLink inside a single node is the recommended starting point.
For models above 200B parameters, use TP=8 within each NVLink-connected node and PP across nodes. Set the number of PP stages equal to the number of nodes minus one to account for pipeline bubble. Use SP for contexts above 32,768 tokens. This configuration is supported by vLLM, TensorRT-LLM, and NVIDIA Dynamo. ClusterBid can provision clusters with the exact NVLink domain sizes needed for your target model parallelism strategy.
