All essays
TechnicalDEEP DIVEFEB 2026

Multi-Node GPU Inference: Tensor Parallelism vs Pipeline Parallelism for Deploying Models Across Multiple GPUs

Technical deep dive comparing tensor parallelism, pipeline parallelism, and sequence parallelism for distributing LLM inference across multi-GPU clusters. Latency, throughput, and memory tradeoffs with benchmarks.

01

The Parallelism Landscape

Deploying large language models across multiple GPUs requires distributing the model across devices. Three primary strategies exist: tensor parallelism (TP), pipeline parallelism (PP), and sequence parallelism (SP). Each makes different tradeoffs between inter-device communication, memory utilization, and end-to-end latency.

The choice of parallelism strategy is the single largest factor in inference performance for models that do not fit on a single GPU. A 405B parameter Llama 4 model requires 810 GB of FP16 weights, exceeding the 288 GB HBM3e capacity of a single B300 GPU. Even at FP4 precision, models above 200B parameters require multi-GPU distribution for both memory and compute throughput.

02

Tensor Parallelism Deep Dive

Tensor parallelism splits individual weight matrices across GPUs. In a 2-way TP configuration, a single matrix multiplication is divided: each GPU holds half the rows or columns and performs a partial matmul, followed by an all-reduce to combine results. TP is the most communication-intensive strategy, requiring an all-reduce for every transformer layer.

TP works best within a single node where NVLink provides 900 GB/s to 1.8 TB/s of inter-GPU bandwidth. Across nodes, TP degrades rapidly because NCCL all-reduce over InfiniBand or Ethernet (50 GB/s per link) becomes the bottleneck. A B200 with 8-way TP over NVLink achieves 2.3x throughput versus 8-way TP split across two nodes over 400 Gbps Ethernet, due entirely to the all-reduce bottleneck.

TP DegreeInterconnectAll-Reduce TimeThroughput (tok/s)Scaling Efficiency
TP=2NVLink 5 (1.8 TB/s)12 us2,45098%
TP=4NVLink 5 (1.8 TB/s)24 us4,68094%
TP=8 (1 node)NVLink 5 (1.8 TB/s)48 us8,10082%
TP=8 (2 nodes)400 Gbps Ethernet810 us3,20032%
TP=4 (4 nodes)200 Gbps InfiniBand940 us1,85019%
03

Pipeline Parallelism Deep Dive

Pipeline parallelism places different transformer layers on different GPUs. Input tokens flow through GPU 0 layers 1-10, then GPU 1 layers 11-20, and so on. PP requires far less inter-GPU communication: only the activations passed between pipeline stages. Communication is proportional to hidden dimension times sequence length per microbatch, not to the full model parameter count.

The primary cost of PP is the pipeline bubble: idle GPU time during the warm-up and cool-down phases when not all stages are processing. With 4 pipeline stages and 8 microbatches, pipeline utilization is approximately 75%. Increasing microbatches to 64 raises utilization to 95%. PP also requires higher total batch sizes, which increases end-to-end latency for interactive applications.

04

Sequence Parallelism and Hybrid Strategies

Sequence parallelism splits the sequence dimension across GPUs, allowing longer context windows than any single GPU can process. SP is often combined with TP (SP + TP = SPT), where the sequence dimension is split for attention computations while weights are split for feed-forward layers. This hybrid approach is the standard deployment configuration for models processing 128K+ token contexts.

The most common production configuration for 100B+ parameter models is TP + PP + SP combined. Typical allocation: TP=8 within a node (using NVLink), PP across 2 or more nodes, and SP for long-context sequences. This three-dimensional parallelism achieves 75-80% scaling efficiency for 1T+ parameter models across 16 B300 GPUs, compared to 40-50% for TP-only or PP-only approaches.

05

Network Requirements for Each Strategy

TP demands the highest inter-device bandwidth and benefits most from NVLink domains. B300 clusters with NVLink 5 at 1.8 TB/s per GPU can sustain TP=8 within a node with minimal overhead. For clusters using Ethernet or InfiniBand only, keeping TP below 4 is essential to avoid all-reduce becoming the bottleneck. Each 1 us of all-reduce latency at TP=8 adds approximately 0.5 ms to per-token latency on a 70B model.

PP and SP are far more network-tolerant. A pipeline parallel deployment across 4 nodes requires only 10-20 GB/s per link for activation transfer, well within the capability of a single 200 Gbps InfiniBand connection. SP requires gradient-quality communication for attention reduction but can use the same relatively modest bandwidth. This makes PP+SP the recommended approach for multi-rack deployments without NVLink.

StrategyBandwidth RequiredLatency SensitivityBest InterconnectMax Practical Scale
Tensor (TP)500 GB/s+Very highNVLink 58 GPUs/node
Pipeline (PP)10-20 GB/sLowEthernet/IBO(100) GPUs
Sequence (SP)20-50 GB/sMediumEthernet/IBO(100) GPUs
TP+PP+SP500 GB/s (TP) + 20 GB/s (PP)MixedNVLink + IBO(1,000) GPUs
06

Practical Deployment Guidance

For teams deploying models below 70B parameters, single-GPU inference with quantization is almost always optimal. The B200 with FP4 can run a 70B model entirely on one GPU, eliminating parallelism overhead entirely. For models between 70B and 200B, TP=4 with NVLink inside a single node is the recommended starting point.

For models above 200B parameters, use TP=8 within each NVLink-connected node and PP across nodes. Set the number of PP stages equal to the number of nodes minus one to account for pipeline bubble. Use SP for contexts above 32,768 tokens. This configuration is supported by vLLM, TensorRT-LLM, and NVIDIA Dynamo. ClusterBid can provision clusters with the exact NVLink domain sizes needed for your target model parallelism strategy.

Filed under
tensor parallelismpipeline parallelismsequence parallelismmulti-GPU inferenceLLM deploymentNVLinkNCCL all-reducevLLM