THE PARALLELISM FRAMEWORK FOR LARGE MODELS
Large models require combining multiple parallelism dimensions because no single dimension scales efficiently beyond 8 GPUs. Tensor parallelism (TP) splits individual layer weights across GPUs, reducing per-GPU VRAM but adding all-reduce communication at every layer. Pipeline parallelism (PP) partitions layer groups across GPUs, adding pipeline bubble idle time. Data parallelism (DP) replicates the model across GPUs and synchronizes gradients, using ZeRO stages to shard optimizer states, gradients, and parameters. Sequence parallelism (SP) splits the sequence dimension across GPUs, critical for long-context workloads. For a 405B model in training, the optimal configuration is TP=8, PP=4, DP=16, SP=2 for 1,024 GPUs total, balancing VRAM, communication, and bubble overhead.
The VRAM budget determines the minimum GPU count. A 405B model at BF16: 810 GB weights + 1,620 GB Adam optimizer (2 states per parameter) + 810 GB gradients = 3,240 GB. With ZeRO-3 sharding across DP, each GPU holds 3,240/DP GB for optimizer states. Additionally, activation memory at micro-batch size 1 with 8K context is approximately 120 GB per layer, with 126 layers total. Activation checkpointing reduces this to 12 GB per layer by recomputing during backward pass. Total VRAM per GPU: 240 GB / (TP * DP) for optimizer states + 12 GB for activations = manageable within H100's 80 GB at TP=8, DP=16 (3,240/128 = 25 GB + 12 GB = 37 GB).
| Parallelism Type | Communication Pattern | Comm Volume per Step | GPU Limit (H100) | When to Use |
|---|---|---|---|---|
| Tensor (TP) | All-reduce | 2x layer params (2-8 GB) | 2-8 GPUs | Within node (NVLink required) |
| Pipeline (PP) | P2P send/recv | Activation size (1-4 GB) | 4-32 GPUs | Across nodes (saves BW) |
| Data (DP + ZeRO-3) | All-gather + reduce-scatter | Gradient size (2-16 GB) | 8-512 GPUs | Main scaling dimension |
| Sequence (SP) | All-to-all | Attention maps (1-8 GB) | 2-8 GPUs | Long context > 32K |
| Expert (EP) | All-to-all | Token hidden states (1-4 GB) | 2-64 GPUs | MoE models only |
INTERCONNECT ANALYSIS: NVLink VS INFINIBAND VS ETHERNET
Tensor parallelism is the most communication-intensive dimension and requires the highest bandwidth interconnect. A single TP all-reduce for Llama 3.1 405B at BF16 transfers 810 GB across GPUs per all-reduce. On H100 NVLink (900 GB/s bidirectional per GPU), this takes 0.9 seconds per training step. On PCIe 5.0 x16 (64 GB/s), it takes 12.7 seconds. On 400 Gbps InfiniBand NDR (50 GB/s per link), with 8 links per GPU, it takes 2.0 seconds. This is why TP is always intra-node with NVLink: PCIe adds 14x overhead and InfiniBand adds 2.2x overhead versus NVLink, directly translating to 14x or 2.2x slower training.
Pipeline parallelism communication is the reverse: it sends activations between pipeline stages rather than sharded weights. Each activation transfer is 4-8 GB per micro-batch per layer boundary. On InfiniBand NDR400 with 8 links per GPU, this takes 2-4 ms per transfer with 6-12 transfers per step, totaling 12-48 ms communication overhead. On RoCEv2 (200 Gbps Ethernet with RDMA), transfer time is 4-8 ms per transfer, totaling 24-96 ms. Data parallelism with ZeRO-3 uses all-gather for parameters and reduce-scatter for gradients. For 405B training with DP=16, the all-gather of 810 GB in FP8 training takes 10 seconds on InfiniBand NDR400 and 16 seconds on RoCEv2. The interconnect choice directly determines the fraction of training time spent communicating versus computing.
| Interconnect | BW per GPU | TP All-Reduce 405B | PP Activation 405B | ZeRO-3 All-Gather 405B | Relative Cost |
|---|---|---|---|---|---|
| NVLink 4.0 (H100) | 900 GB/s bidirectional | 0.9s | 4-9 ms | 0.9s | Included in GPU |
| NVLink 5.0 (B200) | 1,800 GB/s bidirectional | 0.45s | 2-5 ms | 0.45s | Included in GPU |
| InfiniBand NDR400 8x | 50 GB/s per link | 2.0s | 12-48 ms | 10s | $2,000/GPU |
| InfiniBand XDR 8x | 100 GB/s per link | 1.0s | 6-24 ms | 5s | $3,500/GPU |
| RoCEv2 200G 8x | 25 GB/s per link | 4.0s | 24-96 ms | 16s | $800/GPU |
| PCIe 5.0 x16 | 64 GB/s | 12.7s | N/A (not viable) | N/A | Free (on mobo) |
INFERENCE CLUSTER DESIGN FOR 70B AND 405B MODELS
Inference cluster topology differs fundamentally from training. For Llama 3.1 70B inference, the model fits on 1-2 H100s. Single-H100 deployment with AWQ (38 GB + 12 GB KV cache at 8K context = 50 GB) leaves 30 GB for batch processing, serving 380 tok/s at batch size 16. For higher throughput, 2x H100 with TP=2 doubles to 720 tok/s at batch size 32. The interconnect requirement is minimal: TP=2 communication is 2 GB per all-reduce per layer, taking 5.5 ms on PCIe 5.0, 0.4 ms on NVLink. For 70B inference, PCIe 5.0 between 2 GPUs is sufficient; NVLink provides marginal benefit. The cluster design is a high-density configuration: 8x 2-GPU inference nodes per DGX server, each serving one model replica.
For 405B inference, the equation changes. At FP8 on H100, 405B requires 405 GB + 128 GB KV cache at 8K context = 533 GB, requiring 7 H100s minimum (560 GB). With TP=8 across 8 GPUs, each GPU holds 67 GB + 16 GB KV cache = 83 GB, just over H100's 80 GB. AWQ W4A16 reduces to 205 GB + 128 GB = 333 GB, fitting on 4 H100s (320 GB = 4x80) with slight memory pressure. The recommended production configuration: AWQ 405B on 4x H100 with TP=4, EP=1, serving 140 tok/s at batch size 8. For latency-critical deployments at batch size 1, 8x H100 with TP=8 achieves 22 tok/s at sub-1s TTFT with 4K input. The cluster interconnect for 4-GPU inference nodes can use PCIe 5.0; for 8-GPU, NVLink is strongly recommended to avoid communication bottlenecks in the attention all-reduce.
| Model | Precision | GPUs | Parallelism | Peak tok/s | Cost/hr | Cost/1M tok |
|---|---|---|---|---|---|---|
| Llama 3.1 70B | AWQ W4A16 | 1x H100 | TP=1 | 380 | $2.50 | $0.81 |
| Llama 3.1 70B | FP8 | 2x H100 | TP=2 | 720 | $5.00 | $0.93 |
| Llama 3.1 405B | AWQ W4A16 | 4x H100 | TP=4 | 140 | $10.00 | $4.46 |
| Llama 3.1 405B | FP8 | 8x H100 | TP=8 | 190 | $20.00 | $6.58 |
| Qwen3-235B MoE | BF16 | 4x H100 | TP=2, EP=4 | 310 | $10.00 | $1.79 |
| DeepSeek-V3 685B | FP8 | 8x H100 | TP=8, EP=8 | 280 | $20.00 | $4.96 |
CLUSTER BUILD-OUT COST AND SCALING
The cost to build a large-model inference cluster spans two tiers. Tier 1: mid-scale 64-GPU cluster for serving 70B models at 20,000+ tok/s aggregate. Using H100 SXM with NVLink switches, the cluster costs $2.4-3.2M (8x DGX H100 nodes at $300-400K each). With 64 GPUs running 32x 2-GPU 70B replicas, aggregate throughput is 32 * 720 = 23,040 tok/s for Llama 70B FP8 TP=2. Cost per 1M tokens at 100% utilization: $160/hr / 23,040 tok/s = $0.002 per 1M tokens. Monthly serving cost for 100B tokens: $200. The interconnect for intra-node is NVLink (included in DGX), inter-node is 400 Gbps InfiniBand for control and model loading.
Tier 2: large-scale 512-GPU cluster for 405B and MoE models. Using 64x DGX H100 nodes (8 GPUs each) at $300K each = $19.2M hardware. With AWQ 405B on 4-GPU nodes, 128 model replicas at 140 tok/s = 17,920 tok/s aggregate. Cost per 1M tokens: $1,280/hr / 17,920 tok/s = $0.020 per 1M tokens, 10x higher than the 70B cost due to the 4-GPU requirement per replica. For DeepSeek-V3 685B on 8-GPU nodes at TP=8, EP=8, 64 replicas at 280 tok/s = 17,920 tok/s aggregate at the same $1,280/hr = $0.020 per 1M tokens. The 405B+ MoE models converge on similar cost per token because the 8-GPU-per-replica overhead offsets the higher throughput per GPU. GPU efficiency (tokens per GPU per second) remains the critical metric: 70B delivers 360 tok/s/GPU, 405B delivers 35 tok/s/GPU, MoE delivers 45 tok/s/GPU.
