WHY PIPELINE PARALLELISM IS CRITICAL
Models exceeding 30B parameters cannot fit on a single GPU. Pipeline parallelism splits layers across GPUs. A 70B Llama 3 in BF16 needs 420 GB total (params + optimizers + gradients), requiring at least 5-6 H100 80GB GPUs.
The pipeline bubble with naive scheduling is 28.6 percent at 8 stages, 48.5 percent at 32 stages. Modern techniques reduce this to under 10 percent.
| Schedule | Bubble 8 stages | Bubble 32 stages | Memory/GPU | Complexity | Best For |
|---|---|---|---|---|---|
| Naive GPipe | 28.6% | 48.5% | Low | Very simple | Prototyping |
| 1F1B | 14.3% | 24.2% | Low | Simple | Production |
| Interleaved 1F1B | 7.1% | 12.1% | Low | Moderate | Medium clusters |
| Zero-Bubble ZB1 | 3.5% | 6.4% | Moderate | Complex | Large-scale |
| Zero-Bubble ZB2 | 1.8% | 3.5% | High | Very complex | Max throughput |
LAYER PARTITIONING HEURISTICS
Different layers have different compute-to-memory ratios. Attention layers are memory-bandwidth-bound, MLP layers compute-bound. For Llama with 80 layers and 8 stages, naive 10 layers/stage produces 15-22 percent imbalance.
Profile-driven partitioning improves throughput by 12-18 percent. For heterogeneous clusters (H100 + A100 mixed), assign 60 percent layers to H100 stages for 23 percent higher throughput and 18 percent lower cost.
| Method | Throughput vs Naive | Setup Time | Memory Variance | Sensitivity |
|---|---|---|---|---|
| Equal (naive) | Baseline | None | +-12% | Low |
| Manual tuning | +8-15% | 2-4 hours | +-5% | High |
| Profile-driven | +12-18% | 30-60 min | +-3% | Medium |
| Heterogeneous-aware | +18-25% | Profile+10min | +-5% | Medium |
MEMORY AND ACTIVATION CHECKPOINTING
Activation memory for a 70B model with seq 8,192, batch 512, 8 stages: 40 GB per GPU with FlashAttention-2, 85 GB without. Checkpointing reduces to 12 GB at 20-30 percent compute overhead.
Optimal point for 70B on H100 80GB: 8-16 stages with selective checkpointing on attention only, achieving 92-95 percent peak throughput. Long-context training (128K tokens) should reduce pipeline depth by 30-50 percent.
PRACTICAL DEPLOYMENT PATTERNS
Common 70B config: TP=4, PP=8, DP=4 on 128 H100 GPUs achieves 850 tokens/sec/GPU, 109,000 tokens/sec total. Frontier models (500B): TP=8, PP=16, DP=16 on 2,048 H100 GPUs achieves 450 tokens/sec/GPU, 920,000 tokens/sec. At $3.50/GPU-hr, a 2T token training run costs $4.3M.
Pipeline depth should be tuned per run, not fixed. Dynamic partitioning across jobs achieves 15-25 percent higher cluster-wide utilization.
