All essays
TechnicalDEEP DIVEFEB 2026

Pipeline Parallelism for LLM Training: Optimal Layer Partitioning Strategies

Optimal layer partitioning for pipeline parallelism in LLM training. Compare schedules with throughput and memory benchmarks across H100 clusters.

01

WHY PIPELINE PARALLELISM IS CRITICAL

Models exceeding 30B parameters cannot fit on a single GPU. Pipeline parallelism splits layers across GPUs. A 70B Llama 3 in BF16 needs 420 GB total (params + optimizers + gradients), requiring at least 5-6 H100 80GB GPUs.

The pipeline bubble with naive scheduling is 28.6 percent at 8 stages, 48.5 percent at 32 stages. Modern techniques reduce this to under 10 percent.

ScheduleBubble 8 stagesBubble 32 stagesMemory/GPUComplexityBest For
Naive GPipe28.6%48.5%LowVery simplePrototyping
1F1B14.3%24.2%LowSimpleProduction
Interleaved 1F1B7.1%12.1%LowModerateMedium clusters
Zero-Bubble ZB13.5%6.4%ModerateComplexLarge-scale
Zero-Bubble ZB21.8%3.5%HighVery complexMax throughput
02

LAYER PARTITIONING HEURISTICS

Different layers have different compute-to-memory ratios. Attention layers are memory-bandwidth-bound, MLP layers compute-bound. For Llama with 80 layers and 8 stages, naive 10 layers/stage produces 15-22 percent imbalance.

Profile-driven partitioning improves throughput by 12-18 percent. For heterogeneous clusters (H100 + A100 mixed), assign 60 percent layers to H100 stages for 23 percent higher throughput and 18 percent lower cost.

MethodThroughput vs NaiveSetup TimeMemory VarianceSensitivity
Equal (naive)BaselineNone+-12%Low
Manual tuning+8-15%2-4 hours+-5%High
Profile-driven+12-18%30-60 min+-3%Medium
Heterogeneous-aware+18-25%Profile+10min+-5%Medium
03

MEMORY AND ACTIVATION CHECKPOINTING

Activation memory for a 70B model with seq 8,192, batch 512, 8 stages: 40 GB per GPU with FlashAttention-2, 85 GB without. Checkpointing reduces to 12 GB at 20-30 percent compute overhead.

Optimal point for 70B on H100 80GB: 8-16 stages with selective checkpointing on attention only, achieving 92-95 percent peak throughput. Long-context training (128K tokens) should reduce pipeline depth by 30-50 percent.

04

PRACTICAL DEPLOYMENT PATTERNS

Common 70B config: TP=4, PP=8, DP=4 on 128 H100 GPUs achieves 850 tokens/sec/GPU, 109,000 tokens/sec total. Frontier models (500B): TP=8, PP=16, DP=16 on 2,048 H100 GPUs achieves 450 tokens/sec/GPU, 920,000 tokens/sec. At $3.50/GPU-hr, a 2T token training run costs $4.3M.

Pipeline depth should be tuned per run, not fixed. Dynamic partitioning across jobs achieves 15-25 percent higher cluster-wide utilization.

Filed under
Pipeline ParallelismLLM TrainingLayer Partitioning1F1B ScheduleModel ParallelismDistributed Training