The RLHF Training Stack
Reinforcement Learning from Human Feedback (RLHF) requires holding multiple models in memory simultaneously. The canonical stack includes: the actor model (being trained), the critic model (value function), the reference model (frozen, for KL divergence computation), and the reward model (frozen, for scoring completions). For PPO, all four models must be loaded simultaneously. GRPO eliminates the critic but keeps actor, reference, and reward. DPO simplifies to just two identical models: the trained policy and the frozen reference.
| Component | Memory per 7B (FP16) | Memory per 13B (FP16) | Memory per 34B (FP16) | Memory per 70B (FP16) |
|---|---|---|---|---|
| Actor (trainable) | 14 GB + 28 GB (opt states) | 26 GB + 52 GB | 68 GB + 136 GB | 140 GB + 280 GB |
| Critic (trainable) | 14 GB + 28 GB (opt states) | 26 GB + 52 GB | 68 GB + 136 GB | 140 GB + 280 GB |
| Reference (frozen) | 14 GB | 26 GB | 68 GB | 140 GB |
| Reward (frozen) | 7-14 GB | 13-26 GB | 34-68 GB | 70-140 GB |
| Total PPO (all four) | 105-119 GB | 195-209 GB | 510-544 GB | 1,050-1,120 GB |
| Total GRPO (no critic) | 49-56 GB | 91-104 GB | 238-272 GB | 490-560 GB |
| Total DPO (actor+ref) | 28 GB + opt | 52 GB + opt | 136 GB + opt | 280 GB + opt |
PPO GPU Memory Math
PPO is the most memory-intensive alignment method because it requires training two models (actor and critic) simultaneously. The optimizer states (first and second moments for AdamW) consume 2x the model weights in FP32, or equivalently 4x the FP16 weights if stored in FP32. For a 7B PPO training run, the actor optimizer states add 56 GB (14 GB weights x 4 for FP32 moments). Combined with critic optimizer states, weights of all four models, and the generation buffer (completions + log probs), total memory exceeds 120 GB per GPU.
In practice, 7B PPO training requires 2x A100 80 GB GPUs with ZeRO-3 sharding, or 4x A100 40 GB GPUs. DeepSpeed ZeRO-3 + LoRA (rank 64) reduces memory by 60-70%, enabling 7B PPO on a single 80 GB GPU. For 70B PPO, the requirement scales to 16-32 H100 80 GB GPUs with ZeRO-3 + activation checkpointing. The critic model adds significant memory pressure for large models because it is the same size as the actor but must be trained, doubling the optimizer state overhead.
GRPO Optimization
GRPO (Group Relative Policy Optimization) removes the critic model, replacing the learned value function with a group-based reward normalization. This eliminates the critic's weight memory (14-140 GB depending on scale) and its optimizer states (28-280 GB). For 7B GRPO, total memory drops to approximately 50-56 GB, fitting comfortably on a single A100 80 GB or H100. For 70B GRPO, memory requirements drop from ~1,100 GB (PPO) to ~500-560 GB, fitting on 8x H100 80 GB with ZeRO-3.
The memory savings come with a compute trade-off. GRPO generates G completions per prompt (typically G=8-64), requiring Gx the generation throughput of PPO. For G=8, GRPO generates 8x more tokens per training step but avoids the critic forward and backward passes. In practice, GRPO achieves 60-80% wall-clock speedup over PPO for the same model scale because generation is compute-bound in a way that is efficiently parallelizable across GPUs, while critic training is memory-bound and scales poorly.
DPO Simplicity
Direct Preference Optimization (DPO) dramatically simplifies the alignment memory equation. DPO requires only two models: the trained policy and a frozen reference model. No critic, no reward model, no online generation. The loss function operates on paired preference data (chosen/rejected completions) and computes the implicit reward from the policy ratio. Total memory for 7B DPO is approximately 28 GB for weights + 56 GB for optimizer states = 84 GB, fitting on a single 80 GB GPU with gradient checkpointing.
The memory efficiency of DPO makes it the alignment method of choice for teams with limited GPU budgets. A 70B DPO run requires approximately 280 GB for weights + 560 GB for optimizer states = 840 GB, fitting on 16x A100 80 GB (54% utilization) or 12x H100 80 GB (88% utilization). With LoRA (rank 128), 70B DPO runs on a single H100 80 GB with 4-bit quantization of the base model. The trade-off is that DPO does not use online exploration, which can lead to less robust alignment for complex reward structures.
Cluster Sizing for Alignment Training
Alignment training clusters require different GPU configurations than pre-training clusters. The memory-proportional nature of RLHF means GPU memory capacity is the binding constraint, not compute throughput. A cluster optimized for 70B PPO needs 32x H100 80 GB with NVLink (4 nodes of 8 GPUs). The same cluster can run 70B GRPO with 16 GPUs, or 70B DPO with 12 GPUs. The cluster should be designed for the PPO worst case while reaping the efficiency benefits of GRPO and DPO for production runs.
Inter-node bandwidth becomes critical for alignment training because of the frequent gradient synchronization across all training models. For 70B PPO on 32 GPUs, each training step synchronizes ~560 MB of gradients (actor + critic) across the cluster. A 400 Gbps InfiniBand fabric handles this in approximately 12 ms per all-reduce. With 200 Gbps Ethernet, the same operation takes 25-30 ms, potentially reducing training throughput by 15-25%. NVLink within nodes and InfiniBand between nodes is the recommended topology for alignment clusters.
Cost Comparison of Alignment Methods
The cost per alignment run varies by 3-5x between methods at the same model scale. For a 70B model trained on 1M preference pairs with 1 epoch: PPO costs approximately $45,000-60,000 in GPU compute (32x H100 for 3-4 days), GRPO costs $18,000-25,000 (16x H100 for 2-3 days), and DPO costs $6,000-9,000 (12x H100 for 1-1.5 days). These costs assume cloud GPU rental at $3.50/GPU/hr. The price-performance sweet spot is GRPO for most production applications: 80% of PPO alignment quality at 40% of the cost.
| Alignment Method | Min GPUs (70B) | Memory per GPU | Training Time | Total GPU Cost | Relative Quality |
|---|---|---|---|---|---|
| PPO (Full FT) | 32x H100 | 35 GB | 3-4 days | $45K-60K | Baseline (1.0x) |
| PPO + LoRA | 8x H100 | 65 GB | 5-7 days | $15K-20K | 0.85-0.90x |
| GRPO (Full FT) | 16x H100 | 35 GB | 2-3 days | $18K-25K | 0.85-0.95x |
| GRPO + LoRA | 4x H100 | 70 GB | 3-5 days | $6K-10K | 0.75-0.85x |
| DPO (Full FT) | 12x H100 | 70 GB | 1-1.5 days | $6K-9K | 0.80-0.90x |
| DPO + LoRA | 1x H100 | 70 GB | 2-3 days | $2K-3K | 0.70-0.80x |
