THE FOUR-MODEL MEMORY WALL IN RLHF
RLHF with PPO requires holding four models in GPU memory simultaneously: the policy model being trained, the reference model for KL divergence computation, the reward model judging output quality, and the value model estimating expected returns. For a 70B policy, this means 4 x 140 GB = 560 GB at FP16, requiring 7 H100 GPUs before any optimizer states, activations, or KV cache. This four-model overhead makes RLHF 3-5x more memory-intensive than supervised fine-tuning of the same model.
The reference model is frozen but must stay in memory because KL divergence is computed at every generation step. The reward model is typically smaller (7-13B parameters, 14-26 GB FP16) but its inference dominates wall-clock time: for each training step, the policy generates responses, then the reward model scores them, then PPO updates the policy. At batch size 32 with 512-token generations, the reward model processes 16,384 tokens per step. At 70B policy scale, RLHF throughput is typically 0.3-0.5 training steps per second per 8-GPU node, versus 2-4 steps per second for SFT.
| Component | 7B Model | 13B Model | 70B Model | 405B Model |
|---|---|---|---|---|
| Policy (trainable) | 14 GB | 26 GB | 140 GB | 810 GB |
| Reference (frozen) | 14 GB | 26 GB | 140 GB | 810 GB |
| Reward Model | 14 GB | 14 GB | 26 GB | 26 GB |
| Value Model | 14 GB | 26 GB | 140 GB | 810 GB |
| Optimizer (Adam) | 28 GB | 52 GB | 280 GB | 1,620 GB |
| Total Weight Memory | 84 GB | 144 GB | 726 GB | 4,076 GB |
| Min H100 GPUs | 2 | 2-4 | 8-16 | 48-64 |
| Peak Activation Mem | 4 GB | 8 GB | 40 GB | 240 GB |
WHY RLHF IS INFERENCE-DOMINATED TRAINING
RLHF is unique among training paradigms because inference dominates compute: the policy generates responses, the reference model computes log-probabilities for KL, and the reward model scores outputs - all inference passes. For each PPO update, the policy generates 256-512 tokens per prompt, the reference model evaluates the same tokens, and the reward model scores them. At batch size 128 with 512-token generations, this is 65,536 tokens of inference per step. The actual PPO weight update is a tiny fraction of total FLOPs.
On an 8x H100 node with a 70B policy under 4-way tensor parallelism and 2-way data parallelism, generation runs at 30-40 tokens per second per replica. A single training step with 128 prompts generating 512 tokens each takes approximately 3-4 seconds for generation, 1-2 seconds for reward scoring, and 0.1-0.3 seconds for the PPO update. Inference accounts for 85-90 percent of step time. This has hardware implications: NVLink bandwidth between GPUs in a TP domain matters more than HBM bandwidth for the weight update, making H100 SXM clusters with 900 GB/s NVLink the practical minimum.
DISTRIBUTED RLHF: ACTOR-CRITIC CLUSTER ARCHITECTURE
Production RLHF at the 70B+ scale separates the actor (policy + reference) and critic (reward + value) into independent GPU pools. The actor pool runs generation and policy updates; the critic pool runs reward scoring and value estimation. Responses flow through a shared high-bandwidth memory store (Redis or similar) between pools, decoupling the generation and scoring pipelines.
With an actor pool of 32 H100 GPUs (4 nodes, 8-way TP) and a critic pool of 8 H100 GPUs (1 node, 7B reward), throughput reaches 8-12 steps per minute with rollout buffer size 1024. Total GPU cost at $3.00/GPU-hour: $480/hour for the full cluster, training a 70B RLHF model to convergence (50K steps) costs $200,000-350,000.
Current practice combines DeepSpeed-Chat or OpenRLHF with ZeRO-3 for memory optimization across GPUs. The RLHF-specific ZeRO extension shares the reference model across data-parallel replicas (ZeRO-3 parameters) while keeping the policy optimizer states distributed. This reduces per-GPU memory by 30-40 percent compared to naive sharding.
| Scale | Actor GPUs | Critic GPUs | Rollout Buffer | Steps/Min | Training Cost |
|---|---|---|---|---|---|
| 7B prototype | 8 H100 (2xTP4, 2xDP) | 1-2 H100 | 256 | 15-25 | $8,000-12,000 |
| 13B production | 16 H100 (4xTP4, 4xDP) | 2-4 H100 | 512 | 10-18 | $30,000-50,000 |
| 70B full-scale | 32 H100 (8xTP8, 4xDP) | 4-8 H100 | 1024 | 8-12 | $200,000-350,000 |
| 405B frontier | 256+ H100 | 16-32 H100 | 4096 | 2-4 | $1.5M-3M |
KL MANAGEMENT AND REWARD NORMALIZATION OVERHEAD
KL divergence computation requires the reference model's log-probabilities for every generated token. For a 512-token generation, this is 512 forward passes through the reference model's transformer with teacher forcing, not generation. Each forward pass processes the full sequence minus one token, creating a 512-element tensor of log-prob differences. At batch size 128, this is 65,536 log-prob pairs per step consuming approximately 8 GB of temporary GPU memory for the KL penalty term computation.
Reward normalization adds another GPU cost center. Raw reward model outputs must be normalized per batch to maintain stable PPO training. Batch-level z-score normalization requires computing mean and variance of up to 128 reward scores per step - a trivial CPU operation - but the more expensive component is reward whitelisting: running a secondary smaller reward model (typically a DistilBERT classifier) to flag reward hacking on every generation. This secondary reward classifier adds 5-10 percent overhead to the critic pool's compute requirements.
ALTERNATIVE PARADIGMS: DPO, GRPO, AND THEIR GPU PROFILES
Direct Preference Optimization (DPO) eliminates the reward and value models entirely, reformulating RLHF as a supervised loss over preference pairs. Training requires only the policy and reference models - two instead of four. For a 70B model, this drops GPU requirements from 16 H100 to 4-8 H100 and training time by 50-60 percent. The tradeoff: DPO cannot use a separate, more capable reward model because the implicit reward emerges from the policy itself.
Group Relative Policy Optimization (GRPO), introduced with DeepSeek-R1, removes the value model by computing advantages from a group of responses rather than a learned value function. This cuts the model count to three (policy, reference, reward) and reduces critic-pool GPUs. GRPO's advantage computation aggregates reward scores across N=8-64 samples per prompt, trading increased generation compute for reduced critic memory. At N=64 on a 70B model, generation cost increases 64x (requiring more actor GPUs) but critic GPUs drop from 8 to 2-4. The total GPU count often matches or exceeds PPO, but the elimination of value model training instability makes GRPO the preferred choice for frontier models above 100B parameters.
| Method | Models in Memory | 70B Min GPUs | Training Time | Reward Quality | Instability Risk |
|---|---|---|---|---|---|
| PPO | 4 (policy, ref, reward, value) | 16 H100 | Baseline | Best (separate RM) | High |
| DPO | 2 (policy, ref) | 4-8 H100 | 40-50% less | Limited (implicit) | Low |
| GRPO (N=8) | 3 (policy, ref, reward) | 12-20 H100 | 10-20% less | Good (separate RM) | Medium |
| GRPO (N=64) | 3 (policy, ref, reward) | 24-40 H100 | Similar | Best (group based) | Low |
| Reinforce++ | 3 (policy, ref, reward) | 12-16 H100 | 20-30% less | Good | Medium |
B200 AND RLHF: THE MEMORY DENSITY SOLUTION
B200's 192 GB HBM3e transforms RLHF by fitting a 70B policy + reference + optimizer states on a single GPU (140 GB weights + 80 GB optimizer = 220 GB still exceeds one GPU, but with ZeRO-3 distributing optimizer states, 2 GPUs suffice versus 4-6 on H100). The four-model requirement drops from 16 H100 GPUs to 6-8 B200 GPUs for 70B PPO, a 50-60 percent reduction in inter-GPU communication overhead.
The critical improvement is in the generation phase: NVLink bandwidth between B200 GPUs in TP configurations is 1.8 TB/s (3rd-gen NVLink), double H100's 900 GB/s. At 70B model size with 4-way TP, reduced communication overhead translates to 1.5-1.8x generation throughput, directly accelerating the inference-dominated RLHF loop. For DPO training, a 70B model fits on a single B200 with ZeRO-3, compared to 4 H100, making DPO-style preference tuning accessible to teams with single-node GPU budgets.
