All essays
TechnicalDEEP DIVEFEB 2026

RLHF GPU Infrastructure: Scaling Reinforcement Learning from Human Feedback

GPU infrastructure deep dive for RLHF training: PPO memory requirements for policy, reference, reward, and value models. Inference overhead during training, distributed actor-critic GPU cluster design, and cost analysis for scaling RLHF from 7B to 405B parameter models.

01

THE FOUR-MODEL MEMORY WALL IN RLHF

RLHF with PPO requires holding four models in GPU memory simultaneously: the policy model being trained, the reference model for KL divergence computation, the reward model judging output quality, and the value model estimating expected returns. For a 70B policy, this means 4 x 140 GB = 560 GB at FP16, requiring 7 H100 GPUs before any optimizer states, activations, or KV cache. This four-model overhead makes RLHF 3-5x more memory-intensive than supervised fine-tuning of the same model.

The reference model is frozen but must stay in memory because KL divergence is computed at every generation step. The reward model is typically smaller (7-13B parameters, 14-26 GB FP16) but its inference dominates wall-clock time: for each training step, the policy generates responses, then the reward model scores them, then PPO updates the policy. At batch size 32 with 512-token generations, the reward model processes 16,384 tokens per step. At 70B policy scale, RLHF throughput is typically 0.3-0.5 training steps per second per 8-GPU node, versus 2-4 steps per second for SFT.

Component7B Model13B Model70B Model405B Model
Policy (trainable)14 GB26 GB140 GB810 GB
Reference (frozen)14 GB26 GB140 GB810 GB
Reward Model14 GB14 GB26 GB26 GB
Value Model14 GB26 GB140 GB810 GB
Optimizer (Adam)28 GB52 GB280 GB1,620 GB
Total Weight Memory84 GB144 GB726 GB4,076 GB
Min H100 GPUs22-48-1648-64
Peak Activation Mem4 GB8 GB40 GB240 GB
02

WHY RLHF IS INFERENCE-DOMINATED TRAINING

RLHF is unique among training paradigms because inference dominates compute: the policy generates responses, the reference model computes log-probabilities for KL, and the reward model scores outputs - all inference passes. For each PPO update, the policy generates 256-512 tokens per prompt, the reference model evaluates the same tokens, and the reward model scores them. At batch size 128 with 512-token generations, this is 65,536 tokens of inference per step. The actual PPO weight update is a tiny fraction of total FLOPs.

On an 8x H100 node with a 70B policy under 4-way tensor parallelism and 2-way data parallelism, generation runs at 30-40 tokens per second per replica. A single training step with 128 prompts generating 512 tokens each takes approximately 3-4 seconds for generation, 1-2 seconds for reward scoring, and 0.1-0.3 seconds for the PPO update. Inference accounts for 85-90 percent of step time. This has hardware implications: NVLink bandwidth between GPUs in a TP domain matters more than HBM bandwidth for the weight update, making H100 SXM clusters with 900 GB/s NVLink the practical minimum.

03

DISTRIBUTED RLHF: ACTOR-CRITIC CLUSTER ARCHITECTURE

Production RLHF at the 70B+ scale separates the actor (policy + reference) and critic (reward + value) into independent GPU pools. The actor pool runs generation and policy updates; the critic pool runs reward scoring and value estimation. Responses flow through a shared high-bandwidth memory store (Redis or similar) between pools, decoupling the generation and scoring pipelines.

With an actor pool of 32 H100 GPUs (4 nodes, 8-way TP) and a critic pool of 8 H100 GPUs (1 node, 7B reward), throughput reaches 8-12 steps per minute with rollout buffer size 1024. Total GPU cost at $3.00/GPU-hour: $480/hour for the full cluster, training a 70B RLHF model to convergence (50K steps) costs $200,000-350,000.

Current practice combines DeepSpeed-Chat or OpenRLHF with ZeRO-3 for memory optimization across GPUs. The RLHF-specific ZeRO extension shares the reference model across data-parallel replicas (ZeRO-3 parameters) while keeping the policy optimizer states distributed. This reduces per-GPU memory by 30-40 percent compared to naive sharding.

ScaleActor GPUsCritic GPUsRollout BufferSteps/MinTraining Cost
7B prototype8 H100 (2xTP4, 2xDP)1-2 H10025615-25$8,000-12,000
13B production16 H100 (4xTP4, 4xDP)2-4 H10051210-18$30,000-50,000
70B full-scale32 H100 (8xTP8, 4xDP)4-8 H10010248-12$200,000-350,000
405B frontier256+ H10016-32 H10040962-4$1.5M-3M
04

KL MANAGEMENT AND REWARD NORMALIZATION OVERHEAD

KL divergence computation requires the reference model's log-probabilities for every generated token. For a 512-token generation, this is 512 forward passes through the reference model's transformer with teacher forcing, not generation. Each forward pass processes the full sequence minus one token, creating a 512-element tensor of log-prob differences. At batch size 128, this is 65,536 log-prob pairs per step consuming approximately 8 GB of temporary GPU memory for the KL penalty term computation.

Reward normalization adds another GPU cost center. Raw reward model outputs must be normalized per batch to maintain stable PPO training. Batch-level z-score normalization requires computing mean and variance of up to 128 reward scores per step - a trivial CPU operation - but the more expensive component is reward whitelisting: running a secondary smaller reward model (typically a DistilBERT classifier) to flag reward hacking on every generation. This secondary reward classifier adds 5-10 percent overhead to the critic pool's compute requirements.

05

ALTERNATIVE PARADIGMS: DPO, GRPO, AND THEIR GPU PROFILES

Direct Preference Optimization (DPO) eliminates the reward and value models entirely, reformulating RLHF as a supervised loss over preference pairs. Training requires only the policy and reference models - two instead of four. For a 70B model, this drops GPU requirements from 16 H100 to 4-8 H100 and training time by 50-60 percent. The tradeoff: DPO cannot use a separate, more capable reward model because the implicit reward emerges from the policy itself.

Group Relative Policy Optimization (GRPO), introduced with DeepSeek-R1, removes the value model by computing advantages from a group of responses rather than a learned value function. This cuts the model count to three (policy, reference, reward) and reduces critic-pool GPUs. GRPO's advantage computation aggregates reward scores across N=8-64 samples per prompt, trading increased generation compute for reduced critic memory. At N=64 on a 70B model, generation cost increases 64x (requiring more actor GPUs) but critic GPUs drop from 8 to 2-4. The total GPU count often matches or exceeds PPO, but the elimination of value model training instability makes GRPO the preferred choice for frontier models above 100B parameters.

MethodModels in Memory70B Min GPUsTraining TimeReward QualityInstability Risk
PPO4 (policy, ref, reward, value)16 H100BaselineBest (separate RM)High
DPO2 (policy, ref)4-8 H10040-50% lessLimited (implicit)Low
GRPO (N=8)3 (policy, ref, reward)12-20 H10010-20% lessGood (separate RM)Medium
GRPO (N=64)3 (policy, ref, reward)24-40 H100SimilarBest (group based)Low
Reinforce++3 (policy, ref, reward)12-16 H10020-30% lessGoodMedium
06

B200 AND RLHF: THE MEMORY DENSITY SOLUTION

B200's 192 GB HBM3e transforms RLHF by fitting a 70B policy + reference + optimizer states on a single GPU (140 GB weights + 80 GB optimizer = 220 GB still exceeds one GPU, but with ZeRO-3 distributing optimizer states, 2 GPUs suffice versus 4-6 on H100). The four-model requirement drops from 16 H100 GPUs to 6-8 B200 GPUs for 70B PPO, a 50-60 percent reduction in inter-GPU communication overhead.

The critical improvement is in the generation phase: NVLink bandwidth between B200 GPUs in TP configurations is 1.8 TB/s (3rd-gen NVLink), double H100's 900 GB/s. At 70B model size with 4-way TP, reduced communication overhead translates to 1.5-1.8x generation throughput, directly accelerating the inference-dominated RLHF loop. For DPO training, a 70B model fits on a single B200 with ZeRO-3, compared to 4 H100, making DPO-style preference tuning accessible to teams with single-node GPU budgets.

Filed under
RLHF GPU InfrastructurePPO Training GPUReinforcement Learning AIHuman Feedback TrainingDistributed RLHFActor Critic GPU ClusterReward Model Training