SFT VS RLHF OVERVIEW
Supervised fine-tuning adapts a pre-trained model using labeled input-output pairs with standard cross-entropy loss. Reinforcement learning from human feedback requires maintaining multiple model instances-policy, reference, reward, and value models-during training. SFT typically requires 4-8 GPUs for a 70B parameter model, while full PPO RLHF requires 32-64 GPUs for the same parameter count. The GPU memory and interconnect requirements diverge significantly between these approaches.
GPU MEMORY COMPARISON
SFT on a 70B model at BF16 requires approximately 140 GB for model weights, 80 GB for optimizer states with AdamW, and 40-60 GB for activations at moderate batch sizes. With ZeRO-3 sharding across 8 H100s, each GPU needs 35-45 GB. Full PPO RLHF requires 4-6 model instances: policy, reference, reward, and optionally value and critic models. This multiplies memory requirements to 100-140 GB per GPU across 32-64 H100s.
| Configuration | Models in Memory | Min GPUs | GPU Memory/GPU | Interconnect |
|---|---|---|---|---|
| 70B SFT (BF16) | 1 | 4-8 | 35-45 GB | 200 GB/s+ |
| 70B SFT (FP8) | 1 | 2-4 | 20-25 GB | 200 GB/s+ |
| 70B PPO RLHF | 4-6 | 32-64 | 100-140 GB | 800 GB/s NVLink |
| 70B DPO | 2 | 8-16 | 50-70 GB | 400 GB/s+ |
DPO AS A MIDDLE GROUND
Direct Preference Optimization eliminates the need for a separate reward model by training directly on preference pairs. DPO requires only two model instances-policy and reference-reducing GPU requirements by 40-60% compared to PPO. A 70B DPO run fits on 8-16 H100s versus 32-64 for PPO, with comparable alignment quality on standard benchmarks. The trade-off is that DPO is less sample-efficient for complex preference structures.
PPO MULTI-MODEL COST
Full PPO RLHF maintains the policy model, reference model, reward model, and value model simultaneously during training. Each model requires separate forward and backward passes per training step. At 70B scale, PPO consumes approximately 2,500-3,500 GPU-hours on H100 per training run, compared to 400-700 GPU-hours for SFT and 800-1,400 for DPO. The GPU-hour premium for PPO is 3-6x over DPO depending on convergence requirements.
CLUSTER TOPOLOGY REQUIREMENTS
SFT and DPO can run effectively on 8-GPU nodes with 200-400 GB/s inter-node bandwidth using RoCE or InfiniBand. Full PPO benefits significantly from NVLink-connected H100 HGX baseboards with 8 GPUs per node at 900 GB/s intra-node bandwidth. PPO training on 32+ GPUs requires at least 400 GB/s inter-node bandwidth to avoid communication bottlenecks, pushing teams toward InfiniBand or Spectrum-X fabrics.
BUILD VS API DECISION
Teams spending under $50,000 per month on alignment training should use API-based solutions from Together AI, Fireworks, or Anyscale. At $50,000-$200,000 per month, reserved H100 clusters with DPO provide better cost efficiency. Above $200,000 per month, building a dedicated alignment cluster with PPO capability becomes economical. The breakeven point shifts 20-30% lower for teams already owning H100 hardware.
