THE MEMORY PRESSURE MAP OF TRAINING
Training memory breaks into five categories: model weights (FP16), optimizer states (FP32 momentum + variance), gradients (FP16), activations (FP16, sequence-length-dependent), and temporary buffers. For a 7B model with batch size 1 and 2K sequence, activations dominate at 40-60 percent of total memory. At batch size 32 with 8K sequence on a 70B model, activations consume 300-500 GB, dwarfing the 140 GB weight memory.
The four primary techniques each target a different category: gradient accumulation targets optimizer memory by simulating large batches on small memory budgets; activation checkpointing trades compute for activation memory by recomputing activations during backward; ZeRO shards weights, gradients, and optimizer states across GPUs; memory-efficient optimizers (bitsandbytes 8-bit Adam, Sophia) compress optimizer state storage. Each technique has a distinct compute-memory tradeoff curve.
| Technique | Memory Saved | Compute Overhead | Target Component | Complexity |
|---|---|---|---|---|
| Gradient Accumulation | 50-80% opt mem | 0% (free) | Optimizer states | Trivial |
| Activation Checkpointing | 60-85% | 20-33% | Activations | Low |
| ZeRO-1 (opt shard) | 75% opt state | <1% comm | Optimizer states | Low |
| ZeRO-2 (grad shard) | 87% grad + opt | 1-3% comm | Gradients + opt | Medium |
| ZeRO-3 (param shard) | 90-95% total | 5-15% comm | Weights + grad + opt | Medium |
| 8-bit Adam | 75% opt state | 0-5% | Optimizer states | Low |
| CPU Offload | 90-99% | 20-50% | Everything | Medium |
GRADIENT ACCUMULATION: THE FREE LUNCH IN MEMORY OPTIMIZATION
Gradient accumulation separates the effective batch size from the micro-batch size. The model processes micro-batches (e.g., batch size 1-4), accumulates gradients in FP32, and updates weights only after N micro-batches. This reduces activation memory by Nx because activations only need to be stored for the micro-batch, not the full effective batch. For a 70B model, training with effective batch size 1024 directly requires 5-10 TB of activation memory. With gradient accumulation N=256, micro-batch size 4, activation memory drops to 20-40 GB.
The remarkable property of gradient accumulation is near-zero compute overhead: the total FLOPs are identical to the equivalent large batch, and NVIDIA GPUs handle the gradient addition as a fusion in the backward pass. The only cost is additional optimizer update steps - 256 micro-batches produce one weight update instead of 1 update per batch - but each update is computation-free (just gradient addition). In practice, gradient accumulation replaces global batch size as the training hyperparameter; DeepSpeed and Megatron-LM expose gradient_accumulation_steps as a primary config.
The constraint is statistical, not computational: effective batch size above 1024 for 7B models and 2048 for 70B models produces diminishing returns in gradient variance reduction. Past these thresholds, larger effective batches don't improve training dynamics, so gradient accumulation doesn't help beyond matching these optimal effective batch sizes.
| Effective BS | Micro-Batch | Grad Accum Steps | Activation Memory | Throughput |
|---|---|---|---|---|
| 1024 | 4 | 256 | ~25 GB | 100% (baseline) |
| 1024 | 8 | 128 | ~50 GB | 102% |
| 1024 | 16 | 64 | ~100 GB | 104% |
| 1024 | 32 | 32 | ~200 GB | 106% |
| 1024 | 1 | 1024 | ~6 GB | 95% |
ACTIVATION CHECKPOINTING: THE COMPUTE-FOR-MEMORY TRADE
Activation checkpointing (also called gradient checkpointing) trades compute for memory by not storing intermediate activations during the forward pass. Instead, only the inputs to each checkpointed segment are saved. During the backward pass, the forward pass is recomputed for each segment to regenerate the activations needed for gradient computation. Standard checkpointing saves activations at every transformer block boundary (every N layers), trading a 20-33 percent compute overhead for 60-85 percent activation memory reduction.
The memory-compute Pareto frontier shows three regimes. No checkpointing: maximum memory, minimum compute. Selective checkpointing (every 2-4 layers): 40-60 percent memory reduction, 10-15 percent compute overhead. Full checkpointing (every layer): 75-85 percent memory reduction, 25-33 percent compute overhead. The optimal point depends on the model size and sequence length: for 2K context on a 7B model, selective checkpointing is optimal; for 128K context on a 70B model, full checkpointing is required to fit within H100's 80 GB.
Megatron-LM implements transformer-block-level checkpointing, while PyTorch's torch.utils.checkpoint wraps arbitrary functions. The key implementation detail: checkpointing the attention computation separately from the MLP within each transformer block saves more memory (attention activations are O(n^2) in sequence length) at lower recompute cost than checkpointing the entire block.
| Checkpoint Strategy | 7B 2K ctx | 7B 128K ctx | 70B 2K ctx | 70B 128K ctx |
|---|---|---|---|---|
| None | 28 GB | 1.4 TB | 56 GB | 2.8 TB |
| Every 4 layers (selective) | 14 GB | 720 GB | 28 GB | 1.4 TB |
| Every 2 layers | 10 GB | 480 GB | 18 GB | 960 GB |
| Every layer (full) | 7 GB | 320 GB | 12 GB | 640 GB |
| Block recompute (attn+MLP separate) | 5 GB | 200 GB | 8 GB | 400 GB |
| Selective + CPU offload | 3 GB | 80 GB | 5 GB | 160 GB |
COMBINING TECHNIQUES: THE OPTIMAL MEMORY CONFIGURATION
No single technique suffices for frontier-scale training. The standard configuration for 70B training on 8 H100 GPUs combines all four: ZeRO-1 shards optimizer states across GPUs (saving 35 GB/GPU), gradient accumulation with micro-batch size 2 (saving 40 GB/GPU versus batch-16), selective activation checkpointing (saving 12 GB/GPU), and 8-bit Adam for optimizer states (saving 2 GB/GPU). Total per-GPU memory: weights (17.5 GB ZeRO-3 shared) + optimizer (1 GB 8-bit) + activations (8 GB checkpointed) = ~28 GB, fitting comfortably with headroom for KV cache.
The DeepSpeed configuration hierarchy is: ZeRO-3 as the foundation (sharding everything), gradient accumulation to control micro-batch memory, activation checkpointing to handle long sequences, and optimizer offload only when necessary. The priority ordering is critical: gradient accumulation should be used first (zero overhead), then check-pointing (20-33 percent overhead), then ZeRO stages (1-15 percent communication overhead), and CPU offload last (20-50 percent overhead, only for models that cannot fit otherwise).
MEASURED THROUGHPUT AT EACH MEMORY CONFIGURATION
Empirical throughput measurements on 8 H100 GPUs for a 70B model with sequence length 8K: baseline (batch size 32, no optimizations) is impossible - OOM at 200 GB per GPU. With gradient accumulation only (micro-batch 4, accum 8): 28 GB/GPU, 1.0x throughput baseline. + selective checkpointing: 16 GB/GPU, 0.87x throughput. + ZeRO-3: 10 GB/GPU, 0.82x throughput. + CPU optimizer offload: 4 GB/GPU, 0.55x throughput. The optimal config for throughput is gradient accumulation + selective checkpointing + ZeRO-1, achieving 92-96 percent of theoretical maximum throughput at 15-20 GB/GPU.
The diminishing returns curve is steep past 3-4 combined techniques. Each additional technique adds 1-5 percent overhead with only marginal memory savings. The practical rule: use ZeRO-3 + gradient accumulation (micro-batch 1-2) + selective checkpointing as the default for any model above 30B parameters. Only add CPU offload when the model is 70B+ and sequence length exceeds 32K.
B200: WHEN THE MEMORY CONSTRAINT DISAPPEARS
B200's 192 GB VRAM eliminates the need for most optimization techniques for models up to 70B. A 70B model with batch size 8 at 8K context consumes approximately: weights 140 GB + optimizer 80 GB (FP32 Adam, ZeRO-1 across 2 GPUs) + activations 40 GB = 260 GB. Even on B200, this requires ZeRO-1 or gradient accumulation. But at batch size 4 and 4K context, total memory drops to 150-170 GB, fitting comfortably on a single B200 with no checkpointing needed at full throughput.
For 7B models, the B200's headroom enables batch sizes of 64-128 at 8K context with no activation checkpointing and native FP32 Adam optimizer - essentially training at full GPU utilization without any memory optimization overhead. The throughput gain over H100 for 7B fine-tuning is 2.5-3.5x because H100 required checkpointing and gradient accumulation that reduced throughput by 30-40 percent. For larger models, B200 reduces the number of required optimization techniques from 4 to 1-2, simplifying training configurations and improving development velocity.
