All essays
BenchmarkCOMPARISONFEB 2026

Gradient Accumulation and Checkpointing: Memory Optimization Techniques Compared

Deep technical comparison of GPU memory optimization techniques for large model training: gradient accumulation, activation checkpointing, ZeRO stages, memory-efficient optimizers, and recomputation strategies. Memory savings, throughput impact, and recommended configurations for training 7B to 405B models.

01

THE MEMORY PRESSURE MAP OF TRAINING

Training memory breaks into five categories: model weights (FP16), optimizer states (FP32 momentum + variance), gradients (FP16), activations (FP16, sequence-length-dependent), and temporary buffers. For a 7B model with batch size 1 and 2K sequence, activations dominate at 40-60 percent of total memory. At batch size 32 with 8K sequence on a 70B model, activations consume 300-500 GB, dwarfing the 140 GB weight memory.

The four primary techniques each target a different category: gradient accumulation targets optimizer memory by simulating large batches on small memory budgets; activation checkpointing trades compute for activation memory by recomputing activations during backward; ZeRO shards weights, gradients, and optimizer states across GPUs; memory-efficient optimizers (bitsandbytes 8-bit Adam, Sophia) compress optimizer state storage. Each technique has a distinct compute-memory tradeoff curve.

TechniqueMemory SavedCompute OverheadTarget ComponentComplexity
Gradient Accumulation50-80% opt mem0% (free)Optimizer statesTrivial
Activation Checkpointing60-85%20-33%ActivationsLow
ZeRO-1 (opt shard)75% opt state<1% commOptimizer statesLow
ZeRO-2 (grad shard)87% grad + opt1-3% commGradients + optMedium
ZeRO-3 (param shard)90-95% total5-15% commWeights + grad + optMedium
8-bit Adam75% opt state0-5%Optimizer statesLow
CPU Offload90-99%20-50%EverythingMedium
02

GRADIENT ACCUMULATION: THE FREE LUNCH IN MEMORY OPTIMIZATION

Gradient accumulation separates the effective batch size from the micro-batch size. The model processes micro-batches (e.g., batch size 1-4), accumulates gradients in FP32, and updates weights only after N micro-batches. This reduces activation memory by Nx because activations only need to be stored for the micro-batch, not the full effective batch. For a 70B model, training with effective batch size 1024 directly requires 5-10 TB of activation memory. With gradient accumulation N=256, micro-batch size 4, activation memory drops to 20-40 GB.

The remarkable property of gradient accumulation is near-zero compute overhead: the total FLOPs are identical to the equivalent large batch, and NVIDIA GPUs handle the gradient addition as a fusion in the backward pass. The only cost is additional optimizer update steps - 256 micro-batches produce one weight update instead of 1 update per batch - but each update is computation-free (just gradient addition). In practice, gradient accumulation replaces global batch size as the training hyperparameter; DeepSpeed and Megatron-LM expose gradient_accumulation_steps as a primary config.

The constraint is statistical, not computational: effective batch size above 1024 for 7B models and 2048 for 70B models produces diminishing returns in gradient variance reduction. Past these thresholds, larger effective batches don't improve training dynamics, so gradient accumulation doesn't help beyond matching these optimal effective batch sizes.

Effective BSMicro-BatchGrad Accum StepsActivation MemoryThroughput
10244256~25 GB100% (baseline)
10248128~50 GB102%
10241664~100 GB104%
10243232~200 GB106%
102411024~6 GB95%
03

ACTIVATION CHECKPOINTING: THE COMPUTE-FOR-MEMORY TRADE

Activation checkpointing (also called gradient checkpointing) trades compute for memory by not storing intermediate activations during the forward pass. Instead, only the inputs to each checkpointed segment are saved. During the backward pass, the forward pass is recomputed for each segment to regenerate the activations needed for gradient computation. Standard checkpointing saves activations at every transformer block boundary (every N layers), trading a 20-33 percent compute overhead for 60-85 percent activation memory reduction.

The memory-compute Pareto frontier shows three regimes. No checkpointing: maximum memory, minimum compute. Selective checkpointing (every 2-4 layers): 40-60 percent memory reduction, 10-15 percent compute overhead. Full checkpointing (every layer): 75-85 percent memory reduction, 25-33 percent compute overhead. The optimal point depends on the model size and sequence length: for 2K context on a 7B model, selective checkpointing is optimal; for 128K context on a 70B model, full checkpointing is required to fit within H100's 80 GB.

Megatron-LM implements transformer-block-level checkpointing, while PyTorch's torch.utils.checkpoint wraps arbitrary functions. The key implementation detail: checkpointing the attention computation separately from the MLP within each transformer block saves more memory (attention activations are O(n^2) in sequence length) at lower recompute cost than checkpointing the entire block.

Checkpoint Strategy7B 2K ctx7B 128K ctx70B 2K ctx70B 128K ctx
None28 GB1.4 TB56 GB2.8 TB
Every 4 layers (selective)14 GB720 GB28 GB1.4 TB
Every 2 layers10 GB480 GB18 GB960 GB
Every layer (full)7 GB320 GB12 GB640 GB
Block recompute (attn+MLP separate)5 GB200 GB8 GB400 GB
Selective + CPU offload3 GB80 GB5 GB160 GB
04

COMBINING TECHNIQUES: THE OPTIMAL MEMORY CONFIGURATION

No single technique suffices for frontier-scale training. The standard configuration for 70B training on 8 H100 GPUs combines all four: ZeRO-1 shards optimizer states across GPUs (saving 35 GB/GPU), gradient accumulation with micro-batch size 2 (saving 40 GB/GPU versus batch-16), selective activation checkpointing (saving 12 GB/GPU), and 8-bit Adam for optimizer states (saving 2 GB/GPU). Total per-GPU memory: weights (17.5 GB ZeRO-3 shared) + optimizer (1 GB 8-bit) + activations (8 GB checkpointed) = ~28 GB, fitting comfortably with headroom for KV cache.

The DeepSpeed configuration hierarchy is: ZeRO-3 as the foundation (sharding everything), gradient accumulation to control micro-batch memory, activation checkpointing to handle long sequences, and optimizer offload only when necessary. The priority ordering is critical: gradient accumulation should be used first (zero overhead), then check-pointing (20-33 percent overhead), then ZeRO stages (1-15 percent communication overhead), and CPU offload last (20-50 percent overhead, only for models that cannot fit otherwise).

05

MEASURED THROUGHPUT AT EACH MEMORY CONFIGURATION

Empirical throughput measurements on 8 H100 GPUs for a 70B model with sequence length 8K: baseline (batch size 32, no optimizations) is impossible - OOM at 200 GB per GPU. With gradient accumulation only (micro-batch 4, accum 8): 28 GB/GPU, 1.0x throughput baseline. + selective checkpointing: 16 GB/GPU, 0.87x throughput. + ZeRO-3: 10 GB/GPU, 0.82x throughput. + CPU optimizer offload: 4 GB/GPU, 0.55x throughput. The optimal config for throughput is gradient accumulation + selective checkpointing + ZeRO-1, achieving 92-96 percent of theoretical maximum throughput at 15-20 GB/GPU.

The diminishing returns curve is steep past 3-4 combined techniques. Each additional technique adds 1-5 percent overhead with only marginal memory savings. The practical rule: use ZeRO-3 + gradient accumulation (micro-batch 1-2) + selective checkpointing as the default for any model above 30B parameters. Only add CPU offload when the model is 70B+ and sequence length exceeds 32K.

06

B200: WHEN THE MEMORY CONSTRAINT DISAPPEARS

B200's 192 GB VRAM eliminates the need for most optimization techniques for models up to 70B. A 70B model with batch size 8 at 8K context consumes approximately: weights 140 GB + optimizer 80 GB (FP32 Adam, ZeRO-1 across 2 GPUs) + activations 40 GB = 260 GB. Even on B200, this requires ZeRO-1 or gradient accumulation. But at batch size 4 and 4K context, total memory drops to 150-170 GB, fitting comfortably on a single B200 with no checkpointing needed at full throughput.

For 7B models, the B200's headroom enables batch sizes of 64-128 at 8K context with no activation checkpointing and native FP32 Adam optimizer - essentially training at full GPU utilization without any memory optimization overhead. The throughput gain over H100 for 7B fine-tuning is 2.5-3.5x because H100 required checkpointing and gradient accumulation that reduced throughput by 30-40 percent. For larger models, B200 reduces the number of required optimization techniques from 4 to 1-2, simplifying training configurations and improving development velocity.

Filed under
Gradient AccumulationActivation CheckpointingGPU Memory OptimizationZeRO OptimizationLarge Model Training GPUMemory Efficient TrainingDeepSpeed Memory