THE OPTIMIZER MEMORY TAX
AdamW stores two FP32 values per parameter: the first moment (momentum) and second moment (variance). For a 70B model, this is 70 billion x 2 x 4 bytes = 560 GB for optimizer states alone - more than the 140 GB for the FP16 weights themselves. This 4:1 optimizer-to-weight memory ratio means the optimizer state, not the model weights, is the primary memory constraint in distributed training. On an 8-GPU H100 node, 560 GB of optimizer states requires 70 GB per GPU before any weights, activations, or gradients are accounted for.
Every optimizer variant reduces this tax differently: 8-bit Adam compresses each optimizer state to 1 byte (87.5 percent reduction, 70 GB total for 70B), Sophia eliminates the second moment by using a Hutchinson trace estimator of the Hessian (50 percent reduction, 280 GB total), LOMO eliminates optimizer states entirely by performing SGD-style updates during the backward pass (100 percent reduction, 0 GB), and GaLore reduces optimizer memory by projecting gradients to a low-rank subspace before applying Adam.
| Optimizer | State per Param | 70B Opt Memory | 7B Opt Memory | Algorithm | Memory Redux |
|---|---|---|---|---|---|
| AdamW | 8 bytes (2x FP32) | 560 GB | 56 GB | Momentum + variance | Baseline |
| AdamW (FP32 weights) | 12 bytes | 840 GB | 84 GB | FP32 weights + states | -50% |
| 8-bit Adam | 2 bytes (2x INT8) | 70 GB | 7 GB | Block-wise quantized | 87.5% |
| Sophia | 4 bytes (FP32 1st) | 280 GB | 28 GB | Hessian trace estimate | 50% |
| Sophia (8-bit) | 1 byte (1st INT8) | 35 GB | 3.5 GB | Quantized Sophia | 93.75% |
| LOMO | 0 bytes | 0 GB | 0 GB | SGD in backward pass | 100% |
| GaLore (r=256) | <1 byte avg | 8-16 GB | 1-2 GB | Low-rank gradient proj | 97-98% |
8-BIT ADAM: THE PRACTICAL DEFAULT FOR MOST TRAINING
8-bit Adam (bitsandbytes) uses block-wise quantization: each optimizer state is divided into 2048-element blocks, each block is quantized independently to INT8 with dynamic range per block, and dequantized on-the-fly during the optimizer step. The quantization error averages 0.1-0.5 percent of the full-precision value and does not accumulate across steps because each step re-quantizes from the updated FP32 values. The 87.5 percent memory reduction is effectively free: convergence quality and training speed are indistinguishable from full AdamW across all tested model sizes.
The throughput impact is 0-5 percent depending on GPU architecture. On H100 with FP8 tensor cores, the quantization overhead is absorbed by the mixed-precision pipeline with <1 percent throughput loss. On A100, the CPU-based quantization in bitsandbytes adds 3-5 percent overhead per optimizer step. For a 70B model with gradient accumulation steps = 256, the optimizer step runs every 256 micro-batches, so even 5 percent overhead on the optimizer step translates to <0.02 percent total overhead. The recommended default: 8-bit Adam for any model above 1B parameters.
SOPHIA: SECOND-ORDER OPTIMIZATION WITH GPU TRADEOFFS
Sophia replaces the variance term in Adam with a Hessian trace estimate computed via Hutchinson's method, which requires one additional forward-backward pass every N steps (typically every 10-20 steps). The Hessian-backward pass doubles the per-step compute for those steps but provides a better preconditioner that reduces total steps to convergence by 30-50 percent. For a 70B model training for 200K steps, Sophia-H completes in 120K-140K steps versus 200K for AdamW, saving 60K-80K GPU-hours at the cost of extra computation on 10K of those steps.
The GPU infrastructure requirement is the doubled backward pass on Hessian-estimation steps, which requires 2x activation memory for those steps. On H100 with 80 GB, this constraint forces a 50 percent reduction in micro-batch size or sequence length on Hessian-estimation steps, or the use of activation checkpointing to free memory. In practice, Sophia-H achieves 380-420 teraFLOP/s per H100 versus 440-460 for AdamW, but requires 30-50 percent fewer total FLOPs to target loss - a net win of 10-20 percent in wall-clock time.
The memory savings from removing the second moment are valuable for large models. For a 405B model, Sophia uses 810 GB for optimizer states (FP32 first moment only) versus 1,620 GB for AdamW (both moments). Combined with ZeRO-3, the per-GPU optimizer memory drops from 20 GB to 10 GB for a 64-GPU cluster.
| Metric | AdamW | 8-bit Adam | Sophia-H | Sophia-H (8-bit) | LOMO |
|---|---|---|---|---|---|
| Steps to target loss (7B) | 100K | 100K | 60-70K | 60-70K | 100-120K |
| Peak TFLOPS/s per H100 | 460 | 460 | 380-420 | 380-420 | 480-500 |
| 70B Opt Memory | 560 GB | 70 GB | 280 GB | 35 GB | 0 GB |
| Wall-Clock Speedup | Baseline | 1.0x | 1.1-1.2x | 1.1-1.2x | 0.7-0.85x |
| Final Quality (MMLU) | 72.0% | 71.8% | 72.5% | 72.3% | 69.5-71.0% |
| Complexity | Trivial | Low | Medium | Medium | High |
LOMO AND GALORE: ZERO-OPTIMIZER-STATE APPROACHES
LOMO (Low-Memory Optimizer) fuses the gradient computation with the parameter update: as each parameter's gradient is computed during backpropagation, LOMO immediately applies the SGD-style update without storing the gradient tensor. This eliminates the need for both optimizer states and gradient storage, reducing memory to just the model weights and activations. For a 7B model, LOMO trains on a single 24 GB GPU (RTX 4090) that cannot even load AdamW's optimizer states. The tradeoff is convergence quality: LOMO's SGD-like updates lack adaptive learning rates and moment estimation, requiring 10-20 percent more steps to reach equivalent loss on most benchmarks.
GaLore (Gradient Low-Rank Projection) projects the gradient matrix into a low-rank subspace (typically rank 256-1024) before applying AdamW. The optimizer states operate on the low-rank projection rather than the full gradient. For a 70B model, a linear layer with weight shape [4096, 4096] has a 4096x4096 = 16.8M gradient element. At rank 256, the projected gradient is 256+256 = 512 dimensions, a 32,768x reduction. Combined with the optimizer state savings, GaLore reduces total memory for the 70B model from 700 GB (weights + AdamW states) to 180 GB (weights + low-rank states), enabling training on 4-8 H100 GPUs versus 16-24 for full AdamW. Quality is within 0.5-1.0 percent of full-rank training at rank 1024.
| Configuration | Total Memory 7B | Total Memory 70B | Min GPUs 70B | Quality vs AdamW | Training Time |
|---|---|---|---|---|---|
| AdamW FP32 | 120 GB | 1.2 TB | 16-24 H100 | 100% | 1.0x |
| 8-bit Adam | 28 GB | 280 GB | 4-8 H100 | 99.8% | 1.0x |
| Sophia-H | 56 GB | 560 GB | 8-16 H100 | 100.5% | 0.85x |
| GaLore r=256 | 18 GB | 180 GB | 2-4 H100 | 98.5-99% | 1.05-1.1x |
| GaLore r=1024 | 32 GB | 320 GB | 4-8 H100 | 99-99.5% | 1.02-1.05x |
| LOMO | 14 GB | 140 GB | 2-4 H100 | 96-98% | 0.7-0.85x |
DISTRIBUTED OPTIMIZER PATTERNS: COMMUNICATION OVERHEAD
In distributed training, the optimizer interacts with parallelism in three ways: FSDP/ZerO-3 requires an all-gather before each forward pass and a reduce-scatter during backward, both independent of optimizer choice. ZeRO-1 (optimizer state sharding) is optimizer-dependent: AdamW shards 8 bytes per parameter, 8-bit Adam shards 2 bytes, Sophia shards 4 bytes, and GaLore shards just the low-rank components (bytes per parameter = projection_dim / original_dim * 8). Lower per-parameter communication reduces the all-reduce volume during optimizer steps.
For a 70B model on 64 GPUs with ZeRO-1, the optimizer all-reduce volume per step is: AdamW 560 GB / 64 = 8.75 GB per GPU, 8-bit Adam 70 GB / 64 = 1.1 GB per GPU, LOMO 0 GB / 64 = 0 GB. For LOMO, the elimination of the optimizer all-reduce saves 1-2 ms per training step in communication overhead. At 50K steps, this saves 50-100 seconds of total training time - modest, but additive with other savings. GaLore's communication advantage is larger: the low-rank projection reduces gradient communication by 100-1000x depending on rank, which is particularly valuable for cross-node training where network bandwidth (typically 25-200 Gbps) is the bottleneck.
B200: MAKING ADAMW AFFORDABLE AGAIN
B200's 192 GB VRAM eliminates the need for most optimizer memory optimizations for models up to 13B. A 13B model with full AdamW takes: weights 26 GB + optimizer 104 GB + activations at batch size 4 with 4K context 24 GB = 154 GB, fitting on a single B200. On H100 (80 GB), this configuration requires 8-bit Adam or ZeRO-2 across 4 GPUs. The simplification for 13B-and-below training means teams can use standard AdamW with zero configuration complexity, reducing training setup time from days to hours.
For 70B models, B200 reduces the number of GPUs required for AdamW from 16-24 H100 to 8-12 B200, because each GPU holds 192 GB versus 80 GB. Combined with ZeRO-1 (not ZeRO-3), per-GPU memory for 70B AdamW becomes: weights (140 GB / 12 GPUs = 11.7 GB) + optimizer (560 GB / 12 = 46.7 GB) + activations (40 GB / 12 = 3.3 GB) = 61.7 GB per GPU on B200 - fitting comfortably without ZeRO-3's communication overhead. The recommended B200 configuration for 70B: ZeRO-1 + standard AdamW + selective checkpointing = 95-98 percent of peak FLOPs utilization.
