The VRAM Wall in Mid-2026
The gap between model size and GPU VRAM continues to widen in mid-2026. The largest commercially available consumer GPU, the Nvidia RTX 6090, offers 48 GB of VRAM. The largest data center GPU, the B200, offers 192 GB. State-of-the-art open models like Llama 4 Behemoth (2T parameters, FP16) require approximately 3,700 GB of VRAM for a single copy of the weights and at least that much again for optimizer states and activations during training. Even with aggressive quantization to FP8, a 2T parameter model requires 1,850 GB of VRAM just for weights, making it impossible to fit on fewer than 10 B200s. Techniques to squeeze models into available VRAM are not academic optimizations - they are prerequisites for working with frontier-scale models on any cluster smaller than 1,024 GPUs.
The memory breakdown for a typical 70B-parameter model in BF16 on a single B200 shows where the VRAM goes: model weights consume 140 GB (1,024 * 70B * 2 bytes), optimizer states (Adam) consume 280 GB (2 states * 140 GB), gradients consume 140 GB, and activations consume 30 to 60 GB depending on sequence length and batch size. Total: 590 to 620 GB for a single replica, far exceeding the 192 GB available on a single B200. Distributed across 4 B200s with ZeRO-3, each GPU holds 148 to 155 GB, which fits with 37 to 44 GB to spare. The margin is thin, and every optimization technique matters.
Activation Checkpointing (Gradient Checkpointing)
Activation checkpointing, also called gradient checkpointing, trades compute for memory by not storing intermediate activations during the forward pass. Instead, only a subset of activations at checkpoints (typically every 4 to 12 transformer layers) are saved, and the discarded activations are recomputed on the fly during the backward pass. The standard implementation in PyTorch (torch.utils.checkpoint) supports both selectable and heuristic checkpointing strategies.
The VRAM savings are substantial. For a 70B Llama model with sequence length 8,192 and batch size 32, activation memory drops from 58 GB to 11 GB with checkpointing every 4 layers, a reduction of 81%. The tradeoff is a 20% to 30% training throughput penalty because the backward pass must recompute the activations. With heuristic checkpointing, where the checkpoint interval is automatically selected by a memory budget, the savings range from 55% to 75% with a 15% to 22% throughput penalty. The optimal checkpoint interval is model-specific: for Llama-architecture models, every 4 to 6 layers provides the best memory-compute tradeoff, while for deeper models like GPT-4-style architectures, every 8 to 12 layers is more efficient.
Selective checkpointing is an emerging refinement that checkpoints only specific operations with high activation memory footprints, such as attention softmax and layer normalization, while keeping cheap operations like linear projections uncheckpointed. The DeepSpeed implementation reduces the throughput penalty from 25% to 11% while maintaining 70% of the memory savings compared to full activation checkpointing. As of mid-2026, selective checkpointing is supported in PyTorch 2.8, DeepSpeed 0.18, and FSDP2, and is recommended over full checkpointing for any model larger than 30B parameters.
| Checkpoint Strategy | VRAM (70B, seq 8192, bs 32) | Savings vs Baseline | Throughput Penalty | Recommended For |
|---|---|---|---|---|
| No checkpointing | 58 GB | 0% | 0% | Sub-7B, batch size <8 |
| Full (every layer) | 6 GB | 90% | 30-35% | Oversubscribed VRAM |
| Every 4 layers | 11 GB | 81% | 20-25% | 70B+, recommended default |
| Every 8 layers | 18 GB | 69% | 15-18% | 30-70B, throughput priority |
| Selective (DeepSpeed) | 23 GB | 60% | 11-14% | Best balance for 70B+ |
| Heuristic (auto) | 19 GB | 67% | 15-22% | Uncertain memory requirements |
Mixed Precision: FP8, FP16, and FP32 Tradeoffs
Mixed precision training reduces VRAM usage by storing weights and activations in lower-precision formats while maintaining a master copy of weights in higher precision for numeric stability. The standard regimes in mid-2026 are FP16 (2 bytes), BF16 (2 bytes), FP8 (1 byte, with E4M3 and E5M2 variants), and FP4 (0.5 bytes, experimental). Each step down the precision ladder approximately halves VRAM requirements for weights but introduces accuracy risks that must be validated per model.
FP16 and BF16 are the safe defaults. Both use 2 bytes per parameter, cutting weight memory in half compared to FP32. BF16 has the same exponent range as FP32 (8 bits vs 5 bits for FP16) and is the preferred format for training because it eliminates the overflow issues that FP16 encounters during gradient accumulation. For a 70B model, BF16 weights consume 140 GB versus 280 GB in FP32. The optimizer states (Adam) remain in FP32, so the total VRAM for weights + optimizer + gradients is 140 GB (FP16 weights) + 280 GB (Adam FP32 states) + 140 GB (FP16 gradients) = 560 GB, compared to 840 GB with full FP32 training.
FP8 training (H100 and B200 native) reduces weight storage by another 50% relative to BF16. A 70B model in FP8 consumes 70 GB for weights, 280 GB for optimizer states (still FP32), and 70 GB for gradients (FP8), totaling 420 GB. The throughput gain from FP8 tensor cores is 1.6x to 2x compared to BF16 on H200 and B200 GPUs. However, FP8 training introduces quantization noise that degrades model quality by 0.2% to 1.5% on downstream benchmarks depending on the model architecture and training stability. The recommendation in mid-2026 is to use FP8 for pre-training runs where the quality gap can be compensated with additional tokens, and to use BF16 for fine-tuning where every quality point matters. FP4 is not production-ready outside of Nvidia's internal testing pipeline and should not be used for production training as of mid-2026.
| Precision | Bytes/Param | 70B Weight VRAM | Total VRAM (70B, 4-GPU) | Throughput vs BF16 | Quality Impact |
|---|---|---|---|---|---|
| FP32 | 4 B | 280 GB | N/A (does not fit) | 0.5x | Baseline |
| BF16 | 2 B | 140 GB | 155 GB/GPU (4x) | 1.0x | Baseline (no loss) |
| FP16 | 2 B | 140 GB | 155 GB/GPU (4x) | 1.1x | Overflow risk (gradients) |
| FP8 (E4M3) | 1 B | 70 GB | 105 GB/GPU (4x) | 1.8x | -0.5 to -1.5% accuracy |
| FP8+FP16 master | 1/2 B | 70 GB | 140 GB/GPU (4x) | 1.5x | -0.2 to -0.8% accuracy |
| FP4 (experimental) | 0.5 B | 35 GB | 87 GB/GPU (4x) | 2.5x | -2 to -5% accuracy |
ZeRO Optimization Stages 1, 2, and 3
The ZeRO (Zero Redundancy Optimizer) family of optimizations eliminates memory redundancy across data-parallel GPUs by partitioning optimizer states, gradients, and model parameters across the data-parallel group. ZeRO-1 partitions only optimizer states. ZeRO-2 partitions optimizer states and gradients. ZeRO-3 partitions optimizer states, gradients, and model parameters. Each stage reduces per-GPU memory at the cost of increased communication volume.
ZeRO-1 reduces per-GPU VRAM by roughly the optimizer state fraction divided by the data-parallel degree. For a 70B model on 8 GPUs with Adam optimizer, ZeRO-1 reduces optimizer state memory from 280 GB (total across all GPUs) to 35 GB per GPU (280/8), saving 35 GB per GPU compared to no partitioning. ZeRO-2 additionally partitions gradients, saving another 17.5 GB per GPU (140/8). ZeRO-3 partitions everything, including model weights, bringing weight memory from 140 GB to 17.5 GB per GPU (140/8). The cumulative per-GPU savings from ZeRO-3 on 8 GPUs is approximately 110 GB compared to DDP with full replication.
The communication overhead of ZeRO scales with the degree of partitioning. ZeRO-1 adds negligible communication because optimizer state updates are infrequent. ZeRO-2 adds a gradient all-gather at each step, increasing communication volume by 30% to 50% compared to DDP. ZeRO-3 adds a parameter all-gather before the forward and backward passes, which increases communication volume by 200% to 300% and introduces a latency bubble at the beginning of each training step. On an 8-GPU node with NVLink, the ZeRO-3 overhead is 10% to 15%. Across nodes with InfiniBand, it is 25% to 40%. The recommendation as of mid-2026 is to use ZeRO-2 for single-node training and ZeRO-3 for multi-node training where model size exceeds available VRAM. ZeRO-1 is rarely used alone because ZeRO-2's additional savings cost very little extra communication.
| ZeRO Stage | Partitioned | 70B VRAM/GPU (8 GPUs) | Comm Overhead vs DDP | Best Used When |
|---|---|---|---|---|
| DDP (no ZeRO) | None | ~155 GB | 0% | Model fits without partitioning |
| ZeRO-1 | Optimizer states | ~138 GB | <5% | Training small models |
| ZeRO-2 | Opt states + gradients | ~120 GB | 30-50% | Single-node, up to H100 |
| ZeRO-3 | Opt states + grads + params | ~19 GB | 200-300% | Multi-node, models >70B |
| FSDP2 (ZeRO-3 variant) | Same as ZeRO-3 | ~19 GB | 150-250% | PyTorch native, better perf |
CPU and NVMe Offloading
When even ZeRO-3 is insufficient to fit a model into available GPU VRAM, the next recourse is offloading model states to CPU memory or NVMe storage. The DeepSpeed Offload library supports offloading optimizer states and gradients to CPU memory (DeepSpeed-CPU), and offloading all model states including parameters to NVMe (DeepSpeed-NVMe). The tradeoff is compute throughput for memory capacity.
CPU offloading moves optimizer states from GPU VRAM to CPU DRAM, which is typically 256 GB to 1,024 GB on a GPU server. For a 70B model with Adam optimizer, offloading optimizer states (280 GB) to CPU frees 140 GB of GPU VRAM (the 2x gradient copy stays on GPU). The throughput penalty for CPU offloading is 15% to 25% because the optimizer step now includes CPU-GPU transfers over PCIe Gen 5 (approximately 32 GB/s per direction for H200). On B200 with PCIe Gen 5 x16 (64 GB/s theoretical), the transfer of 280 GB adds roughly 4.4 seconds per step, compared to a typical 2-second training step, making CPU offloading viable only for the optimizer step, not for every forward-backward pass.
NVMe offloading is the nuclear option, moving all model states to NVMe SSDs through DeepSpeed-NVMe or the newer torch.distributed.checkpoint offload API. The effective capacity becomes the size of the NVMe disk (3.2 TB to 30 TB per node), but the bandwidth bottleneck is severe: PCIe Gen 5 x4 to NVMe provides approximately 7 GB/s throughput, compared to 3,300 GB/s for GPU VRAM bandwidth. Training throughput with full NVMe offloading drops to 2% to 5% of baseline, which means what would be a 2-day training run in VRAM takes 40 to 100 days with NVMe offloading. The practical use case for NVMe offloading is not training but batch inference or model evaluation where latency is not critical, or fine-tuning very small subsets of weights (LoRA) on a model that is otherwise offloaded.
| Offload Target | Freed VRAM (70B) | Throughput Impact | Per-Step Time (70B, 8-GPU) | Use Case |
|---|---|---|---|---|
| CPU (optimizer only) | ~70 GB | -15% to -25% | 2.3 sec (vs 2.0) | Training, moderate penalty |
| CPU (optimizer + gradients) | ~105 GB | -30% to -40% | 2.8 sec (vs 2.0) | Tight VRAM training |
| NVMe (full offload) | ~155 GB | -95% to -98% | 40-100 sec (vs 2.0) | Inference, eval, LoRA |
| NVMe + CPU (hybrid) | ~130 GB | -60% to -80% | 6-10 sec (vs 2.0) | Fine-tuning with high patience |
Compute-in-Memory and Emerging Techniques
Compute-in-memory is an emerging technique where model operations are fused into in-memory operations, avoiding the bandwidth cost of moving activations between GPU VRAM and compute units. Nvidia's Tensor Memory Accelerator (TMA) in the Blackwell architecture (GB200, B200) implements hardware-level support for asynchronous data movement that overlaps memory transfers with computation, reducing the effective memory footprint of intermediate activations by 15% to 30% depending on the model architecture. The TMA unit handles tensor copy operations independently of the CUDA cores and tensor cores, freeing compute resources for kernel execution while data is moved in the background.
Memory-efficient attention variants like Flash Attention 3.2 (integrated into PyTorch 2.8) and vLLM's PagedAttention further reduce activation memory by not materializing the full attention score matrix. Flash Attention 3.2 tiles the attention computation into blocks that fit in on-chip SRAM, avoiding large intermediate writes to HBM. For a 70B model with seq length 32,768, Flash Attention reduces activation memory from 42 GB to 8 GB, a reduction of 81%. Combined with activation checkpointing and mixed precision, an 8-GPU B200 cluster can train a 70B model with global batch size 64 at 82% MFU, which was impossible with the same hardware using the standard PyTorch attention implementation in 2024.
The practical takeaway for infrastructure teams in mid-2026 is that these techniques compose multiplicatively, not additively. A single B200 serving a 70B model cannot run it with full BF16 weights. But BF16 weights + ZeRO-3 across 4 GPUs + activation checkpointing every 4 layers + Flash Attention 3.2 + FP8 gradient communication reduces per-GPU VRAM from an impossible 155 GB to 43 GB, well within the 192 GB available on B200. The throughput is 75% of what a fully replicated 8-GPU setup would achieve, but the cost is half the GPUs. For inference, vLLM with FP8 quantization and PagedAttention can serve a 70B model on a single B200 with 8,192-token context at 1,200 tokens per second, making the most cost-effective inference configuration not a cluster at all but a single high-memory GPU with the right optimization stack.
| Technique Stack | Per-GPU VRAM (70B) | GPUs Required | Relative Throughput | GPU Cost/Month |
|---|---|---|---|---|
| Full BF16, no optimizations | 155 GB | N/A (does not fit on 8) | N/A | N/A |
| BF16 + ZeRO-3 (8 GPUs) | 19 GB | 8 | 1.0x | $49,000 |
| BF16 + ZeRO-3 + act ckpt (8 GPUs) | 11 GB | 8 | 0.82x | $49,000 |
| FP8 + ZeRO-3 + act ckpt (4 GPUs) | 43 GB | 4 | 0.75x (of 8-GPU) | $24,500 |
| FP8 + ZeRO-3 + act ckpt + FA3 (4 GPUs) | 35 GB | 4 | 0.80x (of 8-GPU) | $24,500 |
| FP8 + vLLM (inference, 1 GPU) | 70 GB | 1 | N/A (inference) | $6,100 |
Decision Framework for Infrastructure Teams
The choice of memory optimization technique depends on the workload type, the model architecture, and the available GPU configuration. For training models that already fit on the available GPU count (e.g., 70B on 4 B200s with ZeRO-3 and activation checkpointing), the priority is to maximize throughput by minimizing optimization overhead. Use selective activation checkpointing, BF16 precision, Flash Attention 3.2, and ZeRO-3 with the DeepSpeed communication overlap optimization. Avoid CPU offloading, NVMe offloading, and FP8 quantization unless VRAM headroom is below 10%. For training models that are one generation ahead of what the GPU count can comfortably fit, use FP8 mixed precision, aggressive activation checkpointing, and ZeRO-3. Accept a 15% to 25% throughput reduction versus the ideal configuration.
For inference, the optimization goal is different: maximize the number of concurrent request slots per GPU while maintaining latency targets. vLLM with PagedAttention and FP8 quantization is the baseline. Adding MIG partitioning for models under 30B parameters increases concurrent serving capacity by 2x to 4x per GPU. Using speculative decoding reduces per-request compute by 20% to 35%, increasing throughput proportionally. For models that exceed a single GPU's VRAM, use tensor parallelism with ZeRO-3 inference across 2 to 8 GPUs, accepting that tensor-parallel inference across GPUs adds 5% to 15% latency overhead depending on the inter-GPU bandwidth.
