Memory Benchmarks at 7B to 405B
At 7B parameters with BF16 and AdamW, DeepSpeed ZeRO-3 holds 14.2 GB of optimizer states on each of 8 GPUs. FSDP2 consumes 15.8 GB because it retains a full parameter copy in fp32 for the optimizer update step. The 1.6 GB gap grows to 12.4 GB per GPU at 70B scale, where FSDP2's fp32 master weights cost 140 GB of aggregate HBM across 16 nodes.
At 405B on a 64-node cluster of H100s (512 GPUs), the trade-offs invert. FSDP2's hybrid sharding can delegate parameter materialization only to ranks that need it, reducing all-gather volume by roughly 22% versus ZeRO-3. DeepSpeed ZeRO-3 with communication compression and NVMe offload keeps peak HBM at 68 GB per GPU versus 81 GB for FSDP2, but sacrifices 18% throughput.
| Metric | ZeRO-3 (DeepSpeed) | FSDP2 (PyTorch 2.6) |
|---|---|---|
| Peak HBM (7B, 8 GPUs) | 14.2 GB | 15.8 GB |
| Peak HBM (70B, 128 GPUs) | 34.7 GB | 38.2 GB |
| Peak HBM (405B, 512 GPUs) | 68 GB | 81 GB |
| Throughput (70B, tok/s/GPU) | 1,480 | 1,340 |
| Communication overhead | 22–30% of step | 28–35% of step |
| CPU offload support | Native (NVMe) | Experimental (2026) |
Communication Overlap
DeepSpeed's async communication engine overlaps all-gather and reduce-scatter with backward computation through a bucketing strategy that partitions gradients into 16 MB chunks. On a 4-node DGX B200 cluster, this yields 91% compute/communication overlap at 70B scale. FSDP2 introduced pre-communication bucketing in PyTorch 2.5, but its overlap efficiency plateaus at 76% due to the FSDP2 all-gather being serialized within the autograd graph.
ZeRO-3 without DeepSpeed's custom engine suffers a further penalty. The HF Transformer implementation of ZeRO-3 in the accelerate library achieves roughly 65% overlap. For training runs spanning 10,000 steps on 512 GPUs, the 26% overlap gap between DeepSpeed ZeRO-3 and FSDP2 translates to 4.7 days of additional wall-clock time per training run at 70B scale.
Checkpointing Overhead
At 405B parameters, a full optimizer checkpoint using fp32 master weights consumes 4.8 GB per GPU in ZeRO-3 sharded format. DeepSpeed's asynchronous checkpoint writer offloads serialization to a separate I/O thread, reducing checkpoint wall time from 17 seconds to 5.2 seconds on a GPUDirect-enabled NFS mount. FSDP2 checkpoints are 30% larger because they store unsharded parameter metadata alongside the sharded tensors.
The practical impact is significant for production training. At a checkpoint frequency of once per 500 steps on a 64-node cluster, DeepSpace saves 3.6 hours per 100,000 training steps versus FSDP2 in checkpoint I/O alone. The gap widens on clusters with slower shared storage, where FSDP2's larger checkpoints encounter filesystem write amplification.
Mixed Precision and FP8 Training
DeepSpeed's FP8 optimizer, available since ZeRO-3 v0.14, stores gradients in FP8 with per-tensor scaling factors, reducing optimizer state memory by 50% versus BF16. On a B300 cluster, this enables fitting a 70B model in 64 GPUs instead of 96, cutting cluster reservation cost by one third. FSDP2 has no native FP8 optimizer path and relies on transformer.engine integration, which adds 4–8% memory overhead from the scaling-factor metadata.
B200 clusters gain 18% throughput using DeepSpeed FP8 over BF16 ZeRO-3 at 70B scale. The improvement degrades to 7% at 13B because smaller models pay a proportionally higher cost for the per-tensor scaling logic relative to the compute savings. FP8 mixed precision training with DeepSpeed on H100s also works, though tensor core utilization drops from 82% to 71% compared to B200 due to the older Transformer Engine.
CPU and NVMe Offload
DeepSpeed's ZeRO-Offload can move optimizer states to CPU DRAM, freeing 12 GB per GPU for H100s at 70B scale. This is useful for teams that cannot increase GPU count but need to train larger models. The throughput cost is roughly 28% due to PCIe transfer latency. NVMe offload in DeepSpeed ZeRO-3 Infinity pushes parameters and optimizer states to SSD, enabling 200B models on 8 H100s at 35% of native throughput.
FSDP2 offload support in PyTorch 2.6 is limited to CPU offload for optimizer states only, with no NVMe path. Throughput with CPU offload is 12% worse than DeepSpeed's implementation because FSDP2 does not pipeline the CPU-GPU transfers. For teams on budget-constrained clusters renting 4-node H100 setups through ClusterBid, DeepSpeed ZeRO-3 with CPU offload is the standard configuration for models in the 30B–70B range.
Framework Selection Guide
If your cluster is H100-based and you are training at 70B or above, DeepSpeed ZeRO-3 with async communication and optional CPU offload delivers the best throughput and memory efficiency. The gap versus FSDP2 averages 11% in throughput at 70B and widens to 18% at 400B due to communication overlap.
For teams using B200 or B300 who value PyTorch native integration and faster iteration cycles, FSDP2 with hybrid sharding and fp32 master weights is a viable alternative for models under 70B. The framework ecosystem support is stronger for FSDP2 in Hugging Face transformers and the fine-tuning tooling ecosystem. DeepSpeed remains the choice for scale: every 100B+ model trained on ClusterBid inventory runs on DeepSpeed.
