THROUGHPUT AND MEMORY EFFICIENCY BENCHMARKS
Benchmarking Llama 70B training on 8x H100 SXM with NVLink (1 node, FSDP data parallel) shows DeepSpeed ZeRO-3 achieving 2,840 tokens/second per GPU with 72 GB peak memory, FSDP2 achieving 2,910 tokens/second with 68 GB peak memory, and FairScale achieving 2,100 tokens/second with 74 GB peak memory. FSDP2's 2.5% throughput advantage over DeepSpeed comes from tighter integration with `torch.compile`'s inductor backend, which fuses communication ops with computation more efficiently. FairScale's 26% throughput deficit reflects its older ZeRO implementation without overlapping communication, effectively serializing the all-gather and compute steps. At 64x H100 (8 nodes with InfiniBand), DeepSpeed ZeRO-3 with gradient partitioning achieves 22,400 tokens/second aggregate, FSDP2 achieves 22,100 tokens/second, and FairScale achieves 14,300 tokens/second.
Memory efficiency benchmarks show ZeRO-3 and FSDP2 converging. For Llama 70B with Adam optimizer, theoretical per-GPU memory is 35 GB for parameters, gradients, and optimizer states at full sharding on 8 GPUs, plus approximately 12 GB for activations at batch size 1 per GPU with 4K sequence length. DeepSpeed's activation checkpointing (`deepspeed.ops.activation_checkpointing`) uses 8 GB for checkpoints versus FSDP2's activation offloading which uses 6 GB. The total per-GPU memory: DeepSpeed 55 GB, FSDP2 53 GB, FairScale 62 GB. The 9 GB difference between FSDP2 and FairScale represents FairScale's inability to overlap activation recomputation with gradient computation, requiring larger activation buffers. For memory-constrained scenarios like fine-tuning 70B on a single 80 GB H100, DeepSpeed ZeRO-3 with CPU offload (`offload_optimizer_device="cpu"`) fits training at batch size 2 with 78 GB peak, while FSDP2 with `cpu_offload=torch.distributed.fsdp.CPUOffload(offload_params=True)` fits at 75 GB peak.
| Configuration | DeepSpeed ZeRO-3 | FSDP2 | FairScale |
|---|---|---|---|
| Llama 70B, 8x H100, BS/GPU=1 | 2,840 tok/s/GPU | 2,910 tok/s/GPU | 2,100 tok/s/GPU |
| Llama 70B, 64x H100, BS/GPU=1 | 22,400 tok/s total | 22,100 tok/s total | 14,300 tok/s total |
| Peak Memory (70B, 8 GPU) | 55 GB | 53 GB | 62 GB |
| Peak Memory (70B, 1 GPU + CPU offload) | 78 GB | 75 GB | OOM |
| Llama 8B, 8x H100, BS/GPU=4 | 12,400 tok/s/GPU | 12,800 tok/s/GPU | 9,800 tok/s/GPU |
| Comm Overhead (% of step) | 8-12% | 7-10% | 18-25% |
ADVANCED FEATURES: FP8, CPU OFFLOAD, AND HETEROGENEOUS CLUSTERS
DeepSpeed leads in advanced features with ZeRO++ and FP8 ZeRO. ZeRO++ introduces quantized communication for cross-node all-gather operations, quantizing gradients from FP16 to INT8 before all-reduce and dequantizing afterward, reducing communication bandwidth by 50% with less than 0.1% accuracy loss. The `zero_quantized_gradients` option in the DeepSpeed config: `"communication_data_type": "fp8"` enables INT8 gradient compression. For Llama 70B training on 64 GPUs across 8 nodes with 200 Gbps InfiniBand, ZeRO++ reduces all-reduce time from 18ms to 9ms per layer, improving throughput by 12%. DeepSpeed also supports NVMe offload (`offload_param_device="nvme"`) that swaps parameters to SSD, enabling fine-tuning of 175B-scale models on a single H100 with batch size 1.
FSDP2's unique advantage is `torch.compile` integration. The `@torch.compile(backend="inductor", fullgraph=True)` decorator on the training step function compiles the entire forward-backward-update loop into a single fused CUDA graph, reducing Python overhead and communication latency. `torch.compile` with FSDP2 achieves 1.15-1.25x throughput improvement over eager-mode FSDP2 on Llama 70B. FairScale's development has effectively ceased: the repository was archived in December 2023 with a recommendation to migrate to PyTorch FSDP. For new GPU cluster deployments in 2026, the choice is between DeepSpeed (advanced features, ZeRO++, CPU/NVMe offload) and FSDP2 (native PyTorch integration, torch.compile, simpler API). The default recommendation is FSDP2 for most workloads, switching to DeepSpeed when specific features like ZeRO++ gradient compression or NVMe offload are required.
PRODUCTION DEPLOYMENT PATTERNS AND HYBRID SHARDING
Production GPU training deployments use hybrid sharding that combines data parallelism (FSDP/ZeRO within a node) with tensor parallelism (TP across nodes). The standard configuration for 70B on 32x H100 (4 nodes x 8 GPUs) uses TP=8 per node (leveraging NVLink within DGX), FSDP=4 across nodes (4-way data parallel). This hybrid configuration reduces all-reduce communication across InfiniBand links by 75% compared to pure FSDP=32. DeepSpeed supports hybrid sharding natively via `deepspeed.initialize(model=model, config=ds_config)` with `"tensor_parallel": {"enabled": true, "tp_size": 8}`. FSDP2 requires manual wrapping: wrap the model in `tensor_parallel` first using Megatron-LM or `torch.distributed.tensor.parallel`, then apply `fully_shard` for the FSDP dimension.
The cost-optimized cluster configuration on ClusterBid for 70B training uses 8x H100 nodes ($20/hr) with 400 Gbps InfiniBand. DeepSpeed ZeRO-3 with activation checkpointing and BF16 mixed precision trains Llama 70B at 2,840 tok/s/GPU (22,720 aggregate). The estimated training time for 1 trillion tokens on this configuration: 1T / (22,720 tok/s) = 44,014,085 seconds = 12,226 hours = 509 days. At $20/hr for the 8-GPU node, the total compute cost is $244,520. Using FP8 ZeRO reduces total cost by approximately 40% to $146,712 due to 1.8x throughput improvement. The GPU cluster cost dominates the training economics: selecting the optimal distributed training library (FSDP2 or DeepSpeed) can reduce training time and cost by 10-30%. The recommended workflow benchmarks each library on the target cluster for 100 training steps before committing to a full training run.
