THE DATA LOADING BOTTLENECK
An H100 consumes 15-25 GB/s preprocessed data. NVMe delivers 3-7 GB/s. S3 delivers 125-625 MB/s. The gap is 25-200x. GPU utilization drops to 40-60 percent when data pipeline is bottlenecked. On 64 H100 GPUs at $3,840/hr, 40 percent wastes $1,536/hr.
Video training with 50-500 MB samples makes this worse: GPUs spend 50-70 percent of time waiting for data decompression and augmentation.
| Storage Layer | Bandwidth | Capacity | Cost/TB/Month | Latency | Best For |
|---|---|---|---|---|---|
| GPU HBM3 | 3.35 TB/s | 80 GB | $0 | ~10 ns | Current batch |
| CPU DRAM | 200-500 GB/s | 1-2 TB | $5-10 | ~100 ns | In-flight batches |
| Local NVMe | 3-7 GB/s | 4-15 TB | $15-30 | ~10 us | Shard cache |
| Distributed FS | 50-200 GB/s | 100+ TB | $50-100 | ~200 us | Shared datasets |
| Object Store | 1-25 Gbps | Unlimited | $20-25 | ~10 ms | Cold storage |
WEBDATASET SHARDING
Converting from individual files to WebDataset shards (256 MB tar archives) improves read bandwidth from 30 MB/s to 3 GB/s - a 100x improvement. For a 10TB dataset of 100M image-text pairs, use 40,000 shards at 256 MB each.
Prefetch_factor=8 achieves 92-95 percent GPU utilization vs 55-65 percent without prefetching. Each GPU reads from 4-8 shards concurrently with num_workers=4-8.
| Format | File Count | Effective BW | GPU Utilization | Preprocess/Batch | Framework |
|---|---|---|---|---|---|
| Individual files | 10M | 30-50 MB/s | 40-55% | 800 ms | Native |
| WebDataset 128MB | 80K shards | 6-8 GB/s | 82-88% | 120 ms | wds+DALI |
| WebDataset 256MB | 40K shards | 8-12 GB/s | 88-93% | 80 ms | wds+DALI |
| WebDataset 512MB | 20K shards | 10-14 GB/s | 90-95% | 60 ms | wds+DALI |
TIERED CACHING
Four tiers: object store (cold), distributed cache like Alluxio (warm), local NVMe (hot), GPU memory (extreme). Cache-warm pre-fetches first 30 min of data before training starts. Background threads keep hot cache hit rate above 95%.
Cache infra for 64-node cluster: $5,200-5,500/month. Benefit: GPU utilization improves from 55% to 92%, recovering $38-52K/month in GPU time. Cache pays for itself 7-10x.
PREPROCESSING PIPELINE DESIGN
Optimal GPU-to-CPU ratio depends on task: text tokenization needs 8-16 CPU cores per GPU, image needs 16-32, video needs 32-64. NVIDIA DALI reduces batch preprocessing from 150 ms to 8 ms.
FFCV library compressed a 1.4B CLIP training from 14 days to 9 days (36% reduction). Profile before optimizing: if GPU kernel launches >1-2 ms apart, data pipeline is the bottleneck.
