All essays
InfrastructureINFRASTRUCTUREFEB 2026

Training Data Caching to Optimize GPU Loading and Preprocessing

Training data caching strategies for GPU pipelines. Reduce idle GPU time with NVMe cache, webdataset sharding, and data prefetching for H100 training clusters.

01

THE DATA LOADING BOTTLENECK

An H100 consumes 15-25 GB/s preprocessed data. NVMe delivers 3-7 GB/s. S3 delivers 125-625 MB/s. The gap is 25-200x. GPU utilization drops to 40-60 percent when data pipeline is bottlenecked. On 64 H100 GPUs at $3,840/hr, 40 percent wastes $1,536/hr.

Video training with 50-500 MB samples makes this worse: GPUs spend 50-70 percent of time waiting for data decompression and augmentation.

Storage LayerBandwidthCapacityCost/TB/MonthLatencyBest For
GPU HBM33.35 TB/s80 GB$0~10 nsCurrent batch
CPU DRAM200-500 GB/s1-2 TB$5-10~100 nsIn-flight batches
Local NVMe3-7 GB/s4-15 TB$15-30~10 usShard cache
Distributed FS50-200 GB/s100+ TB$50-100~200 usShared datasets
Object Store1-25 GbpsUnlimited$20-25~10 msCold storage
02

WEBDATASET SHARDING

Converting from individual files to WebDataset shards (256 MB tar archives) improves read bandwidth from 30 MB/s to 3 GB/s - a 100x improvement. For a 10TB dataset of 100M image-text pairs, use 40,000 shards at 256 MB each.

Prefetch_factor=8 achieves 92-95 percent GPU utilization vs 55-65 percent without prefetching. Each GPU reads from 4-8 shards concurrently with num_workers=4-8.

FormatFile CountEffective BWGPU UtilizationPreprocess/BatchFramework
Individual files10M30-50 MB/s40-55%800 msNative
WebDataset 128MB80K shards6-8 GB/s82-88%120 mswds+DALI
WebDataset 256MB40K shards8-12 GB/s88-93%80 mswds+DALI
WebDataset 512MB20K shards10-14 GB/s90-95%60 mswds+DALI
03

TIERED CACHING

Four tiers: object store (cold), distributed cache like Alluxio (warm), local NVMe (hot), GPU memory (extreme). Cache-warm pre-fetches first 30 min of data before training starts. Background threads keep hot cache hit rate above 95%.

Cache infra for 64-node cluster: $5,200-5,500/month. Benefit: GPU utilization improves from 55% to 92%, recovering $38-52K/month in GPU time. Cache pays for itself 7-10x.

04

PREPROCESSING PIPELINE DESIGN

Optimal GPU-to-CPU ratio depends on task: text tokenization needs 8-16 CPU cores per GPU, image needs 16-32, video needs 32-64. NVIDIA DALI reduces batch preprocessing from 150 ms to 8 ms.

FFCV library compressed a 1.4B CLIP training from 14 days to 9 days (36% reduction). Profile before optimizing: if GPU kernel launches >1-2 ms apart, data pipeline is the bottleneck.

Filed under
Data CachingGPU TrainingData LoadingWebDatasetNVMe CacheData PipelinePreprocessing Optimization