All essays
InfrastructureINFRASTRUCTUREFEB 2026

GPU Storage Tiering: Hot, Warm, and Cold Data for AI Workloads

GPU storage tiering strategies for AI workloads. Hot (local NVMe), warm (network SSD), and cold (object storage) data tiers for training and inference pipelines.

01

Why GPU Storage Architecture Determines Training Throughput

GPU storage is the most underestimated bottleneck in AI infrastructure. A single H100 SXM5 GPU can process data at 3.35 TB/s from HBM during training. If the data loader cannot feed the GPU at a rate that keeps it busy, the GPU stalls. This is the I/O bottleneck: the GPU finishes computing the current batch and waits for the next batch to arrive from storage. At H100 rental rates of $1.15/hr on ClusterBid, a GPU waiting 10% of the time adds $0.115/hr to the effective cost - a 10% efficiency loss across a 256-GPU cluster costing $294/hr wastes $29.40/hr.

The storage hierarchy for AI workloads spans five tiers with latency differences spanning five orders of magnitude. L1 is GPU HBM (3.35 TB/s, ~200-400ns latency). L2 is the GPU's L2 cache (50MB on H100, ~50-100ns). L3 is local NVMe SSD on the GPU host (7-14 GB/s, ~5-10 microseconds). L4 is network-attached SSD storage (1-5 GB/s per node, ~100-500 microseconds). L5 is object storage like S3 (0.5-2 GB/s per connection, ~5-50 milliseconds). Each tier is 3-10x slower and 5-20x cheaper per GB than the tier above it. Designing the data pipeline to keep the hottest data in the fastest tiers is the storage architect's primary task.

The data pipeline for AI training has distinct phases with different storage requirements: dataset storage (terabytes of training data, accessed sequentially during each epoch), shuffle buffer (randomized read order, requires fast random access), preprocessing cache (transformed data that should be cached between epochs), checkpoint storage (model weight snapshots, write-heavy, high reliability required), and logging/output storage (metrics, sample outputs, tensorboard data). Each phase benefits from a different storage tier configuration, and the optimal architecture matches each data type to its appropriate tier.

02

The Hot Tier: Local NVMe SSD for Training Data Streaming

Local NVMe SSD storage on the GPU host node is the hot tier for AI training. A DGX H100 node includes 4-8 x 3.84TB NVMe SSDs in RAID0 or JBOD configuration, providing 15-30TB of usable capacity with read speeds of 14-28 GB/s depending on RAID striping. This capacity is sufficient to hold the working dataset for most training runs - typically 1-10TB for NLP datasets and 5-50TB for vision datasets. The hot tier should hold the current training dataset in its entirety during the training run.

The hot tier performance requirement for GPU training is driven by the GPU's data consumption rate. A single H100 training a 70B model with Flash Attention consumes approximately 500-800MB/s of training data (tokens, images, or other input) depending on batch size and sequence length. An 8-GPU node consuming 4-6 GB/s of training data requires a local NVMe RAID array capable of at least this throughput. A single PCIe 4.0 NVMe SSD at 5-7 GB/s per drive meets the requirement for a single GPU node. The DGX H100's RAID0 of 4-8 SSDs provides headroom for multiple concurrent training jobs or data preprocessing on the same node.

The hot tier configuration advice: use 70-80% of local NVMe capacity for the training dataset, reserving 20-30% for checkpoint storage and temporary buffers. Pre-stage the training dataset to the hot tier before the training job starts using a data staging script that runs during the GPU provisioning phase. The staging time for a 5TB dataset over a 25 Gbps network connection is roughly 25 minutes - a cost of roughly $4.80 for a 256-GPU cluster ($1.15/hr per GPU, 25 minutes of 8 GPUs preprocessing) that is recovered within the first 2-3 minutes of training through avoided I/O waits.

03

The Warm Tier: Network-Attached SSD for Dataset Repositories

The warm storage tier is a network-attached SSD filesystem (Lustre, GPUDirect Storage-compatible filesystem, or NFS over RDMA) shared across the GPU cluster. It holds the full dataset repository - all datasets the team may use, not just the currently active training dataset. The warm tier capacity typically ranges from 50TB to 500TB for an active AI team, at a cost of $0.15-0.35/GB/month for managed filesystems (Amazon FSx for Lustre, GCP Filestore with SSD, or a self-hosted GPUDirect Storage cluster).

The warm tier is also the primary checkpoint storage target for most training runs. Checkpoints are written to the warm tier during training and retained for analysis, rollback, and model governance. The warm tier must support high write throughput for concurrent checkpointing from multiple training jobs. A 256-GPU cluster with 4 concurrent training jobs may generate checkpoint writes totaling 500-2000 MB/s during checkpoint intervals. The warm tier should support at least 10 GB/s of write throughput with sub-second consistency for checkpoint atomicity.

GPUDirect Storage (GDS) enables direct data transfer between GPU memory and NVMe storage without passing through the CPU host memory. On H100 with GDS-compatible storage (Lustre, WekaFS, VAST Data, or basic NVMe over Fabrics), data loading throughput improves by 30-50% for large dataset reads compared to CPU-buffered I/O. GDS reduces CPU utilization for data loading from 15-25% of a CPU core per GPU to 2-5%, freeing CPU resources for preprocessing and data augmentation. Enable GDS in your storage configuration if your warm tier storage system supports it: the configuration is a single additional mount option and typically improves training throughput by 5-10% for I/O-bound workloads.

TierMediaThroughputLatencyCost/GB/Month
HotLocal NVMe RAID14-28 GB/s per node5-10us$0.06-0.10/GB
WarmNetwork SSD (Lustre/NFS)1-10 GB/s per cluster100-500us$0.15-0.35/GB
ColdObject Storage (S3/R2)0.5-2 GB/s per prefix5-50ms$0.01-0.02/GB
04

The Cold Tier: Object Storage for Dataset Archives and Long-Term Retention

The cold storage tier for AI workloads is object storage (AWS S3, Cloudflare R2, GCP GCS, or MinIO) used for long-term dataset archives, completed training run outputs, and infrequently accessed data. Object storage costs $0.01-0.02/GB/month for standard access and $0.001-0.004/GB/month for archival access (Glacier, S3 Glacier Deep Archive). The latency of 5-50 milliseconds for the first byte makes object storage unsuitable for direct training data access but ideal for infrequent reads and write-once data.

The data migration strategy between tiers: cold-to-warm migration occurs when a dataset is scheduled for training. The data moves from object storage to the warm tier using the provider's high-speed transfer tool (aws s3 sync, gsutil rsync, rclone) 1-2 hours before the training start. Warm-to-hot migration stages the active training dataset to local NVMe during the GPU provisioning phase. The migration should happen asynchronously in the background while the pod starts, using a data preparation init container that completes before the training container starts. Hot-to-cold migration occurs after training: completed checkpoints and outputs move from warm tier to cold tier for archival within 24 hours of job completion.

Cloudflare R2 is an increasingly popular choice for the cold tier in GPU clusters because it charges zero egress fees, unlike AWS S3 ($0.09/GB for internet egress) or GCP GCS ($0.12/GB). For teams that frequently move data between GPU providers (multi-cloud strategy), the egress fees from S3 or GCS can exceed the storage costs by 10-100x. A team that generates 500GB of checkpoint data per day and stores it in S3 with an occasional move to a different provider for evaluation pays $1,350/month in egress for a single data movement, versus $0 with R2.

05

Checkpoint Storage: The Special Storage Tier for Model Persistence

Checkpoint storage deserves its own design consideration within the tiering hierarchy. Checkpoints have different access patterns than training data: they are written infrequently (every 250-500 training steps, typically every 5-30 minutes), written as large sequential files (130-280GB for a 70B model with FSDP), and read primarily for resume or evaluation. The write throughput requirement is high: a 200GB checkpoint written within 30 seconds requires 6.7 GB/s of sustained write throughput. The hot tier (local NVMe) is the designated checkpoint write target for speed, with asynchronous copy to the warm tier for persistence.

The checkpoint storage strategy: write checkpoints to local NVMe for the fastest write and fastest resume (no network latency). After each checkpoint write, start an asynchronous copy to the warm tier object store. The copy runs in the background while training continues. If the training job fails before the copy completes, the checkpoint is still available on local NVMe for immediate resume. If the job fails after the copy completes, the checkpoint is available on the warm tier for resume even if the local NVMe is lost. This two-tier checkpoint strategy provides both fast writes and durable persistence with zero additional latency to the training loop.

Checkpoint compression reduces storage costs and copy time. Zstandard compression (zstd level 3-6) compresses checkpoint files by 20-35% with negligible CPU overhead (1-2% of a CPU core during compression). The compressed checkpoint is smaller on local NVMe (freeing space for more checkpoints), faster to copy to the warm tier (20-35% less data to transfer), and cheaper to store (20-35% less warm and cold tier cost). The decompression during resume adds roughly 5-15 seconds to the resume time, which is acceptable for most use cases and can be overlapped with GPU initialization.

06

Intelligent Data Caching: Avoiding the Data Upload Tax on Every Epoch

Training data caching is the highest-impact storage optimization for multi-epoch training. The naive approach reads the training dataset from the warm tier on every epoch, incurring the full data load overhead (100-500 microseconds latency, limited throughput) for each epoch. The optimized approach caches the dataset in the hot tier after the first epoch: the first epoch reads from the warm tier (slower), transforms and loads the data, and the preprocessed data remains in the hot tier's page cache for subsequent epochs. For a 10-epoch training run, the first epoch is 2-3x slower due to cold cache, and the remaining 9 epochs read from the hot tier cache at full NVMe speed.

Intelligent caching goes beyond simple page caching with a GPU-aware caching layer that prefetches the next batch while the GPU computes the current batch. The prefetch thread reads data from the hot tier into a pinned memory buffer that the GPU can DMA directly. With a double-buffered prefetch (one buffer for the current batch, one buffer prefetching the next), the GPU never waits for data as long as the prefetch completes within one training step. For typical training step times of 300-800ms on H100 for 70B models, the prefetch must complete within 200-600ms - easily achievable with local NVMe at 14 GB/s throughput.

The LRU-like eviction policy for the data cache: the hot tier retains the most recently used datasets across training runs. When a new dataset needs more space than available, the oldest dataset (by last access time) is evicted. The evicted dataset's data is preserved in the warm tier and can be re-cached when next needed. The eviction policy should prioritize keeping the working dataset (the one currently being trained on) and the next scheduled dataset (prefetched based on the job queue). This cache policy reduces average data load time across consecutive training runs by 60-80% versus reading from warm tier on every run.

07

Architecting the GPU Storage Stack: Implementation Decisions

The storage architecture decision for a GPU cluster follows a clear pattern by cluster size. Small clusters (8-32 GPUs, single node or small multi-node): use local NVMe for hot tier, cloud object storage for warm and cold tiers. The warm tier can be skipped for small clusters - the hot tier is large enough to hold the dataset, and the cold tier provides persistence. The architecture simplicity (no network filesystem) reduces operational complexity and cost for small deployments.

Medium clusters (32-128 GPUs, multi-node): add a warm tier network filesystem shared across nodes. The warm tier enables sharing datasets across training jobs, storing checkpoints from any node, and running data preprocessing that writes to a location accessible from all nodes. Deploy a GPUDirect Storage-compatible filesystem (Lustre on AWS FSx, VAST Data, or WekaFS) with at least 10 GB/s aggregate throughput. The warm tier cost of $0.15-0.35/GB/month is justified by the elimination of data silos and the checkpoint access flexibility.

Large clusters (128+ GPUs, multi-rack): implement the full three-tier storage architecture with local NVMe hot tier, GPUDirect Storage warm tier with 50+ GB/s throughput, and object storage cold tier with lifecycle policies for automatic data migration between storage classes. The warm tier should be deployed as a dedicated storage cluster (separate from the compute cluster) with InfiniBand or high-speed Ethernet connectivity to all GPU nodes. The storage cluster should be sized independently: 1-2 storage nodes per 8-16 GPU nodes provides balanced throughput. The full architecture adds 10-20% to total cluster cost but prevents the I/O bottlenecks that reduce GPU utilization below 60% in clusters with inadequate storage.

Filed under
Storage TieringGPU StorageNVMeObject StorageData PipelineTraining I/OCheckpoint Storage