WHY GPU TRAINING BREAKS TRADITIONAL STORAGE
A single NVIDIA H100 GPU can sustain 3.3 TB/s of memory bandwidth internally, but training at scale requires feeding data from external storage to 128, 512, or 1024 GPUs simultaneously. At 128 GPUs reading 256 KB samples at a throughput of 40 GB/s aggregate, a single training job can saturate a 400 Gbps InfiniBand link with checkpoint writes every 10-30 minutes. Traditional NFS servers with a single 25 GbE link and kernel NFSd bottleneck at 1-2 GB/s before CPU core exhaustion becomes the limiting factor, creating a mismatch of roughly 20-40x between what GPUs demand and what conventional NAS can supply.
The solution is parallel filesystems designed for concurrent access from thousands of clients. These systems stripe data across multiple storage nodes and present a single POSIX-compliant namespace, enabling aggregate throughput that scales linearly with node count. For GPU training clusters of 256 H100 GPUs and larger, sustained read throughput of 50-200 GB/s is the minimum viable target. Achieving this requires careful design across four dimensions: metadata server architecture, data path parallelism, network fabric integration (InfiniBand vs RoCE v2), and GPU-direct storage (GDS) support.
LUSTRE: THE OPEN-SOURCE STANDARD WITH OPERATIONAL COMPLEXITY
Lustre remains the most widely deployed parallel filesystem in HPC and GPU training environments, powering approximately 60 percent of the TOP500 supercomputers and a comparable share of large-scale AI clusters. The architecture separates metadata and data paths through distinct servers: Metadata Servers (MDS) manage the namespace and file layout, typically running on 2-4 nodes in active-standby or active-active configuration, while Object Storage Servers (OSS) manage 10-200+ Object Storage Targets (OSTs) that store file data. Clients mount the Lustre filesystem via the lnet kernel module and communicate over RDMA-capable fabric.
For GPU training clusters, a typical Lustre deployment uses 8-24 OSS nodes, each with 8-24 OSTs backed by NVMe SSDs, connected via InfiniBand NDR400 or HDR200. A well-tuned 16-OSS Lustre deployment with 128 OSTs on NVMe can sustain 120-180 GB/s read throughput and 60-100 GB/s write throughput on 8+4 MB stripe sizes. The tuning knobs are extensive: stripe count, stripe size, OST thread count, RPC coalescing, and ldlm lock management. The `lfs setstripe -c -1 -S 4M /lustre/dataset` command stripes a file across all OSTs with 4 MB chunks. The operational cost of tuning these parameters for each workload is Lustre's primary drawback versus turnkey alternatives.
| Parameter | Small Dataset (<10 TB) | Medium (10-100 TB) | Large (>100 TB) |
|---|---|---|---|
| OST Count | 8-16 | 32-64 | 64-256 |
| Stripe Size | 1 MB | 4 MB | 8 MB |
| OSS Nodes | 2-4 | 8-16 | 16-32 |
| MDS RAM | 128 GB | 256 GB | 512 GB+ NVMe-backed |
| Network | HDR100 RoCE v2 | HDR200 InfiniBand | NDR400 InfiniBand |
| Estimated Read Throughput | 15-30 GB/s | 60-120 GB/s | 160-400 GB/s |
| Lustre Version | 2.15.x LTS | 2.15.x LTS | FSD (exascale) |
WEKA: THE GPU-FIRST SOFTWARE-DEFINED FILESYSTEM
WEKA has become the most popular parallel filesystem for commercial GPU training clusters, deployed at CoreWeave, Lambda, and multiple top-tier AI labs. Its architecture eliminates the Lustre metadata/data path separation in favor of a unified user-space data plane running on commodity x86 servers with NVMe SSDs. Each WEKA node runs a containerized IO server process that handles both data and metadata operations, using a distributed map and non-blocking RDMA for internode communication. The system scales in increments of 6-48 nodes per filesystem, with each node contributing throughput and capacity.
The key technical advantage of WEKA for GPU training is its native integration with NVIDIA GPUDirect Storage (GDS). GDS enables data to flow directly from storage to GPU memory without passing through the CPU host buffer, reducing latency from 50-80 microseconds per operation to 5-10 microseconds. In benchmarks from the MLPerf storage track, a 24-node WEKA cluster with 48x H100 GPUs achieved 3.2 TB/s aggregate read bandwidth on the UNet3D benchmark, with 92 percent GPU utilization versus 64 percent with traditional NFS. WEKA's data placement engine also supports erasure coding (8+2 or 12+2) rather than replication, delivering 77 percent usable capacity versus 33 percent with 3x replication, at a performance overhead of 3-8 percent on mixed workloads.
| Configuration | WEKA 6-Node | WEKA 12-Node | WEKA 24-Node |
|---|---|---|---|
| NVMe per Node | 6 x 7.68 TB | 12 x 7.68 TB | 24 x 7.68 TB |
| Raw Capacity | 276 TB | 1.1 PB | 4.4 PB |
| Usable (EC 8+2) | 207 TB | 845 TB | 3.4 PB |
| Sustained Read | 45 GB/s | 90 GB/s | 180 GB/s |
| Sustained Write | 25 GB/s | 50 GB/s | 100 GB/s |
| GDS Latency (GPU→NVMe) | 12 microseconds | 12 microseconds | 12 microseconds |
| Estimated Cost/TB (5yr TCO) | $4,500-5,500 | $3,200-4,200 | $2,800-3,500 |
DAOS: INTEL'S DISAGGREGATED STORAGE FOR EXASCALE AI
DAOS (Distributed Asynchronous Object Storage) represents a architectural departure from both Lustre and WEKA. Developed by Intel for the Aurora exascale supercomputer and now available open source, DAOS presents a POSIX-compatible namespace but stores data as versioned objects rather than inodes. The architecture ships metadata and data on physically separate NVMe pools, with metadata stored in Intel Optane Persistent Memory (or DRAM-only mode on newer deployments) for sub-microsecond operations. DAOS targets 100 microseconds per IO operation end-to-end at the 99.9th percentile, compared to 200-300 microseconds for WEKA and 500-1000 microseconds for Lustre.
DAOS's internal benchmark results on a 16-server cluster with dual-port NDR400 InfiniBand show 284 GB/s read throughput on 64 MB I/O with 4 KB random read at 16.2 million IOPS. For GPU training, DAOS supports the DAOS Client Library (libdaos) for native integration with PyTorch and TensorFlow dataloaders, bypassing the POSIX VFS layer entirely and achieving 2.1 GB/s per GPU on checkpoint write operations from 256 H100 GPUs. However, DAOS adoption outside of DOE labs and a few hyperscale AI clusters remains limited because the deployment requires specific hardware (NVMe SSDs with PLP, Optane or large DRAM pools) and the community ecosystem is smaller than Lustre's.
VAST DATA: DISAGGREGATED SHARED-EVERYTHING ARCHITECTURE
VAST Data has emerged as a compelling option for GPU storage by combining a disaggregated shared-everything architecture with QLC NVMe flash economics. The VAST Universal Storage platform uses a DASE (Disaggregated Shared Everything) design where compute and storage are on separate nodes connected via 100/200 GbE NVMe-oF. The storage nodes run QLC SSDs with 3D XPoint or SLC-class write buffers to mitigate the native write endurance limitation of QLC. VAST's proprietary data reduction engine claims 5:1 to 10:1 effective compression on training datasets (mixed formats), delivering usable capacity at $0.10-0.15 per GB-hour total cost of operation.
For GPU training, VAST's key differentiator is its unified protocol support: NFS, SMB, S3, and its own proprietary NVMe-oF protocol all access the same data store without data duplication. This simplifies the AI pipeline architecture because preprocessing scripts writing to S3 buckets and training jobs reading via NFS or NVMe-oF operate on the same data. In published benchmarks, an 8-node VAST cluster serving 128 H100 GPUs over 8x 200 GbE links delivered 105 GB/s sustained read throughput with 3.8x single-blob throughput improvement over an equivalent-capacity WEKA deployment for large file sequential reads, though WEKA outperformed on metadata-heavy workloads by 2.1x.
| Dimension | Lustre | WEKA | DAOS | VAST Data |
|---|---|---|---|---|
| License Model | Open Source (GPL) | Proprietary / SaaS | Open Source (Apache 2.0) | Proprietary / Hardware |
| Metadata Architecture | Separate MDS + MGS | Distributed (unified) | Versioned objects (Optane) | Disaggregated DASE |
| GDS Support | Via lnet (limited) | Native (GPU→NVMe) | DAOS Client Library | Proprietary NVMe-oF |
| Max Aggregate Read | 400 GB/s (256 OST) | 3.2 TB/s (48 nodes) | 284 GB/s (16 nodes) | 200+ GB/s (16 nodes) |
| Read Latency (P99.9) | 500-1000 microseconds | 200-300 microseconds | 50-100 microseconds | 150-250 microseconds |
| Min Recommended Nodes | 6 (2 MDS + 4 OSS) | 6 | 8 | 6 (4 storage + 2 compute) |
| Maturity / Commercial Support | Very High / DDN, HPE, others | High / WEKA | Medium / Intel (transitioning) | Medium-High / VAST Data |
STORAGE NETWORK INTEGRATION AND CHECKPOINT OPTIMIZATION
The storage network architecture must be carefully integrated with the GPU compute fabric. The standard pattern for NVIDIA DGX-based clusters uses two separate networks: a compute fabric (InfiniBand NDR400 or Spectrum-X Ethernet for NCCL traffic) and a storage fabric (InfiniBand or RoCE v2 for storage traffic). The separation prevents checkpoint writes and dataset reads from interfering with gradient synchronization during training. On DGX H100 systems, 8x NDR400 ConnectX-7 HBAs typically split as 4x for compute fabric and 4x for storage fabric, with CUDA-aware MPI routing storage traffic through the storage-specific ports.
Checkpoint optimization deserves specific attention. Training runs on 256-GPU clusters typically write 5-50 GB checkpoints every 10-30 minutes. A 5 GB checkpoint write to Lustre at 10 GB/s takes 500ms, during which the training step pauses (global barrier on all ranks). Asynchronous checkpointing using PyTorch Distributed's `torch.save` on a separate thread with CUDA stream synchronization can reduce wall-clock checkpoint overhead from 2-3 seconds to 200-400ms. Combined with WEKA's GDS path and 8 MB aligned checkpoint file writes, aggregate checkpoint overhead can be held below 2 percent of total training time for clusters of up to 512 GPUs.
