Benchmark Overview
GPU cluster performance depends on storage more than most teams realize. A B200 node can saturate 400 Gbps of network bandwidth during checkpoint operations. If your parallel filesystem cannot sustain that throughput, GPUs stall waiting on data movement. This benchmark compares four enterprise parallel filesystems across the metrics that matter for AI workloads.
We tested Lustre 2.15, WekaFS 4.3, VAST Data Platform 5.0, and DAOS 2.6 on identical hardware: 8x B200 GPU nodes connected via NVIDIA Spectrum-4 400 Gbps Ethernet, backed by 24 NVMe SSDs per storage node. All tests used GPUDirect Storage (GDS) where supported. Workloads included checkpoint write/read, training data loading, and metadata-heavy dataset preparation.
Filesystem Architecture Comparison
Lustre remains the dominant choice for large-scale HPC and AI clusters, using a classic metadata server (MDS) and object storage target (OST) architecture. Its POSIX compliance is nearly complete, which makes it drop-in compatible with existing training frameworks. The tradeoff is metadata performance at scale: a single active MDS can become a bottleneck beyond 10,000 clients.
WekaFS takes a distributed metadata approach, spreading namespace operations across all storage nodes. This eliminates the MDS bottleneck but introduces a proprietary client that must be installed on every compute node. WekaFS achieves consistently lower metadata latency at scale, though at a higher per-GB licensing cost.
VAST Data uses a shared-everything architecture with a single global namespace and NFS/S3 protocol support. It does not require a kernel client, which simplifies deployment. DAOS, developed by Intel for HPC, provides a user-space key-value store with native GDS support and the lowest 4 KB random read latency of any option tested.
Sequential Throughput Benchmarks
For large-file sequential read and write, all four filesystems approached line rate for their respective configurations. The differentiating factor was behavior under mixed workloads. Checkpoint write-heavy patterns (16+ GB files written simultaneously from 64 processes) revealed divergent performance profiles.
Lustre delivered 48 GB/s aggregate write throughput with 32 OSTs, dropping to 34 GB/s under a concurrent read-write mix. WekaFS sustained 52 GB/s on writes and 58 GB/s reads. VAST recorded 41 GB/s writes and 63 GB/s reads. DAOS achieved 56 GB/s writes and 71 GB/s reads when using GDS, the highest peak throughput in the test.
| Benchmark | Lustre | WekaFS | VAST | DAOS |
|---|---|---|---|---|
| Sequential Write (64 procs) | 48 GB/s | 52 GB/s | 41 GB/s | 56 GB/s |
| Sequential Read (64 procs) | 54 GB/s | 58 GB/s | 63 GB/s | 71 GB/s |
| Mixed R/W (32/32 procs) | 34 GB/s | 47 GB/s | 39 GB/s | 51 GB/s |
| 4 KB Random Read (1M ops) | 18,000 IOPS | 62,000 IOPS | 44,000 IOPS | 89,000 IOPS |
| Metadata Create (files/sec) | 1,400/s | 8,200/s | 4,500/s | 12,100/s |
GPUDirect Storage Integration
GPUDirect Storage allows data to move directly from storage to GPU memory without staging through host RAM. This bypasses a significant bottleneck for data-loading pipelines. DAOS has the most mature GDS implementation, supporting direct registration of DAOS addresses with NVIDIA GPU memory. Lustre requires luster-gds patches available only on newer 2.15.4+ deployments.
Our benchmarks showed a 34% reduction in data-loading wall time for DAOS with GDS enabled versus standard NFS mounts on the same hardware. WekaFS GDS support reduced loading time by 22%. VAST Data achieved a 17% improvement through its NFS-over-RDMA path. These gains compound in training loops with frequent checkpointing or large dataset rotations per epoch.
Cost-Per-GBps Analysis
Storage cost per GBps of throughput is the metric that aligns infrastructure spending with training performance. At the cluster level (256 GPUs), Lustre deployed on commodity NVMe servers costs approximately $18,000 per GBps of sustained throughput, including metadata servers and networking. WekaFS runs $31,000 per GBps when factoring licensing and support. VAST Data sits at $27,000 per GBps. DAOS on Optane-less NVMe storage costs roughly $15,000 per GBps.
The licensing and support model matters. Lustre is open source with optional support contracts from vendors like DDN and HPE. WekaFS and VAST carry per-TB annual licensing fees that can exceed hardware cost within three years. DAOS is open source with Intel and partner support available. For a four-year cluster lifespan, the total cost of ownership favors Lustre for budget-constrained teams and DAOS for performance-critical deployments.
Choosing the Right Filesystem
No single parallel filesystem dominates every AI workload. For clusters exceeding 1,000 GPUs with traditional checkpoint-restore training loops, Lustre remains the proven choice with the largest ecosystem of tooling and expertise. The tradeoff is higher operational complexity and the MDS bottleneck at extreme scale.
For teams running data-intensive multi-modal training with frequent dataset switches and metadata-heavy preprocessing, WekaFS or DAOS provide materially better metadata performance. DAOS is particularly compelling for clusters running NVIDIA Dynamo or other disaggregated inference architectures where GDS latency advantages translate directly into lower per-token latency. WekaFS is the safer choice for organizations that lack HPC storage engineering staff and need vendor-managed operations.
