All essays
TechnicalDEEP DIVEFEB 2026

The Parallel Filesystem Tax: How Lustre, WekaFS, DAOS, and VAST Choices Quietly Add 15-25% to Every GPU Training Run in 2026

The parallel filesystem for AI training quietly burns 15-25% of every GPU hour. Lustre vs WekaFS vs DAOS vs VAST, real GPU starvation telemetry, and how to fix it.

01

How A Bad Parallel Filesystem Starves A $4M B200 Cluster

The single biggest hidden cost on a 2026 GPU training run is not the GPU. It is the parallel filesystem for AI training sitting under it. We have watched 64-node B200 clusters running Llama-4-class pretraining sit at 58 to 62 percent SM utilization for hours at a time, while the same model on the same hardware against a properly tuned WekaFS or VAST tier holds 88 to 92 percent. The delta is 25 to 30 points of GPU efficiency on a cluster billing $180,000 to $220,000 per month per 8-node pod. That is the filesystem tax and nobody puts it on the invoice.

The telemetry pattern is consistent. You watch GPU iteration time on a Llama-4 405B-class run. The forward and backward passes are fine. Then you hit a checkpoint write or a global shuffle reload and the cluster goes quiet. nvidia-smi shows the SMs idle. The InfiniBand counters look healthy. The bottleneck is upstream: the filesystem cannot ingest the 5 to 8 terabyte checkpoint inside the iteration window, so the next forward pass blocks waiting for the optimizer state to land on stable storage. On a 1,024-GPU run, every five minutes of starvation is roughly 85 GPU-hours of waste, and at $2.02 per H200-hour or $3.36 per B200-hour, the math gets ugly fast.

The frustrating part is that this is almost always diagnosed late. Teams benchmark their model on a single node, get good numbers, scale to 64 nodes and lose 20 to 30 percent of throughput somewhere they cannot find. The training storage bottleneck shows up as GPU starvation, not as a filesystem error. There is no log line that says "WekaFS metadata server is rate-limiting your dataloader." The cluster just gets slower the more nodes you add and the team chalks it up to scaling overhead. It is not. It is storage.

Workload patternStorage requirementCommon bottleneck
Dense pretraining (sequential reads)1-2 GB/s per GPU readAggregate read bandwidth
Checkpoint write (every 30-60 min)200-400 GB/s burst writeBurst buffer / metadata
MoE training (random sparse)Low-latency random readsIOPS, metadata scaling
RAG indexing (small files)5-50M IOPS small readsMetadata server saturation
LoRA fine-tuning (small checkpoints)Modest throughput, low latencySnapshot overhead
02

Lustre vs WekaFS vs DAOS vs VAST: How The Architectures Actually Differ

Most procurement decks treat these four as interchangeable parallel filesystems. They are not. The architecture choices behind each one push you toward very different workloads, ops budgets, and unit economics. The Lustre vs WekaFS vs DAOS choice is the most consequential storage decision a training team makes after picking the GPU.

Lustre is the open-source baseline. It split metadata (MDS) and object storage (OSS) into separate servers years before that was fashionable, and it scales sequential read bandwidth almost arbitrarily by adding OSTs. AWS FSx for Lustre ships persistent SSD tiers that quote 1,000 MB/s per TiB of provisioned throughput and burst higher, and Azure Managed Lustre lands in a similar band. The hard truth about open-source Lustre on bare metal is the ops burden. You run an MDS failover pair, you tune stripe counts per dataset, you deal with the OST imbalance that creeps in over months, and you keep a Lustre kernel engineer on payroll or you eat downtime. For teams that have one, Lustre is the cheapest dollar per GB per second in the market. For teams that do not, the all-in cost looks very different.

WekaFS is the commercial answer to "we cannot keep a Lustre engineer." It is NVMe-native, runs metadata and data on the same nodes, and exposes POSIX, NFS, SMB, and S3 from the same namespace. The throughput numbers Weka publishes are real: production deployments routinely sustain 720 GB/s read across 8 nodes and burst to 7.2 TB/s aggregate at the cluster level, with native GPU Direct Storage support that bypasses the host CPU on the read path. The cost is the per-TB license, which sits in the high single digits to low double digits per usable TB per month at typical deal sizes, on top of the underlying hardware.

DAOS is the Intel-originated, SCM-tiered design developed starting in 2015 for DOE exascale programs and prominently deployed in the Aurora supercomputer. It pushes metadata into Optane-class persistent memory (or its successors) and treats the storage stack as a key-value store rather than block-and-inode. The latency numbers are extraordinary for small random IO and that maps directly to MoE training and RAG indexing patterns. The catch is the ecosystem. The supported hardware stack is narrow, the operator base is small, and the future of Optane has been turbulent enough that several teams paused DAOS rollouts in 2024 and 2025 while they figured out which SCM technology to bet on. The teams that committed have very fast clusters.

VAST Data is the architectural outlier. Their DASE design (Disaggregated, Shared-Everything) puts QLC flash and storage class memory behind a fabric that lets any compute node see any storage device, with no fixed shard ownership. The practical result is that the namespace scales without metadata partitioning headaches and the QLC tier brings price-per-TB down to a band that competes with hybrid disk arrays at all-flash speeds. VAST has been winning AI deals through 2025 and into 2026 because the unit economics on a 10-petabyte training corpus look very different from WekaFS, and the GPU Direct Storage support is mature.

FilesystemArchitectureSweet spot workloadOps burden
Lustre (OSS)MDS + OST splitDense pretraining, big sequentialHigh
FSx LustreManaged LustreSpiky training, AWS-nativeLow
WekaFSNVMe shared-nothingMixed training and inferenceMedium
DAOSSCM + NVMe KV storeMoE, RAG, small random IOVery high
VAST DataDASE, QLC + SCMLarge corpora, mixed accessLow to medium
03

Matching Your Workload To The Right Filesystem Stack

The honest answer to "which filesystem should we run" is "what does your access pattern actually look like." Most teams do not know. The data scientists know the model. The infra team knows the GPU spec. Nobody has the heat map of read sizes, read locality, and write bursts the model actually generates, which is why the choice ends up being made by whichever vendor sent the most polished deck.

Dense pretraining (Llama-4 dense, Mistral, classic transformer architectures) is the easiest case. The access pattern is large sequential reads of the tokenized dataset, with periodic large sequential checkpoint writes. Aggregate read bandwidth is what matters. Lustre, WekaFS, and VAST all handle this well. DAOS works but you are paying for IOPS you do not use. The deciding factor here is checkpoint write speed, because a 5 to 8 TB optimizer state from a 405B-parameter model has to land in 60 to 90 seconds to fit inside a reasonable iteration cadence. That implies 80 to 130 GB/s of write bandwidth from the cluster to the storage, sustained for the duration of the write.

MoE training (Llama-4 Maverick-class with 128 experts, Mixtral, DeepSeek-MoE) flips the pattern. Expert routing means active parameters change per token, and the data loader pattern becomes more random as different experts pull from different parts of the dataset. The metadata server stops being decorative and starts being the hot path. DAOS shines here, VAST holds up well because of the disaggregated metadata design, WekaFS is competent, and Lustre struggles unless you partition aggressively. We have seen MoE training runs on stock Lustre lose 18 to 24 percent of GPU throughput just to metadata contention.

RAG indexing and fine-tuning are the small-file problem. You are reading millions of 4-to-50 KB chunks. The IOPS ceiling matters far more than bandwidth. WekaFS publishes 18 million IOPS at the cluster level on the WEKApod Nitro 150 and Nitro 180 appliances, DAOS goes higher per node, and VAST is competitive. Lustre is the wrong answer for this workload regardless of how big you build it. LoRA fine-tuning is a different small-file problem because the checkpoints are tiny (single-digit GB) but the iteration cadence is fast, so the bottleneck is snapshot and version overhead more than raw throughput.

04

What Each Tier Actually Costs Per Petabyte In 2026

Storage cost analysis is where the marketing finally meets the invoice. The number you care about is total cost per usable petabyte per month, including hardware, licenses, support, power, and the operators required to keep it healthy. Most published vendor numbers leave out at least two of those.

AWS FSx for Lustre Persistent SSD lists at roughly $0.300 per GB-month for the 1,000 MB/s/TiB tier in us-east-1 as of Q2 2026, which is roughly $307,000 per usable petabyte per month at list price before any enterprise discount. That includes the management plane, the hardware, and Amazon's margin. There is no separate license, no support contract, no operator cost. For teams that do not want to run storage, this is the floor.

On-prem WekaFS lands in a different band. The hardware (NVMe servers, NICs, switches) runs $400,000 to $700,000 per petabyte of usable capacity at current SSD pricing, amortized over 5 years that is $7,000 to $12,000 per month. The Weka license is the next layer and is negotiable but typically sits in the $5,000 to $10,000 per usable PB per month band for AI workloads at the multi-PB scale. Add support and operator overhead and you are at $15,000 to $30,000 per usable PB per month all-in. Far cheaper than FSx at steady state, materially more expensive in operator burden.

VAST tends to come in at the lowest steady-state cost per PB at the multi-PB tier because the QLC flash design pushes hardware cost down. We see all-in numbers landing $10,000 to $20,000 per usable PB per month on multi-PB deployments. DAOS bare metal numbers are harder to publish cleanly because the SCM tier is the swing variable; on Optane-class hardware the cost is roughly comparable to WekaFS, but on the post-Optane SCM alternatives shipping in 2026 the cost band is still settling.

The honest summary is that pure storage cost is not the line item that decides the question. A 25 percent GPU utilization difference on a $200,000 per month cluster is $50,000 per month of GPU waste. That dwarfs any reasonable filesystem price delta. The right question is which filesystem gets your GPUs to 90 percent utilization, then optimize storage cost within that constraint.

Filesystem optionAll-in $/usable PB/monthOperator load
AWS FSx Lustre (persistent SSD)$300K-$310KNone
Azure Managed Lustre$120K-$160KNone
Lustre on bare metal (open-source)$8K-$15K + 1-2 FTEVery high
WekaFS on-prem$15K-$30KMedium
VAST Data multi-PB$10K-$20KLow to medium
DAOS bare metal$15K-$25K + 1-2 FTEVery high
05

GPU Direct Storage And The CPU Bypass Most Teams Skip

GPU Direct Storage (GDS) is the part of the storage stack that determines whether your H200 or B200 cluster actually sees the throughput the storage vendor sells you. The mechanism is straightforward. Without GDS, every byte read from storage hops through the host CPU and main memory before reaching the GPU. With GDS, the GPU pulls bytes from NVMe or RDMA storage directly into its HBM with no CPU involvement. The CPU savings are real but the bigger win is latency and predictability under load.

Native GPU Direct Storage support in 2026 is uneven across the filesystem options. WekaFS shipped GDS support early and the implementation is mature. VAST has supported GDS since the 2021 GA launch and has hardened it through multiple production deployments. DAOS has had GDS-style direct paths from the start because the architecture was designed for it. Lustre's GDS support varies by distribution; the upstream open-source path requires recent kernel work, while the managed offerings from AWS and Azure ship it enabled by default. The teams skipping this are usually running open-source Lustre on a stack where GDS was never enabled and they do not know.

The throughput numbers with and without GDS are not subtle. NVIDIA's reference benchmarks show 2x to 3x improvement in storage read throughput when GDS is active, which translates to meaningful dataloader throughput gains on I/O-bound training runs. On a B200 cluster, that is the difference between feeding the GPUs at 90 percent utilization and feeding them at 50 percent. If you are paying the B200 hourly rate and you have not verified GDS is actually active end-to-end, you are likely leaving half your throughput on the floor.

The verification is also non-obvious. nvidia-smi will not tell you GDS is active. You need to check the cufile.json config, run the gdscheck tool from the NVIDIA repo, and confirm the actual storage path is registered as a GDS-capable mount. On managed cloud filesystems this is documented; on bare-metal Lustre or DAOS deployments the team has to verify it themselves. We have seen multi-million-dollar clusters running for weeks with GDS misconfigured because nobody on the team thought to check.

06

Beyond Throughput: What To Actually Verify Before Signing

The bandwidth and IOPS numbers in vendor decks are necessary but never sufficient. The variables that decide whether a parallel filesystem will work for a real training cluster sit one layer below the headline performance numbers, and they are the ones teams catch only after the deployment is in production.

Data tiering policy. Most modern AI corpora are 10 to 50 PB and a small fraction is hot at any time. You want the filesystem to tier cold data to cheaper media without manual intervention and without breaking the global namespace. WekaFS tiers to S3 transparently. VAST manages its own QLC tier internally so there is no tiering to configure. Lustre and DAOS rely on HSM (Hierarchical Storage Management) policies that you write yourself. Ask the vendor for a real, documented example of a customer running their tiering policy at your scale, not the slide deck claim.

Snapshot cost and frequency. Training teams snapshot checkpoints and intermediate artifacts constantly. A filesystem that charges 1.5x or 2x storage cost for snapshot data is going to bend your budget in ways nobody plans for. VAST is aggressive on this with its similarity reduction and the effective snapshot overhead is low. WekaFS snapshots are space-efficient but eventually compound. Bare-metal Lustre has no native snapshot story and most teams roll their own via ZFS on the OSTs, which works but adds complexity.

Multi-tenant isolation. If the cluster is shared across multiple teams (which is the common case for GPU clusters costing seven figures per month), you need quota enforcement, per-tenant throughput caps, and namespace isolation. WekaFS and VAST handle this natively. Lustre has project quotas but the per-tenant throughput control is weaker. DAOS isolates at the pool level which is clean but requires upfront capacity planning. The wrong answer here is discovering at month four that one team's RAG indexing job is starving the pretraining cluster of metadata IOPS.

Burst buffer behavior. Checkpoint writes are bursty. Most filesystems have an effective write cache that absorbs the burst and drains it to the persistent tier over the next several minutes. Ask for the burst cache size and the drain rate. We have seen WekaFS configurations that absorb 8 TB checkpoints in 40 seconds and then take three minutes to drain, which is fine, versus FSx Lustre configurations on smaller throughput tiers where the burst cache is the bottleneck and the GPUs sit idle until the previous checkpoint clears.

07

When To Bring Your Own Filesystem vs Use The Data Center's Tier

The choice between bringing your own storage stack and using your GPU provider's managed tier is the decision most procurement teams handle backwards. The default assumption is that BYO is cheaper because you control the hardware. At small scale and short contract terms, that is wrong. At large scale and long terms, it can be right but only if you have the operator headcount to back it up.

Use the provider's managed storage tier when the deployment is under 12 months, the cluster is under 256 GPUs, or the team does not already have a Lustre or Weka engineer. The all-in cost is higher per PB but the time to first training job is days instead of months, and the provider absorbs the operational risk. For Series A and Series B AI teams running their first multi-node training, this is almost always the right answer.

Bring your own storage when the deployment is multi-year, multi-petabyte, and the team has the operators to run it. The break-even on a properly sized WekaFS or VAST deployment versus FSx Lustre or Azure Managed Lustre lands around 18 to 24 months at the multi-PB scale - the TCO framework for this decision is covered in detail in the bare-metal vs cloud comparison. The savings compound at larger scale, which is why every frontier lab runs their own storage tier even though they could afford to outsource it.

The middle path that most teams miss is sourcing GPU capacity at a data center that already runs the storage tier you need. A B200 cluster at a DC that already operates a multi-PB VAST or WekaFS deployment is materially cheaper than BYO storage at a DC that was not designed for it, because the storage capex is amortized across many tenants and the operator cost is shared. ClusterBid's vetting process explicitly checks parallel filesystem options, throughput per node, and bandwidth between storage and GPU racks at each DC, which is why the 15-question DC vetting framework we publish covers the storage adjacencies (power, fabric, cooling, SLA, egress) you need to interrogate alongside the filesystem itself.

08

How To Source A Cluster With The Right Storage Tier

Most teams renting GPUs forget to ask the storage question until the cluster is running and the training is slow. The right question goes in the RFQ. What parallel filesystem do you operate, what is the aggregate read and write bandwidth between storage and the GPU racks, is GPU Direct Storage enabled by default, and what is the burst buffer size for checkpoint writes. If the provider cannot answer all four cleanly, that is the answer.

The data center spread on storage tier in 2026 is wider than the spread on GPU pricing itself. Some DCs run mature WekaFS or VAST tiers and quote 800 GB/s read between storage and GPUs by default. Others quote a single NFS share at 12 GB/s and assume the buyer will figure it out. The brokered model we use at ClusterBid maps the customer's data pattern (random vs sequential, checkpoint cadence, dataset size, model architecture) to a DC that already runs the right filesystem stack, instead of forcing a team to BYO storage at a place that was not designed for it. Pricing on H100, H200, B200, and B300 capacity with verified storage tiers is published live on our inventory page and updated as operators add capacity. Pricing reflects current GPU availability as of May 2026 and can fluctuate based on supply and demand.

The filesystem tax is real, it is 15 to 25 percent of every GPU hour on a poorly matched stack, and it compounds every month for the life of the contract. The teams that take this seriously in 2026 will be running their B200 and B300 fleets at 90 percent utilization while their competitors burn the same hardware at 65 percent and wonder why their loss curves are slower. If you want a routed quote that takes the filesystem topology into account, the sourcing desk handles the matchmaking and the contract review at no cost to the buyer, and it pairs naturally with the commissioning timeline work we published recently.

Filed under
Parallel filesystem for AI trainingLustre vs WekaFS vs DAOSVAST Data DASEGPU Direct StorageGPU starvationTraining cluster storage tierCheckpoint throughput