All essays
GuideGUIDEFEB 2026

NFS over RDMA for GPU Clusters: 2026 Complete Guide - Architecture, Performance Benchmarks and Cost Analysis

Complete guide to NFS over RDMA for AI training and inference clusters. Type: shared NAS. Performance: 5-15 GB/s per export. Access: NFS v4.1/v4.2 over RDMA (NFSoRDMA). Cost: $40-$120/TB-month. Platforms: NFS server + RDMA NIC, NetApp, Dell PowerScale. Architecture, benchmarks, deployment, and cost optimization.

01

NFS over RDMA Architecture

NFS over RDMA is a shared NAS storage system providing NFS v4.1/v4.2 over RDMA (NFSoRDMA). The storage cluster delivers 5-15 GB/s per export of throughput through shared NAS architecture. Key design features include: NFS v4.1/v4.2 over RDMA (NFSoRDMA) protocol support, erasure coding or replication for data durability, and integration with major GPU cluster orchestration platforms.

02

Performance Benchmarks for AI

Performance benchmarks for NFS over RDMA: sequential read throughput of 5-15 GB/s per export with 1M+ IOPS; metadata operations at 50K-500K ops/second depending on namespace design; checkpoint write throughput of 30-70% of peak read due to parity/replication overhead. For checkpointing a 70B model (140 GB) with 32 GPUs: checkpoint save time of 2-10 seconds with GPUDirect Storage vs 20-60 seconds without.

03

Cost Analysis and Budgeting

NFS over RDMA costs $40-$120/TB-month on NFS server + RDMA NIC, NetApp, Dell PowerScale. For a 64-GPU cluster with 100 TB of active storage for training data and checkpoints: monthly storage cost of $2,000-$25,000 depending on performance tier. Total storage as percentage of cluster TCO: 10-30% for parallel filesystems, 5-15% for object storage, 2-5% for local NVMe.

04

GPU Cluster Integration

Integration with GPU clusters: Kubernetes CSI driver for container-native storage access; Slurm job integration with filesystem mounts; GPUDirect Storage support for direct GPU<->filesystem DMA; and data lifecycle policies for hot/warm/cold tiering. Training frameworks (PyTorch, NeMo, Megatron-LM) access NFS over RDMA via standard filesystem interfaces or optimized data loaders.

05

Deployment Patterns

Deployment of NFS over RDMA for AI clusters: dedicated storage cluster with GPU-direct networking; all-flash NVMe storage nodes for maximum throughput; storage nodes connected via InfiniBand or high-speed Ethernet to compute cluster; and data staging from object store to parallel filesystem for training jobs. For multi-cluster deployments, global namespace and cross-region data replication are critical.

06

Recommendations and Best Practices

Best practices for NFS over RDMA: provision 30-50% more throughput than peak checkpoint load to avoid training stalls; separate checkpoint and dataset storage for independent scaling; implement data caching on local NVMe to reduce parallel filesystem load; use GPUDirect Storage for 2-5x checkpoint speedup; monitor filesystem metadata performance separately from data throughput; and plan for 50% annual capacity growth.

Filed under
NFS over RDMA GPU StorageAI Storage NFS over RDMAGPU Cluster Storage NFS over RDMATraining Storage NFS over RDMAParallel Filesystem NFS over RDMA