All essays
TechnicalDEEP DIVEFEB 2026

GPU Memory Pooling & Disaggregation: Architectures for AI Clusters

GPU memory pooling and disaggregation architectures for AI clusters. Unified memory fabrics, NVLink memory sharing, and remote GPU memory access.

01

The GPU Memory Crisis: Why 80GB HBM Is No Longer Enough

In mid-2026, the most widely deployed GPU for AI training and inference is the H100 SXM5 with 80GB of HBM3. When that hardware was designed in 2022, 80GB per GPU was generous for the models of the era - GPT-3 175B at FP16 needed roughly 350GB of total memory, fitting across 5-8 H100 GPUs with tensor parallelism. Two model generations later, Llama 4 Maverick (400B parameters), Llama 3.1 405B, and Qwen 3 300B push per-GPU memory requirements far beyond 80GB. A single Llama 3.1 405B inference instance at FP16 requires roughly 810GB of total GPU memory - at least 11 H100 GPUs, requiring pipeline parallelism across multiple nodes. The memory capacity gap between available HBM and model requirements is widening with each model release.

The market response has been predictable: NVIDIA filled the gap with H200 (141GB HBM3e, launched as a Hopper refresh) and B200 (192GB HBM3e on Blackwell), giving teams the option of larger per-GPU memory without changing their cluster architecture. But these upgrades are expensive. H200 carries a 20-30% premium over H100 on most markets, and B200 is roughly 3x the per-hour cost. Memory pooling and disaggregation offer an alternative: instead of buying next-generation GPUs with larger HBM, connect existing H100 GPUs to a shared memory fabric that any GPU can access, effectively pooling their individual 80GB HBM into a larger virtual memory space.

Memory disaggregation takes this concept further by physically separating GPU compute from GPU memory. In a disaggregated architecture, GPU compute units and HBM pools reside in separate chassis connected by a high-bandwidth fabric. Compute GPUs access memory from remote HBM pools as needed, with the fabric providing transparent remote memory access. This enables independent scaling of compute and memory capacity - a team needing 160GB of usable GPU memory for a large model but only moderate compute throughout can provision one compute GPU and two memory pools, rather than two full compute GPUs with underutilized compute units.

03

Disaggregated GPU Memory: The Economics and Architecture

A disaggregated GPU memory architecture consists of compute pods (GPU compute units with minimal local HBM, perhaps 20-40GB per GPU) connected through a high-bandwidth memory fabric to memory pods (dedicated HBM pools without compute units, each providing 100-500GB of HBM accessible over the fabric). The compute-to-memory ratio can be tuned per workload: a memory-intensive inference workload might connect to a 4:1 memory-to-compute ratio (4 memory pods for every compute pod), while a compute-intensive training workload might use a 1:2 ratio.

The economic case for disaggregation rests on HBM utilization. In a traditional GPU cluster, each GPU's HBM is dedicated to the workload running on that GPU. Utilization varies widely: a training job using 40GB of 80GB available leaves 50% of HBM idle. Across a 256-GPU cluster, the aggregate stranded memory capacity can reach 5-10 TB - equivalent to 25-50 additional H100 GPUs worth of HBM. Disaggregation pools that stranded memory and makes it available to workloads on any compute GPU, improving aggregate memory utilization to 80-90%.

The technical challenge of disaggregation is the fabric latency overhead. A disaggregated memory access over a fabric (CXL, NVLink over optical, or custom interconnects) adds 1-10 microseconds of latency compared to local HBM access at 200-400 nanoseconds. For inference workloads where memory access is on the critical path for every decode step, a 5us additional latency per memory access can degrade ITL by 15-30%. Disaggregation is practical primarily for workloads where memory access latency is not the bottleneck: training workloads with large batch sizes, batch inference, and data preprocessing.

04

CXL and Emerging Standards for Memory Disaggregation

CXL (Compute Express Link) is the leading industry standard for cache-coherent memory disaggregation. CXL 3.0 supports memory pooling across up to 4,096 devices with sub-microsecond latency over PCIe 6.0 physical links. CXL memory expanders - devices that provide DRAM or persistent memory accessible over CXL - are entering the market in 2025-2026, offering 256GB-2TB of memory per expander at $2-5/GB, significantly cheaper than HBM at roughly $20-40/GB for H100-class GPUs.

The CXL approach to GPU memory disaggregation: use CXL memory expanders as a slow-but-large memory tier that backs the GPU's fast-but-small HBM. The GPU's local HBM functions as a cache for the larger CXL-attached memory pool. Hot pages (frequently accessed model weights, KV cache entries) reside in HBM for fast access. Cold pages (infrequently accessed parameters, long-tail KV cache entries) migrate to the CXL memory tier. The cache coherence protocol ensures that the GPU sees a consistent view of memory regardless of where the physical page resides.

The latency hierarchy in a disaggregated GPU cluster: L1 is GPU register file and shared memory (access in 10-30 cycles), L2 is local HBM (200-400ns, ~800-1,600 cycles), L3 is remote GPU HBM over NVLink (3-5us, ~12,000-20,000 cycles), L4 is CXL-attached memory expander (5-15us, ~20,000-60,000 cycles), and L5 is CPU system memory (100-200ns for local NUMA, 200-500ns for remote NUMA, but the PCIe bridge adds latency for GPU access). Effective memory pooling requires workload-aware page placement across this hierarchy, with migration policies that move pages between tiers based on access frequency.

05

Which Workloads Benefit from Memory Pooling and Which Do Not

Memory pooling benefits are workload-dependent. The workloads that benefit most are those with high peak memory requirements but low average memory utilization: large model inference with variable batch sizes, multi-tenant inference where models are loaded and unloaded dynamically, and training runs that use memory-intensive techniques like gradient checkpointing or large KV caches. For these workloads, pooled memory reduces the number of GPUs needed to handle peak memory requirements.

Workloads that do not benefit from memory pooling: single-GPU inference of models that fit in a single GPU's HBM (models up to 40B parameters at FP8 on H100, or up to 100B parameters on B200), training runs that fully utilize each GPU's HBM continuously (a well-tuned FSDP training run saturates HBM with model parameters, optimizer states, and gradients), and latency-sensitive inference where any additional memory access latency degrades the user experience. For these workloads, the cost of the memory fabric interconnect does not pay for itself.

The decision framework for memory pooling: if your workload's peak memory requirement exceeds a single GPU's HBM by more than 50%, memory pooling may reduce the total GPU count needed. If your workload's memory utilization across GPUs varies by more than 30% (some GPUs near capacity while others are half-empty), pooling can level the utilization. If your workload has predictable memory access patterns (sequential layer-by-layer model execution, flat KV cache access), software pooling via DeepSpeed ZeRO or UVM is sufficient without custom hardware. If your access pattern is random and latency-sensitive, only hardware-transparent NVLink 5 pooling or similar is viable.

06

Implementing GPU Memory Pooling: Practical Configurations

The simplest memory pooling configuration for H100 clusters: use CUDA unified memory with advice flags (cudaMemAdviseSetAccessedBy and cudaMemAdviseSetPreferredLocation) to guide the UVM driver's page migration decisions. Configure the UVM fault handler to prefer remote GPU HBM over CPU system memory for secondary pages. This software-only approach requires no additional hardware and works on any H100 cluster with NVLink connectivity. The performance benefit is workload-dependent but typically reduces peak memory usage by 30-50% for inference workloads with variable batch sizes.

For B200 clusters with NVLink 5 fabric memory, the configuration is hardware-transparent. Enable fabric-attached memory through CUDA's memory pool API. Configure the pool size as a percentage of the total domain memory (e.g., 50% of total 13.8 TB in a 72-GPU domain reduces to 6.9 TB of poolable memory). Set the allocation policy to favor local GPU HBM first and allocate from the pool only when local HBM is exhausted. Monitor pool allocation rates to determine whether the pool size configuration matches workload requirements.

The advanced configuration for CXL-based disaggregation: deploy CXL memory expanders on the same PCIe fabric as the GPU compute nodes. Configure the operating system's memory tiering subsystem (Linux kernel 6.5+ with tiered memory support) to define the HBM as the fast tier and CXL memory as the slow tier. Set the page promotion threshold (hot page migration to HBM) to trigger when a page is accessed more than 5 times per second, and the demotion threshold to trigger when a page is not accessed for 10 seconds. Monitor the promotion and demotion rates to verify that the tiering policy matches workload access patterns. If promotion rates exceed 1,000 pages per second per GPU, increase the HBM cache size or reduce the workload's working set.

07

The Future of GPU Memory: Pooled, Disaggregated, and Composable

The trajectory of GPU memory architecture is toward fully composable systems where compute, memory, and storage are independent resource pools connected by a unified fabric. NVIDIA's roadmap beyond Blackwell points to a multi-year transition where GPU compute modules and HBM stacks are physically separated and connected through optical interconnects. The GB200 NVL72 rack-scale design is the first step toward this architecture, consolidating 72 GPUs and their 13.8 TB of memory into a single logical domain.

The impact on GPU cluster design will be significant. Instead of buying fixed-ratio GPU configurations (8 GPUs per node with 80GB each, rigidly coupled), teams will specify their compute requirements (TFLOPS) and memory requirements (GB) independently and the composable fabric will assemble the appropriate configuration dynamically. A cluster with 1,000 compute modules and 5,000 memory modules could support workloads ranging from a single 4,000-parameter model requiring 2 PB of GPU memory to 200 small inference jobs each using 1 compute module and 1 memory module.

The timeline for mainstream disaggregated GPU memory is 2027-2029. For teams planning GPU infrastructure in mid-2026, the practical steps: adopt GPU memory pooling (NVLink 5 or UVM) to get comfortable with pooled memory architectures, invest in workload profiling that measures per-GPU memory utilization independently from compute utilization, and design cluster networking with the bandwidth and latency characteristics needed for future disaggregation. The cluster networking decisions made in 2026 (fabric topology, cable plant, switch selection) will determine whether composable memory architectures can be adopted in 2028 without a networking overhaul.

Filed under
Memory PoolingDisaggregated ArchitectureNVLink MemoryUnified MemoryGPU MemoryHBM CapacityInference Optimization