All essays
TechnicalDEEP DIVEFEB 2026

GPU Memory Pooling and Disaggregation: CXL and NVLink Shared Memory Architectures

A technical deep dive into GPU memory pooling and disaggregation using CXL and NVLink, covering shared memory pool architectures, performance trade-offs, and deployment use cases.

01

The Memory Wall Problem

GPU compute capacity has grown faster than GPU memory capacity. An H100 SXM delivers 1979 TFLOPS of FP8 compute with 141GB HBM3e. A B300 delivers 5482 TFLOPS FP8 with 288GB HBM3e. Over three GPU generations, compute throughput has increased 6.2x while memory capacity has only increased 2x. This growing imbalance means large model inference is increasingly memory-bound: GPUs have compute cycles that cannot be utilized because model weights and KV cache cannot fit in local HBM.

Memory pooling and disaggregation address this by allowing GPUs to access memory beyond their local HBM. The two dominant technologies are CXL (Compute Express Link), an open industry standard for cache-coherent memory expansion over PCIe, and NVLink, NVIDIA's proprietary high-bandwidth GPU interconnect. Both enable shared memory pools, but they operate at different bandwidth and latency regimes suitable for different workload types.

02

CXL Memory Pooling

CXL 3.0 enables pooled memory systems where multiple GPUs share a common memory pool over PCIe Gen 6 x16 links. Each link provides approximately 64 GB/s of bandwidth with sub-microsecond latency. A CXL memory expander with 8 DRAM slots can host up to 2TB of DDR5 memory, shared across 4-16 GPU nodes. The pooled memory appears as coherent system memory, allowing GPUs to fault in pages on demand from the shared pool.

The practical bandwidth limitation of CXL (64 GB/s vs GPU HBM bandwidth of 4-8 TB/s) restricts CXL pool use to cold memory: infrequently accessed model weights, long-tail KV cache entries, and checkpoints. NVIDIA's Dynamo framework uses CXL for tier-2 KV cache offloading, achieving 85% hit rates on the local HBM tier while the remaining 15% is served from CXL-pooled DRAM at 200-500 microsecond latency overhead. For inference workloads with prefix-heavy traffic patterns, this adds only 3-5% to P99 latency.

04

Disaggregated Memory Architectures

The most advanced deployments combine CXL and NVLink in a three-tier memory hierarchy. Tier 1 is local HBM (288GB per B300, 8 TB/s bandwidth). Tier 2 is NVLink-pooled HBM from neighboring GPUs (2TB+ across 8 GPUs, 900 GB/s per GPU). Tier 3 is CXL-pooled DDR5 (16TB+ per rack, 64 GB/s per CXL link). Software frameworks automatically page memory between tiers based on access frequency, with migration triggered at cache-line granularity.

This three-tier hierarchy effectively extends the memory-constrained GPU footprint. A B300 node with 8 GPUs traditionally has 2.3TB of total HBM. With NVLink pooling, the same 8 GPUs can address the full 2.3TB as a unified pool. With CXL expansion, the total addressable memory reaches 10-20TB per node. For inference of models above 500B parameters, this disaggregated approach reduces the GPU count needed by 40-60% compared to a non-pooled deployment.

Memory TierTechnologyCapacity per NodeBandwidthLatencyUse Case
Tier 1 (hot)HBM3e (local GPU)288 GB8 TB/s~100 nsActive weights, KV cache
Tier 2 (warm)NVLink-pooled HBM2.3 TB900 GB/s~1 usModel shards, optimizer states
Tier 3 (cold)CXL-pooled DDR516 TB64 GB/s~300 nsCold KV cache, checkpoints
Tier 4 (archive)NVMe over fabric1 PB+8 GB/s~10 usDataset storage, model versions
05

Performance Trade-offs

Memory pooling introduces NUMA effects even within a single node. A GPU accessing local HBM (Tier 1) gets full 8 TB/s bandwidth. Accessing a neighbor GPU's HBM via NVLink (Tier 2) gets 450-900 GB/s, a 5-10x bandwidth reduction. Accessing CXL-pooled memory (Tier 3) gets 64 GB/s, a 125x reduction versus local HBM. Workloads must be aware of these tiers to place memory accesses appropriately.

The practical solution is page migration at the OS or driver level. Linux kernel 6.8 introduced CXL memory tiering support, and NVIDIA's CUDA 13 driver stack includes automatic page migration between NVLink and local HBM. In production testing, automatic tiering achieves 92-96% of the performance of manually optimized memory placement, suggesting that most workloads can adopt memory pooling without application code changes. The remaining 4-8% performance gap comes from cold-start page faults that access CXL memory before the hot tier warms up.

06

Use Cases

Long-context inference is the killer application for GPU memory pooling. Serving 1M context windows on Llama 4 Maverick requires 1.5TB of KV cache per request. Without pooling, this requires 6 B300s dedicated to KV cache storage. With NVLink+CXL pooling, the KV cache expands into Tier 2 and Tier 3 memory, reducing the GPU count to 2 B300s while maintaining 100ms P99 token generation. The cost savings are 60-70% for long-context inference workloads.

Training checkpointing also benefits significantly. A checkpoint for a 70B model at FP8 weights with optimizer states is approximately 400GB. Writing to NVMe takes 20-30 seconds. Writing to CXL-pooled DDR5 takes under 1 second. Asynchronous checkpointing to CXL memory with background flush to NVMe reduces training idle time from 30 seconds to under 2 seconds per checkpoint, recovering 3-5% of training throughput for models that checkpoint every 30 minutes.

07

Market Timeline

CXL 3.0 memory pooling is shipping in production today. AMD's EPYC Turin and Intel's Granite Rapids support CXL 3.0 with PCIe Gen 5 x16, and CXL memory expanders from Samsung, Micron, and SK Hynix are available in volume. NVIDIA's NVLink domain pooling is available on B300 NVL systems with NVLink Switch. The first fully disaggregated memory systems combining CXL and NVLink are being deployed in early 2026 by major cloud providers.

The next milestone is CXL 4.0 (expected 2027), which increases per-link bandwidth to 128 GB/s and adds support for CXL switches enabling rack-scale memory pools shared across 100+ nodes. NVIDIA's Rubin architecture is also expected to introduce NVLink 6 with 2.5 TB/s per GPU, further blurring the line between local and remote GPU memory. By 2028, memory pooling and disaggregation will be a standard feature of all GPU clusters above 100 GPUs, not a premium add-on.

Filed under
CXL memory poolingNVLink shared memorymemory disaggregationGPU memory fabricHBM3eCXL 3.0NVLink Switchdisaggregated infrastructure