All essays
TechnicalDEEP DIVEFEB 2026

GPU Memory Pooling and Disaggregation Architectures

GPU memory pooling and disaggregation architectures. Compare NVLink, CXL, and InfiniBand-based memory sharing for H100 and B100 GPU clusters.

01

THE VRAM WALL

An H100 SXM has 80 GB HBM3. A 175B model in FP16 needs 350 GB, requiring 5 H100s. Even FP8 needs 3. Model sizes outpace per-GPU memory - MoE models at 1T+ parameters need 2 TB+. Memory pooling lets GPUs transparently access remote memory with 200-500 ns added latency.

Key technologies: NVLink/NVSwitch (NVIDIA, 900 GB/s), CXL (open standard, 64 GB/s), InfiniBand with GPUDirect RDMA (50 GB/s).

TechnologyBandwidth/LinkLatency LocalLatency RemoteMax Pool SizeAvailability
NVLink 4.0 H100900 GB/s<100 ns200-300 ns8 GPUsShipping
NVLink 5.0 B1001.8 TB/s<100 ns150-250 ns8-16 GPUs2025+
CXL 3.064-128 GB/s~100 ns500-1000 nsThousandsEarly 2025
InfiniBand NDR40050 GB/s~100 ns1-3 usUnlimitedShipping
03

CXL 3.0: OPEN STANDARD MEMORY POOLING

CXL 3.0 is an open standard (Intel, AMD, ARM) providing 64 GB/s per link. CXL memory costs $2-3/GB/month versus $20-30/GB for HBM3. Suitable for KV cache overflow (40-80 GB for 128K context) but not for model weights due to 3-10x latency premium.

First production platforms arrived late 2025. Reported latency: 300-500 ns local CXL, 600-1000 ns pooled. Still niche for AI as of early 2026.

04

DISAGGREGATED INFERENCE

Separating pre-fill (compute-bound) and decode (memory-bandwidth-bound) phases yields 1.6x improvement on same GPUs. A 4-H100 deployment achieves 65 tok/s disaggregated vs 40 tok/s standard, with zero additional GPU cost.

TensorRT-LLM 0.10+ supports this pattern. Engineering cost: 4-6 weeks.

ArchitectureGPU Count (70B)Tokens/secCost/1M TokensKV Cache/GPUPre-fill Latency
Standard TP=44 H10040$2.1010 GB180 ms
Disaggregated 2+24 H10065$1.295 GB90 ms
Standard + CXL KV4 H100+512GB CXL45$1.952 GB+CXL180 ms
Memory-pooled 8 GPU8 H100 NVLink85$0.994 GB shared85 ms
05

FUTURE LANDSCAPE

HBM4 (2026-2027) will provide 256 GB per GPU at 2.0+ TB/s. NVLink 6.0 targets 3.6 TB/s for 64-GPU pools. CXL 4.0 (2027) targets 256 GB/s with sub-200 ns latency.

Early adopters report 20-35 percent per-inference cost reduction for large-context models. Pooled memory expected to become default by 2028.

Filed under
GPU Memory PoolingMemory DisaggregationNVLinkCXLH100 MemoryB100 GPU