THE VRAM WALL
An H100 SXM has 80 GB HBM3. A 175B model in FP16 needs 350 GB, requiring 5 H100s. Even FP8 needs 3. Model sizes outpace per-GPU memory - MoE models at 1T+ parameters need 2 TB+. Memory pooling lets GPUs transparently access remote memory with 200-500 ns added latency.
Key technologies: NVLink/NVSwitch (NVIDIA, 900 GB/s), CXL (open standard, 64 GB/s), InfiniBand with GPUDirect RDMA (50 GB/s).
| Technology | Bandwidth/Link | Latency Local | Latency Remote | Max Pool Size | Availability |
|---|---|---|---|---|---|
| NVLink 4.0 H100 | 900 GB/s | <100 ns | 200-300 ns | 8 GPUs | Shipping |
| NVLink 5.0 B100 | 1.8 TB/s | <100 ns | 150-250 ns | 8-16 GPUs | 2025+ |
| CXL 3.0 | 64-128 GB/s | ~100 ns | 500-1000 ns | Thousands | Early 2025 |
| InfiniBand NDR400 | 50 GB/s | ~100 ns | 1-3 us | Unlimited | Shipping |
NVIDIA NVLink-BASED DISAGGREGATION
NVLink 4.0 provides 900 GB/s within 8-GPU nodes. GPUDirect extends via InfiniBand but bandwidth drops to 50 GB/s. NVLink 5.0 on B100 promises 1.8 TB/s with cross-node connectivity without InfiniBand.
GH200 Grace Hopper delivers 1.4x inference throughput for 175B models versus H100+x86 due to 576 GB/s unified memory eliminating PCIe KV cache transfer.
CXL 3.0: OPEN STANDARD MEMORY POOLING
CXL 3.0 is an open standard (Intel, AMD, ARM) providing 64 GB/s per link. CXL memory costs $2-3/GB/month versus $20-30/GB for HBM3. Suitable for KV cache overflow (40-80 GB for 128K context) but not for model weights due to 3-10x latency premium.
First production platforms arrived late 2025. Reported latency: 300-500 ns local CXL, 600-1000 ns pooled. Still niche for AI as of early 2026.
DISAGGREGATED INFERENCE
Separating pre-fill (compute-bound) and decode (memory-bandwidth-bound) phases yields 1.6x improvement on same GPUs. A 4-H100 deployment achieves 65 tok/s disaggregated vs 40 tok/s standard, with zero additional GPU cost.
TensorRT-LLM 0.10+ supports this pattern. Engineering cost: 4-6 weeks.
| Architecture | GPU Count (70B) | Tokens/sec | Cost/1M Tokens | KV Cache/GPU | Pre-fill Latency |
|---|---|---|---|---|---|
| Standard TP=4 | 4 H100 | 40 | $2.10 | 10 GB | 180 ms |
| Disaggregated 2+2 | 4 H100 | 65 | $1.29 | 5 GB | 90 ms |
| Standard + CXL KV | 4 H100+512GB CXL | 45 | $1.95 | 2 GB+CXL | 180 ms |
| Memory-pooled 8 GPU | 8 H100 NVLink | 85 | $0.99 | 4 GB shared | 85 ms |
FUTURE LANDSCAPE
HBM4 (2026-2027) will provide 256 GB per GPU at 2.0+ TB/s. NVLink 6.0 targets 3.6 TB/s for 64-GPU pools. CXL 4.0 (2027) targets 256 GB/s with sub-200 ns latency.
Early adopters report 20-35 percent per-inference cost reduction for large-context models. Pooled memory expected to become default by 2028.
