All essays
TechnicalDEEP DIVEFEB 2026

Grace Blackwell GH200: When Unified CPU-GPU Memory Changes Your AI Architecture

GH200 combines 72-core Grace CPU and H100 GPU with 900GB/s NVLink-C2C and 624GB unified memory. Which workloads benefit from this architecture.

01

GH200 ARCHITECTURE OVERVIEW

NVIDIA Grace Blackwell GH200 pairs a 72-core ARM Grace CPU with an H100 GPU connected via NVLink-C2C at 900 GB/s, providing 624 GB of coherent unified memory. The Grace CPU contributes 480 GB of LPDDR5X memory at 500 GB/s bandwidth, while the H100 contributes 144 GB of HBM3 at 3.35 TB/s. Coherent memory access across the NVLink-C2C link means CPU and GPU can transparently access each other's memory without explicit data copies.

02

WORKLOAD FIT ANALYSIS

GH200 excels at workloads that require frequent CPU-GPU data movement or large in-memory datasets that exceed H100's 80 GB HBM capacity. Graph neural networks, genome analysis, and recommendation systems that process multi-gigabyte embedding tables benefit significantly. Pure training workloads with well-batched small-batch sizes see minimal benefit since the HBM on H100 already provides sufficient capacity. The 2-5x advantage appears in memory-bound workloads where CPU-GPU data transfer is the bottleneck.

03

INFERENCE USE CASES

GH200 shines in large-model inference where KV cache demands exceed HBM capacity. Serving Llama 4 405B with 64K context requires approximately 180 GB for KV cache under batching, which cannot fit on a single H100 but fits comfortably in GH200 unified memory. The 900 GB/s NVLink-C2C provides sufficient bandwidth for KV cache spills to Grace's LPDDR5X memory. Early benchmarks show GH200 achieves 95-98% of H100 per-token latency for 405B parameter models while supporting 2-3x larger batch sizes.

04

TRAINING APPLICATIONS

For training, GH200's value proposition is running models that exceed single-GPU HBM capacity. Models with large embedding tables, such as recommendation transformers with 200B+ parameter embeddings, train 30-50% faster on GH200 than on CPU-GPU systems with PCIe interconnects. However, for standard transformer pre-training, a DGX H100 with NVLink-connected H100s remains more cost-effective. GH200 shines specifically when memory capacity, not compute, is the primary constraint.

05

COST COMPARISON

GH200 systems cost approximately $2.50-3.50 per hour on the spot market versus $2.00-2.80 for standard H100. The premium of 25-35% is justified when unified memory enables workloads that cannot run on standard H100. On a per-GB-of-memory basis, GH200 at $5.61/GB-hr is substantially cheaper than H100 at $35.00/GB-hr when considering total system memory. For memory-bound workloads, GH200 often delivers 2-3x better cost efficiency.

ConfigurationGPU MemoryCPU MemoryTotal Unified$/hr$/GB-hr
H100 80 GB SXM80 GB0 GB80 GB$2.50$31.25
A100 80 GB80 GB0 GB80 GB$1.50$18.75
GH200144 GB480 GB624 GB$3.50$5.61
B200192 GB0 GB192 GB$4.50$23.44
06

PROCUREMENT GUIDANCE

Reserve GH200 only for workloads that demonstrably benefit from unified memory. Run a 7-day benchmark comparing GH200 versus standard H100 for your specific workload before committing to multi-month reservations. GH200 is ideal for recommendation systems, large embedding models, and extreme-context-length inference. For standard LLM training and inference, reserve H100 or B200 instead and route GH200 savings toward memory-bound workloads.

Filed under
GH200Grace BlackwellUnified MemoryNVLink C2CCPU-GPUH100ARM AI