GH200 ARCHITECTURE OVERVIEW
NVIDIA Grace Blackwell GH200 pairs a 72-core ARM Grace CPU with an H100 GPU connected via NVLink-C2C at 900 GB/s, providing 624 GB of coherent unified memory. The Grace CPU contributes 480 GB of LPDDR5X memory at 500 GB/s bandwidth, while the H100 contributes 144 GB of HBM3 at 3.35 TB/s. Coherent memory access across the NVLink-C2C link means CPU and GPU can transparently access each other's memory without explicit data copies.
WORKLOAD FIT ANALYSIS
GH200 excels at workloads that require frequent CPU-GPU data movement or large in-memory datasets that exceed H100's 80 GB HBM capacity. Graph neural networks, genome analysis, and recommendation systems that process multi-gigabyte embedding tables benefit significantly. Pure training workloads with well-batched small-batch sizes see minimal benefit since the HBM on H100 already provides sufficient capacity. The 2-5x advantage appears in memory-bound workloads where CPU-GPU data transfer is the bottleneck.
INFERENCE USE CASES
GH200 shines in large-model inference where KV cache demands exceed HBM capacity. Serving Llama 4 405B with 64K context requires approximately 180 GB for KV cache under batching, which cannot fit on a single H100 but fits comfortably in GH200 unified memory. The 900 GB/s NVLink-C2C provides sufficient bandwidth for KV cache spills to Grace's LPDDR5X memory. Early benchmarks show GH200 achieves 95-98% of H100 per-token latency for 405B parameter models while supporting 2-3x larger batch sizes.
TRAINING APPLICATIONS
For training, GH200's value proposition is running models that exceed single-GPU HBM capacity. Models with large embedding tables, such as recommendation transformers with 200B+ parameter embeddings, train 30-50% faster on GH200 than on CPU-GPU systems with PCIe interconnects. However, for standard transformer pre-training, a DGX H100 with NVLink-connected H100s remains more cost-effective. GH200 shines specifically when memory capacity, not compute, is the primary constraint.
COST COMPARISON
GH200 systems cost approximately $2.50-3.50 per hour on the spot market versus $2.00-2.80 for standard H100. The premium of 25-35% is justified when unified memory enables workloads that cannot run on standard H100. On a per-GB-of-memory basis, GH200 at $5.61/GB-hr is substantially cheaper than H100 at $35.00/GB-hr when considering total system memory. For memory-bound workloads, GH200 often delivers 2-3x better cost efficiency.
| Configuration | GPU Memory | CPU Memory | Total Unified | $/hr | $/GB-hr |
|---|---|---|---|---|---|
| H100 80 GB SXM | 80 GB | 0 GB | 80 GB | $2.50 | $31.25 |
| A100 80 GB | 80 GB | 0 GB | 80 GB | $1.50 | $18.75 |
| GH200 | 144 GB | 480 GB | 624 GB | $3.50 | $5.61 |
| B200 | 192 GB | 0 GB | 192 GB | $4.50 | $23.44 |
PROCUREMENT GUIDANCE
Reserve GH200 only for workloads that demonstrably benefit from unified memory. Run a 7-day benchmark comparing GH200 versus standard H100 for your specific workload before committing to multi-month reservations. GH200 is ideal for recommendation systems, large embedding models, and extreme-context-length inference. For standard LLM training and inference, reserve H100 or B200 instead and route GH200 savings toward memory-bound workloads.
