What a GPU Actually Does for AI
A GPU is a processor designed for parallel computation. Where a CPU has 8-32 powerful cores optimized for sequential logic, a GPU packs thousands of simpler cores that execute the same instruction across many data points simultaneously. This maps directly to the matrix multiplications at the heart of neural network training and inference.
The distinction matters because AI workloads are not general-purpose computing. A single forward pass through a 70B-parameter model requires roughly 140 petaflops of matrix math. A CPU would take hours. A modern GPU does it in seconds.
VRAM, Memory Bandwidth, and TFLOPS
Three numbers dominate every GPU spec sheet. VRAM is the amount of high-speed memory on the card, measured in GB. It determines the largest model you can load without offloading to CPU or NVMe. Memory bandwidth, measured in GB/s, determines how fast the GPU can feed data to its compute cores. TFLOPS measures the raw math throughput of those cores.
These three interact directly. A GPU with 80 GB VRAM can hold a 70B model in FP16 (roughly 140 GB of parameters and KV cache), but if memory bandwidth is low, the cores starve and utilization drops. The balance between capacity, bandwidth, and compute is what separates usable GPUs from impressive paper specs.
| Spec | What it measures | Why it matters |
|---|---|---|
| VRAM (GB) | On-card memory capacity | Largest model you can fit |
| Memory bandwidth (GB/s) | Data transfer rate to cores | How fast the GPU stays fed |
| TFLOPS (FP16) | Peak math throughput | Raw compute speed |
NVIDIA GPU Generations for AI
Four NVIDIA GPU generations dominate AI infrastructure in 2026. The A100 (Ampere, 2020) started the wave with 80 GB HBM2e at 2 TB/s and 312 TFLOPS FP16. The H100 (Hopper, 2022) jumped to 80 GB HBM3 at 3.35 TB/s and 989 TFLOPS FP16 with FP8 and Transformer Engine. The H200 (Hopper refresh, 2024) kept the same architecture but swapped to 141 GB HBM3e at 4.8 TB/s. The B200 (Blackwell, 2025) delivers 192 GB HBM3e at 8 TB/s and 2.25 PFLOPS FP8.
Each generation roughly doubles compute throughput and improves memory bandwidth by 40-60%. The A100 is now the budget option for inference-only workloads that can tolerate FP16. The H100 remains the workhorse for training. The B200 is the premium choice for large-scale pre-training where every day of GPU time costs tens of thousands of dollars.
| GPU | VRAM | Bandwidth | FP16 TFLOPS | Release |
|---|---|---|---|---|
| A100 | 80 GB HBM2e | 2.0 TB/s | 312 | 2020 |
| H100 | 80 GB HBM3 | 3.35 TB/s | 989 | 2022 |
| H200 | 141 GB HBM3e | 4.8 TB/s | 989 | 2024 |
| B200 | 192 GB HBM3e | 8.0 TB/s | 2,250 | 2025 |
Tensor Cores and Precision Formats
NVIDIA GPUs contain dedicated tensor core hardware that performs fused multiply-add operations on 4x4 matrices in a single clock cycle. These tensor cores support multiple precision formats including FP32, TF32, FP16, BF16, FP8, and INT8. The choice of precision directly trades accuracy for speed: FP8 runs roughly 4x faster than FP16 on the same hardware but requires careful scaling to avoid numerical drift.
The B200 adds FP4 and FP6 support, enabling inference at precision levels that were previously only achievable with quantization after training. This means a B200 can serve a 70B model at 4-bit precision from a single GPU, achieving throughput that would have required an 8-GPU H100 cluster in 2024.
Training vs. Inference GPU Requirements
Training demands sustained compute across hundreds or thousands of GPUs for weeks. The bottleneck is typically memory bandwidth during forward and backward passes, and inter-GPU communication bandwidth during gradient synchronization. Training benefits most from high memory bandwidth and fast interconnects like NVLink and InfiniBand.
Inference is more varied. Latency-sensitive applications (chat, real-time API) need fast single-GPU throughput with minimal batching. Throughput-optimized inference (batch processing, offline) benefits from large VRAM to maximize batch sizes and reduce per-token cost. Most inference workloads are memory-bandwidth-bound, not compute-bound.
How GPUs Talk to Each Other
A single GPU is rarely enough for modern models. Multi-GPU systems use NVLink for intra-node GPU communication (900 GB/s on H100 SXM, 1.8 TB/s on B200) and InfiniBand or Spectrum-X Ethernet for inter-node communication (400-800 Gb/s per link). The interconnect topology determines how efficiently you can scale training across nodes.
The rule of thumb: each GPU in a training cluster should be able to communicate with every other GPU at a latency under 10 microseconds per message. Any slower, and the gradient synchronization overhead eats into compute utilization. This is why SXM form factors with full NVLink dominate training clusters while PCIe GPUs with Ethernet-only connectivity are limited to inference and small-scale training.
The GPU Procurement Reality in 2026
In 2026, lead times for B200 clusters range from 36 to 52 weeks. H100 and H200 are available on secondary markets at a 30-50% discount from their 2024 peak prices, but the supply of B200 is constrained by CoWoS packaging capacity and HBM3e allocation. Renting from hyperscalers or GPU brokers is the only way to get B200 capacity in weeks instead of months.
The practical takeaway: match GPU generation to workload criticality. If you are pre-training a frontier model, B200 is worth the wait and the premium. If you are fine-tuning or serving inference, H100 and H200 clusters are available now and deliver 80-90% of the performance at 50-60% of the cost. The best GPU is the one you can actually get running today.
