GOOGLE CLOUD'S DUAL-ACCELERATOR STRATEGY
Google Cloud is the only major cloud provider that operates both NVIDIA GPUs and its own custom TPUs (Tensor Processing Units). This dual-accelerator strategy is both a strength and a source of complexity. TPUs are designed for Google's internal workloads (Search, YouTube, Gemini training) and are optimized for large-scale, synchronous distributed training. GPUs serve the broader market and provide compatibility with the full PyTorch/TensorFlow ecosystem. Google's bet is that customers will use TPUs for training (where JAX and TensorFlow XLA compilation provide 1.5-2x better price-performance than GPUs) and GPUs for inference (where the wider model ecosystem matters more than raw throughput).
As of Q2 2026, Google Cloud operates approximately 50,000 NVIDIA H100 GPUs and 15,000 H200 GPUs, plus an estimated 8,000 TPU v5p chips (each TPU v5p pod has 8,960 chips, so approximately one pod worth). GPU capacity is available in 14 regions, TPU capacity in 4 regions (us-central1, us-east1, europe-west4, and asia-east1). The fleet makes Google Cloud the second-largest GPU cloud by volume (behind AWS) and the largest TPU cloud by a wide margin.
| Accelerator | Type | Compute (FP16) | Memory | Interconnect | Best For | Price/hr |
|---|---|---|---|---|---|---|
| TPU v5p | Custom ASIC | 459 TFLOPS bf16 | 95 GB HBM2e | ICI (Inter-Core Interconnect) | Large-scale training (JAX/Flax) | $4.50 (per chip) |
| TPU v5e | Custom ASIC | 197 TFLOPS bf16 | 16 GB HBM2e | ICI | Mid-size training, inference | $1.20 (per chip) |
| H100 SXM | NVIDIA GPU | 1,979 TFLOPS sp | 80 GB HBM3 | NVLink + InfiniBand | Training + inference (PyTorch) | $3.55-$4.00 |
| H200 SXM | NVIDIA GPU | 1,979 TFLOPS sp | 141 GB HBM3e | NVLink + InfiniBand | Large context inference | $4.00-$4.50 |
| B200 SXM | NVIDIA GPU | 4,510 TFLOPS sp | 192 GB HBM3e | NVLink + InfiniBand | Next-gen training/inference | $5.50-$6.00 |
| L4 | NVIDIA GPU | 121 TFLOPS sp | 24 GB GDDR6 | PCIe Gen4 | Cost-effective inference | $0.60-$0.80 |
GKE FOR AI: KUBERNETES AS THE AI ORCHESTRATION LAYER
Google Cloud's orchestrator strategy centers on GKE (Google Kubernetes Engine) as the universal control plane for AI workloads. GKE supports GPU node pools with automatic GPU driver installation, node auto-scaling for GPU workloads, and Node Auto-Repair for GPU health monitoring. The service is deeply integrated with Google's Vertex AI, which provides managed ML tools that run on GKE under the hood. For teams that prefer to manage their own stack, GKE GPU clusters start from a standard GKE cluster with `--accelerator` flags for GPU nodes.
GKE's GPU support extends to advanced scheduling features including GPUs as an extended resource, MIG (Multi-Instance GPU) partition scheduling for A100/H100, and dynamic GPU allocation via the NVIDIA GPU Operator. GKE also supports GPUDirect-TCPX for optimized network performance between GPU pods. For teams that need to run multi-node distributed training, GKE supports the Kubeflow Training Operator (which manages PyTorchJob and TFJob custom resources) and Kueue for quota-based scheduling with job queuing. This stack is mature: Google has been running internal ML workloads on Kubernetes for longer than any other company.
TPU VS GPU: THE DECISION FRAMEWORK
The decision between TPUs and GPUs on Google Cloud depends on three factors: framework, scale, and model architecture. TPUs require JAX or TensorFlow: if your stack is PyTorch-first, TPUs add significant migration friction. TPUs excel at scale: a TPU v5p pod with 8,960 chips achieves 95+ percent scaling efficiency for large Transformer training at 100B+ parameters, while GPU clusters typically see 70-85 percent efficiency at comparable scale due to Ethernet-based inter-node communication overhead. However, TPUs underperform on inference for most model architectures because the JIT compilation overhead (3-10 minutes for initial compilation) makes dynamic batching and model swapping impractical.
For most teams, the practical recommendation is: use TPUs for training runs that exceed 1,000 GPU-hours per job and use JAX; use GPUs for everything else. Google's own Gemini models are trained entirely on TPUs, but their inference serving infrastructure uses GPUs. This internal precedent is informative: TPUs for training efficiency, GPUs for serving flexibility. Google's pricing reflects this: TPU v5p at $4.50/chip-hour is 20-30 percent more expensive per chip than H100, but achieves 30-50 percent more tokens per dollar for large JAX-based training runs.
| Workload | Recommended Accelerator | Rationale | Cost/Token (Training) | Cost/Token (Inference) |
|---|---|---|---|---|
| Small model training (<7B) | GPU (H100/L4) | Framework flexibility, fast iteration | $1.00-2.50/Mtok | $0.15-0.50/Mtok |
| Large model training (>70B) | TPU v5p | Superior scaling efficiency in JAX | $0.50-1.20/Mtok | N/A |
| MoE model training (8x22B+) | TPU v5p | TPU expert parallelism is more mature | $0.60-1.00/Mtok | N/A |
| Real-time inference | GPU (H100/L4) | Dynamic batching, ecosystem support | N/A | $0.10-0.80/Mtok |
| Batch inference | GPU (H200) | Large VRAM for batch processing | N/A | $0.05-0.30/Mtok |
PRICING MODEL AND COMMITMENT OPTIONS
Google Cloud's GPU pricing is similar to AWS and Azure for on-demand instances (H100 at $3.55-4.00/GPU-hour) but offers more flexible commitment options. Committed Use Contracts (1-year or 3-year) provide 20-40 percent discounts on GPU instances. Preemptible instances (equivalent to AWS Spot) offer 60-80 percent discounts with 24-hour max runtime. Google also offers Sustained Use Discounts: if you run an instance for more than 25 percent of the month, you automatically receive a discount that scales up to 30 percent for full-month usage, without any commitment.
The key pricing differentiator is custom machine types. GCP allows users to configure GPU nodes with custom CPU/RAM ratios rather than fixed instance shapes. For inference workloads where GPUs are the bottleneck, a custom node with minimum CPUs and RAM can reduce total cost by 15-25 percent versus predefined instance types. This flexibility is unique to GCP among the hyperscalers and is particularly valuable for high GPU-to-CPU workloads like batch inference and embedding generation.
| Commitment Type | H100 Discount | TPU v5p Discount | Min Term | Cancelable | Best For |
|---|---|---|---|---|---|
| On-Demand | 0% | 0% | N/A | Yes (per second) | Short-term, experiments |
| Sustained Use (auto) | Up to 30% | Up to 30% | 1+ month of usage | Auto-applied | Steady workloads |
| 1-Year CUD | 20-25% | 20-25% | 1 year | No | Moderate commitment |
| 3-Year CUD | 40-45% | 40-45% | 3 years | No | Production workloads |
| Preemptible | 60-80% | N/A (no preempt TPUs) | 1 sec | Provider can terminate | Fault-tolerant training |
REGIONAL GPU CAPACITY AND AVAILABILITY DYNAMICS
Google Cloud's GPU capacity is concentrated in us-central1 (Iowa), which hosts approximately 30 percent of GCP's GPU fleet. us-east1 (South Carolina) and us-west1 (Oregon) host another 30 percent combined. European regions (europe-west4 Netherlands, europe-west1 Belgium) host 20 percent, and APAC regions (asia-east1 Taiwan, asia-southeast1 Singapore) host the remaining 20 percent. GPU availability is highly correlated with region popularity-us-central1 often has capacity for urgent H100 workloads, while asia-southeast1 frequently shows limited availability for large GPU clusters.
During the 2023-2024 GPU shortage, GCP implemented a GPU quota increase request system that became a bottleneck: requests for 64+ GPUs required 2-4 weeks review. As of 2026, this has improved to 1-2 days for H100 quotas under 256 GPUs, but 512+ GPU allocations still require a sales conversation. GCP also reserves GPU capacity for Google's internal AI workloads (Gemini, DeepMind), which can compete with customer allocations during peak usage periods. Google has been transparent about this dynamic: they prioritize internal workload allocation at the beginning of each quarter, releasing remaining capacity to customers mid-quarter.
COMPETITIVE POSITION AND MARKET FIT
Google Cloud's GPU offering is strongest for teams that use JAX/TensorFlow, want tight integration with GKE and Vertex AI, or need TPU access for large-scale training. The custom machine type flexibility provides a unique cost optimization lever. Google's commitment to open-source AI infrastructure (Kubernetes, Kubeflow, TensorFlow, JAX) makes them the preferred cloud for teams that build their own AI platforms rather than using managed services.
The challenges are Google's higher list prices (matching AWS and Azure but above OCI and CoreWeave), the complexity of the TPU vs. GPU decision which can confuse less experienced teams, and capacity competition with Google's internal AI workloads. Google Cloud's GPU business is estimated at $2-3 billion in annual revenue, growing approximately 40-50 percent year-over-year, making it a significant but not dominant player in a market where AWS leads in absolute GPU volume.
