All essays
MarketMARKET REPORTFEB 2026

Google Cloud GPU Infrastructure Deep Dive: TPU + NVIDIA Strategy, GKE for AI, and Pricing Model

Comprehensive analysis of Google Cloud

01

GOOGLE CLOUD'S DUAL-ACCELERATOR STRATEGY

Google Cloud is the only major cloud provider that operates both NVIDIA GPUs and its own custom TPUs (Tensor Processing Units). This dual-accelerator strategy is both a strength and a source of complexity. TPUs are designed for Google's internal workloads (Search, YouTube, Gemini training) and are optimized for large-scale, synchronous distributed training. GPUs serve the broader market and provide compatibility with the full PyTorch/TensorFlow ecosystem. Google's bet is that customers will use TPUs for training (where JAX and TensorFlow XLA compilation provide 1.5-2x better price-performance than GPUs) and GPUs for inference (where the wider model ecosystem matters more than raw throughput).

As of Q2 2026, Google Cloud operates approximately 50,000 NVIDIA H100 GPUs and 15,000 H200 GPUs, plus an estimated 8,000 TPU v5p chips (each TPU v5p pod has 8,960 chips, so approximately one pod worth). GPU capacity is available in 14 regions, TPU capacity in 4 regions (us-central1, us-east1, europe-west4, and asia-east1). The fleet makes Google Cloud the second-largest GPU cloud by volume (behind AWS) and the largest TPU cloud by a wide margin.

AcceleratorTypeCompute (FP16)MemoryInterconnectBest ForPrice/hr
TPU v5pCustom ASIC459 TFLOPS bf1695 GB HBM2eICI (Inter-Core Interconnect)Large-scale training (JAX/Flax)$4.50 (per chip)
TPU v5eCustom ASIC197 TFLOPS bf1616 GB HBM2eICIMid-size training, inference$1.20 (per chip)
H100 SXMNVIDIA GPU1,979 TFLOPS sp80 GB HBM3NVLink + InfiniBandTraining + inference (PyTorch)$3.55-$4.00
H200 SXMNVIDIA GPU1,979 TFLOPS sp141 GB HBM3eNVLink + InfiniBandLarge context inference$4.00-$4.50
B200 SXMNVIDIA GPU4,510 TFLOPS sp192 GB HBM3eNVLink + InfiniBandNext-gen training/inference$5.50-$6.00
L4NVIDIA GPU121 TFLOPS sp24 GB GDDR6PCIe Gen4Cost-effective inference$0.60-$0.80
02

GKE FOR AI: KUBERNETES AS THE AI ORCHESTRATION LAYER

Google Cloud's orchestrator strategy centers on GKE (Google Kubernetes Engine) as the universal control plane for AI workloads. GKE supports GPU node pools with automatic GPU driver installation, node auto-scaling for GPU workloads, and Node Auto-Repair for GPU health monitoring. The service is deeply integrated with Google's Vertex AI, which provides managed ML tools that run on GKE under the hood. For teams that prefer to manage their own stack, GKE GPU clusters start from a standard GKE cluster with `--accelerator` flags for GPU nodes.

GKE's GPU support extends to advanced scheduling features including GPUs as an extended resource, MIG (Multi-Instance GPU) partition scheduling for A100/H100, and dynamic GPU allocation via the NVIDIA GPU Operator. GKE also supports GPUDirect-TCPX for optimized network performance between GPU pods. For teams that need to run multi-node distributed training, GKE supports the Kubeflow Training Operator (which manages PyTorchJob and TFJob custom resources) and Kueue for quota-based scheduling with job queuing. This stack is mature: Google has been running internal ML workloads on Kubernetes for longer than any other company.

03

TPU VS GPU: THE DECISION FRAMEWORK

The decision between TPUs and GPUs on Google Cloud depends on three factors: framework, scale, and model architecture. TPUs require JAX or TensorFlow: if your stack is PyTorch-first, TPUs add significant migration friction. TPUs excel at scale: a TPU v5p pod with 8,960 chips achieves 95+ percent scaling efficiency for large Transformer training at 100B+ parameters, while GPU clusters typically see 70-85 percent efficiency at comparable scale due to Ethernet-based inter-node communication overhead. However, TPUs underperform on inference for most model architectures because the JIT compilation overhead (3-10 minutes for initial compilation) makes dynamic batching and model swapping impractical.

For most teams, the practical recommendation is: use TPUs for training runs that exceed 1,000 GPU-hours per job and use JAX; use GPUs for everything else. Google's own Gemini models are trained entirely on TPUs, but their inference serving infrastructure uses GPUs. This internal precedent is informative: TPUs for training efficiency, GPUs for serving flexibility. Google's pricing reflects this: TPU v5p at $4.50/chip-hour is 20-30 percent more expensive per chip than H100, but achieves 30-50 percent more tokens per dollar for large JAX-based training runs.

WorkloadRecommended AcceleratorRationaleCost/Token (Training)Cost/Token (Inference)
Small model training (<7B)GPU (H100/L4)Framework flexibility, fast iteration$1.00-2.50/Mtok$0.15-0.50/Mtok
Large model training (>70B)TPU v5pSuperior scaling efficiency in JAX$0.50-1.20/MtokN/A
MoE model training (8x22B+)TPU v5pTPU expert parallelism is more mature$0.60-1.00/MtokN/A
Real-time inferenceGPU (H100/L4)Dynamic batching, ecosystem supportN/A$0.10-0.80/Mtok
Batch inferenceGPU (H200)Large VRAM for batch processingN/A$0.05-0.30/Mtok
04

PRICING MODEL AND COMMITMENT OPTIONS

Google Cloud's GPU pricing is similar to AWS and Azure for on-demand instances (H100 at $3.55-4.00/GPU-hour) but offers more flexible commitment options. Committed Use Contracts (1-year or 3-year) provide 20-40 percent discounts on GPU instances. Preemptible instances (equivalent to AWS Spot) offer 60-80 percent discounts with 24-hour max runtime. Google also offers Sustained Use Discounts: if you run an instance for more than 25 percent of the month, you automatically receive a discount that scales up to 30 percent for full-month usage, without any commitment.

The key pricing differentiator is custom machine types. GCP allows users to configure GPU nodes with custom CPU/RAM ratios rather than fixed instance shapes. For inference workloads where GPUs are the bottleneck, a custom node with minimum CPUs and RAM can reduce total cost by 15-25 percent versus predefined instance types. This flexibility is unique to GCP among the hyperscalers and is particularly valuable for high GPU-to-CPU workloads like batch inference and embedding generation.

Commitment TypeH100 DiscountTPU v5p DiscountMin TermCancelableBest For
On-Demand0%0%N/AYes (per second)Short-term, experiments
Sustained Use (auto)Up to 30%Up to 30%1+ month of usageAuto-appliedSteady workloads
1-Year CUD20-25%20-25%1 yearNoModerate commitment
3-Year CUD40-45%40-45%3 yearsNoProduction workloads
Preemptible60-80%N/A (no preempt TPUs)1 secProvider can terminateFault-tolerant training
05

REGIONAL GPU CAPACITY AND AVAILABILITY DYNAMICS

Google Cloud's GPU capacity is concentrated in us-central1 (Iowa), which hosts approximately 30 percent of GCP's GPU fleet. us-east1 (South Carolina) and us-west1 (Oregon) host another 30 percent combined. European regions (europe-west4 Netherlands, europe-west1 Belgium) host 20 percent, and APAC regions (asia-east1 Taiwan, asia-southeast1 Singapore) host the remaining 20 percent. GPU availability is highly correlated with region popularity-us-central1 often has capacity for urgent H100 workloads, while asia-southeast1 frequently shows limited availability for large GPU clusters.

During the 2023-2024 GPU shortage, GCP implemented a GPU quota increase request system that became a bottleneck: requests for 64+ GPUs required 2-4 weeks review. As of 2026, this has improved to 1-2 days for H100 quotas under 256 GPUs, but 512+ GPU allocations still require a sales conversation. GCP also reserves GPU capacity for Google's internal AI workloads (Gemini, DeepMind), which can compete with customer allocations during peak usage periods. Google has been transparent about this dynamic: they prioritize internal workload allocation at the beginning of each quarter, releasing remaining capacity to customers mid-quarter.

06

COMPETITIVE POSITION AND MARKET FIT

Google Cloud's GPU offering is strongest for teams that use JAX/TensorFlow, want tight integration with GKE and Vertex AI, or need TPU access for large-scale training. The custom machine type flexibility provides a unique cost optimization lever. Google's commitment to open-source AI infrastructure (Kubernetes, Kubeflow, TensorFlow, JAX) makes them the preferred cloud for teams that build their own AI platforms rather than using managed services.

The challenges are Google's higher list prices (matching AWS and Azure but above OCI and CoreWeave), the complexity of the TPU vs. GPU decision which can confuse less experienced teams, and capacity competition with Google's internal AI workloads. Google Cloud's GPU business is estimated at $2-3 billion in annual revenue, growing approximately 40-50 percent year-over-year, making it a significant but not dominant player in a market where AWS leads in absolute GPU volume.

Filed under
Google Cloud GPUTPU v5pGKE AIGoogle Cloud AIGCP GPU PricingNVIDIA on GCPCloud TPU Strategy