THE AWS GPU INSTANCE LINEUP IN 2026
AWS offers the broadest set of GPU instance families among public cloud providers. As of 2026, the high-performance compute tier is led by the P5e instances (H100 SXM3 141 GB with NVLink 4.0) in us-east-1, us-west-2, eu-west-1, and ap-southeast-1. The P5 family continues with the original P5 (H100 SXM 80 GB) and P5e (H100 SXM3 141 GB), with the latter delivering 2 TFLOPS more FP8 throughput per GPU due to the higher-bandwidth HBM3e memory. Below the P5 line, the P4d and P4de instances (A100 80 GB) remain relevant for inference workloads where H100 oversupply is unnecessary, particularly for serving under 30B parameter models.
On the inference-optimized side, the G5 family (A10G 24 GB) and G6 instances (L40S 48 GB) cover the mid-range. The Trn1 and Trn2 instances using AWS custom Trainium2 chips compete with H100 for training workloads at a 30-40 percent lower effective cost when AWS Neuron SDK compatibility is satisfied. Inf2 instances (Inferentia2) target low-cost inference for Transformer models below 10B parameters at pricing that undercuts equivalent GPU-based instances by 40-50 percent.
| Instance Family | GPU Type | GPUs per Node | GPU VRAM | On-Demand $/hr | 1yr Reserved $/hr | Spot Avg $/hr |
|---|---|---|---|---|---|---|
| p5e.48xlarge | H100 SXM3 141GB | 8 | 1,128 GB | $38.22 | $24.54 | $9.80-12.40 |
| p5.48xlarge | H100 SXM 80GB | 8 | 640 GB | $32.77 | $21.06 | $8.20-10.50 |
| p4de.24xlarge | A100 80GB | 8 | 640 GB | $22.64 | $14.35 | $5.40-7.10 |
| p4d.24xlarge | A100 40GB | 8 | 320 GB | $19.67 | $12.58 | $4.80-6.20 |
| g6.12xlarge | L40S 48GB | 4 | 192 GB | $6.98 | $4.52 | $1.80-2.40 |
| g5.12xlarge | A10G 24GB | 4 | 96 GB | $5.67 | $3.72 | $1.40-1.90 |
| trn1.32xlarge | Trainium2 | 16 | N/A (HBM) | $18.45 | $11.83 | $5.20-6.80 |
| inf2.48xlarge | Inferentia2 | 12 | 48 GB shared | $9.68 | $6.36 | $2.80-3.60 |
REGIONAL AVAILABILITY AND CAPACITY CONSTRAINTS
AWS GPU availability is heavily regional. P5e H100 instances are available in us-east-1 (9 AZs), us-west-2 (4 AZs), eu-west-1 (3 AZs), and ap-southeast-1 (2 AZs). Us-east-1 has the deepest pool, with AWS reporting over 20,000 P5e GPUs deployed across its Northern Virginia availability zones as of mid-2026. Ap-southeast-1 (Singapore) remains capacity-constrained with typical wait times of 2-4 weeks for 8-node (64 GPU) P5e clusters.
Trn1 Trainium2 instances have broader regional coverage including us-east-1, us-west-2, eu-west-1, eu-central-1, ap-northeast-1, and ap-southeast-1. AWS has invested heavily in Trainium capacity, making it the most available high-performance accelerator regionally. The G5 (A10G) family is available in all 26 AWS commercial regions, making it the default choice for multi-region inference deployments.
| Region | P5e H100 | P5 H100 | P4de A100 | Trn1 Trn2 | Availability Score |
|---|---|---|---|---|---|
| us-east-1 (N.Virginia) | Yes (9 AZs) | Yes | Yes | Yes | Excellent |
| us-west-2 (Oregon) | Yes (4 AZs) | Yes | Yes | Yes | Very Good |
| eu-west-1 (Ireland) | Yes (3 AZs) | Yes | Yes | Yes | Very Good |
| eu-central-1 (Frankfurt) | No | Yes | Yes | Yes | Moderate |
| ap-southeast-1 (Singapore) | Yes (2 AZs) | Yes | Yes | Yes | Constrained |
| ap-northeast-1 (Tokyo) | No | Yes | Yes | No | Limited |
| sa-east-1 (Sao Paulo) | No | No | Yes | Yes | Limited |
ON-DEMAND VERSUS SPOT VERSUS RESERVED: COST OPTIMIZATION STRATEGIES
Spot GPU pricing on AWS fluctuates between 70-85 percent below on-demand for P5 instances during non-peak hours (weekends, nighttime in the primary region). However, spot termination rates for P5e in us-east-1 average 5-12 percent per hour during business hours, making spot viable only for fault-tolerant training jobs using checkpointing every 15-30 minutes. G5 spot instances see termination rates below 3 percent per hour, supporting near-production inference workloads at 65-75 percent below on-demand pricing.
One-year reserved instances for P5e offer a 36 percent discount over on-demand at $24.54/hr for the full 8-GPU instance. Three-year reservations with partial upfront reach $18.12/hr, 53 percent below on-demand. AWS Convertible Reservations allow migration between GPU instance types, useful for teams upgrading from P4d to P5. The break-even point: at 6+ hours of daily GPU usage, one-year reserved is cheaper than on-demand; at 12+ hours daily, three-year reserved wins.
CLUSTER NETWORKING AND MULTI-NODE PERFORMANCE
P5e and Trn1 instances support Elastic Fabric Adapter (EFA) with 3,200 Gbps aggregate bandwidth per node, essential for multi-node training. P5e uses NVLink 4.0 with 900 GB/s GPU-to-GPU bandwidth within a node, and EFA for inter-node communication. Multi-node training on 64+ P5e instances achieves 90+ percent scaling efficiency for GPT-class models when using NeMo Megatron or AWS SageMaker distributed training libraries.
A notable differentiator: Trn1 instances use NeuronLink (AWS proprietary interconnect) that delivers 768 GB/s per accelerator, comparable to NVLink within the node, with 1,600 Gbps EFA for inter-node. In practice, Trn2 at 64 nodes achieves 85-90 percent of H100 cluster throughput for TensorFlow and PyTorch training jobs optimized with Neuron SDK, at 30-40 percent lower total cost.
| Feature | P5e 8x H100 | P4de 8x A100 | Trn1 16x Trn2 | G5 4x A10G |
|---|---|---|---|---|
| Intra-node Fabric | NVLink 900 GB/s | NVLink 600 GB/s | NeuronLink 768 GB/s | N/A (PCIe 4) |
| Inter-node BW | 3,200 Gbps EFA | 1,600 Gbps EFA | 1,600 Gbps EFA | 100 Gbps ENA |
| GPU-to-GPU Lat | <2 microsec | <3 microsec | <3 microsec | >10 microsec |
| Multi-Node Scaling | 90%+ at 64 nodes | 85%+ at 64 nodes | 85-90% at 64 nodes | 60-70% at 16 nodes |
| Best For | Large training | Fine-tuning | Training only | Inference |
COST PER GPU-HOUR: AGGREGATED COMPARISON
Breaking down AWS GPU costs to a per-GPU basis reveals major savings opportunities. P5e at $38.22/hr for 8 GPUs yields $4.78 per GPU-hour on-demand and approximately $1.23-1.55 per GPU-hour on spot. This is competitive with GCP's A3 instances at $4.62 per A100 GPU-hour and Azure's ND H100 v5 at $4.95 per GPU-hour on-demand. AWS G5 A10G instances at $1.42 per GPU-hour on-demand ($0.35-0.48 spot) lead the mid-range price segment.
The Trainium2 per-chip price of $1.15/hr on-demand (Trn1 at $18.45/16 chips) makes it the cheapest high-performance training accelerator on AWS, but the software stack maturity and limited model support narrow the practical use cases. For teams already using PyTorch with Neuron SDK compilation, the savings are substantial: 60-70 percent lower training cost versus P5e for most NLP model architectures.
WHEN AWS GPU IS NOT THE ANSWER
AWS GPU instances carry a premium over bare-metal GPU providers and secondary markets. A comparable H100 configuration on Lambda GPU Cloud or TensorDock costs 40-60 percent less than AWS P5 on-demand pricing. However, AWS's advantage lies in ecosystem integration: VPC networking, IAM security, CloudWatch monitoring, SageMaker integration, and the broadest regional footprint for latency-sensitive inference workloads.
For workloads with predictable long-running GPU demand (3+ month training runs, continuous inference serving), AWS reserved instances or savings plans reduce the premium to 15-25 percent above bare-metal providers. The decision matrix: use AWS when you need the ecosystem or must deploy inference globally; use specialized GPU providers when training cost is the primary concern and the software stack is portable.
