THE AZURE GPU PORTFOLIO IN 2026
Microsoft Azure segments its GPU offerings into three families in 2026. The ND-series is the flagship training lineup: ND H100 v5 (8x H100 80 GB SXM), ND H100 v5 Boost (8x H100 141 GB SXM3), and ND A100 v4 (8x A100 80 GB). The NC-series targets inference and HPC: NC A100 v4 (4x A100 80 GB), NC H100 v5 (4x H100), and the entry-level NCads H100 v5. The third category is Azure's differentiated offering: confidential GPU instances with AMD MI300X and NVIDIA H100 trusted execution environments (TEE) for regulated workloads requiring data encryption in-use.
Azure's pricing positions slightly above AWS and GCP for on-demand H100 instances. The ND H100 v5 Boost at $39.59/hr (8 GPU) is $1.37/hr more than AWS P5e and $2.48/hr more than GCP A3 Mega. However, Azure offers the deepest reserved instance discounts: 3-year reserved with upfront payment reaches 58 percent off on-demand, bringing ND H100 v5 to $16.63/hr, below both AWS 3-year reserved P5e ($18.12/hr) and GCP 3-year CUD A3 Mega ($17.35/hr).
| Azure Instance | GPU Type | GPUs | VRAM | On-Demand $/hr | 3yr Reserved $/hr | Spot $/hr |
|---|---|---|---|---|---|---|
| ND H100 v5 Boost | H100 SXM3 141GB | 8 | 1,128 GB | $39.59 | $16.63 | $10.80-13.50 |
| ND H100 v5 | H100 SXM 80GB | 8 | 640 GB | $34.12 | $14.33 | $8.90-11.20 |
| ND A100 v4 | A100 80GB | 8 | 640 GB | $24.80 | $10.42 | $6.50-8.20 |
| NC H100 v5 | H100 80GB | 4 | 320 GB | $19.42 | $8.16 | $4.90-6.30 |
| NC A100 v4 | A100 80GB | 4 | 320 GB | $13.49 | $5.67 | $3.40-4.50 |
| NCads H100 v5 | H100 PCIe 80GB | 1 | 80 GB | $5.01 | $2.10 | $1.30-1.70 |
| NCads A100 v4 | A100 PCIe 40GB | 1 | 40 GB | $3.42 | $1.44 | $0.85-1.12 |
CONFIDENTIAL GPU: AZURE'S DIFFERENTIATED OFFERING
Azure's confidential GPU instances use NVIDIA H100 TEE (Trusted Execution Environment) to encrypt data in-use via GPU hardware attestation. This is the only hyperscaler GPU offering that meets the strictest data sovereignty requirements (HIPAA, GDPR Article 46, FedRAMP High) for AI inference on protected health information and financial data. The H100 TEE encrypts model weights, intermediate activations, and output data at the hardware level, preventing access by host OS, hypervisor, or cloud provider personnel.
Confidential NC H100 v5 instances command a 35-40 percent premium over standard instances: $6.78/hr for a single H100 80 GB TEE versus $5.01/hr standard. Initial adoption is concentrated in healthcare (medical imaging AI, clinical NLP) and financial services (fraud detection, credit scoring models). For workloads without compliance requirements, the standard instances offer identical compute at lower cost.
Azure also offers confidential AMD MI300X instances for customers who prefer AMD GPU architecture. The MI300X with 192 GB HBM3 runs at $28.52/hr (8 GPU) with confidential computing at no additional premium. AMD's ROCm software stack maturity has improved significantly through 2025-2026, though PyTorch compilation through ROCm still shows 5-10 percent overhead versus CUDA for some model architectures.
| GPU Type | Standard $/hr | Confidential $/hr | Premium | Certification | Use Case |
|---|---|---|---|---|---|
| H100 80 GB (single) | $5.01 | $6.78 | 35% | HIPAA FedRAMP | Healthcare LLM |
| H100 80 GB (4-GPU) | $19.42 | $25.89 | 33% | HIPAA FedRAMP | Finserv AI |
| A100 80 GB (single) | $3.42 | $4.55 | 33% | HIPAA | Medical imaging |
| MI300X 192 GB (8-GPU) | $28.52 | $28.52 | 0% | FedRAMP High | Gov workloads |
REGIONAL DEPLOYMENT AND CAPACITY STRATEGY
Azure has the most extensive regional GPU deployment among the three major providers, driven by Microsoft's $50+ billion infrastructure investment commitment. ND H100 v5 instances are available in 16 Azure regions including East US, West US 2, West Europe, North Europe, Southeast Asia, Australia East, and newer regions like Qatar Central and Poland Central. This breadth exceeds AWS P5e (4 regions) and GCP A3 (3 regions) by a wide margin.
Capacity allocation uses Azure's quota system with per-VM-family limits. Standard subscriptions start at 8-16 ND-series vCPU quota. Enterprise customers with Microsoft AI agreements can access 1,000+ GPU clusters with regional proximity guarantees. Wait times for large H100 allocations in West US 2 or East US typically run 1-2 weeks versus 2-4 weeks in European or Asian Azure regions.
AZURE ML, COPILOT, AND ECOSYSTEM INTEGRATION
Azure Machine Learning provides managed GPU training and inference with automated hyperparameter tuning, data versioning, and experiment tracking. The primary advantage is integration with the Microsoft AI ecosystem: OpenAI models (GPT-4o, o3) accessible through Azure OpenAI Service on the same network as custom GPU training jobs; GitHub Copilot extending to infrastructure-as-code GPU provisioning; and Microsoft Fabric data pipelines feeding directly into AKS GPU inference clusters.
The Azure Kubernetes Service (AKS) GPU node pool supports the full GPU portfolio with KEDA-based autoscaling that adjusts GPU nodes based on pending pod queue depth. AKS GPU clusters can mix ND-series (training) and NC-series (inference) node pools in a single cluster, enabling a unified K8s control plane for the entire ML pipeline. Azure's GPU monitoring via Container Insights provides per-GPU utilization, memory bandwidth, and NVLink metrics natively.
| Integration Feature | Azure | AWS | GCP | Differentiator |
|---|---|---|---|---|
| Managed ML | Azure ML | SageMaker | Vertex AI | Tight OpenAI integration |
| Serverless GPU | AKS + KEDA | EKS + Karpenter | GKE Autopilot | Deep KEDA integration |
| Inference Endpoint | Azure ML Endpoints | SageMaker Endpoints | Vertex AI Endpoints | Copilot IaC support |
| Cost Management | Azure Cost Mgmt | Cost Explorer | Billing Reports | EA + MACC credits |
TOTAL COST OF AZURE GPU: BEYOND INSTANCE PRICING
Azure GPU costs extend beyond compute rates. Data egress from Azure GPU instances to the internet is $0.05-0.19/GB depending on region and cumulative volume, plus managed disk costs ($0.08/GB-month for Premium SSD) and Azure ML workspace overhead ($50-200/month). A training run on ND H100 v5 at 50% utilization for 30 days: compute $4,900 (spot) or $12,274 (on-demand), egress (2 TB checkpoint) $100-380, storage $80-160, management overhead $100-200.
Enterprise customers with Microsoft Azure Consumption Commitment (MACC) agreements can effectively reduce GPU costs by 15-35 percent by applying committed spend against GPU instances. Microsoft's OpenAI partnership also provides Azure credits for GPU usage tied to Azure OpenAI consumption, making Azure the most cost-effective option for customers already using GPT-4o or other Microsoft AI services.
WHERE AZURE WINS AND LOSES IN GPU
Azure dominates in three scenarios: multi-national inference deployment requiring GPU access in 16+ regions (no other provider matches this breadth); regulated industries needing confidential GPU computing on H100 TEE; and Microsoft enterprise shops that can leverage MACC agreements to reduce effective GPU rates below AWS or GCP pricing.
Azure lags on per-GPU raw pricing (5-10 percent above competitors on-demand), spot instance reliability (Azure spot termination rates for ND-series run 8-15 percent per hour versus 5-12 percent on AWS), and startup-friendliness (Azure's quota and reservation systems are more complex than AWS Launch Wizard or GCP's straightforward CUD model). For non-Microsoft shops without EA agreements, AWS or GCP generally deliver lower total GPU cost.
