AZURE: THE LARGEST GPU CLOUD BY REVENUE
Microsoft Azure has become the largest GPU cloud by revenue, driven primarily by two factors: the OpenAI relationship (OpenAI's training infrastructure runs exclusively on Azure) and Microsoft's aggressive investment in GPU capacity. As of Q2 2026, Microsoft operates an estimated 150,000+ NVIDIA H100 GPUs across 30+ Azure regions, plus growing fleets of H200 and B200. The exact count is not disclosed, but analyst estimates place Azure's GPU fleet at 50-80 percent larger than AWS's, making it the single largest GPU operator in the world.
Azure's GPU infrastructure serves two distinct markets. The OpenAI partnership requires massive contiguous GPU clusters (10,000+ GPUs per training run) with InfiniBand fabric and custom networking. Enterprise customers use the same infrastructure for smaller workloads (8-512 GPUs) through Azure's standard VM offerings. The economies of scale from OpenAI's massive commitment (estimated at $5-10 billion in GPU spend over 5 years) allow Azure to offer competitive pricing for enterprise customers while maintaining healthy margins.
ND H100 V5 SERIES: THE FLAGSHIP AI INSTANCE
The ND H100 v5 series is Azure's flagship AI training instance. Each ND96isr_H100_v5 VM provides 8x H100 SXM GPUs (80 GB each, NVLink-connected), 96 vCPUs (AMD Genoa), 1.9 TB RAM, and 3.2 TB/s NVMe local storage. The key differentiator is the InfiniBand networking: each VM has 8x 400 Gbps NVIDIA Quantum-2 InfiniBand connections, providing a total of 3.2 Tbps of inter-node bandwidth. Azure's InfiniBand fabric is designed for the highest-scale distributed training workloads, enabling linear scaling up to 1,000+ GPU jobs with 90+ percent scaling efficiency for standard Transformer architectures.
Azure also offers the ND H100 v5 in a 'Super' configuration (ND960isr_H100_v5) with 8 GPUs and 8 InfiniBand NICs, optimized for the largest model parallelism scenarios. The Super variant is the same hardware with enhanced firmware tuning for lower latency, reducing inter-node communication overhead by an additional 5-10 percent. This configuration is primarily used for OpenAI's internal workloads and is available to select enterprise customers through a private preview process.
| Instance | GPU | vCPUs | RAM | Inter-node | Disk | On-Demand/hr | 1-Yr Reserved/hr |
|---|---|---|---|---|---|---|---|
| ND96isr_H100_v5 | 8x H100 (80 GB) | 96 (AMD) | 1.9 TB | 8x 400 Gbps InfiniBand | 3.2 TB NVMe | $26.80 ($3.35/GPU) | $19.30 ($2.41/GPU) |
| ND96asr_A100_v4 | 8x A100-80G | 96 (AMD) | 900 GB | 8x 200 Gbps InfiniBand | 1.9 TB NVMe | $15.00 ($1.88/GPU) | $10.80 ($1.35/GPU) |
| NC96ads_A100_v4 | 4x A100-80G | 96 (AMD) | 880 GB | 100 Gbps Ethernet | 1.9 TB NVMe | $9.80 ($2.45/GPU) | $7.40 ($1.85/GPU) |
| NC4ads_L40S_v2 | 1x L40S | 4 (AMD) | 32 GB | 25 Gbps Ethernet | 128 GB SSD | $0.65 | $0.45 |
| ND960isr_H100_v5 (Super) | 8x H100 (80 GB) | 96 (AMD) | 1.9 TB | 8x 400 Gbps InfiniBand (tuned) | 3.2 TB NVMe | $29.50 ($3.69/GPU) | Private preview |
THE OPENAI FACTOR: HOW OPENAI SHAPED AZURE'S GPU ARCHITECTURE
OpenAI's relationship with Azure has been the most consequential factor in shaping Azure's GPU infrastructure strategy. OpenAI's requirements pushed Azure to deploy InfiniBand at hyperscale earlier than any other cloud provider. The GPT-4 training run required 10,000+ GPU clusters with sub-3 microsecond inter-node latency, which forced Azure to develop custom networking software (the Azure HPC CNI plugin, GPUDirect-NetX integration) and to work with NVIDIA to build the Quantum-2 InfiniBand platform at cloud scale. These optimizations, built for OpenAI, are now available to all Azure H100 customers.
The scale of OpenAI's GPU consumption is difficult to overstate. OpenAI is estimated to consume 30-50 percent of Azure's total GPU capacity (30K-60K H100s continuously, with periodic spikes to 100K+ for training runs). This creates a double-edged dynamic for other Azure GPU customers: capacity is more constrained than on AWS or GCP because OpenAI gets priority allocation, but the infrastructure quality and optimization are better because OpenAI's demands drive continuous improvements. Azure's GPU quota system explicitly separates OpenAI's allocation from customer capacity, but during training runs at OpenAI's scale, the InfiniBand fabric experiences congestion that can affect adjacent customer jobs.
PRICING STRATEGY: AZURE HYBRID BENEFIT AND ENTERPRISE DEALS
Azure's GPU pricing strategy is built on a foundation of enterprise licensing. Azure Hybrid Benefit allows customers with existing Windows Server and SQL Server licenses to apply them to Azure GPU VMs, reducing the compute portion of the bill by 40-60 percent. For customers with Enterprise Agreements, GPU pricing is negotiated individually-typically 15-40 percent below list price for 3-year commitments. This makes Azure's effective pricing highly variable: a large enterprise customer with EA pricing and Azure Hybrid Benefit might pay $1.80-2.20 per H100-hour, while a startup paying by credit card sees $3.35 per H100-hour on-demand.
Azure also offers Spot GPU VMs (up to 80 percent discount) but with the shortest eviction notice among hyperscalers–30 seconds versus GCP's 30-second notice and AWS's 2-minute notice. The tight eviction window makes Azure Spot challenging for training workloads but acceptable for batch inference and parameter-sweep experiments. Azure's spot GPU capacity is more volatile than AWS's: availability fluctuates by 40-60 percent day-to-day depending on OpenAI's variable demand, making it less reliable for production workloads that require guaranteed GPU access.
| Pricing Tier | H100/hr (8-GPU) | Effective/GPU-hr | Requirements | Best For |
|---|---|---|---|---|
| On-Demand (Pay-as-you-go) | $26.80 | $3.35 | None | Short-term, variable workloads |
| 1-Year Reserved | $19.30 | $2.41 | Prepay or monthly commit | Steady-state training |
| 3-Year Reserved | $15.50 | $1.94 | 3-year commit | Production, predictable load |
| Enterprise Agreement (EA) | $14.40-$20.00 | $1.80-$2.50 | EA contract + $500K+ commit | Enterprise, negotiated pricing |
| EA + Azure Hybrid Benefit | $11.00-$15.00 | $1.38-$1.88 | EA + SQL/Windows licenses | Microsoft-first enterprises |
| Spot | $5.36-$8.04 | $0.67-$1.01 | No guarantee, 30s eviction | Batch, experimentation |
REGIONAL GPU DISTRIBUTION AND CAPACITY MANAGEMENT
Azure's GPU capacity is distributed across 30+ regions but concentrated in 5 primary AI regions: East US (Virginia), South Central US (Texas), West US (California/Washington), West Europe (Netherlands), and Southeast Asia (Singapore). East US has the largest GPU concentration (estimated 40 percent of Azure's total), driven by OpenAI's primary training clusters. Europe (West Europe and North Europe combined) holds 25 percent, US West 20 percent, Asia 10 percent, and the remaining 5 percent across South America, Australia, and other regions.
Azure operates a capacity reservation system that is more rigid than AWS's: customers can reserve GPU capacity in a specific region for 1-month to 3-year terms, but reserved capacity that goes unused is still billed. This discourages multi-region GPU reservation as a strategy-most Azure customers focus on their primary region and accept occasional capacity shortages. During the peak GPU demand periods (typically September-November for AI research deadlines), Azure's South Central US and Southeast Asia regions often show 'capacity not available' for large ND-series deployments, forcing customers to use West Europe or East US instead.
COMPETITIVE POSITION AND MARKET FIT
Azure's GPU cloud is best suited for two customer segments: enterprises with existing Microsoft licensing and AI-native teams that need the highest possible scale. For Microsoft-first enterprises (those using Azure AD, Office 365, SQL Server, and Visual Studio), Azure GPU is the path of least resistance: GPU access integrates with existing Azure subscriptions, billing, and compliance frameworks. The Azure Hybrid Benefit makes GPU pricing highly competitive for this segment-effectively cheaper than any competitor when licensing savings are factored in.
For AI-native companies that are not Microsoft-first, Azure is less compelling. The list pricing is in line with AWS and GCP ($3.35/GPU-hr versus $3.55-4.00 on GCP and $3.75-4.13 on AWS), and the capacity prioritization for OpenAI creates availability uncertainty. Azure's managed AI services (Azure ML) are feature-complete but less polished than Vertex AI or SageMaker. The strongest case for Azure in the non-Microsoft segment is scale: if you need a 1,000+ GPU cluster for training, Azure's InfiniBand fabric and custom CNI are the most battle-tested in the cloud.
