The GPU Provider P&L
Every dollar a GPU provider charges for inference compute is split across four buckets: hardware depreciation (40–55%), power and cooling (15–25%), facility and interconnect (10–15%), and operating margin (5–30%). The range in operating margin is enormous - well-run providers with high utilization and long-term power contracts can clear 25–30% gross margins, while thinly capitalized neoclouds operating at 40% utilization may be margin-negative on an all-in basis.
Understanding this P&L is not academic. When a provider's margins are thin, they cut costs in areas that affect your workload: degraded cooling (thermal throttling), oversubscribed network fabric, or reduced staffing for incident response. The providers with sustainable margins are the ones who will respond to your NCCL failure at 2 AM.
Revenue Per GPU-Hour
Revenue per GPU-hour varies dramatically by deployment model. A provider running dedicated reserved instances for a single customer captures $2.80–$3.40/hr per H100 (their highest per-unit revenue), but utilization risks are borne by the customer via minimum commitments. Spot market revenue floats at $1.80–$2.80/hr depending on region and oversubscription, with zero revenue assurance.
The most profitable revenue model for providers is the managed inference API. By wrapping GPU compute with a model-serving runtime and charging per-token, providers earn the equivalent of $4–$8 per GPU-hour on the same hardware that would net $2.50/hr as raw compute. The margin comes from three compounding advantages: multi-tenancy (serving 5–20 customers per GPU via continuous batching), higher utilization (85%+ versus 60–70% for raw rental), and the markup on software abstraction.
Cost Structure Breakdown by Provider Type
The cost to deliver a GPU-hour differs by provider tier. Hyperscalers benefit from volume power purchase agreements ($0.04–$0.06/kWh vs $0.08–$0.12/kWh for midsize operators) and amortized facility costs. Neoclouds compete on lower overhead and thinner margins. The table below estimates all-in cost per H100 GPU-hour for three provider tiers, assuming 70% utilization and 3-year hardware depreciation.
Hardware costs assume H100 SXM at $30,000 MSRP with volume discounts. Facility costs include data center lease, networking equipment, and staffing.
| Cost Component | Hyperscaler | Mid-Size Neocloud | Small Operator |
|---|---|---|---|
| Hardware Depreciation (3yr) | $1.14 | $1.22 | $1.35 |
| Power & Cooling | $0.32 | $0.48 | $0.65 |
| Facility & Network | $0.28 | $0.35 | $0.50 |
| Staffing & Support | $0.18 | $0.22 | $0.30 |
| All-In Cost per GPU-hr | $1.92 | $2.27 | $2.80 |
| Avg. Spot Revenue | $2.60 | $2.40 | $2.10 |
| Gross Margin | 26% | 5% | -25% |
Who Is Profitable and Who Is Not
The GPU provider market in 2026 is experiencing a margin squeeze. The H100 spot price has declined from $3.50/hr in early 2025 to $2.40–$2.80/hr as new capacity from B200 and B300 deployments flooded the market. Providers that financed H100 clusters at peak hardware prices ($35,000–$40,000 per GPU) are now operating at negative margins on spot sales, subsidizing losses with reserved contract revenue.
The winners are providers with three characteristics: hardware purchased at or below MSRP, long-term power contracts at $0.05/kWh or less, and managed inference API revenue that earns $4–$8/GPU-hr equivalent. Providers relying solely on raw GPU rental at spot rates are running 0–10% gross margins in mid-2026. This margin compression will drive consolidation - we expect 20–30% of neocloud providers to exit the market or be acquired in the next 12 months.
What Margins Mean for Buyers
Thin provider margins create buyer risk. A provider operating at 0–5% margin has no cushion for hardware failures, network upgrades, or power price spikes. When their costs go up, service quality goes down before prices go up - slower support response, degraded cooling, deferred maintenance. The cheapest GPU-hour on the spot market often comes from the most distressed provider.
For buyers, the right question is not "What is the lowest GPU-hour price?" but "What is the lowest price a provider can sustainably charge for this hardware?" The answer is roughly $2.20–$2.50/hr for H100 (the hyperscaler cost plus a reasonable margin). Any provider pricing significantly below that range is either subsidizing with venture capital, depreciating hardware over an unrealistic 5-year schedule, or deferring maintenance - all of which create concentration risk for your workload.
Market Outlook Through 2027
The margin trajectory for H100 is clear: continued compression toward $2.00–$2.20/hr all-in cost floor as more Blackwell capacity comes online. H100 becomes the commodity compute tier, and providers will differentiate on reliability, interconnect quality, and support rather than price. B200 and B300 will maintain a margin premium of 25–35% over H100 through early 2027 as demand for high-bandwidth inference outstrips supply.
The structural opportunity for buyers is to treat GPU procurement as a portfolio problem. Lock in a base layer of reserved capacity with a well-capitalized provider at sustainable margins (2-year term, 10–20% below spot). Layer spot capacity from providers that offer transparent interconnect specs and published maintenance windows. Use transfer desks (including ClusterBid's) to exit positions that no longer make sense as the market evolves.
