3-YEAR TCO MODEL FRAMEWORK
The total cost of ownership for GPU compute breaks into capital expenditure (purchase, facility, networking) and operating expenditure (power, cooling, maintenance, labor). For on-premise, CapEx dominates at 60-70% of 3-year TCO. For cloud rental, OpEx is 100% of TCO but with zero upfront commitment. The comparison point is the utilization breakeven: on-premise wins at sustained utilization above 55-70% depending on GPU type.
Our TCO model includes all cost factors often omitted from vendor analysis: facility cooling at $0.05-0.12/kWh add-on, networking at $2,000-5,000/node, maintenance at 3-5% of hardware cost annually, labor at $15,000-30,000/admin/year, and decommissioning at 5% of purchase price. These hidden costs add 25-40% to the headline server price.
PURCHASE: H100 ON-PREMISE TCO
H100 secondary market pricing in mid-2026: $15,000-20,000 for SXM, $12,000-16,000 for PCIe. New H100 SXM HGX baseboard with 8 GPUs costs $140,000-180,000. Full server deployment including InfiniBand networking, storage, and facility integration adds 35-50%: total 8-GPU node CapEx of $200,000-270,000. Annual OpEx for power (700W/GPU + overhead) at $0.12/kWh: $7,350/GPU or $58,800/8-GPU node.
Three-year TCO for an 8-GPU H100 node at 60% utilization: $200,000 (CapEx) + $176,400 (3yr OpEx) = $376,400. Cost per GPU-hour: $376,400 / (8 GPUs x 24hr x 365d x 3yr x 60%) = $376,400 / 126,144 = $2.98/GPU-hr. At 80% utilization: $2.24/GPU-hr. At 40% utilization: $4.48/GPU-hr. Breakeven against cloud rental of $2.50/hr occurs at approximately 55% utilization.
RENTAL: CLOUD GPU TCO
2026 on-demand pricing varies by provider: AWS p5 (H100) at $2.50-3.50/hr, Lambda at $1.75-2.50/hr, RunPod at $1.50-2.20/hr, Vast at $1.20-1.80/hr. Spot pricing reduces these by 50-70%: AWS spot H100 at $0.75-1.20/hr, RunPod spot at $0.50-0.90/hr. Reserved 1-year contracts provide 30-40% discount: $1.50-2.00/hr on neoclouds, $2.00-2.50/hr on AWS.
Three-year cloud TCO for 8 GPUs at on-demand $2.50/hr: $2.50 x 8 x 24 x 365 x 3 = $525,600. No upfront cost, no decommissioning, no facility management. However, cloud costs include data egress ($0.05-0.12/GB), storage ($0.02-0.08/GB-month), and networking ($500-2,000/month/cluster), adding 10-20% overhead. Effective 3-year cost: $578,000-630,000.
| Cost Factor | On-Prem (8x H100) | Cloud On-Demand (8x H100) | Cloud Spot (8x H100) |
|---|---|---|---|
| 3-Year CapEx | $200,000-270,000 | $0 | $0 |
| 3-Year GPU Cost | $0 | $525,600-735,840 | $157,680-378,000 |
| 3-Year Power/Cooling | $176,400 | $0 (included) | $0 (included) |
| 3-Year Labor/Maintenance | $60,000-90,000 | $0 (managed) | $0 (managed) |
| 3-Year Networking | $16,000-40,000 | $18,000-72,000 | $18,000-72,000 |
| Total 3-Year Cost | $452,000-576,400 | $543,600-807,840 | $175,680-450,000 |
| Effective $/GPU-hr (60% util) | $2.98-3.81 | $2.58-3.83 | $0.83-2.14 |
BREAKEVEN ANALYSIS
On-premise breaks even with cloud on-demand at sustained GPU utilization of 55-65% for H100. For B200 at $35,000-50,000 purchase price and $40,000-60,000 fully loaded server cost, breakeven rises to 70-80% utilization due to higher depreciation and faster obsolescence risk. MI300X at $18,000-25,000 purchase: breakeven at 50-60% utilization.
The breakeven utilization decreases with scale. A 128-GPU cluster achieves 60% breakeven vs 70% for an 8-GPU cluster due to shared facility and labor costs. At 512+ GPU scale, on-premise breakeven drops to 45-55%. This explains why only large organizations with >100 GPUs and >60% sustained utilization tend to purchase hardware.
RISK FACTORS AND HIDDEN COSTS
Hardware obsolescence risk is the largest hidden cost of purchasing. Blackwell B200 makes H100 generationally obsolete for FP4 workloads. An H100 purchased in 2024 at $30,000 has 50% residual value in mid-2026; by 2027, residual may drop to 20-30%. Three-year depreciation may significantly exceed standard accounting schedules if Blackwell adoption accelerates.
Cloud rental risks include price increases (AWS raised GPU pricing by 10-25% in 2024-2025), capacity allocation during crunch periods, and egress fees that lock customers into a provider. On-premise risks include facility power constraints (many data centers maxed on GPU power allocations), cooling failures, and NVIDIA GPU lead times of 12-20 weeks for replacement units.
DECISION FRAMEWORK
Purchase on-premise if: sustained utilization > 60%, cluster size > 32 GPUs, deployment horizon > 2 years, power is under $0.10/kWh, and workload is training or high-throughput inference with predictable demand. Rent cloud if: utilization is variable or < 50%, cluster size < 16 GPUs, workload is latency-sensitive inference, or deployment horizon is < 18 months.
The hybrid model is increasingly popular: purchase baseline capacity for predictable load at 70-80% utilization, rent burst capacity from cloud for demand spikes. A 32-GPU baseline with 32-GPU cloud burst capacity achieves effective blended cost of $1.80-2.20/GPU-hr, 25-35% below pure on-prem or pure cloud at the same effective utilization.
