The Enterprise GPU Procurement Challenge
Enterprise GPU procurement in 2026 is fundamentally different from any other IT hardware purchase. GPU supply chains are constrained, lead times stretch 4-52 weeks depending on generation, and pricing is opaque -- no two buyers pay the same rate. The enterprise procurement team must navigate a market where 80% of B200 supply is pre-allocated to hyperscalers, where GPU generations change every 12-18 months, and where the wrong procurement decision can waste millions in excess capacity or delay critical AI timelines by months.
The structured procurement process used by leading enterprises includes: requirements definition (workload profiling, GPU generation requirements, capacity planning), market analysis (vendor landscape, pricing benchmarks, lead time assessment), request for quotation (structured RFQ sent to 5-15 GPU providers), vendor evaluation (scorecard-based comparison of proposals), contract negotiation (pricing, terms, SLAs, flexibility provisions), and ongoing management (vendor performance monitoring, capacity adjustments).
This post provides the templates and frameworks used in each stage, based on procurement processes from 40+ enterprise GPU deployments.
Requirements Definition and Workload Profiling
The requirements document is the foundation of a successful procurement. It translates AI workload characteristics into GPU infrastructure specifications. The key dimensions are: GPU generation (H100, H200, B200, or future requirements), GPU count and form factor (SXM for training, PCIe for inference), memory capacity (80 GB for standard workloads, 192 GB for large models), inter-node networking (InfiniBand vs RoCEv2, bandwidth requirements), storage requirements (parallel filesystem capacity and performance), and location and data residency requirements.
Workload profiling should include at least 30 days of GPU utilisation data if migrating from existing infrastructure, or detailed workload estimates if building new. The profile should cover: peak GPU utilisation percentage, average training job duration, inference request rate and latency requirements, data set sizes and storage growth rate, and model size distribution (parameters, precision).
The workload profile determines the GPU generation and configuration. For example, a profile dominated by 70B+ model inference with 128K context windows requires B200 (192 GB) or multi-GPU H200 configurations. A profile dominated by medium model training (7B-70B) is well-served by H100, with significant cost savings over B200.
RFQ Template: The Essential Sections
A well-structured RFQ enables comparable responses from GPU providers. The RFQ should include: background (company overview, AI workload description, current infrastructure if any), technical requirements (GPU generation and quantity, networking specifications, storage requirements, GPU software stack requirements), terms (contract duration, start date, renewal options, termination notice), pricing structure (GPU-hour pricing, included services, additional charges for networking, storage, support), location preferences (primary and secondary data centre regions), compliance requirements (SOC2, HIPAA, FedRAMP, GDPR), SLAs (uptime guarantee, GPU failure response time, network availability), and vendor qualifications (experience with AI workloads, reference deployments, financial stability).
The RFQ should be sent to 5-15 providers: hyperscalers (AWS, Azure, GCP), GPU-focused neoclouds (CoreWeave, Lambda, RunPod, Vast.ai), data centre operators (Equinix, Digital Realty), and bare-metal GPU providers (Applied Digital, Cirrascale).
Set a firm response deadline (typically 3-4 weeks) and provide a structured response template so responses can be compared apples-to-apples. The RFQ process from distribution to vendor shortlisting typically takes 4-6 weeks.
Vendor Scorecard: Evaluating GPU Providers
Vendor evaluation should use a weighted scorecard. The table below shows the evaluation dimensions and typical weightings based on enterprise procurement best practices. The total score determines vendor ranking and negotiation priority.
| Evaluation Dimension | Weight | Measurement Criteria |
|---|---|---|
| Pricing competitiveness | 25% | GPU-hour rate vs market benchmark, discount for volume/commitment |
| Capacity availability | 20% | GPU lead time, ability to scale with demand, geographic coverage |
| Technical capability | 20% | GPU generation offered, networking options, storage integration, software stack |
| Compliance and security | 15% | SOC2 Type II, HIPAA eligibility, FedRAMP if required, data centre certifications |
| Operational support | 10% | Incident response SLA, on-call availability, GPU replacement time |
| Financial stability | 10% | Funding status, operating history, financial references |
Contract Negotiation: Key Terms and Protections
GPU procurement contracts should include several protections beyond pricing. GPU replacement SLA: maximum 4 hours to detect a failed GPU and 24 hours to replace it. Some providers offer hot-spare GPUs for instant failover. Performance guarantees: minimum NVLink and inter-node bandwidth guarantees, with credits if performance falls below 90% of advertised. Capacity expansion rights: pre-agreed pricing and allocation for expanding GPU count during the contract term, without renegotiation. GPU upgrade rights: ability to upgrade to next-generation GPUs at a predetermined formula (e.g., H100 to B200 at 1.5x the H100 rate).
Termination for convenience: 30-90 day termination clause with minimal penalties (3-6 months GPU payments). This is the most important protection for rapidly evolving AI workloads. Data egress: limits or caps on data egress charges when moving models and data out of the provider's infrastructure.
Price protection: annual price increase cap of 5% or CPI + 2%, whichever is lower. Given declining H100 pricing, buyers should negotiate price reductions rather than guarantees -- a market adjustment clause allowing quarterly repricing for H100 capacity.
Pricing Benchmarks and Market Rates (Mid-2026)
Understanding market pricing is essential for RFQ evaluation. The table below shows mid-2026 pricing benchmarks across contract types and GPU generations. These are based on ClusterBid's transaction data across 40+ providers.
Benchmark ranges vary by provider type: hyperscalers charge a premium (20-40% above neocloud) for integrated services (networking, storage, support). Neoclouds offer the most competitive GPU-only pricing but may have limited add-on services. Bare-metal providers offer the lowest per-GPU prices but require infrastructure management.
| Contract Type | H100 80GB (per GPU-hr) | H200 141GB (per GPU-hr) | B200 192GB (per GPU-hr) |
|---|---|---|---|
| Spot / On-demand | $1.20 - $1.80 | $2.50 - $3.50 | $7.00 - $9.00 |
| 3-month reserved | $1.80 - $2.20 | $2.80 - $3.20 | $6.00 - $7.00 |
| 6-month reserved | $1.50 - $2.00 | $2.50 - $3.00 | $5.00 - $6.00 |
| 12-month reserved | $1.20 - $1.80 | $2.00 - $2.80 | $4.50 - $5.50 |
| 24-month committed | $1.00 - $1.50 | $1.80 - $2.40 | $3.50 - $4.50 |
| 36-month committed | $0.80 - $1.20 | $1.50 - $2.00 | $3.00 - $4.00 |
Post-Contract Vendor Management
After contract signing, ongoing vendor management ensures the relationship delivers value. Monthly capacity reviews track utilisation against committed capacity, alerting to situations where capacity is under-utilised (request capacity reduction before contract lock-in) or over-utilised (trigger capacity expansion rights). Quarterly business reviews (QBRs) review GPU uptime SLAs, support ticket resolution times, and any incidents affecting training or inference workloads.
Performance benchmarking should be conducted quarterly. Run standardised benchmarks (NCCL all-reduce bandwidth, MLPerf training benchmarks, inference latency tests) to verify that the provided infrastructure meets contracted performance levels. If performance degrades, the benchmark results provide evidence for SLA credit claims.
Provider switching capability should be maintained even with a signed contract. Maintain model portability (containerised models that run on any GPU infrastructure) and data portability (model weights and datasets stored in version-controlled repositories accessible from any provider). The switching cost should be a few days of engineering time, not a multi-month migration project. This negotiating leverage ensures continued good service from the current provider.
