UTILIZATION METRICS AND DEMAND MODELS
Primary metric: effective GPU utilization = sum(DCGM_FI_DEV_GPU_UTIL > 10% over 5min) / sum(total GPU-seconds). A cluster at 85% with 12% idle allocation indicates waste. For 256 GPUs: utilization below 60% = no expansion; 60-75% = monitor; 75-85% = plan; above 85% = immediate procurement.
Demand forecasting uses multiplicative decomposition: forecast = base_hours * (1+growth)^months * seasonality * event. AI training demand doubles every 8-12 months. Weekly pattern: research runs Monday-Thursday, weekends drop 40-60%. End-of-quarter model releases cause 30% spikes.
Forecast outputs probability distribution: P50 drives baseline planning, P90 drives maximum procurement. Gap between P50 and P90 defines needed buffer. A cluster at P50=80%, P90=95% needs 15% buffer; 84% is max safe steady-state utilization.
| Metric | Source | Normal Range | Expansion Signal |
|---|---|---|---|
| Effective utilization | GPU active min / total GPU min | 40-75% | >85% for 2 weeks |
| Queue wait time (P50) | avg submit to start | <15 min | >60 min for 1 week |
| Queue depth (pending) | jobs waiting for GPU | <5% of GPUs | >20% of total GPUs |
| Preemption rate | preempted / total | <5% | >15% non-preemptible |
| Peak-to-average | peak/avg GPU-Hr | 1.5-2.5 | >3.5 (fragmentation) |
| Monthly growth | this/last month -1 | 5-15% | >20% for 3 months |
EXPANSION TRIGGERS AND PROCUREMENT TIMELINES
Soft triggers (planning): utilization >75% for 14 days, P99 wait >30 min, growth >15% month-over-month for 3 months. Initiate RFQ, 8-12 weeks. Hard triggers (immediate): utilization >85% for 7 days, queued GPU-hours >25% of total. Expedited procurement 4-8 weeks.
Procurement timeline 2026: H100/H200 lead times 8-16 weeks, B200 12-20 weeks, server integration 2-4 weeks, facility upgrades 8-16 weeks, InfiniBand switches 4-8 weeks. Total: 22-44 weeks from trigger to production. Maintain rolling 18-month capacity reserve.
Expansion in capacity chunks: 8 GPUs (1 leaf port), 32 GPUs (4 leaf ports), 64 GPUs (full rack), 128+ GPUs (new leaf switch). Minimum viable expansion determined by fabric port, power, cooling, and uplink availability.
| Expansion Size | GPU Count | Infrastructure Needs | Lead Time | Estimated Cost |
|---|---|---|---|---|
| Single Node | 8 | 1 leaf port, 7kW, 8 RU | 8-16 weeks | $350-500K |
| Half Rack | 32 (4 nodes) | 4 leaf ports, 28kW, 32 RU | 12-20 weeks | $1.4-2.0M |
| Full Rack | 64 (8 nodes) | 8 leaf ports, 56kW, 48 RU | 16-24 weeks | $2.8-4.0M |
| Multi-Rack | 256 (32 nodes) | New leaf switch, 224kW | 22-36 weeks | $11-16M |
| Cluster Expansion | 512+ | New spine, fabric partition | 28-44 weeks | $22-40M+ |
GPU SKU SELECTION AND FUTURE-PROOFING
H100 SXM 80GB: default for training, $30-35K/GPU. H200 with 141GB HBM3e: 1.4-1.8x speedup on memory-bound workloads at 1.3x price. B200: 2x H100 FP8 at 1.6x price, requires CUDA 12.6+, NCCL 2.22+. Model effective throughput per dollar for workload mix.
Technology refresh schedule: H100 for immediate capacity, B200 for next cycle 12-18 months out, option for B300 (expected 2027). Rotating strategy: new nodes run latest architecture, older nodes relegated to inference, nodes >5 years decommissioned.
Power is often binding constraint: 8 H100 nodes (64 GPUs) draws 56-64 kW requiring liquid cooling. Air-cooled limited to 15-20 kW/rack. Direct-to-chip liquid cooling at $8-12K/rack, power distribution upgrade at $15-25K/rack.
SCENARIO MODELING AND WHAT-IF ANALYSIS
Three scenarios: Base Case (current growth), Upside Case (+30% demand for new model), Downside Case (efficiency improvements). Each models utilization over 24 months. Size expansion for Upside, fund only Base Case incrementally.
What-if: if teams improve utilization from 50% to 75%, effective capacity increases 50% without hardware. If all training adopts FP8, throughput doubles. Utilization efficiency coefficient in model shows: 10% improvement delays $5M expansion by 6 months, saving $500K.
Output: 12-18 month procurement roadmap with decision points. Month 0: RFQ for 64 H100s. Month 6: evaluate utilization, decide option for 64 more. Month 12: B200 evaluation begins. Each point has utilization threshold preventing over/under-provisioning.
BUDGETING AND FINANCIAL METRICS FOR GPU CAPACITY
Annual TCO for 1024-GPU H100 cluster: hardware $8M, facilities $2.2M, operations $1.8M, network $1.2M. Total $13.2M ($12,900/GPU/year, $1.47/GPU-hour at 100%). Budget scenarios: Base $13.2M, Expansion $18.5M (+256 GPUs), Efficiency $11.5M.
Cost per effective GPU-hour: $2.10 at 40% utilization, $1.60 at 80%. Marginal cost for expansion: $5.3M for 256 GPUs providing 1.12M incremental GPU-hours at 50% utilization = $4.73/GPU-hour. Payback period: approximately 13 months.
Annual demand survey to team leads. Teams overestimating >50% two quarters in a row moved to higher-cost on-demand tier. Accurate forecasts retain reserved pricing. Improves forecast accuracy from 60-70% to 85-90% over 4-5 cycles.
| Budget Category | Annual Cost (1024 H100) | Per GPU/Hr | Cost Driver |
|---|---|---|---|
| GPU Hardware (depreciation) | $8,533,333 | $0.95 | 3yr straight-line $25K/GPU |
| Server Hardware (depreciation) | $3,413,333 | $0.38 | CPU, RAM, NVMe, chassis |
| InfiniBand Fabric | $1,200,000 | $0.14 | Switch/Router depreciation |
| Power (14 MW) | $2,012,800 | $0.23 | $0.10/kWh, PUE 1.4 |
| Data Center Space | $168,000 | $0.02 | $150/kW-month |
| Operations + Software | $1,800,000 | $0.21 | 5 FTE, monitoring, orchestration |
| Total Annual TCO | $17,127,466 | $1.91 | At 100% utilization |
| Effective at 40% Util | $17,127,466 | $4.78 | 3,584 effective hours/GPU |
| Effective at 80% Util | $17,127,466 | $2.39 | 7,168 effective hours/GPU |
