All essays
TechnicalDEEP DIVEFEB 2026

GPU Over-Provisioning in 2026: Why AI Teams Waste 35-45% of Their Compute Budget and How to Fix It

GPU over-provisioning in AI teams wastes 35-45% of compute spend. Learn to measure real utilization, rightsize your fleet, and cut costs without sacrificing performance.

01

The 37% Problem: What Utilization Data Actually Shows

GPU over-provisioning in AI teams is not a niche problem - it is the default state. According to FinOps Foundation research across enterprise AI organizations in 2026, average GPU utilization sits at roughly 37%. That means for every three dollars you spend on compute, about two are doing real work. The third one is paying for idle silicon.

The 35-45% waste range is not a worst-case scenario. It is the median band for companies that have not deliberately instrumented and audited their GPU fleet. The number spikes higher - sometimes to 60% waste - at companies that scaled quickly, provisioned ahead of demand, or have multiple teams sharing a cluster without any chargeback discipline.

Where does the idle time actually go? Data loading stalls account for a larger share than most teams expect. A GPU sitting at 0% utilization while PyTorch waits for the next batch from NFS still shows up as reserved and billing. Checkpoint saves to slow storage pause training jobs for minutes at a time. Model compilation on first run (torch.compile, TensorRT export) consumes wall-clock time without doing productive flops. And then there are the gaps between experiments - someone finishes a run at 11 PM, the next run does not start until 10 AM, and the cluster idles for eleven hours at full reserved cost.

02

5 Budget Leaks That AI Teams Consistently Miss

Training clusters that sit idle between runs are the most obvious category, but teams normalize it because it feels like necessary slack. The standard justification is 'we need it available when we need it.' The real answer is spot instances with checkpointing, which eliminates that cost entirely for non-latency-critical training. We have seen teams save $40,000/month just by shutting down a 16xH100 cluster on weekends and resuming from checkpoint Monday morning.

Oversized inference fleets are subtler and more expensive at scale. A team deploying Llama 3.1 8B on H100 SXM5 nodes because 'that is what we use for training' is paying $2.50-3.00/hr for a GPU that an L40S ($1.50-1.80/hr) handles without breaking a sweat for that model size. The H100's 80GB HBM and 3.35 TB/s memory bandwidth is genuinely wasted on an 8B parameter model that fits in 16GB at FP16. Multiply this across a fleet serving sub-30B models and you are looking at 40-50% inference overspend.

Dev and test waste is the category no one wants to talk about. Developers spin up 8xA100 instances to test a training script, find a bug on line 12, and leave the instance running while they fix it. Two hours later, $80 has evaporated. Multiply by a team of 15 ML engineers and a culture of 'just spin one up,' and dev/test waste alone can run $15,000-30,000/month at a mid-size AI company. Failed jobs are a close cousin - a distributed training run that hits an OOM error on step 1 still burns GPU-hours for the full job startup sequence across every node.

03

Measuring What Actually Matters: Beyond the nvidia-smi Lie

Here is the trap that catches most teams: nvidia-smi reports GPU utilization as 'any kernel was running in this sampling interval.' That is not the same as 'the GPU was doing useful compute.' A GPU that runs one tiny kernel every 100ms - maybe a memory copy or a health check - reports 100% utilization in nvidia-smi if the sampling period is short enough. You can have a training job 'at 100% GPU utilization' that is actually spending 70% of wall-clock time waiting for data.

The metrics that actually matter come from NVIDIA's DCGM (Data Center GPU Manager). The key ones are: DCGM_FI_PROF_GR_ENGINE_ACTIVE (Graphics Engine Active - the fraction of time compute engines were running), DCGM_FI_PROF_PIPE_TENSOR_ACTIVE (Tensor Core utilization - the only thing that matters for transformer workloads), and DCGM_FI_PROF_DRAM_ACTIVE (memory bandwidth utilization). A well-optimized training job on H100 should show Tensor Core utilization above 60% and DRAM activity above 70%. If you are seeing 30% Tensor Core utilization on a 'busy' cluster, you have a preprocessing or data loading bottleneck, not a compute bottleneck.

SM (Streaming Multiprocessor) utilization is the metric between nvidia-smi and DCGM - it tells you what fraction of SMs had at least one active warp. Good training workloads hit 85-95% SM utilization. Inference is trickier: at low batch sizes, SM utilization can be 20-30% even when the service is 'running fine' by latency SLOs. That low utilization at the SM level is the signal that you could be serving multiple model replicas on the same GPU - something MIG (Multi-Instance GPU) partitioning enables on H100 and H200 hardware.

04

Which GPU to Use for Which Model: The Inference Rightsizing Map

The single biggest rightsizing win for most AI teams in 2026 comes from matching GPU memory to model size instead of defaulting to 'whatever we have in training.' The math is simple: a 7B parameter model at FP16 needs ~14GB VRAM plus KV cache overhead, putting it comfortably on an L40S (48GB, $1.50-1.80/hr spot). The same workload on an H100 SXM5 (80GB, $2.50-3.00/hr) is paying for 66GB of unused HBM. That is a 40-50% cost premium for headroom you will never use.

H100 SXM is genuinely worth the premium in specific situations: models above 40B parameters at FP16, heavy batching workloads that need the 3.35 TB/s memory bandwidth, or latency-sensitive inference where NVLink interconnect matters for tensor parallelism. For everything smaller, the L40S or even the A100 80GB ($1.80-2.20/hr) delivers better cost efficiency. The A100 PCIe (40GB variant) is too small for most production deployments now, but the 80GB A100 SXM4 remains one of the most underpriced GPUs in the current market.

For quantized models - and most production inference runs at FP8 or INT4 by 2026 - the effective memory footprint drops further. A Llama 3.3 70B model at INT4 AWQ fits in roughly 38GB, which puts it on a single L40S or A100 80GB with room for a healthy KV cache. Running that on an H200 ($3.07-3.16/hr) because 'bigger is better' is a $1.50/hr mistake per GPU, times 24 hours, times however many replicas you are running.

Model SizePrecisionMin VRAMBest GPUCost/hr
7B-13BFP1614-26 GBL40S (48GB)$1.50-1.80
30B-34BFP1660-68 GBA100 80GB SXM$1.80-2.20
70BFP835-40 GBL40S or A100 80GB$1.50-2.20
70BFP16140 GB2x H100 80GB$5.00-6.00
8x7B MoEFP845-50 GBH100 80GB PCIe$2.00-2.50
405B+FP8200 GB+H200 or B200$3.16+
05

2 Utilization Levers Most Teams Have Not Pulled Yet

Spot instances for development and non-latency training are the most impactful underused tool in 2026. H100 spot rates on neoclouds and the secondary market range from $0.34-1.03/hr - against $2.50-3.00/hr on-demand. That is a 65-86% discount. The objection is always 'but we will lose the run if it gets preempted.' The counter is: if you are not checkpointing every 30-60 minutes anyway, you are going to lose runs to hardware failures too. Preemption-resilient training via Megatron-LM checkpointing or FSDP's save_policy is table stakes for any serious training operation. Add spot, save 70% on dev and experimental training costs.

Batch scheduling for non-latency workloads is the second lever. Evaluation pipelines, batch embedding generation, offline re-ranking, and dataset preprocessing do not need dedicated always-on GPUs. They need compute for 2-4 hours a day, at best. Scheduling these as batch jobs against shared capacity - or against spot instances with a queue - can cut the GPU-hours required by 60-70% compared to running a dedicated instance that sits idle while waiting for the next scheduled run. SLURM with a GPU partition or Kubernetes with the `batch.kubernetes.io` job API handles this cleanly.

06

How to Audit Your Current GPU Spend in 3 Days

Day one: collect baseline utilization data. Deploy DCGM Exporter (it is open source, runs as a DaemonSet on Kubernetes) and point it at a Prometheus/Grafana stack. If you are on AWS, Azure, or GCP, their native GPU monitoring surfaces DCGM metrics directly in CloudWatch, Azure Monitor, and Cloud Monitoring. You want 72 hours of data minimum to capture weekend patterns, which is where a lot of idle waste accumulates.

Day two: categorize your spend by workload type. Pull your GPU reservation and usage data - cloud providers expose this via Cost Explorer (AWS), Cost Management (Azure), or Billing (GCP). Tag every GPU reservation with: training (large model), training (small model/fine-tune), inference (production), inference (dev/staging), or dev/test/experimental. Most teams are surprised to find dev/test consuming 20-30% of total GPU spend with <10% actual utilization during business hours and near-zero overnight.

Day three: build the waste map. Cross-reference the utilization data with the spend categories. For each category, compute: total GPU-hours reserved, GPU-hours with >50% SM utilization, and the gap. That gap, multiplied by the hourly rate, is your addressable waste. For inference, also compute the ratio of GPU VRAM to actual model footprint - anything below 50% VRAM utilization per GPU is a candidate for downsizing or MIG partitioning. A $500,000 annual GPU budget with 40% waste represents $200,000 in annual savings that does not require any model performance tradeoff.

07

The Rightsizing Procurement Path: From Audit to Action

The audit tells you what you are wasting and where. The next problem is executing the transition - moving from over-provisioned on-demand H100s to a mix of spot for dev, smaller GPUs for appropriate inference workloads, and reserved capacity only for what genuinely needs it. This is where most teams stall. Cloud provider on-demand pricing for the right GPU tier often costs more than the wrong tier on a reserved contract. The math only works if you are buying or renting the rightsized hardware at competitive rates.

In practice, rightsizing often means leaving the hyperscaler for specific workloads. An L40S on AWS on-demand runs $1.96/hr. The same GPU through a neocloud provider on a 1-month reserved rate runs $1.50-1.65/hr. An H100 SXM5 that you no longer need for inference but still want for training at a lower cost is available on the secondary market and through platforms like ClusterBid at $1.00-1.50/hr spot - fractions of the $2.50-3.00/hr on-demand rate. The savings from rightsizing are only fully realized when the procurement channel matches the compute tier.

Procurement pricing is also dynamic. H100 spot hit $0.34/hr briefly in Q1 2026 as Blackwell upgrade cycles flooded secondary supply. B200 reserved rates held at $3-5/hr through the same period due to HBM3e constraints. The teams who audited their fleets in early 2026 and shifted training workloads to spot H100 captured savings that will not be available when H100 secondary market supply tightens again. Timing matters, and having a sourcing relationship that surfaces those windows is worth more than the occasional discount. ClusterBid's inventory reflects live market rates across neoclouds and secondary market supply - it is the fastest way to see where the arbitrage actually sits right now.

08

What Good GPU Utilization Actually Looks Like

For training clusters, target average SM utilization above 75% during scheduled training windows, with Tensor Core utilization above 55%. Idle periods should be intentional and bounded - 'this cluster is off on weekends and resumes from checkpoint Monday at 8 AM' is a policy, not a failure. Unplanned idle - a job that crashed at 2 AM and the cluster sat unused until the team noticed at 9 AM - is waste. Alerting on idle clusters above a cost threshold ($50 of GPU-hours with <10% utilization for >2 hours) is worth setting up.

For inference, the target depends on the SLO. A latency-sensitive production service is correctly sized when P95 time-to-first-token stays within spec at peak traffic - which often means running at 60-70% average utilization to have headroom for spikes. But dev and staging inference should be running at much higher utilization (70-90%) or sharing capacity across multiple models via MIG. Running a full dedicated H100 for a staging endpoint that serves 10 requests per minute is the kind of waste that piles up invisibly in cost reports.

The honest summary: most AI teams that have not explicitly instrumented and audited their GPU utilization are wasting between one-third and one-half of their compute budget. The good news is that this waste is recoverable without any model performance tradeoff. Smaller GPUs for appropriate inference workloads, spot for development, MIG partitioning for low-QPS staging services, and scheduled shutdown for idle training clusters - these are not exotic optimizations. They are the basics of GPU FinOps that hyperscaler default pricing has allowed teams to skip. At 20-30% of total engineering spend going to AI infrastructure, skipping them is getting expensive.

Filed under
GPU UtilizationFinOpsCost OptimizationDCGM MetricsRightsizingAI InfrastructureCompute Budget