All essays
GuideGUIDEFEB 2026

AI GPU Energy Optimization: DVFS, Power Capping, and Dynamic Voltage Scaling for Reducing AI Infrastructure Costs

Techniques for reducing GPU energy consumption without sacrificing training throughput. DVFS tuning, power capping limits, undervolting results, and per-workload energy optimization across H100, B200, and B300.

01

The GPU Energy Problem at Scale

A single B300 GPU at 1000W TDP consumes 24 kWh per day. A 1,024-GPU cluster draws 1 MW continuously. At $0.12/kWh industrial electricity rates, power alone costs $2,880 per day or $86,400 per month. For colocated clusters, power is typically the second-largest cost after GPU depreciation, often exceeding 30% of total monthly expenditure.

The carbon implications are equally significant. A 1,024-GPU B300 cluster operating at 65% utilization for a year emits approximately 4,200 metric tons of CO2e at the US average grid carbon intensity. Energy optimization techniques can reduce both cost and emissions by 15-35% without reducing training throughput for most workloads.

02

DVFS Fundamentals for AI GPUs

Dynamic Voltage and Frequency Scaling (DVFS) allows the GPU to adjust its clock speed and voltage based on workload demand. NVIDIA GPUs expose multiple performance states (P-states) through nvidia-smi, ranging from P0 (maximum performance) to P12 (minimum power). By default, training workloads run at P0 to maximize throughput.

The key insight is that many AI workloads are not compute-bound. Memory-bound operations like attention softmax, element-wise activations, and data-loading pipelines spend significant time waiting on HBM bandwidth. Reducing core clock frequency for these operations saves power with near-zero throughput impact. The DVFS overhead from state transitions is approximately 50 us, negligible for training steps lasting 100-500 ms.

03

Power Capping Limits and Tradeoffs

nvidia-smi provides a power cap interface that limits maximum GPU power draw. On the H100 SXM, the default power cap is 700W. Reducing it to 600W cuts peak power by 14% while reducing training throughput by only 2-4% for memory-bound models. At 500W, throughput drops by 8-12% with power savings of 29%. The optimal cap depends on the compute-to-memory ratio of the specific model.

For B300 at 1000W default, a 800W cap reduces measured training throughput by 3-6% across standard LLM benchmarks while cutting power by 20%. The B300's larger HBM3e stack and improved memory controllers make it less sensitive to power capping than the B200, where a 20% power reduction costs 8-10% throughput. Clusters running continuous batching inference can often cap at 75% TDP with under 3% per-token latency increase.

GPUDefault TDPCap at 90%Thruput LossCap at 80%Thruput Loss
H100 SXM700W630W1-2%560W4-6%
H200 SXM700W630W1-2%560W3-5%
B200 NVL1000W900W2-3%800W8-10%
B300 NVL1000W900W1-2%800W3-6%
04

Dynamic Voltage Scaling and Undervolting

Dynamic voltage scaling adjusts GPU voltage independent of frequency. NVIDIA does not publicly expose voltage control via nvidia-smi, but platform tools like NVFlash and custom firmware allow voltage curve modification on select SKUs. Undervolting reduces the voltage supplied at each frequency step, cutting power quadratically (P = CV^2f). A 5% voltage reduction yields approximately 10% power reduction at the same frequency.

Production results from large GPU operators show that most H100 and B200 chips can sustain their rated boost clocks at 45-70 mV below the factory voltage curve. This translates to 8-12% power savings with identical throughput. The caveat is chip-specific variance: approximately 15% of chips fail stability tests at reduced voltage and must run at stock settings. GPU operators should validate undervolt stability with a 48-hour matrix multiplication and attention kernel stress test before deploying at scale.

05

Per-Workload Energy Optimization

Different phases of the training pipeline have different energy profiles. The forward pass is typically memory-bound, benefiting from power capping and reduced core clocks. The backward pass and weight update are compute-bound, requiring full performance to maintain throughput. Dynamic power management that adjusts the cap between forward and backward passes can save 12-18% total energy versus static capping.

Inference workloads show even greater optimization potential. Batch sizes vary throughout the day, and GPU utilization for online inference can drop below 30% during off-peak hours. Dynamic power capping tied to utilization levels can reduce idle GPU power from 150W (idle) to under 75W with brief wake latency. For a 256-GPU inference cluster, this dynamic approach saves approximately $4,000 per month in power costs at $0.12/kWh.

Workload TypeDefault PowerOptimized PowerThroughput ImpactMonthly Savings (256 GPUs)
LLM Training (compute-bound)700W H100630W-2%$2,200
LLM Training (memory-bound)700W H100560W-4%$5,200
Online Inference (batch)700W H100560W cap day / 350W night-2% latency$3,800
Multi-Modal Training1000W B300850W-3%$7,400
06

Implementation Guide for Cluster Operators

Start with nvidia-smi power capping, which is fully supported and requires no firmware changes. Set a cluster-wide default cap at 85-90% of TDP and measure per-model throughput for one week. Models with attention-heavy architectures (long context transformers, multi-modal encoders) tolerate larger caps. Models with dense compute (CNNs, MLP-heavy transformers) lose more throughput per watt saved.

For teams ready to go further, undervolting requires vendor support or custom firmware. Most GPU providers contractually prohibit undervolting. If your provider allows it, validate at the node level before cluster-wide deployment. ClusterBid works with providers that support power optimization, and we include per-workload power capping configuration in our cluster management layer.

Filed under
GPU energy optimizationDVFSpower cappingdynamic voltage scalingAI power efficiencyH100 power tuningB200 undervoltinggreen AI