RATIO PLANNING FRAMEWORK
The training-to-inference GPU ratio depends on the organization's ML maturity stage, model volume, and user base. Early-stage teams with 1-3 models in production often need a 70-30 training-to-inference split. Mature organizations with 10+ deployed models and millions of users typically invert to 30-70. The optimal ratio also depends on GPU generation, since older GPUs remain viable for inference longer than for training.
TRAINING WORKLOAD MATH
Training GPU requirements scale with model size, dataset volume, and iteration frequency. A team fine-tuning Llama 4 70B weekly on 50B tokens needs approximately 256 H100-hours per training run, or 1,024 GPU-hours per month. Pre-training a 70B model from scratch requires 2,000-4,000 H100-days. Training workloads are predictable and schedulable, making them ideal for reserved GPU capacity with utilization targets above 85%.
INFERENCE WORKLOAD MATH
Inference GPU requirements scale with daily active users, average query length, and latency targets. A consumer product serving 100M daily tokens at sub-200ms latency needs approximately 8-16 H100s for batch sizes of 32-64. Inference workloads have predictable baseline traffic patterns with unpredictable spikes to 3-5x baseline during peak hours. Inference GPU utilization typically ranges from 30-60%, making it a candidate for dynamic allocation strategies.
DYNAMIC ALLOCATION STRATEGIES
Dynamic allocation requires a GPU orchestration layer that can preempt inference during training priority windows and scale inference during traffic surges. Kubernetes with the Volcano scheduler and GPU metrics server enables this. Key metrics for auto-scaling decisions include GPU memory pressure, inference queue depth, and training checkpoint urgency. Most mature teams use a 3-tier priority system: production inference, experimental training, preemptible batch jobs.
SCALING DECISIONS FOR 2026
Scale the inference fleet when daily token volume exceeds 70% of sustained capacity for 7 consecutive days. Scale the training fleet when new model training cycles require more than 80% of available compute for more than 2 weeks. Consider GPU generation separation: allocate H100s to training and older A100s to inference for optimal cost. The Vera Rubin generation arriving in late 2026 will likely shift this calculus for early adopters.
