The GPU Scheduling Challenge
GPU clusters are expensive shared resources. A single 256-GPU cluster costs $1-1.5M per month. Maximising utilisation while ensuring that high-priority workloads get timely access is the central challenge of AI infrastructure management. Unlike CPU scheduling, GPU scheduling must account for GPU memory topology, NVLink affinity, NCCL ring placement, and long-running jobs that cannot be easily preempted.
At mid-2026, the three dominant GPU orchestrators -- Slurm (32% of AI clusters), Kubernetes (45%), and Ray (18%) -- each approach scheduling with different trade-offs. Slurm excels at batch job scheduling with preemption support. Kubernetes provides containerised workload management with GPU device plugins. Ray handles distributed training with dynamic scaling. The remaining 5% use custom schedulers or cloud-native solutions.
This post covers scheduling policies, preemption strategies, queue management, and fair-share allocation models for GPU clusters, with implementation guidance for each orchestrator.
Priority Levels and Preemption Policies
Most GPU clusters implement 3-4 priority tiers. Critical (real-time inference and production training runs), High (important training experiments, model fine-tuning with deadlines), Normal (standard development training, evaluation runs), and Low (exploratory experiments, spot-equivalent workloads, CI/CD pipeline jobs). Each priority level has different preemption behaviour and resource guarantees.
Preemption policies define what happens when a high-priority job arrives and no GPU capacity is available. Gang preemption (Slurm) preempts all jobs on the required nodes simultaneously. Graceful preemption (Kubernetes) sends SIGTERM to running pods, allowing checkpointing before termination. Queue-based preemption (Ray) waits for running jobs to complete before dispatching high-priority work.
The table below compares preemption models across the major orchestrators. The choice depends on workload characteristics and the cost tolerance for preempted compute.
| Orchestrator | Preemption Model | Checkpoint Support | Preemption Granularity |
|---|---|---|---|
| Slurm | Gang preemption (suspend/requeue) | Built-in (requeue) | Node-level |
| Kubernetes | Graceful (SIGTERM + priority class) | Application-level required | Pod-level |
| Ray | Queue-based (no hard preemption) | Built-in (object store) | Task-level |
| Custom (e.g., Gavel) | Policy-based (combinatorial) | Native checkpointing | Job-level |
GPU-Aware Job Placement
Job placement on GPU clusters must account for: GPU-to-GPU connectivity (NVLink topology -- GPUs on the same NVSwitch domain communicate faster), fabric affinity (jobs requiring inter-node GPU communication should be placed on the same InfiniBand leaf), NUMA proximity (GPU-to-CPU affinity for data loading performance), and node exclusivity (MIG partitions versus full-node GPU allocation).
The placement constraint that most affects training performance is NVLink topology. An 8-GPU H100 node has a full NVSwitch mesh, meaning any GPU can communicate with any other GPU at the full NVLink bandwidth (900 GB/s). However, in a multi-node setup with 32 GPUs (4 nodes), the inter-node bandwidth drops to InfiniBand (typically 400-800 Gbps, or 50-100 GB/s). The scheduler should attempt to place jobs on the minimum number of nodes to maximise intra-node communication.
Bin-packing algorithms are the standard approach: sort GPU nodes by available GPU count, allocate the most constrained nodes first, and spread load to minimise fragmentation. Open-source schedulers use GPU-aware bin-packing by default, but the defaults often need tuning for specific cluster topologies.
Queue Management and Fair-Share Allocation
Multi-tenant GPU clusters require fair-share allocation to prevent one team from consuming the entire cluster's capacity. The three standard mechanisms are: static partitioning (each team has a guaranteed GPU allocation), dynamic fair-share (teams receive GPU time proportional to their historical usage and queue depth), and auction-based (teams bid for GPU time, the highest bidder gets priority).
Static partitioning is simplest but underutilises the cluster (each partition's idle GPUs cannot be used by other teams). Dynamic fair-share, implemented through Slurm's fair-share tree or Kubernetes priority classes, achieves 85-95% utilisation while maintaining allocation fairness. Auction-based allocation (used by the largest AI labs) extracts maximum value from GPU capacity but requires sophisticated cost accounting.
The fair-share decay factor is critical: a well-tuned decay factor (weighting recent usage more heavily than historical usage) prevents teams from accumulating stale priority. The recommended decay half-life is 7-14 days, balancing responsiveness to current workload with stability for long-running training jobs.
Backfill Scheduling and Utilisation Optimisation
Backfill scheduling fills idle GPU capacity with lower-priority jobs without delaying the start of higher-priority jobs. When a high-priority job is queued waiting for resources (e.g., 64 GPUs), the backfill scheduler identifies jobs that can run on the currently available GPUs and will finish before the high-priority job's start time. This dramatically improves utilisation.
In practice, backfill increases GPU cluster utilisation from 60-70% to 80-90% by filling the idle gaps between large job allocations. The challenge is predicting how long currently running jobs will occupy their GPUs, which is straightforward for fixed-duration training but challenging for variable-length inference workloads.
The implementation pattern in Kubernetes uses the Descheduler and priority-based preemption with PodDisruptionBudgets. In Slurm, backfill is built in with configurable parameters for job time limits and reservation policies. The key configuration parameter is the backfill window (typically 30-60 minutes), which limits how far ahead backfill scheduling considers idle resources.
Preemption-Aware Training: Checkpointing and Elasticity
If your cluster uses preemptive scheduling, your training jobs must be preemption-aware. This means: frequent checkpointing (every 15-30 minutes for preemptible jobs), checkpoint path registration in a shared filesystem accessible from any node, and automatic job restart on any available GPU (not requiring the same GPU or node).
Elastic training frameworks (DeepSpeed, PyTorch Elastic, Horovod Elastic) support dynamic GPU membership changes. When a job is preempted on 8 of its 64 GPUs, the remaining 56 GPUs can continue training at reduced throughput while the scheduler allocates replacement GPUs. The training loop handles GPU membership changes through NCCL communicator re-initialisation.
The cost of preemption for a checkpointed job is approximately 2-3% of total compute (the lost compute between the last checkpoint and preemption). For a non-checkpointed job, preemption wastes all compute since the last full save. Teams operating on preemptible GPU instances report that checkpointing overhead of 5-10% is acceptable given the 60-80% cost savings versus reserved pricing.
Scheduling for Inference vs Training
Inference and training have fundamentally different scheduling requirements. Inference jobs need fixed GPU capacity with low latency variance. Training jobs need bulk GPU capacity with high throughput and tolerance for scheduling delays. A well-architected GPU cluster partitions the fleet into an inference pool (20-40% of GPUs, fixed allocation, low preemption, latency-optimised) and a training pool (60-80% of GPUs, dynamic allocation, preemption allowed, throughput-optimised).
The inference pool should use Kubernetes with resource reservations and minimum guaranteed GPU allocations per model instance. The training pool should use Slurm or Ray with backfill scheduling, preemption, and fair-share allocation. Both pools share the same GPU hardware but use different scheduling policies.
At mid-2026, the leading practice is a unified fleet pool where GPUs can shift between inference and training based on demand patterns. The cluster management platform dynamically allocates the split based on queue depth and latency SLAs. This hybrid approach achieves 88-94% overall GPU utilisation while maintaining inference latency SLAs -- significantly better than the 70-80% typical of statically partitioned fleets.
