All essays
TechnicalDEEP DIVEFEB 2026

GPU Job Scheduling: SLURM vs Kubernetes for AI Workloads

Comparison of GPU job scheduling approaches covering SLURM, Kubernetes with volcano/Kueue, and hybrid architectures for training and inference workload orchestration.

01

THE GPU JOB SCHEDULING LANDSCAPE

GPU job scheduling divides into two primary camps. SLURM dominates traditional HPC and research environments with 88 percent adoption among top-500 supercomputers. Kubernetes with AI extensions (Volcano, Kueve, Run:AI) dominates enterprise deployments with 72 percent adoption for production inference. The convergence trend shows 34 percent of organizations running hybrid schedulers as of 2025, up from 12 percent in 2023.

The fundamental difference is workload philosophy. SLURM schedules batch jobs with fixed resource requirements into queues with priority policies. Kubernetes schedules containers with dynamic scaling and microservice orchestration. Training workloads favor SLURM while inference workloads favor Kubernetes, driving hybrid adoption for AI platform teams.

FeatureSLURMKubernetes + KueueHybrid Architecture
Job queuingNative (priority/partition)Via queue/cohort pluginsBoth
GPU fractional allocationVia GRES/gpuDevice plugin (1 GPU min)Unified
Gang schedulingNative (--gres)Volcano gang schedulingBoth
Auto-scalingManual (sinfo/dynamic)Cluster autoscaler (native)K8s for elastic
Topology awarenessNative (slurm topology)NUMA-aware pluginsExtended NRI
PreemptionNative (preempt)Priority/preemption pluginsCombined
02

TRAINING WORKLOAD SCHEDULING

Training workloads demand gang scheduling (all GPUs available simultaneously) and topology-aware placement. SLURM excels here with native gang scheduling and topology-aware allocation using the hwloc library. A 256-GPU training job on SLURM achieves consistent allocation within 30-120 seconds. On Kubernetes, Volcano enables gang scheduling but adds 60-180 seconds of scheduling overhead.

Node-level GPU allocation differs significantly. SLURM allocates GPUs at job granularity with exclusive node mode preventing co-location by default. Kubernetes supports GPU oversubscription with quality-of-service classes, enabling inference jobs to share nodes with training workloads during off-peak hours, improving cluster utilization from 55 percent to 75 percent.

MetricSLURM (100% Dedicated)K8s (100% Dedicated)K8s Shared Cluster
Avg scheduling latency45 sec120 sec180 sec
Cluster utilization55-65%50-60%72-82%
Preemption time5-15 sec30-90 sec60-180 sec
Topology placementOptimal (hwloc)Good (NRI)Good (NRI)
Job fragmentationLowMediumMedium
03

INFERENCE WORKLOAD SCHEDULING

Kubernetes dominates inference scheduling due to native support for autoscaling, rolling updates, and service discovery. Horizontal Pod Autoscaler with custom metrics (queue depth, request latency) scales inference deployments from 1 to 100 replicas within 60-120 seconds. SLURM lacks native service scaling, requiring external tools to manage inference deployments.

GPU memory oversubscription via MIG or MPS enables higher inference density. Kubernetes with NVIDIA MIG support partitions A100/H100 GPUs into up to 7 instances. A single A100 can serve 3 Llama 3 8B replicas simultaneously with MIG, achieving 3x inference density versus exclusive GPU allocation.

04

HYBRID SCHEDULING ARCHITECTURE

The optimal approach for organizations running both training and inference is a hybrid architecture. SLURM manages training workloads with dedicated GPU partitions for job duration. Kubernetes manages inference workloads with dynamic scaling across shared GPU pools. A federation layer maps GPU resources between schedulers based on demand, achieving 78-85 percent cluster utilization.

Implementation requires a resource accounting layer to prevent conflicts. Run:AI and CoScale provide scheduler-agnostic GPU allocation with policy-based partitioning. Default configuration reserves 60 percent of GPUs for training and 40 percent for inference, with dynamic rebalancing every 15 minutes. Organizations using hybrid scheduling report 28 percent lower GPU spend.

Filed under
SLURMKubernetesJob SchedulingGPU OrchestrationVolcanoKueueWorkload Management