All essays
BenchmarkCOMPARISONFEB 2026

GPU Workload Orchestration: Kubernetes vs Slurm vs Ray for AI at Scale

Compare Kubernetes, Slurm, and Ray for GPU workload orchestration. Decision framework covering scheduling, autoscaling, job prioritization, data locality, and hybrid deployments for H100/B200 clusters.

01

THE GPU ORCHESTRATOR TRILEMMA

GPU workload orchestration requires choosing between schedulers that were designed for fundamentally different use cases. Kubernetes was designed for stateless microservices and extended for batch GPU scheduling through Volcano and Koordinator. Slurm was designed for HPC batch jobs with fixed resource allocations and gang scheduling. Ray was designed for distributed Python workloads with dynamic task graphs and elastic scaling. Each makes different tradeoffs in scheduling latency (milliseconds vs seconds vs minutes), autoscaling speed (sub-second pod scaling vs per-job queue wait), resource utilization efficiency (bin-packing vs exclusive node allocation vs elastic scaling), and GPU topology awareness.

The practical consequence is that no single orchestrator is optimal for all GPU workloads. Hyperparameter search jobs (thousands of short-lived trials) benefit from Ray's elastic scaling and task-level parallelism. Large-scale LLM training (hundreds of GPUs for weeks) needs Slurm's gang scheduling and deterministic resource allocation. Model inference serving (latency-sensitive, variable traffic) requires Kubernetes with HPA/VPA and GPU-aware request routing. The most sophisticated GPU clusters use all three in a federated orchestration plane, with a common scheduler hierarchy that routes workloads to the appropriate orchestrator based on resource requirements, duration, and SLAs.

CharacteristicKubernetes (+Volcano)SlurmRay
Scheduling UnitPod (container)Job (batch step)Task (Python function)
Min Scheduling Latency10-50 ms (pod)100-500 ms (job)1-10 ms (task)
Gang SchedulingVia Volcano / KoordinatorNative (salloc --exclusive)Via Placement Group
GPU Topology AwarenessNVIDIA device plugin + NFDslurm.conf Gres + topologyAuto-detected by Ray core
AutoscalingCluster Autoscaler, KarpenterManual (no built-in autoscale)Ray Autoscaler (node-level)
Job PrioritizationPriorityClass + PreemptionPriority + Fairshare (multi-factor)Task queue priorities
Multi-Node NCCL SupportVia Kubeflow Training OperatorNative (sbatch -N 8 --gres=gpu:8)Via Ray Train + NCCL backend
Stateful/Inference ServingNative (Deployment + Service)Not designed for itVia Ray Serve
Ecosystem MaturityBroadest (Helm, Istio, Argo)Deep HPC integrationsPython/ML ecosystem native
02

KUBERNETES: MICROSERVICES HERITAGE WITH BATCH GPU EXTENSIONS

Kubernetes handles GPU workloads through the device plugin framework, which advertises `nvidia.com/gpu` resources across 1-8 count depending on GPU density. The vanilla Kubernetes scheduler (kube-scheduler) treats GPUs as discrete countable resources: a pod requesting `nvidia.com/gpu: 8` must land on a node with 8 available GPUs, and the scheduler uses a best-fit algorithm to minimize fragmentation. However, vanilla Kubernetes lacks gang scheduling (all pods in a distributed training job must start simultaneously for NCCL initialization), GPU topology awareness (NVLink domains within a DGX node require GPUs to be scheduled on the same NUMA node), and batch job queue management.

Volcano (formerly kube-batch) addresses these gaps as an alternative Kubernetes scheduler. Volcano's `Queue` and `PodGroup` CRDs implement gang scheduling: all pods in a training job are held until the scheduler can place them simultaneously, then released for parallel startup. Volcano also supports preemption (higher-priority job preempts running pods with PriorityClass), fair-sharing across queues (e.g., research team gets 40 percent of GPU capacity, production gets 60 percent), and resource reservation through the `Reservation` CRD. The `volcano-bj` scheduler plugin adds NUMA-aware GPU placement: when a pod requests 4 GPUs, Volcano ensures they are on the same NVSwitch domain rather than spanning across DGX nodes. The overhead of gang scheduling adds 200-800ms to pod scheduling time but prevents a 30-second NCCL timeout that would occur when 7 of 8 pods start and the 8th is delayed.

K8s GPU FeatureVanilla KubernetesWith Volcano SchedulerWith Koordinator
GPU SchedulingFit (first available node)Gang + Fair-share queueElastic bin-packing + NUMA
NCCL Multi-NodeManual (Kubeflow Training Op)PyTorchJob CRD + gang scheduleSame as Volcano
PreemptionPriorityClass (preempt lower priority)Queue-based preemption + SLAsColocation with CPU/memory tiers
GPU OversubscriptionNot supportedNot supportedShared GPU + reclaimable resources
Autoscaling GPU nodesCluster Autoscaler (+15-60 sec)Volcano + CA (gang-aware)Koordinator + CA (colocation-aware)
Maturity / Production ReadinessProven for inferenceProven (Ant Group, Huawei, others)Proven (Alibaba, 10k+ node clusters)
03

SLURM: HPC-GRADE GPU JOB SCHEDULING WITH DETERMINISTIC ALLOCATION

Slurm's design for HPC workloads aligns well with large-scale GPU training: jobs declare resource requirements upfront (`--gres=gpu:8 --nodes=8`), Slurm's backfill scheduler finds placements in the queue that maximize utilization without starving large jobs, and allocations are exclusive and deterministic. The slurm.conf GPU configuration specifies both GRES (Generic Resource Scheduling) for GPU counting and GPU topology via `GresTypes=gpu` and `NodeName=DGX-01 Gres=gpu:8`. Slurm's topology plugin (`TopologyPlugin=topology/tree`) ensures that multi-node GPU jobs are allocated within the same leaf switch to minimize NCCL latency.

Slurm's weakness for GPU clusters is threefold. First, Slurm has no built-in autoscaling: if the cluster has 128 GPUs and a 256-GPU job is submitted, the job stays in the queue indefinitely. GPU clusters using Slurm must integrate with a separate auto-scaling layer (e.g., AWS ParallelCluster or GCP Slurm-GCP) that provisions or deprovisions node groups based on queue depth. Second, Slurm jobs are monolithic: dynamic task spawning (Ray's core use case) requires `srun` within the job allocation, which adds 200-400ms per `srun` call. Third, inference serving is not supported: there is no equivalent to Kubernetes Service for routing requests to GPU pods. Despite these gaps, Slurm remains the dominant orchestrator for clusters where training is the sole or primary workload type, with approximately 40 percent of large >1000-GPU clusters using Slurm in 2026.

04

RAY: DISTRIBUTED PYTHON WITH ELASTIC GPU EXECUTION

Ray's task-level parallelism is uniquely suited to GPU workloads that require dynamic execution patterns. Unlike Slurm's fixed allocation or Kubernetes' pod-level scheduling, Ray's scheduler (GCS) allocates GPU resources at the task level from a shared pool of Ray nodes. A root task can spawn 64 child tasks that each request 1 GPU, execute for 10 seconds, and terminate, with the GPUs immediately available for the next wave. This is ideal for hyperparameter optimization (Optuna + Ray Tune), reinforcement learning (multiple environment rollout workers each consuming a GPU for a few seconds), and data preprocessing pipelines that use GPU-accelerated libraries like RAPIDS cuDF.

Ray's elastic autoscaler manages GPU node lifecycle based on task demand: `ray up` creates the initial cluster, and the autoscaler adds or removes nodes based on pending GPU task count. The autoscaler configuration specifies `min_workers`, `max_workers`, and `target_utilization_fraction`. For GPU clusters, `target_utilization_fraction: 0.8` triggers new node provisioning when GPU utilization exceeds 80 percent across the cluster. The autoscaler supports spot/preemptible GPU instance integration with `max_failures` and node provider rate limiting to prevent throttling. Ray's weakness is the lack of gang scheduling: if a training job requires 64 GPUs, Ray assigns them on a first-come-first-served basis, potentially allocating 8 GPUs on 8 different nodes when a single 8-GPU DGX node would be more efficient. Ray Placement Groups (`ray.util.placement_group`) provide a workaround by reserving GPUs before task execution, adding 100-500ms reservation latency.

Workload TypeBest OrchestratorGPU Count RangeKey Metric
LLM Training (large-scale)Slurm128-4096Time-to-train, allocation determinism
Hyperparameter Search (1000s of trials)Ray Tune8-256Trials/hour, GPU utilization
Model Inference ServingKubernetes1-512p99 latency, tokens/sec
Reinforcement Learning (parallel envs)Ray RLlib16-128Env steps/second, GPU utilization
Data Preprocessing (GPU-accelerated)Ray Data / K8s8-64Throughput GB/s per GPU
Hybrid (training + inference on same cluster)Kubernetes + Ray64-1024Resource utilization, scheduling efficiency
05

FEDERATED ORCHESTRATION: HYBRID DEPLOYMENT PATTERNS

Real-world GPU clusters above 512 GPUs rarely use a single orchestrator. The federated orchestration pattern uses Kubernetes as the control plane with Slurm and Ray as co-schedulers for batch workloads. The Kubernetes management cluster runs both Slurm controller and Ray head as stateful workloads, with the GPU nodes partitioned into Slurm and Ray node pools via node labels and taints. A unified admission controller (`gpu-scheduler-admission`) intercepts pod creation and routes the pod to the appropriate scheduler based on annotations: `scheduler.clusterbid.com/scheduler: volcano` for batch training, `scheduler.clusterbid.com/scheduler: ray` for distributed Python tasks, and default Kubernetes for inference services.

This federated approach achieves 85-92 percent cluster-wide GPU utilization, versus 60-75 percent for any single scheduler alone. The key enabling technology is the GPU resource pool abstraction: a resource pool of 256 H100 GPUs is dynamically split across orchestrators based on queue depth and priority. Slurm gets 128 GPUs at 9 AM when batch training jobs flood the queue, Ray gets 192 GPUs when hyperparameter sweep jobs run overnight, and Kubernetes inference pods get priority access to 32 GPUs during business hours with burst capacity from the shared pool. The resource pool rebalancing happens through the cluster autoscaler's expander plugins, with a custom `gpu-priority-expander` that selects the node group based on scheduler queue metrics.

Filed under
Kubernetes GPU OrchestrationSlurm GPU ClusterRay Distributed GPUGPU Workload SchedulingKubernetes vs Slurm vs RayGPU AutoscalingAI Workload Orchestration