All essays
TechnicalDEEP DIVEFEB 2026

Slurm GPU Management: Partition Configuration, Job Scheduling, and GPU Accounting

Slurm GPU management for AI clusters: partition configuration for GPU nodes, GRES GPU plugin setup, job scheduling policies (backfill, gang, topology), GPU accounting and fairshare, and production configuration for H100/B200 partitions.

01

GRES GPU CONFIGURATION IN SLURM

Slurm manages GPU resources through the GRES (Generic Resource Scheduling) plugin. The slurm.conf configuration defines GPU resources at the node level: GresTypes=gpu and per-node NodeName=h100-gpu-[1-32] Gres=gpu:8 for 8-GPU nodes. The gres.conf file specifies GPU device files and MIG configurations: Name=gpu Type=h100 File=/dev/nvidia[0-7]. For MIG-partitioned GPUs, each MIG slice is defined as a separate GRES: Name=gpu Type=h100-mig-1g File=/dev/nvidia-caps/nvidia-cap[0-55]. NVIDIA provides the slurm-gres-nvidia-plugin which auto-detects GPU topology and MIG profiles.

The GRES plugin supports GPU affinity through --gres-flags=enforce-binding which pins jobs to specific GPU devices. On H100 nodes, --gres=gpu:4 combined with --gres-flags=enforce-binding allocates 4 specific GPUs and sets CUDA_VISIBLE_DEVICES to those device IDs. Slurm 24.11+ adds GPU topology-aware allocation: --gres-flags=topology_fit places GPU allocations within the same NVSwitch domain when possible, significantly reducing all-reduce latency.

02

GPU PARTITION DESIGN AND NODE GROUPING

Production Slurm GPU clusters organize nodes into partitions based on GPU type, interconnect, and workload characteristics. A typical H100 cluster partition scheme: training-h100 (8xH100 NVLink nodes, full-node jobs only, InfiniBand interconnect), inference-h100 (8xH100 PCIe nodes, fractional GPU jobs, Ethernet interconnect), mig-h100 (H100 nodes with MIG 1g slices, single-slice containers), and interactive-h100 (2xH100 nodes, short time limit, low priority). For mixed-vendor clusters, separate partitions prevent workload placement on wrong GPU architectures.

Partition configuration in slurm.conf: PartitionName=training-h100 Nodes=h100-trn-[1-128] Default=NO MaxNodes=128 DefaultTime=24:00:00 State=UP. The OverSubscribe=NO directive prevents CPU overcommit on GPU partitions. AMD MI400 nodes use a separate partition with the rocm feature label: PartitionName=amd-mi400 Nodes=mi400-[1-32] Feature=rocm. The Priority parameter controls partition-level scheduling priority.

PartitionNode TypeGPU CountInterconnectJob TypeDefault Time Limit
training-h100H100 SXM8 per nodeNVSwitch + InfiniBand NDR400Full-node training jobs24 hours
inference-h100H100 PCIe8 per nodePCIe Gen5 + Ethernet 100GFractional GPU serving7 days
mig-h100H100 MIG-enabled56 MIG slicesNVLink (shared per GPU)Single-slice inference48 hours
b200-trainingB200 SXM8 per nodeNVSwitch 5 + InfiniBand XDRFull-node training24 hours
amd-mi400MI4008 per nodeInfinity Fabric + IB NDR400ROCm training24 hours
interactiveH100 SXM2 per nodeEthernet 25GDev/debug sessions1 hour
debugH100 SXM8 per nodeNVSwitchQuick testing1 hour
03

JOB SCHEDULING POLICIES FOR GPU WORKLOADS

Slurm's scheduling policies significantly impact GPU cluster utilization. The default FIFO scheduler leaves GPUs idle when large jobs queue behind smaller ones. Backfill scheduling (SchedulerType=sched/backfill) enables small jobs to run while large jobs wait for resources, improving utilization by 15-30% on GPU clusters. For training jobs requiring 8 or more GPUs, gang scheduling (GangSchedType=preempt) allows higher-priority jobs to preempt running jobs with checkpoint resumption.

GPU topology-aware scheduling in Slurm 24.11+ uses the TopologyParam=TopologyTree directive to build a tree of GPU interconnects. Jobs requesting --switches=1@00:10:00 require all allocated GPUs to be within one switch domain. The --gpu-bind=none|per_task|closest flag controls GPU-to-CPU binding. For AI training, --gpu-bind=closest binds each GPU to the nearest NUMA domain, reducing PCIe latency by 8-12%.

Scheduler ConfigurationParameterGPU Cluster ImpactRecommendation
Backfill Schedulingsched/backfill+15-30% utilizationEnable bf_interval=60
Gang Schedulingpreempt-type=gang+20% throughput for mixed jobsEnable with checkpoint
Topology-AwareTopologyTree + switches-25% all-reduce latencyRequired for NVSwitch
FairsharePriorityType=priority/multifactorPrevents team resource hoggingEnable with 3 factors
GPU Exclusive--exclusivePrevents GPU sharing conflictsDefault for training
Job Array--array=1-100%4Parallel hyperparameter tuningLimit concurrent %N
04

GPU ACCOUNTING AND FAIRSHARE

Slurm's accounting infrastructure tracks GPU usage through the sacct and sacctmgr commands. GPU accounting requires AccountingStorageTRES=gpu in slurm.conf and the jobacct_gather/cgroup plugin to report per-job GPU utilization. The sacct --format=JobID,Elapsed,AllocTRES%40,ConsumedEnergyRaw,GPUFrac command shows per-job GPU allocation and utilization fraction. The GPUFrac field, new in Slurm 24.11, reports actual GPU utilization as a fraction of allocated GPU time.

Fairshare scheduling (PriorityType=priority/multifactor) for GPU clusters uses three factors: Fairshare (usage history), Partition Priority (partition-level weighting), and Job Size (large GPU jobs get higher priority). For GPU clusters shared between training and inference teams, configure separate associations with different fairshare weights. GPU chargeback tracks GPU-hours consumed per team via sacctmgr, supporting internal billing at $2.50-5.00/GPU-hour for H100 resources.

05

PRODUCTION SLURM.CONF FOR GPU CLUSTERS

A production slurm.conf for a 128-node H100 GPU cluster includes these critical directives: GresTypes=gpu,mps (enable GPU and MPS sharing), SelectType=select/cons_res (consumable resource selection), SelectTypeParameters=CR_GPU (GPU as a consumable resource), PriorityType=priority/multifactor (fairshare), SchedulerType=sched/backfill (backfill scheduling), PreemptMode=gang (gang preemption), TaskPlugin=task/affinity,task/cgroup (GPU binding), PrologFlags=Contain (containerize prolog/epilog).

Performance-critical Slurm configuration parameters for GPU workloads: MessageTimeout=60 (reduce from default 120 for faster job startup on InfiniBand clusters), TreeWidth=512 (optimized for 128+ node GPU clusters), SchedulerParameters=default_queue_depth=10000,sched_interval=30,bf_interval=60,bf_max_time=60000,bf_resolution=500. The bf_resolution=500 millisecond setting enables sub-second GPU allocation granularity for backfill scheduling. SchedulerParameters=max_gpu_job_launch=8 limits concurrent GPU job launches.

06

SLURM-READY GPU INSTANCES ON CLUSTERBID

ClusterBid lists 3,200+ GPU instances across 14 providers that support Slurm workload manager. The --scheduler slurm filter identifies providers offering pre-installed Slurm or Slurm-on-demand installation. Instances are tagged with Slurm version (slurm-24.11, slurm-24.08), GRES configuration (gres/gpu:8), and MIG support. For dedicated Slurm GPU clusters, ClusterBid partners with 6 providers offering managed Slurm control nodes with automatic partition configuration.

Pricing for Slurm-managed GPU clusters on ClusterBid: 8xH100 Slurm partition at $24-36/hour; 8xMI400 Slurm partition at $18-28/hour; 8xB200 Slurm partition at $40-60/hour. The --slurm-config parameter accepts a custom slurm.conf fragment that ClusterBid applies during provisioning. Slurm-enabled H100 partitions deliver 92-95% scheduling utilization versus 70-80% for shared cloud GPU queues, with 30-45% lower effective cost per GPU-hour.

Filed under
Slurm GPUSlurm GPU AccountingSlurm GRES GPUSlurm Partition ConfigurationGPU Job SchedulingSlurm Fairshare GPUHPC GPU Cluster