All essays
InfrastructureINFRASTRUCTUREFEB 2026

GPU Cluster User Management: Multi-Tenant Isolation, Resource Quotas, and Priority Scheduling

Manage multi-tenant GPU clusters with Slurm QoS, Kubernetes ResourceQuota, fairshare scheduling, priority preemption, and user access controls. Isolate teams, enforce quotas, and optimize GPU allocation fairness.

01

SLURM QOS AND PARTITION DESIGN FOR GPU CLUSTERS

Three QoS levels: high-priority (GrpTRES=gpu=512, MaxWall=168h, Priority=1000), normal (GrpJobs=50/user, GrpTRES=gpu=256, Priority=500), low (GrpTRES=gpu=1024, Priority=100, Preempt=YES). Configured via sacctmgr add qos with respective limits. Partitions map to QoS: gpu-prod, gpu-research, gpu-low.

Physical partitions: gpu-h100-80gb, gpu-a100-80gb, gpu-h100-mig. Logical partitions: gpu-prod (reserved, QoS gpu-prod), gpu-research (shared, QoS gpu-research), gpu-preemptible (shared, QoS gpu-low). Nodes assigned in slurm.conf.

FairShare: each team (account) has FairShare 1-100. Decay half-life 7 days, PriorityWeightAge=0 (no queue time bonus), PriorityFavorSmall=NO. Recalculation every 60s via scontrol reconfig.

Preemption: PreemptMode=REQUEUE, send_user_signal=USR1 at 30s before termination for checkpointing. PreemptType=partition_prio: higher-priority partition preempts lower. USR1 handler in training framework enables seamless resume.

ParameterQoS: gpu-prodQoS: gpu-researchQoS: gpu-low (preempt)
Priority1000500100
GrpTRES=gpu5122561024
MaxWall168:00:00 (7 days)24:00:004:00:00
MaxJobs/user1020100
PreemptNoNoYes
FairShare weight1005010
Charge factor1.0x1.0x0.3x
02

KUBERNETES RESOURCE QUOTAS AND LIMIT RANGES

ResourceQuota per namespace: spec.hard.nvidia.com/gpu: 32 (GPU limit), requests.nvidia.com/mig-1g.10gb: 16 (MIG slices), limits.cpu: 64, count/deployments.apps: 20. Prevents any namespace from exhausting cluster resources.

LimitRange per namespace: default.nvidia.com/gpu: 2, max.nvidia.com/gpu: 8, min.nvidia.com/gpu: 1. default.cpu: 4, default.memory: 16Gi. Ensures GPU pods have adequate host resources.

PriorityClass: gpu-prod-high (1,000,000), gpu-research-mid (500,000), gpu-development-low (100,000). Higher priority preempts lower when resources are insufficient. PDB protects production inference pods.

Resourceprod-team nsresearch-team nsci-cd ns
GPU quota64164
Max GPU per pod842
Default GPU per pod211
CPU quota128 cores32 cores8 cores
Memory quota512 Gi128 Gi32 Gi
Storage quota (SSD)10 Ti2 Ti500 Gi
Priority classprod-high (1M)research-mid (500K)dev-low (100K)
Preemption allowedNo (PDB protected)NoYes
03

USER IDENTITY MANAGEMENT AND ACCESS CONTROL

User management integrates with LDAP/AD for centralized identity. SSH key auth via AuthorizedKeysCommand fetching from LDAP. Kubernetes RBAC: ClusterRole gpu-user for submitting pods, gpu-viewer for read-only. OIDC groups mapping.

User onboarding: LDAP account creation, group assignment (gpu-prod-team, gpu-research-team), Slurm account via sacctmgr, Kubernetes RoleBinding. Job submit plugins validate group membership against partition access.

Offboarding: LDAP deactivation, SSH key removal, OIDC token expiry, Slurm MaxJobs=0. Audit log records event. Running jobs terminated via scancel or kubectl delete pods.

04

PRIORITY SCHEDULING AND PREEMPTION STRATEGIES

Slurm priority: PriorityWeightQOS=1000, PriorityWeightFairShare=500, PriorityWeightPartition=100. QoS is dominant factor, fairshare secondary tiebreaker. PriorityCalcPeriod=60.

Backfill: bf_interval=60, bf_max_time=86400 (schedule up to 24h ahead). Smaller jobs fill gaps left by large jobs waiting for all GPUs. bf_continue keeps backfill running after main scheduling pass.

Gang scheduling (time-slicing): fixed time slices (1h) for interactive workloads. Context switching takes 5-15s for GPU memory load, ~1-3% overhead per hour. Acceptable for development.

Topology-aware placement: select/cons_res + topology/block places jobs on nodes sharing the same leaf switch. topology.conf defines SwitchName hierarchy. Allocates GPUs from minimum switches, improving NCCL bandwidth 15-30%.

StrategyThroughput ImpactFairnessLatencyBest For
FIFOLow (idle GPUs)HighHighBatch with reservations
BackfillHigh (fills gaps)MediumLowMixed 500+ GPUs
FairShare (priority)HighVery highMediumMulti-team clusters
PreemptiveVery highLowVery lowProd + preemptible mix
Gang (time-slice)Medium (overhead)HighVery lowInteractive/dev
Topology-awareHigh (NCCL optimized)NeutralLowMulti-rack clusters
05

USAGE MONITORING, REPORTING, AND ENFORCEMENT

Slurm accounting to Elasticsearch via JobCompType=jobcomp/elasticsearch. Daily aggregate per-user GPU-hours, per-partition utilization, queue wait percentiles. Grafana dashboard with per-user self-service view.

Enforcement: job submit plugin checks GPU request against historical utilization. If user averages <40% utilization for requested count over 7 days, plugin suggests reduction. K8s admission webhook validates quota and per-user limits.

Waste tracking: GPU hours allocated but <5% utilized for >30 minutes. Team waste displayed on dashboard, included in weekly report. >20% waste for 2 consecutive weeks: QoS priority reduced 10%. Reduces waste from 15-25% to 8-12% in 3 months.

Filed under
GPU Multi-TenantSlurm QoS GPUKubernetes ResourceQuotaFairshare Scheduling GPUGPU Priority PreemptionUser GPU Access ControlCluster Partition Management