GPU OPERATOR ARCHITECTURE: NVIDIA VS AMD
The NVIDIA GPU Operator (v24.9) deploys and manages all NVIDIA GPU software components on Kubernetes through a set of operators: the NVIDIADriver daemon handles driver installation; the NVIDIAContainerToolkit daemon set configures the container runtime; the NVIDIADevicePlugin exposes GPUs as extend resources; the NVIDIAMIGManager partitions GPUs into MIG slices; the GFD (GPU Feature Discovery) labels nodes with GPU attributes; and the DCGMExporter exports GPU metrics to Prometheus. The operator is installed via a single Helm chart: helm install gpu-operator nvidia/gpu-operator --namespace gpu-operator --create-namespace.
The AMD GPU Operator (v2.5) follows a similar pattern but with an important architectural difference: AMD splits its operator into the amd-gpu-operator (which manages the ROCm driver, container runtime, and device plugin) and the k8s-device-plugin (which manages SR-IOV virtual functions). AMD's operator uses Node Feature Discovery (NFD) from the Kubernetes ecosystem rather than a custom feature discovery component. The AMD operator's Helm chart is: helm install amd-gpu-operator amd/amd-gpu-operator --namespace amd-gpu --create-namespace.
| Component | NVIDIA GPU Operator 24.9 | AMD GPU Operator 2.5 |
|---|---|---|
| Driver Management | NVIDIADriver daemon (auto-download) | ROCmDriver daemon (kernel module) |
| Container Runtime | nvidia-container-toolkit | rocm-container-runtime |
| Device Plugin | nvidia.com/gpu (MIG-aware) | amd.com/gpu (SR-IOV-aware) |
| GPU Partitioning | MIGManager (MIG profiles) | SR-IOV VF configuration |
| Feature Discovery | GPU Feature Discovery (custom) | Node Feature Discovery (K8s ecosystem) |
| Metrics Export | DCGM Exporter (Prometheus) | AMDSMI Exporter (Prometheus) |
| Validation | GPU health validator daemon | ROCm health check init container |
| Helm Chart Size | 200+ configurable values | 80+ configurable values |
NVIDIA GPU OPERATOR: PRODUCTION SETUP AND VALUES TUNING
A production NVIDIA GPU Operator deployment requires careful values.yaml configuration. The critical parameters: devicePlugin.enabled=true (default, required for GPU scheduling); migManager.enabled=true for MIG-partitioned GPUs; gfd.enabled=true for GPU feature labels; dcgmExporter.enabled=true for Prometheus metrics; toolkit.enabled=true for container runtime configuration. For H100 clusters, set toolkit.version: v1.18.0-ubuntu24.04 for CUDA 13 compatibility. The migManager.migProfileDefaults must match the GPU SKU: H100 SXM supports profiles all-1g.11gb, all-2g.22gb, all-3g.46gb, all-4g.46gb, all-7g.78gb.
The validation command sequence for a healthy NVIDIA GPU Operator deployment: kubectl get pods -n gpu-operator (all pods Running); kubectl describe node | grep nvidia.com/gpu (GPU capacity listed); kubectl logs -n gpu-operator -l app=nvidia-device-plugin-daemon (device plugin registered each GPU). A common production issue is the driver compatibility check failing on RHEL-based nodes with Secure Boot enabled. Workaround: helm upgrade gpu-operator nvidia/gpu-operator --set driver.mksImage=nvidia/mkisofs:latest --set driver.rdma.enabled=true.
| Configuration | Helm Value | Default | Production Recommendation |
|---|---|---|---|
| MIG Manager | migManager.enabled | false | true (for MIG deployment) |
| Default MIG Profile | migManager.defaultGPUsAsMIGEnabled | false | true (if all GPUs use MIG) |
| Container Runtime | toolkit.version | v1.14.0-ubuntu22.04 | v1.18.0-ubuntu24.04 (CUDA 13) |
| Metrics | dcgmExporter.enabled | true | true |
| Feature Discovery | gfd.enabled | true | true |
| GPU Resource Name | devicePlugin.resourceName | nvidia.com/gpu | nvidia.com/gpu (default) |
| Node Feature Labels | gfd.plugins | [gpu, mig, ...] | [gpu, mig, memory, timeout] |
AMD GPU OPERATOR: SR-IOV CONFIGURATION AND ROCM DEPLOYMENT
The AMD GPU Operator configuration centers on SR-IOV virtual function management. The key values.yaml parameters: devicePlugin.enabled=true; sriovDevicePlugin.enabled=true (SR-IOV VF allocation); rocmDriver.enabled=true (ROCm kernel module installation); rocmDriver.version: 6.5.0; amdsmiExporter.enabled=true. The SR-IOV configuration specifies VF count per GPU: sriovDevicePlugin.sriovResources[0].vfAmount: 8 creates 8 virtual functions per MI400 GPU. AMD's device plugin registers each VF as a separate amd.com/gpu resource.
Production AMD GPU Operator validation: kubectl get pods -n amd-gpu (all pods Running); kubectl exec -n amd-gpu deploy/amd-smi-exporter -- rocm-smi --showvf (lists VF topology). AMD's operator supports dynamic SR-IOV reconfiguration: kubectl annotate node <node> amd.com/sriov-vf-count=12 changes VF count without draining workloads. The RuntimeClass for AMD GPU pods must be configured as runtimeClassName: amd-gpu in the pod spec.
MONITORING, ALERTING, AND OBSERVABILITY
Both operators export GPU metrics to Prometheus via their respective exporters. NVIDIA's DCGM Exporter provides 250+ GPU metrics including: DCGM_FI_DEV_GPU_UTIL (GPU utilization 0-100%), DCGM_FI_DEV_MEM_COPY_UTIL (memory bandwidth utilization), DCGM_FI_DEV_POWER_USAGE (power draw in milliwatts), DCGM_FI_DEV_SM_CLOCK (SM clock frequency), DCGM_FI_DEV_NVLINK_BANDWIDTH_TOTAL (total NVLink bandwidth), and DCGM_FI_DEV_GPU_TEMP (temperature in Celsius). The recommended Prometheus recording rules include avg by (node) (rate(DCGM_FI_DEV_GPU_UTIL[5m])) for cluster GPU utilization.
AMD's AMDSMI Exporter provides comparable metrics: amdgpu_utilization_percent, amdgpu_memory_used_bytes, amdgpu_power_draw_watts, amdgpu_temperature_celsius, and amdgpu_vf_count (number of active VFs). AMD's metrics coverage is narrower than NVIDIA's: 80 metrics versus 250, with no SM occupancy or memory bandwidth utilization metrics. The recommended alert thresholds for both operators: GPU utilization < 30% for > 5 minutes, GPU temperature > 85C, GPU power > 95% of TDP for > 10 minutes, and memory utilization > 90%.
| Monitoring Dimension | NVIDIA DCGM Exporter | AMD AMDSMI Exporter |
|---|---|---|
| Total Metrics Exported | 250+ | 80+ |
| GPU Utilization | Yes (per-GPU) | Yes (per-VF) |
| Memory Utilization | Yes (allocated/free/total) | Yes (used/total) |
| Power Draw | Yes (mW) | Yes (W) |
| Temperature | Yes (multiple sensors) | Yes (junction/memory) |
| NVLink Bandwidth | Yes (per-link) | N/A (Infinity Fabric) |
| SM Occupancy | Yes | Not available |
| Memory Bandwidth Utilization | Yes | Not available |
| VF Count / MIG Slices | Yes (per-profile) | Yes (per-VF) |
PRODUCTION DEPLOYMENT PATTERNS AND COMMON PITFALLS
For multi-node GPU training on Kubernetes, both operators support node-affinity-based GPU scheduling. The pattern: label GPU nodes with gpu-type: h100 or gpu-type: mi400, then use nodeSelector and tolerations in pod specs. For NVLink-connected GPU pods, set requiredDuringSchedulingIgnoredDuringExecution node affinity to ensure all GPUs are on the same node. For multi-node training with InfiniBand, both operators support RDMA device plugin registration.
Common pitfalls: First, MIG configuration ordering - NVIDIA GPU Operator requires MIG to be configured BEFORE the device plugin starts, otherwise GPU capacity is double-counted. Fix: helm upgrade --set migManager.enabled=true --set devicePlugin.enabled=false && helm upgrade --set devicePlugin.enabled=true in two steps. Second, AMD SR-IOV VF exhaustion - each VF consumes 1 GB of host memory for DMA buffers; 16 VFs per MI400 on a 32-GPU node consumes 512 GB host RAM. Third, GPU time-slicing conflicts with MIG: if both timeSlicing.enabled and migManager.enabled are true, the time-slicing scheduler can preempt MIG slices.
GPU OPERATOR READY INSTANCES ON CLUSTERBID
ClusterBid's provider network includes GPU instances pre-configured with both NVIDIA GPU Operator and AMD GPU Operator support. The --gpu-operator filter identifies instances with the operator pre-installed or available for installation. NVIDIA GPU Operator-ready instances across 12 providers start at $3.10/hour for a 1x H100 node and $28.00/hour for an 8x H100 NVLink node. AMD GPU Operator-ready MI400 instances start at $2.40/hour for a 1x MI400 node.
The recommended ClusterBid workflow for GPU Operator deployment: (1) Filter by GPU operator readiness and GPU type. (2) Select a provider with documented Helm values. (3) Use ClusterBid's --setup-script feature to automate helm install as a post-provisioning step. (4) Validate with the gpu-operator-validator job. (5) Deploy AI workloads with kubectl apply -f training-job.yaml using nvidia.com/gpu: 8 or amd.com/gpu: 8 resource requests.
