MLOPS PIPELINE FOUNDATIONS FOR GPU CLUSTERS
MLOps pipelines for GPU clusters face constraints that traditional CI/CD systems do not. Training runs on 8-512 GPUs take hours to days, consume $500-$50,000 per run in compute cost, and must be reproducible across environments with different GPU SKUs (H100, A100, H200) and CUDA versions. A production MLOps pipeline must orchestrate five phases: data validation and preprocessing (CPU-bound), distributed training initialization (multi-GPU rendezvous), training execution (GPU-bound), model evaluation and validation (GPU or CPU), and staged deployment with traffic shifting.
The pipeline must also handle GPU-specific failure modes: NCCL timeout (typically 30 seconds default, requires tuning to 60-120 seconds for cross-node training), CUDA out-of-memory errors that require automatic batch size reduction with gradient accumulation scaling, and node failures that trigger elastic training job rescheduling. The reference architecture uses Kubernetes for orchestration with Kubeflow Pipelines or Argo Workflows as the DAG engine, MLflow for experiment tracking and model registry, and KFServing or vLLM for deployment with automatic scale-to-zero for inference endpoints.
CI/CD WITH GITHUB ACTIONS AND GPU SELF-HOSTED RUNNERS
GitHub Actions workflows for GPU projects require self-hosted runners with NVIDIA driver and CUDA toolkits installed. The base setup configures an H100 or A100 GPU host with the actions-runner agent, NVIDIA Container Toolkit for Docker GPU access, and pre-cached base images. The runner labels enable workflow targeting: `self-hosted`, `gpu-h100`, `cuda-12.4`. A typical training CI pipeline triggered on `git push` runs code quality checks on CPU, compiles CUDA kernels on GPU (90-180 seconds), runs single-GPU unit tests with small batch sizes (5-10 minutes), and optionally triggers a full multi-GPU training run as a separate workflow dispatch.
Artifact caching is critical: PyTorch and CUDA download artifacts add 3-5 minutes per run without caching. A `setup-cuda` action that caches `/usr/local/cuda-12.4` and conda environments reduces cold-start overhead from 300 seconds to 15 seconds. For multi-GPU testing, a GitHub Actions matrix strategy with `matrix.gpu-count: [1, 2, 4]` enables parallel validation across configurations. The end-to-end CI pipeline for a 4-GPU training job should complete in under 20 minutes for most model architectures, with the constraint that GitHub Actions has a 6-hour job timeout, sufficient for all but the largest training runs.
| CI Phase | Runner Type | Duration | GPU Count | Trigger | Cache Key |
|---|---|---|---|---|---|
| Code quality (ruff, mypy) | CPU-4core | 45-90 sec | 0 | Pull Request | N/A |
| CUDA kernel compile | GPU-H100 | 90-180 sec | 1 | Pull Request | CUDA 12.4 + NVCC |
| Single-GPU unit tests | GPU-H100 | 5-10 min | 1 | Pull Request | cached conda env |
| Multi-GPU integration | GPU-H100x4 | 15-30 min | 4 | Manual dispatch | cached NCCL tests |
| Full training run | GPU-H100x8+ | 1-24 hrs | 8+ | Merge to main | checkpoint artifacts |
ARGO WORKFLOWS: DAG ORCHESTRATION FOR GPU TRAINING
Argo Workflows is the most widely adopted DAG orchestrator for GPU training pipelines on Kubernetes. Each training step is an Argo template running in a Kubernetes pod with GPU resource requests (e.g., `nvidia.com/gpu: 8`), init containers for dataset synchronization, and sidecar containers for metrics emission to Prometheus. A typical pipeline DAG has parallel branches: data preprocessing runs on a CPU pod with high-memory (64 GB per worker), while a separate branch validates that the preprocessing output is available. After both complete, the training step launches with GPU pods using Volcano or Koordinator for gang scheduling.
The Argo YAML manifest for a 64-GPU training step specifies `parallelism: 8`, each with 8 GPUs across 8 nodes, using the `training-operator` template from Kubeflow. The manifest also defines retry policies with backoff (3 retries, 60-second initial backoff, exponential) for transient NCCL failures, artifact tracking with S3 checkpoint storage, and interdependencies with the model registry step using MLflow's tracking URI as an Argo parameter. Argo's workflow-of-workflows pattern enables nested pipelines: the training workflow can trigger a separate validation workflow on completion, which triggers the deployment workflow via a webhook to ArgoCD.
MLFLOW MODEL REGISTRY WITH GPU-SPECIFIC METADATA
MLflow serves as the central model registry with GPU-specific metadata that distinguishes models by target hardware. Each registered model version stores `gpu_architecture` (H100, A100, H200), `cuda_version` (12.4, 12.6), `precision` (FP32, AMP, BF16, FP8), `minimum_vram_gb` (80, 96, 144), and `nvidia_compute_capability` (9.0 for H100, 8.0 for A100). The registry tags enable the deployment system to automatically select the appropriate model version for the target GPU SKU, preventing deployment of an FP8-optimized model to an A100 cluster that lacks Transformer Engine support.
MLflow's model registry stages (Staging, Production, Archived) map to deployment pipeline gates. A model promoted to Staging triggers a canary deployment to 10 percent of inference traffic on a single H100 node for 1 hour of observability (p50/p99 latency, throughput, error rate). If metrics pass the SLO (p99 < 50ms, error rate < 0.1 percent), an automated promotion to Production deploys to the full 32-node cluster with blue-green rollout. MLflow's `model_version` transitions are enforced via CI checks: a `validate_compatibility` GitHub Actions step confirms that the Staging metrics meet the Production SLO before allowing the transition API call.
| Metadata Field | Example Value | Used By | Validation Gate |
|---|---|---|---|
| gpu_architecture | H100-SXM-80GB | Deployment scheduler | Must match target cluster GPU |
| cuda_version | 12.4 | Container build step | CUDA compat check on deploy |
| precision | BF16 + FP8 | Inference config | FP8 requires Compute Capability >=9.0 |
| min_batch_size | 32 | Load tester | Validated during canary phase |
| model_family | llama-3-70b | Rate limiter | Quota allocation per model family |
| training_dataset_hash | sha256:abc123... | Compliance audit | Provenance check before production |
CANARY DEPLOYMENT AND TRAFFIC SHIFTING FOR GPU INFERENCE
GPU inference deployments use traffic mirroring and canary shifting to validate model performance without impacting user-facing latency. The standard pattern deploys the new model version to a shadow inference pod that receives mirrored traffic alongside the current production version. Traffic mirroring adds 5-10 percent CPU overhead on the production pod for connection replication but requires no additional GPU capacity. The shadow pod outputs are compared to production outputs every 100 requests using cosine similarity, with an alert threshold of similarity < 0.95 triggering automatic rollback. This phase runs for 10-15 minutes with approximately 10,000 mirrored requests.
After shadow validation passes, canary deployment shifts 5 percent of real traffic to the new model version for 30 minutes. During canary, the pipeline monitors p50/p99 GPU inference latency, throughput per GPU (tokens/second), error rate, and KV cache utilization. For text generation models with vLLM, the critical metric is TTFT (Time to First Token, target < 200ms at p99) and ITL (Inter-Token Latency, target < 30ms per token). If all metrics remain green, the canary shifts to 25 percent for 1 hour, then 100 percent. Argo Rollouts manages this traffic shifting with Service Mesh (Istio) integration, providing automated rollback if the canary health check fails.
PIPELINE OBSERVABILITY AND COST TRACKING
MLOps pipeline observability requires tracking both technical metrics (GPU utilization, NCCL ring completion time, data loading I/O wait) and business metrics (cost per training run, cost per inference request, model training efficiency). Prometheus metrics scraped from each pipeline step include `gpu_utilization_percent`, `step_duration_seconds`, `data_loader_wait_seconds` (time GPUs spend waiting for data), and `estimated_cost_usd` computed from GPU-hours multiplied by the cluster GPU pricing rate (e.g., $3.50/hr for H100). These metrics power Grafana dashboards that show pipeline efficiency over time and detect regressions.
A critical metric is the Pipeline GPU Efficiency ratio: `total_step_gpu_seconds / (total_wall_seconds * gpu_count)`. A well-optimized pipeline achieves >85 percent efficiency; values below 70 percent trigger investigation into data loading bottlenecks, NCCL synchronization overhead, or checkpoint write stalls. MLflow automatically logs these efficiency metrics to each run record, enabling comparison across model versions, data versions, and cluster configurations. The combination of technical and cost observability enables infrastructure teams to answer the essential question: "Did this pipeline change improve model quality within the same GPU budget, or did it just consume more compute?"
