Why Containerization Matters for GPU Workloads
GPU workloads are notoriously environment-sensitive. A model trained with CUDA 12.4 and PyTorch 2.4 may fail silently or produce different results when run with CUDA 12.6 or PyTorch 2.5. Docker containers solve this by bundling the exact GPU software stack -- NVIDIA driver (minimum version), CUDA runtime, CuDNN library, inference framework, and model weights -- into a portable artefact that runs identically across development laptops, CI/CD runners, and production GPU clusters.
Containerization also enables multi-tenancy on GPU clusters, where different teams run different models with different software stacks on the same GPU hardware. Each team's container includes only the dependencies their model requires, avoiding the dependency conflicts that plague shared GPU environments.
At mid-2026, approximately 85% of production GPU deployments use containerized workloads, up from 55% in 2024. This post covers the best practices for building, optimising, and deploying GPU containers for AI workloads.
NVIDIA Container Toolkit and GPU Runtime
The NVIDIA Container Toolkit (`nvidia-container-toolkit`) is the standard mechanism for exposing GPUs to Docker containers. It replaces the earlier `nvidia-docker2` and provides: automatic GPU detection and enumeration within containers, CUDA version compatibility validation, GPU driver version checking against container requirements, and MIG device partitioning for multi-instance GPU.
The toolkit works by injecting NVIDIA's container runtime (`nvidia-container-runtime`) into Docker's runtime configuration. When a container requests GPUs via `--gpus all` or `--gpus device=0,1`, the runtime mounts the NVIDIA drivers, CUDA libraries, and device files from the host into the container. The key requirement is that the host's NVIDIA driver version must be equal to or greater than the CUDA version expected by the container.
The practical guideline: use the latest stable NVIDIA driver on the host (at mid-2026, version 570.x), and pin a specific CUDA base image in your Dockerfile (e.g., `nvidia/cuda:12.8-runtime-ubuntu24.04`). Do not use `latest` tags for CUDA images, as driver incompatibility between container and host will cause runtime failures.
Dockerfile Optimization for Large Model Images
GPU model containers are typically very large. A container with a 70B-parameter model in FP8, PyTorch, vLLM, and CUDA runtime can exceed 80 GB. Docker image optimization is essential to reduce storage costs, transfer times, and deployment latency.
Multi-stage builds separate the build environment (CUDA toolkit, compiler tools, source code) from the runtime environment (model weights, inference server, minimal runtime libraries). The final image stage includes only what is needed at runtime. This reduces image size by 40-60% for typical model serving containers.
Model weights should not be embedded in the image. Instead, mount weights from a shared filesystem or download from object storage at container start time. This decouples model updates from container rebuilds, reduces image size by 70-90%, and enables A/B testing with different model versions on the same container infrastructure. The image handles model version as an environment variable that selects the weight path.
GPU Device Plugins for Kubernetes
Kubernetes manages GPU allocation through device plugins. The NVIDIA device plugin (`nvidia-device-plugin`) exposes GPUs as node resources (`nvidia.com/gpu`). Pods request GPUs through resource limits, and the plugin handles GPU device assignment and container injection. At mid-2026, the plugin supports MIG partitioning, GPU time-slicing, and GPU sharing.
MIG partitioning is configured at the node level in the device plugin's configuration file. Each node can expose different MIG profiles (e.g., one node exposes 7x 1g.10gb instances, another exposes 2x 3g.40gb instances). Pods request the appropriate profile through the `nvidia.com/mig.config` resource name. This is the recommended approach for multi-tenant GPU clusters because it provides hardware-enforced isolation.
GPU time-slicing (sharing a GPU among multiple pods) is available through the device plugin's time-slicing feature, but it is not recommended for production inference workloads due to unpredictable latency. Time-slicing works best for CI/CD pipelines, experimentation, and low-priority batch workloads where latency variability is acceptable.
Container Registry and Artifact Management
Large GPU model containers (10-80 GB) create challenges for container registries. Standard Docker registries have image size limits (2 GB default for Docker Hub, 10-50 GB for cloud registries with increased limits). Storing multi-GB model containers in registries is expensive and impractical.
The recommended registry architecture for GPU workloads uses a two-tier approach. Tier 1: Standard container registry (ECR, GCR, or self-hosted Harbor) for the inference server container (1-5 GB, multi-stage built, no model weights). Tier 2: Object storage (S3-compatible) for model weights, mounted or downloaded at deployment time. The container image never changes between model updates -- only the weight path changes.
This separation enables rapid model rollouts without the overhead of container image builds and pushes. A new model version is deployed by updating a Kubernetes ConfigMap or environment variable to point to the new weight path, then issuing a rolling update of the inference pods. Deployment time drops from 10-20 minutes (image build + push + pull) to 2-5 minutes (weight download + pod restart).
GPU Container Security
Container security for GPU workloads adds GPU-specific considerations to standard container security practices. The container runtime must not allow access to other tenants' GPU memory. The NVIDIA container runtime enforces this by default, but misconfigurations can expose device files across containers.
The security checklist for GPU containers includes: run containers as non-root (NVIDIA container runtime supports `--user` with GPU access since CUDA 12.2), use read-only root filesystems for inference containers (model weights mounted from external storage), avoid `privileged: true` in Kubernetes pod security contexts (unnecessary for GPU access), and verify MIG isolation when using partitioned GPUs (tenants in different MIG instances should not be able to communicate via shared memory).
For regulated workloads, container image signing (Cosign or Notary) ensures that only approved container images are deployed on GPU clusters. The container image signature chain includes the base image (NVIDIA CUDA), the inference server build, and the deployment configuration. Combined with Kubernetes admission controllers, this provides a complete software supply chain for GPU workloads.
Performance Considerations for GPU Containers
Containerization introduces minimal overhead for GPU workloads when configured correctly. GPU device passthrough via the NVIDIA container runtime has near-zero performance impact because the container directly accesses the GPU device through the driver -- there is no emulation or virtualization layer for CUDA operations.
However, container networking can introduce performance overhead for distributed training. The default Docker bridge network adds latency and bandwidth constraints. For multi-GPU training across nodes, containers should use host networking mode (`--network host` on Docker or `hostNetwork: true` in Kubernetes) to avoid the overhead of container network address translation and port mapping.
Storage access from GPU containers should use GPUDirect-compatible filesystem mounts when possible. Standard NFS mounts through Docker volumes add latency for checkpoint reads and writes. The configuration recommendation: use hostPath mounts for local NVMe scratch storage, and GPUDirect-compatible parallel filesystem mounts (WekaFS, GPUDirect Storage) for shared model weight and checkpoint access.
