CUDA BASE IMAGE SELECTION: SIZE, COMPATIBILITY, AND RUNTIME TRADEOFFS
Selecting the right CUDA base image is the most impactful optimization for GPU container size and performance. NVIDIA publishes four base image families: `cuda:12.6.0-runtime-ubuntu22.04` (CUDA runtime + libcuda, no development tools, ~1.2 GB), `cuda:12.6.0-base-ubuntu22.04` (minimal CUDA driver stubs only, ~600 MB), `cuda:12.6.0-devel-ubuntu22.04` (full CUDA toolkit including nvcc, cuBLAS, cuDNN headers, ~4.5 GB), and `nvidia/cuda:12.6.0-cudnn-runtime-ubuntu22.04` (runtime + cuDNN, ~2.8 GB). The rule: use runtime images for inference serving, base images + pip-install CUDA Python wheels for training, and devel images only in CI build stages.
Image size differences have real operational impact. Pulling a 4.5 GB devel image across 64 GPU nodes (DGX-8 H100) generates 288 GB of network traffic and takes 3-5 minutes per node at 10 Gbps. The identical application built on the 1.2 GB runtime image pulls in 20-30 seconds per node. For clusters using image caching registries (Harbor, Amazon ECR, GCP Artifact Registry with pull-through cache), the first-pull cost is paid once, but image scan (Trivy, Snyk) takes proportionally longer: a devel image with 800+ packages scans in 4-6 minutes versus 120-180 packages in 30-60 seconds for a runtime image. The smaller attack surface of runtime images also reduces CVEs: NVIDIA CUDA devel images averaged 147 critical/high CVEs in 2025, versus 23 for runtime images of the same CUDA version.
| Image Base | Compressed Size | Uncompressed Size | C/C++ Compiler | Average CVEs | Use Case |
|---|---|---|---|---|---|
| nvidia/cuda:12.6.0-base-ubuntu22.04 | 180 MB | 600 MB | No | 12 | Custom CUDA loader, minimal surface |
| nvidia/cuda:12.6.0-runtime-ubuntu22.04 | 380 MB | 1.2 GB | No | 18 | Inference serving (most common) |
| nvidia/cuda:12.6.0-devel-ubuntu22.04 | 1.4 GB | 4.5 GB | Yes (nvcc, gcc) | 147 | CI build stages only |
| nvidia/cuda:12.6.0-cudnn-runtime-ubuntu22.04 | 520 MB | 2.8 GB | No | 25 | Inference with cuDNN ops |
| nvcr.io/nvidia/pytorch:24.10-py3 | 780 MB | 2.9 GB | Yes (nvcc + Python) | 95 | Training (NGC curated, includes PyTorch) |
MULTI-STAGE BUILD PATTERNS FOR GPU APPLICATIONS
Multi-stage Docker builds for GPU applications separate the compilation environment from the runtime environment. The build stage uses `FROM nvidia/cuda:12.6.0-devel-ubuntu22.04 AS builder` with the full CUDA toolkit for compiling CUDA kernels and C++ extensions (e.g., flash-attention, vLLM custom CUDA ops, xFormers). The runtime stage uses `FROM nvidia/cuda:12.6.0-runtime-ubuntu22.04` with only the artifacts copied from the builder. This reduces the final image from 4.5 GB to 1.7 GB for a typical training image, and from 3.2 GB to 1.1 GB for a typical inference image with PyTorch.
The multi-stage pattern for a training container with custom CUDA ops: stage 1 installs `cuda-toolkit-12-6`, `gcc-11`, `cmake`, and compiles `git clone --depth 1 https://github.com/... && pip wheel .`. Stage 2 copies the `.whl` and installs with `pip install --no-build-isolation *.whl` along with PyTorch's CUDA 12.6 wheels. The `--no-build-isolation` flag prevents pip from redundantly installing CUDA headers from PyPI wheels, which would add 400 MB of duplicate CUDA libraries. For vLLM inference containers, the NVCC compilation of the `_C` extension adds 3-5 minutes to the build stage, which is acceptable in CI but should be cached using Docker layer caching: `docker build --cache-from vllm:cache-$CUDA_VERSION.$ARCH`.
NVIDIA CONTAINER TOOLKIT CONFIGURATION FOR PRODUCTION
The NVIDIA Container Toolkit (formerly nvidia-docker2) provides the bridge between container runtimes and GPU hardware. The toolkit consists of three components: `nvidia-container-runtime` (OCI runtime hook that injects GPU devices and libraries into the container), `nvidia-container-toolkit` (the nvidia-container-cli binary that configures the container's GPU environment), and `libnvidia-container` (the Go library for GPU device discovery). The toolkit version must align with the NVIDIA driver: **driver version 550.x requires nvidia-container-toolkit >=1.16.0, driver 535.x requires >=1.14.0**.
Production GPU container configuration goes beyond the basic `--gpus all` flag. The pod or container spec should set `NVIDIA_VISIBLE_DEVICES: all` (or specific GPU UUIDs for MIG isolation), `NVIDIA_DRIVER_CAPABILITIES: compute,utility` (restricts to compute and utility, omitting display/graphics), `NVIDIA_REQUIRE_CUDA: cuda>=12.0` (prevents container launch on nodes with incompatible CUDA driver). For H100 MIG partitions, `NVIDIA_MIG_CONFIG_DEVICES: all` partitions the GPUs on container start, and `NVIDIA_MIG_CONFIG_DEVICES: 0,1` selects specific GPUs for MIG configuration. The `nvidia-ctk` tool validates runtime configuration: `nvidia-ctk runtime configure --runtime=docker --set-as-default && systemctl restart docker`.
| Config Variable | Setting | Effect | Verification |
|---|---|---|---|
| NVIDIA_VISIBLE_DEVICES | all | Makes all detected GPUs available to container | nvidia-smi inside container shows expected count |
| NVIDIA_DRIVER_CAPABILITIES | compute,utility | Only init CUDA + nvidia-smi (no display) | container startup 200ms faster vs 'all' |
| NVIDIA_REQUIRE_CUDA | cuda>=12.0 | Blocks container launch if driver CUDA < 12.0 | nvidia-container-cli info --requirements |
| NVIDIA_MIG_CONFIG_DEVICES | all | Applies MIG configuration on container start | nvidia-smi -L shows MIG devices |
| NVIDIA_DISABLE_REQUIRE | true (for CI) | Skips driver version check in CI containers | Only use in non-production environments |
| -e NVIDIA_READY_LOG | 1 | Logs 'GPU ready' to stdout when CUDA init completes | Container logs show GPU readiness timestamp |
LAYER CACHING AND BUILD OPTIMIZATION FOR GPU CONTAINERS
GPU container build optimization focuses on maximizing Docker layer cache hits. The standard pattern orders Dockerfile layers from least-changing to most-changing: OS packages (`apt-get install` rarely changes), CUDA runtime (changes with CUDA version updates, approximately quarterly), Python runtime (`pip install` for base packages like numpy, changes monthly), application Python dependencies (changes with code), and application code (changes with every commit). Each cache miss invalidates all subsequent layers, so the goal is to push framework and CUDA layers high in the Dockerfile and application code low.
The `uv` package manager (Python) or `mamba` (conda) provides 3-5x faster dependency resolution than pip or conda alone. A Dockerfile for a training container with PyTorch 2.5 and CUDA 12.6: `RUN --mount=type=cache,target=/root/.cache/uv uv pip install torch==2.5.0 --index-url https://download.pytorch.org/whl/cu126` uses Docker's BuildKit mount cache to avoid re-downloading PyTorch wheels across rebuilds, saving 800 MB of download per build. For `pip` users, `pip install --cache-dir /tmp/pip-cache torch` with Docker's `--mount=type=cache,target=/tmp/pip-cache` achieves the same result. The build time reduction: first build 12 minutes, cached layer rebuild 45 seconds for a code change.
PODMAN AND ROOTLESS GPU CONTAINERS
Podman provides a daemonless alternative to Docker with built-in rootless execution, which is increasingly important for multi-tenant GPU clusters where container escape from a root daemon is a threat vector. Podman rootless GPU containers require additional configuration because GPU device access is traditionally a root operation. Podman 4.9+ implements GPU device access through `podman machine init --cpus 16 --memory 64 --gpus 1 --volume /path/to/data:/data` using the `nvidia-container-toolkit` OCI hook in rootless mode. The rootless GPU container user must be in the `video`, `compute`, and `render` groups for device access.
Rootless GPU containers on H100 have two constraints. First, MIG partitioning is not available in rootless mode because MIG configuration requires NVIDIA Management Library access that is restricted to root. Second, NCCL InfiniBand/RoCE v2 inter-container communication in rootless mode requires `CAP_NET_RAW` and `CAP_NET_ADMIN` to create RDMA-CM IDs and set GIDs, which are not granted to rootless containers by default. Podman's solution is `podman run --device=nvidia.com/gpu=all --group-add=keep-groups --cap-add=NET_RAW --userns=keep-id nvcr.io/nvidia/pytorch:24.10-py3`. Rootless GPU container adoption is growing: 25 percent of new GPU clusters deployed in 2026 use Podman instead of Docker for the security advantages, despite the MIG and RDMA limitations.
