NVIDIA CONTAINER TOOLKIT: ARCHITECTURE AND INSTALLATION
The NVIDIA Container Toolkit (v1.18) is the standard GPU container runtime for Linux. Its architecture comprises three layers: the nvidia-container-toolkit shared library (a runc pre-start hook that mounts GPU devices, driver libraries, and CUDA runtime into the container); the nvidia-container-runtime (an OCI runtime wrapper that transparently injects the toolkit hook); and the nvidia-container-cli (a command-line tool for querying GPU access). Installation via sudo apt install nvidia-container-toolkit configures Docker automatically.
The toolkit's ldconfig hook mounts the NVIDIA driver libraries from the host into the container. The device mount hook maps /dev/nvidia0, /dev/nvidiactl, /dev/nvidia-uvm, and /dev/nvidia-uvm-tools into the container. The NVIDIA_VISIBLE_DEVICES environment variable controls GPU access: all, 0,1,2 (specific indices), or GPU-<UUID>. For production: docker run --gpus all -e NVIDIA_VISIBLE_DEVICES=0,1 -e NVIDIA_DRIVER_CAPABILITIES=all nvidia/cuda:13.0-base.
| Component | Function | Configuration File | Key Parameters |
|---|---|---|---|
| nvidia-container-toolkit | OCI pre-start hook library | /etc/nvidia-container-runtime/config.toml | swarm-resource = DOCKER_RESOURCE_GPU |
| nvidia-container-runtime | OCI runtime wrapper | /etc/docker/daemon.json | runtimes: {nvidia: ...} |
| nvidia-container-cli | CLI tool for GPU config | N/A (CLI flags) | nvidia-ctk runtime configure --runtime=docker |
| nvidia-ctk CDI hook | CDI spec generator | /etc/cdi/nvidia.yaml | nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml |
| libnvidia-container | Go library for GPU mount | N/A (programmatic API) | Used by containerd, Podman, CRI-O |
| nvidia-docker2 | Docker v1 plugin (legacy) | /etc/docker/daemon.json | Deprecated |
CONTAINERD GPU RUNTIME CONFIGURATION
Containerd, the container runtime used by Kubernetes, requires explicit GPU runtime configuration. The nvidia-container-toolkit provides containerd integration via sudo nvidia-ctk runtime configure --runtime=containerd, which adds the nvidia runtime to /etc/containerd/config.toml. The resulting configuration adds a runtime section with runtime_type = io.containerd.runc.v2 and appropriate options. Containerd GPU pods are launched with runtimeClassName: nvidia in the pod spec.
The containerd GPU runtime supports MIG-partitioned GPUs through the NVIDIA_MIG_CONFIG_DEVICES environment variable. For MIG, set NVIDIA_VISIBLE_DEVICES=MIG-<UUID>. For production Kubernetes, the ctr command tests GPU access: ctr run --gpus 0 --rm nvidia/cuda:13.0-base nvidia-smi. Performance impact of containerd's GPU runtime is negligible: 0.2-0.5 milliseconds additional startup time versus direct runc.
PODMAN GPU SUPPORT WITH CDI
Podman 5.0+ uses the Container Device Interface (CDI) specification for GPU access. The CDI specification for NVIDIA GPUs describes each GPU as a device with specific mounts, device node permissions, and environment variables. The NVIDIA Container Toolkit generates the CDI specification automatically: sudo nvidia-ctk cdi generate --output=/etc/cdi/nvidia.yaml. Podman accesses GPUs via podman run --device nvidia.com/gpu=all nvidia/cuda:13.0-base nvidia-smi.
CDI offers several advantages over Docker's --gpus flag. CDI is runtime-agnostic: the same /etc/cdi/nvidia.yaml works with Podman, containerd, CRI-O, and SingularityCE. CDI supports multiple device vendors simultaneously: --device nvidia.com/gpu=0 --device amd.com/gpu=0. CDI enables MIG partition naming: each MIG slice appears as a separate device. Podman's GPU performance is within 0.1% of native because CDI does not add runtime overhead.
| Runtime Capability | Docker + nvidia-container-toolkit | Containerd + nvidia-runtime | Podman + CDI |
|---|---|---|---|
| GPU Access Method | --gpus flag | runtimeClassName: nvidia | --device nvidia.com/gpu=all |
| Configuration File | /etc/docker/daemon.json | /etc/containerd/config.toml | /etc/cdi/nvidia.yaml |
| MIG Support | NVIDIA_VISIBLE_DEVICES=MIG-UUID | NVIDIA_MIG_CONFIG_DEVICES | CDI device name per MIG slice |
| Multi-Vendor GPU | No (NVIDIA only) | No (NVIDIA only) | Yes (NVIDIA + AMD + Intel) |
| Startup Overhead vs Native | ~0.5ms | ~0.5ms | ~0.1ms |
| Rootless GPU Access | No (requires root) | No | Yes (Podman 5.0+ rootless) |
| Kubernetes Native | Supported (CRI-O/Docker shim) | Native (CRI plugin) | Via CRI-O CDI support |
OCI RUNTIME SPEC AND GPU DEVICE MOUNTS
The OCI Runtime Specification defines how devices are exposed to containers through the linux.devices array. A GPU device mount in the OCI spec specifies: path (/dev/nvidia0), type (c for character device), major (195 for NVIDIA GPU), minor (0-255 for device instance), fileMode (0666), and uid/gid (0 for root). The NVIDIA Container Toolkit's pre-start hook reads these fields and injects the correct device nodes before the container process starts.
For MIG devices (CUDA 13+), the OCI spec must include /dev/nvidia-caps/nvidia-cap<N>, /dev/nvidia-caps/nvidia-cap-mig<N>, and /dev/nvidiactl. Each MIG slice is a separate character device. The nvidia-container-cli info command displays the device map. Container runtimes that implement the OCI spec (runc, crun, youki) all support NVIDIA GPU access through the pre-start hook mechanism.
GPU CONTAINER RUNTIME FOR KUBERNETES
Kubernetes GPU support has evolved through three phases. Phase 1 (2018-2022): Docker-in-the-middle with nvidia-docker2. Phase 2 (2022-2025): Containerd native with nvidia-container-runtime. Phase 3 (2026+): CDI-based device allocation with CRI-O or containerd. The current recommended stack is containerd with the NVIDIA Container Toolkit v1.18+ and the NVIDIA GPU Operator v24.9+.
The RuntimeClass resource defines the GPU runtime for pods: apiVersion: node.k8s.io/v1; kind: RuntimeClass; metadata: name: nvidia; handler: nvidia. For CDI-based Kubernetes (K8s 1.28+), the RuntimeClass handler is cdi and pods request GPUs via CDI device names. The crictl runp --runtime nvidia command validates GPU runtime configuration.
GPU CONTAINER READY INSTANCES ON CLUSTERBID
ClusterBid's GPU instances are pre-configured with the appropriate container runtime for each provider. The --container-runtime filter selects instances by runtime type: containerd+nvidia (standard for Kubernetes), docker+nvidia (development), podman+cdi (rootless GPU), and cri-o+cdi (lightweight Kubernetes). Over 3,800 GPU instances on ClusterBid ship with nvidia-container-toolkit v1.18+ pre-installed.
Recommended testing workflow on ClusterBid instances: (1) Provision a GPU instance with the desired runtime. (2) Run sudo nvidia-ctk runtime configure --runtime=$RUNTIME. (3) Test GPU access: docker run --rm --gpus all nvidia/cuda:13.0-base nvidia-smi (Docker) or ctr run --gpus 0 --rm nvidia/cuda:12.6-base nvidia-smi (containerd). (4) Verify MIG access: docker run --rm --gpus 'device=MIG-<UUID>' nvidia/cuda:13.0-base nvidia-smi -L.
