All essays
TechnicalDEEP DIVEFEB 2026

Docker GPU Ecosystem: NGC Catalog, DeepOps, and Kubeflow GPU Images Guide

Production guide for Docker GPU images: NVIDIA NGC catalog optimization, DeepOps GPU cluster deployment, Kubeflow GPU training images, container build patterns, and registry management for H100/B200 clusters.

01

NVIDIA NGC CATALOG: OPTIMIZING BASE IMAGE SELECTION

NVIDIA's NGC (NVIDIA GPU Cloud) catalog provides the canonical set of Docker images for GPU computing. The catalog is organized by software stack: `nvcr.io/nvidia/pytorch:24.12-py3` (PyTorch + CUDA + cuDNN + NCCL), `nvcr.io/nvidia/tensorflow:24.12-tf2-py3`, and `nvcr.io/nvidia/tritonserver:24.12-py3` for inference serving. Each image version corresponds to the monthly container release schedule, with `24.12` representing December 2024. The image size is the primary concern: the PyTorch NGC image is 14.2 GB compressed (28.4 GB on disk), driven by the full CUDA toolkit (8.2 GB), cuDNN libraries (2.1 GB), NCCL (1.8 GB), and TensorRT (1.5 GB). For GPU cluster deployments with hundreds of nodes, pulling this image on each node consumes 3-5 minutes per node over a 25 GbE network, adding 5-8 hours to the rollout of a 100-node cluster.

The production optimization is to build a custom slim image from the NGC base using multi-stage Docker builds. The pattern: `FROM nvcr.io/nvidia/pytorch:24.12-py3 AS base`, then copy only the necessary CUDA libraries and the training script into a second stage using `ubuntu:22.04` as the base. The slim image reduces to 4.8 GB compressed (9.6 GB on disk) by excluding CUDA samples, documentation, and redundant libraries. The Dockerfile adds `--from=base /usr/local/cuda-12.4 /usr/local/cuda-12.4` and specific `.so` files for NCCL, cuDNN, and TensorRT. For clusters with heterogeneous GPUs (H100 and A100 mixed), the NGC image should use the `cuda:12.4.1-runtime-ubuntu22.04` base and install the GPU-specific libraries at container startup via volume mounts, avoiding per-GPU-architecture image builds.

Image TypeBase ImageCompressed SizeDisk SizePull Time (25 GbE)GPU Coverage
NGC Full PyTorchnvcr.io/nvidia/pytorch:24.1214.2 GB28.4 GB3-5 minAll NVIDIA
NGC Slim PyTorchCustom multi-stage4.8 GB9.6 GB1-2 minAll NVIDIA
NGC Triton Servernvcr.io/nvidia/tritonserver:24.128.4 GB16.8 GB2-3 minAll NVIDIA
Kubeflow Trainingcustom-kubeflow:24.126.2 GB12.4 GB1.5-2.5 minAll NVIDIA
Minimal CUDA Runtimenvidia/cuda:12.4.1-runtime1.2 GB2.4 GB20-30 secCUDA GPUs
02

NVIDIA CONTAINER TOOLKIT AND GPU OPERATOR CONFIGURATION

The NVIDIA Container Toolkit (nvidia-docker2 or the newer `nvidia-container-toolkit` package) is the runtime hook that enables GPU access from Docker containers. The `nvidia-ctk runtime configure --runtime=docker` command configures the Docker daemon to use the `nvidia` runtime. The key configuration file is `/etc/nvidia-container-runtime/config.toml`, which controls: `accept-nvidia-visible-devices-as-volume-mounts = true` (required for Kubernetes GPU passthrough), `swarm-resource = "DOCKER_RESOURCE_GPU"` (for Docker Swarm GPU scheduling), and `disable-require = false` (ensures NVIDIA drivers are checked before container start). The `--gpus all` flag (Docker) or `nvidia.com/gpu` resource (Kubernetes) triggers the runtime to mount `/usr/local/nvidia/lib64` from the host into the container and set `NVIDIA_VISIBLE_DEVICES` and `CUDA_VISIBLE_DEVICES` environment variables.

The NVIDIA GPU Operator automates this on Kubernetes. Version 24.9+ deploys DaemonSets for the NVIDIA driver, container toolkit, DCGM exporter, MIG manager, and GPU feature discovery as an integrated Helm chart. The operator's `values.yaml` configures `driver.enabled=true` to install the 550.54.15 driver on the host, `toolkit.enabled=true` for the container runtime hook, and `migManager.enabled=true` for H100 MIG partitioning. The GPU Operator eliminates manual driver installation on each GPU node: when a new GPU node joins the cluster, the operator detects the GPU hardware and installs the appropriate driver version automatically. This is the standard deployment pattern for GPU clusters on Kubernetes, used by CoreWeave, Lambda, and most neocloud providers. The GPU Operator adds approximately 3 GB per node in container images and 500 MB of driver artifacts.

03

DEEPOPS: DEPLOYING GPU CLUSTERS WITH ANSIBLE PLAYBOOKS

NVIDIA DeepOps is an open-source Ansible-based deployment toolkit for GPU clusters, supporting Slurm and Kubernetes. The `deepops/slurm-cluster/` playbook installs Slurm with GPU accounting (`slurm.conf` `GresTypes=gpu`), the `nvml` plugin for GPU health checks, and Pyxis/Enroot for containerized GPU job execution. The playbook also configures NCCL topology files and InfiniBand partitions for optimal GPU communication. A `config.yml` specifies the cluster topology: `slurm_node_type: "gpu"`, `nvidia_driver_version: "550.54.15"`, `nvidia_k8s_device_plugin: true`. For a 32-node H100 cluster, DeepOps deploys the complete software stack in 45-60 minutes, compared to 4-8 hours for manual installation.

DeepOps Slurm integration with containerized GPU workloads uses Pyxis (a Slurm SPANK plugin) and Enroot (a container runtime). The `srun --container-image=nvcr.io/nvidia/pytorch:24.12-py3 --container-mounts=/mnt/data:/data training_script.py` command launches a containerized training job with Slurm GPU resource management. Enroot converts the Docker image to a SquashFS file, reducing container startup time from 30-45 seconds (podman/docker pull + run) to 2-5 seconds (SquashFS mount). For production clusters, DeepOps+Pyxis+Enroot is the standard DeepSpeed/Megatron-LM training stack at NVIDIA DGX SuperPOD deployments. DeepOps is less common on neocloud providers (which prefer Kubernetes), but remains the standard for bare-metal GPU clusters in enterprise HPC environments.

Deployment ApproachDeployment Time (32-node)Container RuntimeGPU SchedulingNetworking Config
DeepOps + Slurm45-60 minEnroot (SquashFS)Slurm GRES gpuInfiniBand partition + NCCL
Kubeflow + K8s2-4 hourscontainerd + nvidia-toolkitnvidia.com/gpu device pluginCNI + RDMA plugin
GPU Operator + K8s1-2 hourscontainerd + operator-managednvidia.com/gpu + MIGCNI + GPU Operator
Manual Deployment4-8 hoursDocker + nvidia-runtimeManual GPU assignmentManual IB config
04

KUBEFLOW: GPU TRAINING IMAGES AND PIPELINES

Kubeflow provides the standard machine learning toolkit for Kubernetes GPU clusters. The `kubeflow/training-operator` (formerly `tf-operator` and `pytorch-operator`) defines `PyTorchJob`, `TFJob`, and `MPIJob` custom resources that handle multi-GPU training orchestration. A `PyTorchJob` YAML spec for 8-GPU training specifies `pytorchReplicaSpecs.Worker.template.spec.containers.resources.limits["nvidia.com/gpu"]: 8` along with the NGC-based container image. The training operator manages NCCL rendezvous across Pods by setting `MASTER_ADDR`, `MASTER_PORT`, `WORLD_SIZE`, and `RANK` environment variables automatically. Kubeflow 1.9+ adds Volcano as the default gang scheduler, ensuring all GPUs for a training job are allocated simultaneously before any Pod starts (preventing deadlock from partial GPU allocation).

The `kubeflow/notebooks` component provides JupyterLab images with GPU support: `kubeflownotebookswg/jupyter-tensorflow-full:v1.9.0` and `kubeflownotebookswg/jupyter-pytorch-full:v1.9.0`. These images include CUDA toolkit, cuDNN, and common ML libraries in a 6-8 GB image. For production, teams typically build custom training images extending the NGC base with project-specific dependencies. The recommended CI/CD pattern builds the Docker image in CI, pushes to a private registry (Harbor or ECR), and references the image in the `PyTorchJob` spec. Image pull secrets are configured on the GPU nodes via Kubernetes `kubelet` config `imagePullSecrets`. On ClusterBid, the standard GPU container image for PyTorch training is ~5 GB and pulls in 30-60 seconds on provisioned H100 nodes, compared to 2-4 minutes for the full NGC image.

05

IMAGE CACHING, REGISTRY MANAGEMENT, AND PULL THROUGHPUT

At GPU cluster scale (32-256 nodes), Docker image pull latency becomes a significant deployment bottleneck. A single 5 GB training image pulled on 128 nodes simultaneously generates 640 GB of registry egress, consuming 50-100% of the registry's bandwidth and taking 5-15 minutes for all nodes to complete. The standard solution is a distributed registry cache: Harbor or Sonatype Nexus deployed as a pull-through cache in the same network as the GPU nodes. Each node's `containerd` is configured with `mirrors:"docker.io": endpoint:["http://harbor-cluster:5000"]` to redirect pulls to the local cache. After the first pull on any node, subsequent pulls on all other nodes serve from the local cache at 10-40 Gbps, reducing pull time to 10-30 seconds.

A second optimization is node-level image caching via the `kubelet` image pull policy. Setting `imagePullPolicy: IfNotPresent` on GPU Pods skips the pull if the image tag already exists on the node. For frequent deployments, a systemd timer on each GPU node runs `nerdctl pull` (for containerd) or `crictl pull` (for CRI-O) during off-peak hours to pre-cache the latest images. The cache size is a concern: 10 GPU images at 5 GB each consume 50 GB per node. Setting `kubelet` `--image-gc-high-threshold=80 --image-gc-low-threshold=60` ensures garbage collection removes images when disk exceeds 80% usage. For production GPU clusters, a shared NFS mount at `/var/lib/containerd` caches images across node restarts, preventing the 5-15 minute cache warmup after a node reboot.

Pull OptimizationPull Time (128 nodes, 5 GB image)Bandwidth ConsumptionInfrastructure Required
Direct Registry (Docker Hub)8-15 min640 GBNone (internet)
Pull-Through Cache (Harbor)30-60 sec (cached)5 GB (first pull only)Harbor cluster (3 nodes)
Node-Level Pre-Caching10-30 sec (warm cache)5 GB per node (scheduled)Systemd timer + storage
Shared /var/lib/containerd NFS5-15 sec (persistent cache)5 GB (NFS share)NFS server + network storage
Read-Only RootFS + Overlay5-10 sec (default)Non-persistentImmutable infrastructure
Filed under
Docker GPUNGC CatalogNVIDIA DeepOpsKubeflow GPU ImagesGPU Container BuildNVIDIA Container ToolkitCUDA Docker Image