All essays
InfrastructureINFRASTRUCTUREFEB 2026

GPU Cluster Security Operations: Zero-Trust Architecture for Multi-Tenant Clusters

Zero-trust security for multi-tenant GPU clusters. Tenant isolation (MIG, vGPU, K8s security), network segmentation, workload attestation, GPU memory confidentiality, and audit logging at scale.

01

THE GPU MULTI-TENANT THREAT MODEL

Multi-tenant GPU clusters introduce attack surfaces absent in single-tenant HPC environments. The GPU threat model includes four categories: GPU memory leakage (a tenant's model weights or training data leaking through GPU framebuffer residue to a co-resident tenant), side-channel attacks (GPU utilization timing attacks that infer training dataset characteristics), NCCL interception (a malicious tenant's code injecting into the NCCL communicator to read gradient data during all-reduce), and management plane compromise (unauthorized GPU configuration changes via the NVIDIA management API). The 2024 GPU.zip attack demonstrated that GPU memory compression side channels can leak pixel data from co-resident GPU processes at 95 percent accuracy, validating that GPU-level isolation is necessary, not optional.

The zero-trust GPU architecture treats every tenant, workload, and node as untrusted until proven otherwise. The control plane authenticates and authorizes every GPU allocation request, the data plane encrypts all GPU-to-GPU and GPU-to-storage traffic (even within the same rack), and the audit plane logs every GPU API call, CUDA kernel launch, and memory allocation with tenant attribution. The following sections detail the implementation of each plane for a production multi-tenant GPU cluster with 10+ tenants sharing a 256-H100 cluster.

02

GPU HARDWARE ISOLATION: MIG, VGPU, AND TIME-SLICING

NVIDIA H100 GPUs support three isolation models with different security properties. Multi-Instance GPU (MIG) partitions the GPU at the hardware level, creating up to 7 fully isolated instances (1g.10gb through 7g.80gb profiles) with dedicated framebuffer, cache, and memory controllers. MIG instances are isolated by hardware: one tenant cannot access another tenant's framebuffer even if the CUDA driver is compromised. For H100 with 80 GB, the `nvidia-smi mig -cgi 1g.10gb,1g.10gb,3g.40gb,3g.40gb` creates 4 MIG instances. MIG is appropriate for inference workloads where per-instance GPU requirements are known and fixed.

vGPU (NVIDIA Virtual GPU) provides VM-level GPU passthrough with vGPU Manager interposing on GPU commands for memory protection. vGPU supports larger maximum GPU partitions (40-80 GB) than MIG but adds a virtualization layer that introduces 3-8 percent performance overhead on CUDA kernels. Time-slicing (Kubernetes GPU sharing) provides the weakest isolation: tenant A's kernel executing on GPU 0 can be evicted at any point by the GPU scheduler to run tenant B's kernel, and though framebuffer is allocated exclusively, the shared compute pipeline creates transient state leakage risks. For security-sensitive multi-tenant environments, the recommended configuration is MIG for inference workloads with strict SLAs and full-GPU non-MIG for training workloads with NCCL isolation at the fabric level (per-tenant VLANs on InfiniBand).

Isolation ModelHardware IsolationMax Partitions/GPUPerformance OverheadGPU Memory ProtectionUse Case
MIG (H100)Yes (full hardware)7<1%Full hardware isolationInference, multi-tenant inference serving
MIG (A100)Yes (full hardware)7<1%Full hardware isolationInference, regulated workloads
vGPUYes (hypervisor mediated)Up to GPU memory limit3-8%VMM-backed protectionVM-based GPU workloads
Time-SlicingNo (software scheduler)Unlimited (device plugin)0% when alone, variable sharingNoneDev/test, burstable inference
Full GPU (non-MIG)Yes (single tenant)10%Dedicated framebuffer onlyTraining, single-tenant clusters
03

NETWORK SEGMENTATION AND NCCL TRAFFIC ENCRYPTION

Multi-tenant NCCL traffic requires network-level isolation to prevent tenant A from reading tenant B's gradient data during all-reduce operations (which would leak training data composition and model architecture). The standard approach is per-tenant InfiniBand partitions (PKeys). Each InfiniBand partition defines a set of ports that can communicate, creating a virtual network within the physical fabric. PKey 0x8001 is assigned to tenant A, PKey 0x8002 to tenant B, and so on. The OpenSM subnet manager configures partitions: `partitions config partition_tenant_a pkey=0x8001 ipoib=0 flags=0x2`. All NCCL traffic runs within the tenant's partition, preventing cross-tenant RDMA access even on shared InfiniBand switches.

For RoCE v2 clusters, NCCL traffic encryption is achieved through IPsec or MACsec at the link layer. NVIDIA ConnectX-7 and above support inline IPsec acceleration at wire speed with no throughput penalty: `mlxconfig -d /dev/mst/mt4123_pciconf0 set IP_OVER_IB=0 IP_OVER_ETH=1 IPSEC_OFFLOAD=1`. The NCCL communicator is configured with `NCCL_IB_GID_INDEX=3` to use the encrypted GID. Storage traffic encryption uses NVMe/TCP with TLS 1.3 or NVMe-oF with pre-shared keys. Auditing NCCL traffic isolation is validated by `ibdiagnet -r -p` which traces all partition memberships and reports cross-partition communication attempts as errors.

04

GPU MEMORY CONFIDENTIALITY AND CONFIDENTIAL COMPUTING

GPU memory confidentiality protects model weights, training data, and inference outputs from host system compromise. NVIDIA Confidential Computing for H100 (now generally available in CUDA 12.6+) uses a TEE (Trusted Execution Environment) that encrypts GPU framebuffer memory using a per-tenant hardware key generated during GPU boot. The GPU TEE ensures that even if the host operating system or hypervisor is compromised, the tenant's GPU memory remains encrypted and the host cannot issue DMA reads to the GPU framebuffer. Enabling CC requires `nvidia-smi -cc 1` on the GPU and the CUDA application to link against `libcuda_cc.so`.

The performance impact of GPU Confidential Computing is workload-dependent. In published benchmarks, H100 confidential computing adds 2-5 percent overhead for inference (GEMM-bound, additional TEE context switch per kernel launch) and 5-12 percent for training (higher kernel launch rate with backward pass requires more TEE transitions). The overhead drops to 3-7 percent when using CUDA Graphs to pre-compile kernel launch sequences. The NCCl collective communicator must be configured with `NCCL_PROTO=Simple` when running in confidential computing mode because the NVLink direct memory access paths are restricted. These constraints mean that confidential computing is currently deployed primarily for regulated industries (healthcare HIPAA, financial services PCI-DSS, government) where the compliance requirement outweighs the performance cost.

Workload TypeNon-CC PerformanceCC Enabled PerformanceOverheadUse CC?
LLM Inference (batch 1, H100)2,400 tokens/sec2,280 tokens/sec5%Healthcare/Finance yes, others no
LLM Training (BF16, 64-GPU)45 TFLOPS/GPU40 TFLOPS/GPU11%Regulated data only
Vision Training (FP32, 8-GPU)8.5 GB/s throughput7.7 GB/s throughput9%MedTech / government only
RNN Inference6,800 samples/sec6,324 samples/sec7%If PCI-DSS requirement
CUDA Graph optimized (training)40 TFLOPS/GPU37 TFLOPS/GPU7%Regulated + CUDA Graph capable
05

AUDIT LOGGING AND COMPLIANCE FOR GPU WORKLOADS

GPU audit logging captures all GPU-level events with tenant attribution for compliance and forensics. The NVIDIA GPU Operator enables audit logging through the `nvidia-operator-validator` that writes GPU allocation events (pod-to-GPU mapping, MIG configuration changes, driver load/unload) to Kubernetes audit logs. Each GPU allocation generates a structured log entry containing: `pod_name`, `namespace` (tenant), `gpu_uuid`, `mig_instance` (if MIG), `timestamp`, and `action` (schedule/unschedule/mig-config-change). These logs are forwarded to a SIEM (Splunk, Elastic, or Grafana Loki) via Fluentd with a 1-minute delivery latency SLA.

GPU-specific compliance controls include framebuffer zeroing between tenant allocations: before GPU reassignment, the CUDA driver writes zeros to all framebuffer pages. This is enabled by `nvidia-fabricmanager -gpumem -z` or via the DCGM API `dcgmFieldGroupCreate` with `DCGM_FI_DEV_FB_FREE`. The GPU Operator also enforces Pod Security Standards (PSS): only Restricted-level pods can access GPU resources in multi-tenant mode, banning privileged containers, host networking, and host PID namespaces for GPU pods. Periodic compliance scanning with `kube-bench` and `trivy` validates that GPU node configurations match CIS benchmarks for Kubernetes and NVIDIA GPU Operator, and any drift is automatically remediated through the GitOps reconciliation loop.

06

GPU-SPECIFIC INCIDENT RESPONSE PROCEDURES

When a GPU security incident is reported (suspected cross-tenant memory leakage, unauthorized GPU configuration change, or NCCL eavesdropping), the incident response procedure has GPU-specific steps beyond standard cloud IR. Step 1: isolate the affected GPU node by cordoning the Kubernetes node (`kubectl cordon gpu-node-42`) and evicting all running GPU pods to prevent further data exposure. Step 2: collect GPU forensics including current framebuffer contents via `nvidia-smi -q -d MEMORY` (VRAM allocation snapshot), DCGM event log retrieval via `dcgmi events -v`, and NCCL communicator state via `ncclCommGetState` if the workload is still accessible.

Step 3: memory capture for forensic analysis. NVIDIA's GPU memory dump capability (`nvidia-smi -g 0 --gpu-reset-mode=2 && nvidia-smi -pm 0`) forces GPU reset and clears framebuffer, but for forensic preservation, use the `nvidia-peermem` module to issue DMA reads of GPU memory from the host before reset. Step 4: review audit logs for the affected GPU's allocation history over the past 30 days to identify all tenants that accessed the hardware. Step 5: cross-reference NCCL partition configurations to verify that InfiniBand PKeys or RoCE v2 ACLs are correctly configured. A post-mortem report documents the root cause, which typically falls into one of: GPU Operator misconfiguration allowing pod-to-pod GPU sharing without isolation, NCCL fabric partition misconfiguration, or a node-level compromise bypassing GPU-level controls.

Filed under
GPU Security ArchitectureZero Trust GPU ClusterMulti-Tenant GPU IsolationMIG SecurityGPU Memory ConfidentialityNVIDIA Confidential ComputingKubernetes GPU Pod Security