All essays
TechnicalDEEP DIVEFEB 2026

GPU Cluster Security: Isolation, Multi-Tenancy, and Data Protection for AI Workloads

GPU memory isolation, MIG security boundaries, confidential computing, network isolation, and encryption for multi-tenant GPU clusters. SOC2 and HIPAA compliance frameworks for AI infrastructure at mid-2026.

01

The Multi-Tenant GPU Threat Model

GPU clusters are increasingly shared across teams, workloads, and even organizations. Unlike CPU virtualization where decades of isolation hardening exist, GPU multi-tenancy is still maturing. The primary threat vectors include cross-tenant GPU memory inspection, side-channel attacks via shared NVLink/NVSwitch fabric, and data exposure through incomplete GPU memory reset between workloads.

A 2025 study from researchers at UIUC demonstrated that uninitialized GPU memory in MIG instances could leak data from co-resident tenants. NVIDIA addressed this with MIG memory encryption in the Hopper architecture, but Blackwell and earlier generations have different security profiles that operators must understand.

This post covers the security mechanisms available at each GPU generation, their effectiveness against real attack vectors, and the compliance requirements that drive security architecture decisions.

02

GPU Memory Isolation Mechanisms

GPU memory isolation operates at multiple hardware and software levels. At the hardware level, NVIDIA's MIG (Multi-Instance GPU) partitions H100's 80GB HBM3 into up to 7 isolated instances with hardware-enforced memory and cache partitioning. Each MIG instance has dedicated DRAM banks, L2 cache slices, and memory controllers with no cross-instance access path.

At the software level, CUDA MPS (Multi-Process Service) provides lighter-weight memory isolation without hardware partitioning, using address translation and access control lists. MPS offers better GPU utilization for bursty workloads but provides weaker isolation guarantees compared to MIG.

The practical choice depends on your threat model. For multi-tenant environments where tenants do not trust each other, MIG is the minimum acceptable isolation level. For single-tenant multi-workload scenarios, MPS or even time-sliced sharing may be sufficient.

03

MIG Security Boundaries

NVIDIA's MIG implementation provides hardware-enforced isolation across GPU memory, cache, and compute units. Each MIG instance operates as an independent GPU with its own DRAM, L2 cache, and compute slice. The table below shows MIG partitioning options for H100 80GB and B200 and the security properties of each.

MIG in Hopper also introduced streaming multiprocessor (SM) isolation, preventing co-resident instances from affecting each other's compute performance. This is critical for SLAs in multi-tenant GPU environments.

GPU ConfigurationMIG PartitionsMemory per SliceSM Isolation
H100 80GB SXMUp to 710-80 GBFull hardware isolation
H100 80GB SXM (1g.10gb)7 instances10 GB1 SM per instance
H100 80GB SXM (3g.40gb)2 instances40 GB3 SMs per instance
H100 80GB SXM (7g.80gb)1 instance80 GB7 SMs per instance
B200 192GB SXMUp to 724-192 GBFull hardware isolation
A100 80GB SXMUp to 710-80 GBFull hardware isolation
04

Confidential Computing on GPUs

Confidential computing extends memory encryption to GPU workloads, protecting data in use from host access, hypervisor compromise, and physical memory attacks. NVIDIA's confidential computing support on H100 and B200 uses hardware-based trusted execution environments (TEEs) that encrypt GPU memory with CPU-attested keys.

NVIDIA Hopper introduced GPU TEE with a hardware root of trust that attests the GPU firmware and establishes encrypted communication channels with the CPU. Blackwell extends this with full GPU memory encryption covering HBM3e, NVLink, and PCIe traffic. At mid-2026, confidential GPU compute adds approximately 3-8% performance overhead depending on workload memory access patterns.

AWS Nitro Enclaves and Azure confidential VMs both support GPU confidential computing on H100 instances. The minimum setup for a confidential GPU workload requires a compatible hypervisor (Nitro or Hyper-V), a confidential computing CPU, and the NVIDIA GPU TEE driver stack.

05

Network Isolation Architectures

Multi-tenant GPU clusters require network isolation at multiple layers: inter-node GPU communication (NVLink/NVSwitch), storage access (NFS/GPUDirect), and inference API endpoints. The most common architecture in mid-2026 uses Kubernetes network policies combined with MacVLAN or SR-IOV for GPU-direct network access.

NVLink fabric isolation is more complex. In an 8-GPU HGX baseboard, all GPUs share the same NVSwitch domain. Multi-tenant isolation at the GPU interconnect level requires physical separation -- separate HGX boards for separate tenants -- or NVIDIA's upcoming NVLink switch partitioning in Blackwell Ultra (due Q4 2026).

For cluster operators, the pragmatic approach is tenant-dedicated GPU nodes with Kubernetes namespace isolation and network policies. Node-level isolation eliminates NVLink cross-tenant concerns at the cost of GPU utilization efficiency, typically reducing packing efficiency from 85% to 60-70%.

06

Data Encryption at Rest and In Transit

GPU workloads introduce unique data encryption challenges because model weights and intermediate activations must be accessible in GPU DRAM for computation. Traditional storage encryption (AES-256 at the filesystem or device level) protects data at rest but does not cover data while being transferred to GPU memory or during computation.

For data in transit, NVLink traffic between GPUs in the same node is not encrypted by default on Hopper. Blackwell introduced NVLink encryption with AES-256-GCM at line rate with no measurable throughput impact. For inter-node GPU communication, InfiniBand with IPsec or RoCEv2 with MACsec provide wire-level encryption.

The practical guidance in mid-2026 is to encrypt all inter-node GPU traffic (InfiniBand/RoCE), use AES-256 encrypted storage (local NVMe and NFS), and rely on confidential GPU compute for data-in-use protection on Blackwell. For Hopper-based clusters without confidential computing, ensure adequate physical security in the data center.

07

Compliance Frameworks: SOC2 and HIPAA

GPU cluster operators serving regulated industries must demonstrate compliance with SOC2 (Type II) and/or HIPAA requirements. The table below maps specific control requirements to GPU infrastructure configurations. The key distinction is data-in-use protection, which SOC2 recommends but HIPAA requires for protected health information (PHI).

Most GPU providers in mid-2026 achieve SOC2 Type II compliance through a combination of node-level tenant isolation, encrypted storage, and access controls. HIPAA-compliant GPU compute requires business associate agreements (BAAs), confidential computing or equivalent data-in-use protection, and audit logging covering all GPU job execution. As of June 2026, fewer than 15 GPU providers offer HIPAA-compliant multi-tenant GPU clusters.

Control AreaSOC2 RequirementHIPAA Requirement
Tenant isolationLogical separationHardware-enforced (MIG or dedicated node)
Encryption at restAES-256 recommendedAES-256 required
Encryption in transitTLS 1.2+TLS 1.2+ for all PHI
Data-in-use protectionRecommendedRequired (confidential computing or equivalent)
Audit loggingAccess logsAccess + job execution + data access logs
Incident response72-hour notification60-day breach notification
Filed under
GPU SecurityMulti-TenancyMIGConfidential ComputingSOC2HIPAANetwork IsolationData Protection