The Multi-Tenant GPU Threat Model
GPU clusters are increasingly shared across teams, workloads, and even organizations. Unlike CPU virtualization where decades of isolation hardening exist, GPU multi-tenancy is still maturing. The primary threat vectors include cross-tenant GPU memory inspection, side-channel attacks via shared NVLink/NVSwitch fabric, and data exposure through incomplete GPU memory reset between workloads.
A 2025 study from researchers at UIUC demonstrated that uninitialized GPU memory in MIG instances could leak data from co-resident tenants. NVIDIA addressed this with MIG memory encryption in the Hopper architecture, but Blackwell and earlier generations have different security profiles that operators must understand.
This post covers the security mechanisms available at each GPU generation, their effectiveness against real attack vectors, and the compliance requirements that drive security architecture decisions.
GPU Memory Isolation Mechanisms
GPU memory isolation operates at multiple hardware and software levels. At the hardware level, NVIDIA's MIG (Multi-Instance GPU) partitions H100's 80GB HBM3 into up to 7 isolated instances with hardware-enforced memory and cache partitioning. Each MIG instance has dedicated DRAM banks, L2 cache slices, and memory controllers with no cross-instance access path.
At the software level, CUDA MPS (Multi-Process Service) provides lighter-weight memory isolation without hardware partitioning, using address translation and access control lists. MPS offers better GPU utilization for bursty workloads but provides weaker isolation guarantees compared to MIG.
The practical choice depends on your threat model. For multi-tenant environments where tenants do not trust each other, MIG is the minimum acceptable isolation level. For single-tenant multi-workload scenarios, MPS or even time-sliced sharing may be sufficient.
MIG Security Boundaries
NVIDIA's MIG implementation provides hardware-enforced isolation across GPU memory, cache, and compute units. Each MIG instance operates as an independent GPU with its own DRAM, L2 cache, and compute slice. The table below shows MIG partitioning options for H100 80GB and B200 and the security properties of each.
MIG in Hopper also introduced streaming multiprocessor (SM) isolation, preventing co-resident instances from affecting each other's compute performance. This is critical for SLAs in multi-tenant GPU environments.
| GPU Configuration | MIG Partitions | Memory per Slice | SM Isolation |
|---|---|---|---|
| H100 80GB SXM | Up to 7 | 10-80 GB | Full hardware isolation |
| H100 80GB SXM (1g.10gb) | 7 instances | 10 GB | 1 SM per instance |
| H100 80GB SXM (3g.40gb) | 2 instances | 40 GB | 3 SMs per instance |
| H100 80GB SXM (7g.80gb) | 1 instance | 80 GB | 7 SMs per instance |
| B200 192GB SXM | Up to 7 | 24-192 GB | Full hardware isolation |
| A100 80GB SXM | Up to 7 | 10-80 GB | Full hardware isolation |
Confidential Computing on GPUs
Confidential computing extends memory encryption to GPU workloads, protecting data in use from host access, hypervisor compromise, and physical memory attacks. NVIDIA's confidential computing support on H100 and B200 uses hardware-based trusted execution environments (TEEs) that encrypt GPU memory with CPU-attested keys.
NVIDIA Hopper introduced GPU TEE with a hardware root of trust that attests the GPU firmware and establishes encrypted communication channels with the CPU. Blackwell extends this with full GPU memory encryption covering HBM3e, NVLink, and PCIe traffic. At mid-2026, confidential GPU compute adds approximately 3-8% performance overhead depending on workload memory access patterns.
AWS Nitro Enclaves and Azure confidential VMs both support GPU confidential computing on H100 instances. The minimum setup for a confidential GPU workload requires a compatible hypervisor (Nitro or Hyper-V), a confidential computing CPU, and the NVIDIA GPU TEE driver stack.
Network Isolation Architectures
Multi-tenant GPU clusters require network isolation at multiple layers: inter-node GPU communication (NVLink/NVSwitch), storage access (NFS/GPUDirect), and inference API endpoints. The most common architecture in mid-2026 uses Kubernetes network policies combined with MacVLAN or SR-IOV for GPU-direct network access.
NVLink fabric isolation is more complex. In an 8-GPU HGX baseboard, all GPUs share the same NVSwitch domain. Multi-tenant isolation at the GPU interconnect level requires physical separation -- separate HGX boards for separate tenants -- or NVIDIA's upcoming NVLink switch partitioning in Blackwell Ultra (due Q4 2026).
For cluster operators, the pragmatic approach is tenant-dedicated GPU nodes with Kubernetes namespace isolation and network policies. Node-level isolation eliminates NVLink cross-tenant concerns at the cost of GPU utilization efficiency, typically reducing packing efficiency from 85% to 60-70%.
Data Encryption at Rest and In Transit
GPU workloads introduce unique data encryption challenges because model weights and intermediate activations must be accessible in GPU DRAM for computation. Traditional storage encryption (AES-256 at the filesystem or device level) protects data at rest but does not cover data while being transferred to GPU memory or during computation.
For data in transit, NVLink traffic between GPUs in the same node is not encrypted by default on Hopper. Blackwell introduced NVLink encryption with AES-256-GCM at line rate with no measurable throughput impact. For inter-node GPU communication, InfiniBand with IPsec or RoCEv2 with MACsec provide wire-level encryption.
The practical guidance in mid-2026 is to encrypt all inter-node GPU traffic (InfiniBand/RoCE), use AES-256 encrypted storage (local NVMe and NFS), and rely on confidential GPU compute for data-in-use protection on Blackwell. For Hopper-based clusters without confidential computing, ensure adequate physical security in the data center.
Compliance Frameworks: SOC2 and HIPAA
GPU cluster operators serving regulated industries must demonstrate compliance with SOC2 (Type II) and/or HIPAA requirements. The table below maps specific control requirements to GPU infrastructure configurations. The key distinction is data-in-use protection, which SOC2 recommends but HIPAA requires for protected health information (PHI).
Most GPU providers in mid-2026 achieve SOC2 Type II compliance through a combination of node-level tenant isolation, encrypted storage, and access controls. HIPAA-compliant GPU compute requires business associate agreements (BAAs), confidential computing or equivalent data-in-use protection, and audit logging covering all GPU job execution. As of June 2026, fewer than 15 GPU providers offer HIPAA-compliant multi-tenant GPU clusters.
| Control Area | SOC2 Requirement | HIPAA Requirement |
|---|---|---|
| Tenant isolation | Logical separation | Hardware-enforced (MIG or dedicated node) |
| Encryption at rest | AES-256 recommended | AES-256 required |
| Encryption in transit | TLS 1.2+ | TLS 1.2+ for all PHI |
| Data-in-use protection | Recommended | Required (confidential computing or equivalent) |
| Audit logging | Access logs | Access + job execution + data access logs |
| Incident response | 72-hour notification | 60-day breach notification |
