Why Confidential Inference Matters Now
When you rent GPU compute from a third-party provider, your model weights and inference inputs reside in the provider's memory space. Without hardware-level isolation, a compromised hypervisor, rogue system administrator, or adjacent tenant could extract model parameters or prompt data. For regulated industries handling PHI, PII, or financial data, this is a compliance risk that standard encryption-at-rest does not address.
Confidential AI inference solves this using Trusted Execution Environments (TEEs) that encrypt data in use inside secure GPU enclaves. AMD SEV-SNP, Intel TDX, and NVIDIA Confidential Computing all provide hardware-enforced memory isolation so that even the host operating system cannot access enclave memory. The userverifies this isolation through remote attestation before sending any workload.
TEE Architectures for GPU Workloads
AMD SEV-SNP (Secure Encrypted Virtualization-Secure Nested Paging) encrypts each virtual machine's memory with a unique key held by the AMD secure processor. The hypervisor cannot decrypt VM memory even with physical access. For GPU workloads, AMD's MI300X supports SEV-SNP, allowing entire VM-level enclaves that encompass both CPU and GPU memory. The attestation report proves the VM's encryption key was generated by genuine AMD hardware.
Intel TDX (Trusted Domain Extensions) provides similar VM-level isolation. TDX domains are encrypted with MKTME (Multi-Key Total Memory Encryption) keys managed by the Intel hardware. NVIDIA's H100 and B200 GPUs add GPU-to-GPU encrypted NVLink and PCIe DMA protection so data remains encrypted across the entire compute fabric. The B300 extends this with per-GPU enclave attestation for multi-node inference pipelines.
| Feature | AMD SEV-SNP | Intel TDX | NVIDIA CC |
|---|---|---|---|
| Memory Encryption | AES-256 per-VM | AES-256 per-TD | AES-256 per-GPU |
| GPU Support | MI300X | N/A (CPU only) | H100/B200/B300 |
| Attestation Protocol | AMD ASP (vLEK/ARK) | Intel PCCS (SGX QE3) | NVIDIA NvSwitch CC |
| Interconnect Encryption | Memory only | Memory only | NVLink + PCIe |
| Hypervisor Trust Level | Zero | Zero | Zero |
| Production Readiness | GA (2024) | GA (2024) | Limited GA (2025) |
Remote Attestation: The Verification Flow
Remote attestation is the cryptographic proof that your model is executing inside a genuine TEE on genuine hardware. The flow works in four steps. Step one, the provider's infrastructure generates an attestation report signed by a hardware root of trust (AMD's ASP, Intel's Quoting Enclave, or NVIDIA's secure processor). Step two, the client verifies the report signature against the hardware manufacturer's public key infrastructure.
Step three, the client inspects the report's measurement field (a SHA-384 hash of the initial code and data state) against a known-good reference value. This proves no unauthorized code was injected at boot. Step four, the client establishes an encrypted TLS channel whose key is sealed to that specific attestation report, ensuring only the verified enclave can decrypt the model weights and inference prompts.
Verification Tools and SDKs
AMD provides the SEV Tool (sevtool) for fetching and verifying attestation reports from EPYC processors. The report includes the chip ID, firmware version, platform version, and policy flags. Production deployments must validate the certificate chain from the AMD Ark (Ask Root Key) through the chip endorsement key. Intel's Trust Authority service provides a hosted verification endpoint for TDX and SGX reports.
NVIDIA's confidential computing SDK includes a host-side verifier for GPU attestation reports. The report contains the GPU firmware version, TCB (Trusted Compute Base) status, and platform certificates. The key check is that the TCB is not revoked and that the GPU firmware is on the approved version list. NVIDIA publishes a per-month TCB revocation list that clients must check before establishing the trusted channel.
Attestation vs TLS: Why Both Are Needed
TLS protects data in transit between client and server. It does nothing to protect data once it reaches the GPU's memory. The server terminates TLS, so the plaintext model weights and inference results are visible to the host operating system. Attestation solves the data-in-use problem by ensuring the server is a trusted enclave, not a general-purpose OS that could be compromised.
The production architecture requires both. The client connects over TLS (standard mTLS with certificate pinning), then requests the attestation report over that channel. Only after the report is verified does the client transmit the model key sealed to the enclave measurement. The model weights are decrypted exclusively inside GPU memory that no host process can read.
Practical Deployment Considerations
Performance overhead from TEEs varies by architecture. AMD SEV-SNP adds 3-8% CPU-side overhead for most workloads; the GPU memory encryption on MI300X adds negligible compute latency (under 1% for inference) because the encryption engine runs in parallel with the compute units. Intel TDX has similar CPU overhead. NVIDIA's GPU-side encryption is designed to use dedicated silicon on the H100 SM and B200/B300 Tensor Cores, with sub-2% throughput impact reported in benchmarks.
Memory limits are the more practical constraint. SEV-SNP and TDX each reserve 3-8% of system memory for the secure processor and encryption metadata. On a 2TB EPYC server, you lose roughly 100GB. On the GPU side, NVIDIA's confidential compute mode reserves 1-2GB per GPU for the encryption engine. For a B300 with 288GB, the usable memory becomes 286GB. Batch sizes must be adjusted accordingly.
Provider Verification Checklist
Not all GPU providers who claim confidential computing support actually implement remote attestation. AI teams should verify three things before contracting. First, that the provider exposes the raw attestation report (not a summary) so your tooling can independently validate the certificate chain. Second, that the provider allows you to pin firmware versions and TCB levels in your deployment configuration.
Third, that the provider documents the exact hardware root of trust used for each GPU generation. AMD EPYC Genoa and Turin support SEV-SNP; Intel Xeon 5th/6th Gen support TDX. NVIDIA H100s with the CC feature flag require firmware version 22.7 or newer. A provider that cannot produce per-GPU attestation reports with CPU+GPU coverage is not offering confidential compute regardless of marketing claims.
