CIS BENCHMARKS FOR GPU-ACCELERATED LINUX NODES
CIS Level 1 server profile with GPU-specific exceptions: skip xwindows group (GPU compute nodes don't need X11), disable slub_debug=FP (adds 10-15% memory overhead that can trigger OOM), set kernel.yama.ptrace_scope=0 for nsys profiling. Weekly OpenSCAP scanning with tailored profile, target >85% compliance.
Compliance extends to GPU software stack: NVIDIA packages signed and verified against NVIDIA GPG key, enforced in dnf with localpkg_gpgcheck=1 and repo_gpgcheck=1. Daily security report cross-references installed packages against CVE database.
A fix_cis_gpu.sh systemd oneshot service applies GPU-compatible CIS settings on boot, ensuring nodes rejoining after maintenance meet baseline.
| CIS Control | Standard | GPU Modification | Reason |
|---|---|---|---|
| 1.8.1 GDM removed | package gdm removed | Unchanged | No display needed |
| 3.1.1 dmesg_restrict=1 | kernel.dmesg_restrict=1 | Unchanged | nvidia-persistenced works |
| 3.2.1 source route disabled | accept_source_route=0 | Unchanged | Prevent InfiniBand spoofing |
| 1.5.1 bootloader pw | GRUB2 password | Unchanged | Prevent driver param tampering |
| 3.1.3 kptr_restrict=2 | kernel.kptr_restrict=2 | Modify: set to 1 | nsys needs kernel pointers |
| GPU-specific | NVIDIA GPG verification | localpkg_gpgcheck=1 | Ensures signed drivers |
SELINUX AND APPARMOR POLICIES FOR GPU WORKLOADS
SELinux: nvidia_t domain needs self:capability { sys_admin sys_rawio } for GPU memory mapping and PCIe config access. Custom nvidia.pp module defines nvidia_device_t for /dev/nvidia* device files. Container runtime to nvidia_t domain transition configured in NVIDIA Container Toolkit.
Custom SELinux module rules: allow nvidia_t self:capability { sys_admin sys_rawio ipc_lock }, nvidia_device_t:chr_file rw_file_perms, self:unix_dgram_socket create_socket_perms, hugetlbfs_t:file rw_file_perms.
AppArmor on Ubuntu: define allowed paths for GPU device access, deny unexpected network ports to restrict NCCL. Cache regenerated after driver updates because device file numbers may change. Enforcing mode validated via ausearch and aa-status.
CONTAINER SECURITY FOR GPU WORKLOADS
NVIDIA Container Runtime: no-cgroups=false enables cgroup device control, preventing a pod allocated 2 GPUs from accessing the other 6. This is the primary multi-tenant GPU security boundary. Configured in /etc/nvidia-container-runtime/config.toml.
Pod Security Standards at Baseline: capabilities.drop: [ALL], allowPrivilegeEscalation: false, seccompProfile.type: RuntimeDefault. Compatible with NVIDIA containers. NCCL_SHM_DISABLE=1 bypasses /dev/shm for non-root containers.
No privileged containers needed (device plugin handles GPU access). readOnlyRootFilesystem for inference, seccomp filtering to ~80 syscalls for PyTorch. Kyverno/OPA Gatekeeper enforces as admission webhooks.
| Control | Implementation | GPU Compatibility | Enforcement |
|---|---|---|---|
| cgroup device control | nvidia-container-runtime config | Full (required) | Runtime |
| runAsNonRoot | Pod securityContext | Partial (NCCL_SHM_DISABLE) | Admission Webhook |
| readOnlyRootFilesystem | Pod securityContext | Full (weights on PVC) | Admission Webhook |
| seccomp RuntimeDefault | Pod securityContext | Full | Pod Security Standard |
| NetworkPolicy zero-trust | Cilium | Full (NCCL port range) | NetworkPolicy CRD |
| Image signing (cosign) | cosign + admission | Full | Admission Webhook |
MULTI-TENANT ISOLATION FOR SHARED GPU INFRASTRUCTURE
GPU device isolation: NVIDIA device plugin allocates nvidia.com/gpu resources per pod. MIG on H100 provides hardware partitioning with up to 7 isolated instances per GPU with dedicated memory and compute. MIG sacrifices NVLink, suitable for inference.
Network isolation: NetworkPolicies with Cilium, default-deny, per-job policy allowing NCCL traffic between same-job pods via podSelector. InfiniBand isolation via PKey partitions with separate ipoib subnets per tenant.
Storage isolation: PVCs with per-team StorageClass quotas, Lustre project quotas (lfs setquota). Slurm fairshare: sacctmgr add qos with GrpTRES=gpu=128 limits per team. Prevents any single team from consuming all resources.
| Isolation Layer | Mechanism | Technology | Enforcement |
|---|---|---|---|
| GPU device | cgroup + device plugin | nvidia-container-runtime + MIG | Per-pod allocation |
| GPU memory | MIG hardware partitioning | H100 MIG profiles | Hardware-enforced |
| NCCL / fabric | InfiniBand PKey partitioning | OpenSM partition config | Per-partition PKey |
| Network (K8s) | NetworkPolicy + Cilium | CiliumNetworkPolicy CRD | Admission+dataplane |
| Storage capacity | Project quotas | Lustre project quota | Filesystem IOCTL |
| Scheduler | Slurm QoS / ResourceQuota | sacctmgr / ResourceQuota | Admission + scheduler |
