DIFFERENTIAL PRIVACY TRAINING INFRASTRUCTURE
Differential privacy (DP) training limits the information leakage from individual training examples by adding calibrated noise to gradient updates. The standard approach is DP-SGD (Differentially Private Stochastic Gradient Descent), which clips per-sample gradients to a maximum L2 norm and adds Gaussian noise scaled to the desired privacy budget (epsilon). The privacy budget epsilon controls the tradeoff: lower epsilon values give stronger privacy guarantees but degrade model quality. A typical production deployment targets epsilon = 8 for general-purpose models (matching Apple's reported privacy parameters for on-device ML) and epsilon = 2-4 for high-privacy settings like healthcare or finance models.
The GPU overhead of DP-SGD is substantial. Gradient clipping requires computing per-sample gradients rather than per-batch gradients, which is the standard optimization in non-private training. This increases memory consumption proportionally to batch size. For a 70B model trained with batch size 128, DP-SGD requires 128x the activation memory of non-private training because each sample's gradient must be computed individually before clipping and aggregation. Practical DP training uses micro-batching: compute per-sample gradients on micro-batches of 4-8 samples, clip and accumulate, then update. This adds approximately 15-30% training time overhead for LLMs. The noise injection step is computationally negligible but requires careful scaling. Tools like Opacus (Meta) and JAX Privacy provide DP-SGD implementations optimized for GPU training. On H100 clusters, DP training of a 7B model requires approximately 20-30% more GPU-hours than non-private training for equivalent data throughput.
| Privacy Parameter | Epsilon=8 (Standard) | Epsilon=4 (High Privacy) |
|---|---|---|
| Training overhead | 15-25% more GPU-hours | 30-50% more GPU-hours |
| Memory per batch (70B) | 32-48 GB | 48-64 GB |
| Micro-batch size | 4-8 samples | 2-4 samples |
| Noise std dev | 0.5-1.0 | 1.5-3.0 |
| Quality degradation (perplexity) | 1-3% | 5-10% |
| GPU requirement (7B) | 8x H100 (comparable) | 16x H100 (recommended) |
MACHINE UNLEARNING: DATA DELETION FROM TRAINED MODELS
The right to deletion under GDPR Article 17 (right to erasure) presents a fundamental challenge for AI models: once data is used for training, removing its influence from a trained model requires either exact unlearning or approximate unlearning via model updates. Exact unlearning retrains the model from scratch on the remaining data - computationally prohibitive for frontier models (retraining a 70B model costs $1-2M on H100 clusters). Approximate unlearning techniques reduce the compute cost by fine-tuning the model to forget specific data points or data cohorts.
The infrastructure for unlearning typically combines three strategies. First, data partitioning: train on data shards and use ensemble methods so that forgetting one shard removes only one model in the ensemble. This adds 2-5x training compute but makes deletion point-in-time efficient. Second, SISA (Sharded, Isolated, Sliced, Aggregated) training by Bourtoule et al. (2021) partitions training data into disjoint shards, trains a separate model per shard, and aggregates predictions. Deleting one user's data requires retraining only the affected shard's model. Third, for non-ensemble models, approximate unlearning via fine-tuning with a forget loss that maximizes loss on the target data while maintaining performance on the remaining data. The approximate approach requires 50-200 GPU-hours for a 70B model versus 50,000+ GPU-hours for full retraining. On ClusterBid, teams maintaining GDPR-compliant AI systems can pre-provision the unlearning compute capacity as reserved spot instances, ensuring deletion requests can be processed within the GDPR's 30-day window.
INFERENCE-TIME PRIVACY AND CONSENT ENFORCEMENT
Beyond training data privacy, AI systems must enforce privacy and consent at inference time. This includes: consent-based access control (only generating content for users who have consented to the specific use case), data minimization (processing only the minimum necessary input data), and output privacy (preventing generation of PII or copyrighted content). The consent enforcement infrastructure integrates with the inference pipeline: each request is tagged with the user's consent profile, and the inference engine applies policy filters based on the consent scope.
The infrastructure for inference privacy includes: a consent management service that maintains user consent profiles (stored in a GDPR-compliant database with consent timestamps and scope definitions), a consent-based routing layer that sits between the API gateway and the inference endpoint, and PII scanning models that filter both inputs and outputs for personally identifiable information. PII scanning adds 10-50ms of GPU inference latency per request using models like Presidio (Microsoft) or fine-tuned NER transformers. For text generation, output privacy filters run on the generated text post-inference, costing 5-20ms per generation. For a platform processing 10M inference requests per day with consent enforcement, the privacy pipeline adds approximately 30-80 GPU-hours per day for PII scanning plus the consent management infrastructure (CPU-based, minimal cost).
| Privacy Enforcement Layer | Latency Overhead | GPU Cost per 1M Requests |
|---|---|---|
| Consent profile lookup | 1-5ms | Zero (CPU/Redis) |
| Consent-based routing | 0.5-2ms | Zero (API gateway) |
| Input PII scanning | 10-50ms | $1-5 (NER model) |
| Output PII scanning | 5-20ms | $0.50-3 (NER model) |
| Copyright detection (output) | 20-100ms | $2-10 (embedding + search) |
| Total inference privacy | 35-175ms | $3.50-18 per 1M |
CONSENT MANAGEMENT INFRASTRUCTURE DESIGN
A production consent management system for AI must handle: consent collection (recording user consent for specific AI use cases with granular opt-in), consent storage (maintaining an immutable audit log of all consent events with proof of consent), consent enforcement (checking consent before each AI processing operation), and consent revocation (processing withdrawal requests and propagating to all systems). The infrastructure must satisfy GDPR Article 7 requirements for consent: freely given, specific, informed, unambiguous, and withdrawable at any time. Each consent event must be timestamped and recorded with the exact terms presented to the user.
The consent infrastructure scales with user base and AI use case complexity. A platform with 10M users and 5 AI features needs: 50M consent records (each user x feature), 500M+ consent check events per day (each inference checks consent), and sub-5ms consent lookup latency. The typical architecture uses Redis for consent cache (checking 99% of requests without database query) and PostgreSQL/CockroachDB for the consent audit log (immutable, append-only). The consent cache must support instant invalidation when a user revokes consent. The consent enforcement layer integrates with the inference API gateway, returning 403 Forbidden if the user has not consented to the specific AI feature being invoked. Audit logs must support: proving consent existed at a specific timestamp (for responding to regulatory inquiries) and producing consent statistics (percentage of users consenting to each feature, by region and demographics).
PRIVACY AUDIT AND VERIFICATION INFRASTRUCTURE
Privacy compliance audits verify that AI systems adhere to stated privacy policies and regulatory requirements. The audit infrastructure includes: data flow mapping (tracing data from collection through training through inference to deletion), consent compliance verification (checking that all processed data had valid consent), deletion processing verification (confirming that deletion requests resulted in actual removal), and DP guarantee verification (for systems claiming differential privacy). Each audit dimension requires different infrastructure - data flow mapping uses lineage tracking from the data catalog, while DP verification requires privacy accounting tools like the DP Accountant.
The automated privacy audit pipeline runs on a scheduled basis (monthly for standard systems, weekly for high-privacy systems). It queries: the consent database for consent validity rates, the data deletion logs for deletion completion rates and SLAs, the training pipeline for DP accounting (if applicable), and the inference logs for privacy policy violations (PII generated in outputs). Results are compiled into a privacy compliance report with metrics against defined thresholds. For GDPR compliance, the report must be available for submission to Supervisory Authorities within 72 hours of request. On ClusterBid, teams can store privacy audit data in the same storage tier as GPU evaluation artifacts, creating a unified compliance data lake that supports rapid regulatory response.
| Audit Dimension | Frequency | Infrastructure Components |
|---|---|---|
| Consent compliance | Monthly | Consent DB query + report generator |
| Data flow mapping | Quarterly | Data catalog + lineage system |
| Deletion verification | Per deletion + quarterly | Deletion log + sample check |
| DP guarantee verification | Per training run | Privacy accountant + audit log |
| Inference privacy check | Weekly | Sample inference log review |
| Full privacy audit | Annual | All above + external auditor access |
