THE MULTI-TENANT GPU CHALLENGE UNIQUE TO AI CONSULTING
AI consulting firms and agencies face a GPU infrastructure challenge that’s distinct from either single-company deployments or cloud providers: they must maintain 10-100+ separate model environments-each with different GPU requirements, security postures, data residency constraints, and latency SLAs-on a shared pool of GPUs. A typical mid-size AI consultancy ($5-20M annual revenue) operates 32-256 H100-equivalent GPUs across 15-60 active client projects, each requiring 1-8 GPUs for development and 4-32 GPUs for inference serving. The consulting GPU allocation is rarely static: clients ramp up for model development sprints (4-12 weeks) then drop to 30-50 percent for maintenance inference.
The core technical challenge is multi-tenant GPU scheduling with strong isolation. Unlike cloud providers who can dedicate entire nodes per customer, agencies must share GPU nodes across multiple clients to achieve economic viability. A single H100 node (8 GPUs) costing $20,000-25,000 per month must support 2-4 client projects simultaneously, each running in isolated containers or virtual machines. The isolation requirements include: GPU memory isolation (preventing one client’s model from OOM-killing another’s), network security (VLAN or encrypted VPC per client), data at rest encryption (client-specific encryption keys for model weights), and billing granularity (per-project GPU-hour accounting).
| Agency Size | Annual Revenue | GPU Fleet | Active Clients | Avg GPUs per Client | Infra Team Size |
|---|---|---|---|---|---|
| Boutique | $1M-$5M | 8-32 H100-equiv | 5-15 | 1-4 | 1-2 FTEs |
| Mid-Size | $5M-$20M | 32-256 H100 | 15-60 | 2-8 | 3-6 FTEs |
| Large Agency | $20M-$100M | 256-1,024 H100 | 50-200+ | 4-16 | 8-20 FTEs |
| AI Dev Shop (Captive) | $50M-$300M | 512-4,096 H100 | 10-30 (large) | 8-64 | 15-40 FTEs |
MULTI-TENANT GPU SCHEDULING: KUBERNETES WITH GPU PARTITIONING
The standard approach for agency GPU infrastructure is Kubernetes with node-level GPU partitioning. Each H100 node (8 GPUs) is partitioned using NVIDIA’s MIG (Multi-Instance GPU) technology, creating 7 GPU instances per H100 (3 instances with one GPU slice each for development, 4 instances with two slices for training). MIG provides hardware-level memory and cache isolation: each partition has dedicated VRAM, L2 cache, and memory bandwidth. The trade-off is flexibility-MIG partitions are static and require node reconfiguration to change. An alternative is time-slicing (multiple pods sharing a GPU via NVIDIA’s MPS), which provides finer granularity but no memory isolation: one client’s OOM can crash the shared GPU.
Kubernetes-based agencies implement namespace-per-client RBAC with resource quotas (max GPU-hours per week, max memory per namespace). GPU monitoring uses DCGM exporter with Prometheus to track per-pod GPU utilization, memory, and temperature. The cost allocation model charges clients based on GPU-seconds tracked via the Kubernetes Metrics API, typically marked up 30-50 percent over the agency’s blended GPU cost. A mid-size agency with 128 H100s at $1.50/hour blended cost charges clients $2.00-2.50 per GPU-hour, recovering $100,000-200,000 per month in overhead for GPU management, storage, networking, and infrastructure team salaries.
| GPU Sharing Method | Isolation Level | Max Partitions per H100 | Flexibility | Recommended For |
|---|---|---|---|---|
| MIG (Hardware) | Full memory + cache isolation | 7 (single-GPU slices) | Low (requires reconfig) | Production inference, security-sensitive clients |
| MPS (Time-sliced) | No memory isolation | Up to 48 concurrent | High (dynamic) | Development workloads, batch inference |
| vGPU (Virtualization) | Full GPU isolation (vGPU) | 4-8 per H100 | Medium (license-dependent) | Mixed workloads, enterprise clients |
| Dedicated Node | Complete isolation | 1 per 8-GPU node | None | Large GPU clients, regulated industries |
| K8s Namespace + Quotas | Logical (no hw isolation) | Unlimited (soft limits) | Maximum | Internal dev, non-production workloads |
CLIENT PROJECT GPU LIFECYCLE: FROM PROOF-OF-CONCEPT TO PRODUCTION
AI consulting engagements follow a predictable GPU lifecycle. Phase 1 (proof-of-concept, weeks 1-4): The agency allocates 1-4 GPUs per client for rapid prototyping-finetuning a 7B-13B model on client data, building the inference pipeline. GPU utilization during POC is typically low (30-50 percent) because development is iterative with pauses for client review. Phase 2 (development, weeks 5-12): GPU allocation scales to 4-16 GPUs per client for hyperparameter sweeps, multi-model comparison, and production-hardening. Utilization peaks at 60-80 percent during this sprint.
Phase 3 (production deployment, weeks 13-16): The agency deploys 4-32 GPUs for inference serving, depending on expected query volume. Utilization patterns shift from spiky (development) to steady-state (production). Phase 4 (maintenance, ongoing): GPU allocation drops to 2-8 GPUs for the client, handling model updates (retraining every 2-4 weeks), A/B testing, and inference scaling. The key metric for agencies is GPU utilization across the entire portfolio: the most profitable agencies maintain 65-80 percent fleetwide utilization by carefully timing POC, development, and production phases across their client base. One agency we interviewed targets exactly 75 percent utilization-leaving 25 percent headroom for new client POCs that can start within 48 hours of contract signing.
COST ALLOCATION MODELS: HOW AGENCIES PRICE GPU TO CLIENTS
Three cost allocation models dominate AI agency GPU pricing. The simplest is fixed monthly retainer ($15,000-60,000 per month for 8-32 GPUs), preferred by enterprise clients who need predictable budgets. The most common for mid-size clients is per-GPU-hour pass-through ($1.80-3.00 per H100-hour, marked up 30-70 percent over agency cost), with detailed billing showing GPU-seconds per project phase. The emerging model for sophisticated clients is outcome-based pricing: the agency charges $0.05-0.20 per production inference call (for inference-heavy projects) or a flat success fee ($50,000-200,000) for delivering a production model that meets accuracy KPIs.
The margin analysis for agency GPU pricing is as follows. At $2.50 per GPU-hour billed to clients and $1.50 blended cost (including reserved instance discount, colocation, and networking), gross margin on GPU resale is 40 percent. Subtract infrastructure team costs ($500,000-1,200,000 annually for a mid-size agency’s 3-6 FTE infrastructure team) and storage/network costs ($50,000-150,000 annually), and the net margin on GPU services is 22-30 percent. The most profitable agencies achieve 35-40 percent net margins by operating at 80 percent+ fleet utilization, using spot instances for non-production work, and charging premium rates for compliance-validated GPU environments (HIPAA, SOC 2).
| Pricing Model | Typical Rate | Client Preference | Agency Margin | Best For |
|---|---|---|---|---|
| Fixed Monthly Retainer | $15K-$60K/month | Enterprise (predictable) | 25-40% | Ongoing maintenance clients |
| Per GPU-Hour | $1.80-$3.00/hr | Startups (flexible) | 30-50% | Variable workload clients |
| Per Inference Call | $0.05-$0.20/call | Scale-up (usage-based) | 35-55% | High-volume inference clients |
| Outcome/Success Fee | $50K-$200K per model | Value-oriented clients | 50-80% | One-off model delivery |
| Hybrid (Retainer + Usage) | $10K-$30K base + $1.50-$2.00/hr | Mid-market | 30-45% | Most common, balanced model |
INFRASTRUCTURE PROFILES: THREE AI AGENCIES’ GPU DEPLOYMENT STRATEGIES
Case 1: A boutique NLP agency ($4M revenue, 25 clients) operates 24 H100s at a colocation facility in Dallas, purchased outright at $25,000 each ($600,000 total capex). They use Slurm for job scheduling (preferred over Kubernetes for its simpler GPU accounting) and charge a flat $50/hour per allocated GPU. With 70 percent fleet utilization and $1.40/hour all-in cost (depreciation + power + colo), they generate $1,150/hour revenue on a $336/hour cost base-a 71 percent gross margin. Their biggest operational challenge is client ramp scheduling: when 5 clients simultaneously enter development sprints, they have 24-hour lead time to spin up additional GPU nodes.
Case 2: A mid-size full-stack AI agency ($18M revenue, 60 concurrent clients) uses a hybrid model. They own 128 H100s at a colocation facility for stable client workloads but burst to Lambda and Vast.ai for POCs and spikes. Their Kubernetes-based platform with namespace-per-client and GPU-time-slicing (MPS) allows 15-25 client pods per 8-GPU node. They report that the Kubernetes overhead (API server, etcd, monitoring) consumes approximately 5 percent of total GPU compute and 8 percent of engineering time. Their infrastructure team of 4 engineers manages the platform. Case 3: A large AI dev shop ($85M revenue) operates 1,024 H100s across two colocation facilities (Ashburn and Dallas) and uses a custom multi-tenant GPU scheduler built on top of Slurm, with per-client GPU quotas enforced through cgroups and GPU time banking.
TRENDS SHAPING AI AGENCY GPU INFRASTRUCTURE IN 2026
Three trends are reshaping how AI agencies approach GPU infrastructure. First, the rise of GPU accounting as a service: startups like ClusterBid and GPU.Report provide multi-cloud GPU cost allocation and optimization built for agencies, offering per-client billing, spot instance arbitrage, and automated workload placement across providers. Agency infrastructure teams that adopt these tools report 20-30 percent reduction in GPU costs and 60 percent reduction in billing reconciliation time.
Second, dedicated GPU nodes for regulated clients: as AI deployments in healthcare (HIPAA), finance (SOC 2, PCI), and government (FedRAMP) grow, agencies must maintain air-gapped GPU nodes that never touch the shared scheduling pool. The cost premium for regulated GPU environments is 40-80 percent over standard shared nodes. Third, the build-vs-partner decision is tilting toward partnership: 60 percent of mid-size AI agencies now use a GPU partner (CoreWeave, Lambda, or specialized agency GPU resellers) rather than building colocated clusters, because GPU procurement and management are not core competencies and distraction from client delivery costs more than the 15-25 percent margin premium paid to partners. The agencies that successfully scale are those that treat GPU infrastructure as a managed service layer beneath their client-facing intellectual property-not as a competitive advantage.
