All essays
TechnicalDEEP DIVEFEB 2026

GPU Infrastructure for AI Consulting Firms: How Agencies Serve Multiple AI Clients

How AI consulting firms and agencies manage GPU infrastructure across 10-100+ clients. Multi-tenant GPU scheduling, client isolation, cost allocation, and the agency infrastructure stack for model development and deployment at scale.

01

THE MULTI-TENANT GPU CHALLENGE UNIQUE TO AI CONSULTING

AI consulting firms and agencies face a GPU infrastructure challenge that’s distinct from either single-company deployments or cloud providers: they must maintain 10-100+ separate model environments-each with different GPU requirements, security postures, data residency constraints, and latency SLAs-on a shared pool of GPUs. A typical mid-size AI consultancy ($5-20M annual revenue) operates 32-256 H100-equivalent GPUs across 15-60 active client projects, each requiring 1-8 GPUs for development and 4-32 GPUs for inference serving. The consulting GPU allocation is rarely static: clients ramp up for model development sprints (4-12 weeks) then drop to 30-50 percent for maintenance inference.

The core technical challenge is multi-tenant GPU scheduling with strong isolation. Unlike cloud providers who can dedicate entire nodes per customer, agencies must share GPU nodes across multiple clients to achieve economic viability. A single H100 node (8 GPUs) costing $20,000-25,000 per month must support 2-4 client projects simultaneously, each running in isolated containers or virtual machines. The isolation requirements include: GPU memory isolation (preventing one client’s model from OOM-killing another’s), network security (VLAN or encrypted VPC per client), data at rest encryption (client-specific encryption keys for model weights), and billing granularity (per-project GPU-hour accounting).

Agency SizeAnnual RevenueGPU FleetActive ClientsAvg GPUs per ClientInfra Team Size
Boutique$1M-$5M8-32 H100-equiv5-151-41-2 FTEs
Mid-Size$5M-$20M32-256 H10015-602-83-6 FTEs
Large Agency$20M-$100M256-1,024 H10050-200+4-168-20 FTEs
AI Dev Shop (Captive)$50M-$300M512-4,096 H10010-30 (large)8-6415-40 FTEs
02

MULTI-TENANT GPU SCHEDULING: KUBERNETES WITH GPU PARTITIONING

The standard approach for agency GPU infrastructure is Kubernetes with node-level GPU partitioning. Each H100 node (8 GPUs) is partitioned using NVIDIA’s MIG (Multi-Instance GPU) technology, creating 7 GPU instances per H100 (3 instances with one GPU slice each for development, 4 instances with two slices for training). MIG provides hardware-level memory and cache isolation: each partition has dedicated VRAM, L2 cache, and memory bandwidth. The trade-off is flexibility-MIG partitions are static and require node reconfiguration to change. An alternative is time-slicing (multiple pods sharing a GPU via NVIDIA’s MPS), which provides finer granularity but no memory isolation: one client’s OOM can crash the shared GPU.

Kubernetes-based agencies implement namespace-per-client RBAC with resource quotas (max GPU-hours per week, max memory per namespace). GPU monitoring uses DCGM exporter with Prometheus to track per-pod GPU utilization, memory, and temperature. The cost allocation model charges clients based on GPU-seconds tracked via the Kubernetes Metrics API, typically marked up 30-50 percent over the agency’s blended GPU cost. A mid-size agency with 128 H100s at $1.50/hour blended cost charges clients $2.00-2.50 per GPU-hour, recovering $100,000-200,000 per month in overhead for GPU management, storage, networking, and infrastructure team salaries.

GPU Sharing MethodIsolation LevelMax Partitions per H100FlexibilityRecommended For
MIG (Hardware)Full memory + cache isolation7 (single-GPU slices)Low (requires reconfig)Production inference, security-sensitive clients
MPS (Time-sliced)No memory isolationUp to 48 concurrentHigh (dynamic)Development workloads, batch inference
vGPU (Virtualization)Full GPU isolation (vGPU)4-8 per H100Medium (license-dependent)Mixed workloads, enterprise clients
Dedicated NodeComplete isolation1 per 8-GPU nodeNoneLarge GPU clients, regulated industries
K8s Namespace + QuotasLogical (no hw isolation)Unlimited (soft limits)MaximumInternal dev, non-production workloads
03

CLIENT PROJECT GPU LIFECYCLE: FROM PROOF-OF-CONCEPT TO PRODUCTION

AI consulting engagements follow a predictable GPU lifecycle. Phase 1 (proof-of-concept, weeks 1-4): The agency allocates 1-4 GPUs per client for rapid prototyping-finetuning a 7B-13B model on client data, building the inference pipeline. GPU utilization during POC is typically low (30-50 percent) because development is iterative with pauses for client review. Phase 2 (development, weeks 5-12): GPU allocation scales to 4-16 GPUs per client for hyperparameter sweeps, multi-model comparison, and production-hardening. Utilization peaks at 60-80 percent during this sprint.

Phase 3 (production deployment, weeks 13-16): The agency deploys 4-32 GPUs for inference serving, depending on expected query volume. Utilization patterns shift from spiky (development) to steady-state (production). Phase 4 (maintenance, ongoing): GPU allocation drops to 2-8 GPUs for the client, handling model updates (retraining every 2-4 weeks), A/B testing, and inference scaling. The key metric for agencies is GPU utilization across the entire portfolio: the most profitable agencies maintain 65-80 percent fleetwide utilization by carefully timing POC, development, and production phases across their client base. One agency we interviewed targets exactly 75 percent utilization-leaving 25 percent headroom for new client POCs that can start within 48 hours of contract signing.

04

COST ALLOCATION MODELS: HOW AGENCIES PRICE GPU TO CLIENTS

Three cost allocation models dominate AI agency GPU pricing. The simplest is fixed monthly retainer ($15,000-60,000 per month for 8-32 GPUs), preferred by enterprise clients who need predictable budgets. The most common for mid-size clients is per-GPU-hour pass-through ($1.80-3.00 per H100-hour, marked up 30-70 percent over agency cost), with detailed billing showing GPU-seconds per project phase. The emerging model for sophisticated clients is outcome-based pricing: the agency charges $0.05-0.20 per production inference call (for inference-heavy projects) or a flat success fee ($50,000-200,000) for delivering a production model that meets accuracy KPIs.

The margin analysis for agency GPU pricing is as follows. At $2.50 per GPU-hour billed to clients and $1.50 blended cost (including reserved instance discount, colocation, and networking), gross margin on GPU resale is 40 percent. Subtract infrastructure team costs ($500,000-1,200,000 annually for a mid-size agency’s 3-6 FTE infrastructure team) and storage/network costs ($50,000-150,000 annually), and the net margin on GPU services is 22-30 percent. The most profitable agencies achieve 35-40 percent net margins by operating at 80 percent+ fleet utilization, using spot instances for non-production work, and charging premium rates for compliance-validated GPU environments (HIPAA, SOC 2).

Pricing ModelTypical RateClient PreferenceAgency MarginBest For
Fixed Monthly Retainer$15K-$60K/monthEnterprise (predictable)25-40%Ongoing maintenance clients
Per GPU-Hour$1.80-$3.00/hrStartups (flexible)30-50%Variable workload clients
Per Inference Call$0.05-$0.20/callScale-up (usage-based)35-55%High-volume inference clients
Outcome/Success Fee$50K-$200K per modelValue-oriented clients50-80%One-off model delivery
Hybrid (Retainer + Usage)$10K-$30K base + $1.50-$2.00/hrMid-market30-45%Most common, balanced model
05

INFRASTRUCTURE PROFILES: THREE AI AGENCIES’ GPU DEPLOYMENT STRATEGIES

Case 1: A boutique NLP agency ($4M revenue, 25 clients) operates 24 H100s at a colocation facility in Dallas, purchased outright at $25,000 each ($600,000 total capex). They use Slurm for job scheduling (preferred over Kubernetes for its simpler GPU accounting) and charge a flat $50/hour per allocated GPU. With 70 percent fleet utilization and $1.40/hour all-in cost (depreciation + power + colo), they generate $1,150/hour revenue on a $336/hour cost base-a 71 percent gross margin. Their biggest operational challenge is client ramp scheduling: when 5 clients simultaneously enter development sprints, they have 24-hour lead time to spin up additional GPU nodes.

Case 2: A mid-size full-stack AI agency ($18M revenue, 60 concurrent clients) uses a hybrid model. They own 128 H100s at a colocation facility for stable client workloads but burst to Lambda and Vast.ai for POCs and spikes. Their Kubernetes-based platform with namespace-per-client and GPU-time-slicing (MPS) allows 15-25 client pods per 8-GPU node. They report that the Kubernetes overhead (API server, etcd, monitoring) consumes approximately 5 percent of total GPU compute and 8 percent of engineering time. Their infrastructure team of 4 engineers manages the platform. Case 3: A large AI dev shop ($85M revenue) operates 1,024 H100s across two colocation facilities (Ashburn and Dallas) and uses a custom multi-tenant GPU scheduler built on top of Slurm, with per-client GPU quotas enforced through cgroups and GPU time banking.

Filed under
AI Consulting GPUAgency InfrastructureMulti-tenant GPUGPU Cost AllocationAI Services AgencyClient GPU Management