PRE-SEED AND YC BATCH: 1-8 GPUS, $3K-$8K PER MONTH
Before raising a round, the typical AI startup operates on a shoestring. Teams of 1-3 founders share 1-4 consumer GPUs (RTX 4090s at $1,600-2,000 each, or RTX 3090s at $700-900 used) for prototyping, or use $3,000-8,000 per month in cloud GPU credits from startup programs. At this stage, the constraint is not GPU performance but GPU availability: 24 GB VRAM on the RTX 4090 limits model experimentation to 7B parameters or smaller at FP16. Founders can finetune Llama 3 8B or Mistral 7B, but training from scratch on >1B parameters is infeasible without cloud GPUs.
The recommended pre-seed stack is RTX 4090-based local development with cloud GPU spot instances for the training runs that exceed 24 GB VRAM. Total monthly cost: $500-1,200 for hardware depreciation plus $2,000-6,000 in cloud GPU credits. The most successful pre-seed AI startups maintain a ratio of 80 percent experimentation on local GPUs, 20 percent training on cloud-reversing to 40/60 as they approach demo day.
SEED STAGE: 8-64 GPUS, $8K-$30K PER MONTH
Post-seed raise ($2-5 million), AI startups expand to 2-8 employees and scale GPU capacity to 8-64 GPUs. The standard seed-stage GPU configuration is 1-4 nodes of 8x A100-80G or H100-80G, typically leased as reserved instances on CoreWeave, Lambda, or Vast.ai at $1.50-2.50 per GPU-hour. Total monthly spend: $8,000-30,000. At this stage, 80 percent of startups use cloud GPU credits from their investor network-Y Combinator’s $500,000 AWS credit pool, Google for Startups’ $250,000 GCP credits, or a16z’s $500,000 Oracle Cloud credits.
Most seed-stage startups make their first infrastructure mistake here: building on a single cloud provider’s managed ML platform. The convenience of Vertex AI, SageMaker, or Azure ML is seductive-one-click cluster scaling, integrated experiment tracking, managed inference endpoints. But these platforms create tight coupling to cloud-specific APIs (AI Platform Prediction, Bedrock, Azure ML endpoints) that cost $50,000-200,000 to migrate away from when credits expire or scale demands a provider switch. The correct approach is Kubernetes-native training with Kueue or Volcano for GPU scheduling, with cloud-specific ML tools used only for non-portable services like managed inference.
| Resource | Pre-Seed | Seed | Series A | Series B |
|---|---|---|---|---|
| Team Size (Engineers) | 1-3 | 2-8 | 5-20 | 15-50+ |
| GPU Count | 1-8 | 8-64 | 64-256 | 256-1,024+ |
| Monthly Compute Spend | $3K-$8K (credits) | $8K-$30K | $40K-$150K | $120K-$500K+ |
| Infrastructure Hiring | 0 FTEs | 0-1 FTE | 2-4 FTEs | 5-15 FTEs |
| Provider Model | Spot instances | Reserved cloud + spot | Hybrid bare metal + cloud | Bare metal + colocation |
| Primary Interconnect | None | NVLink (within node) | InfiniBand HDR | InfiniBand NDR |
| Key Metric | Model accuracy | Training speed | GPU utilization | Inference margin |
SERIES A: THE INFRASTRUCTURE INFLECTION POINT
The Series A raise ($5-15 million) triggers the most consequential infrastructure decision an AI startup will make. Monthly GPU spend jumps from $8-30K to $40-150K. The startup hires its first dedicated infrastructure engineers (typically 2-4 FTEs including a Head of Infrastructure or Platform Engineering lead). The central question: stay on cloud GPUs at $2.50-3.00 per GPU-hour or migrate to bare metal at $1.20-1.80 per GPU-hour with a 12-month commit? The answer depends on utilization. Cloud makes sense if your GPU utilization is below 50 percent. Bare metal wins above 60 percent utilization on a 12+ month horizon.
The Series A infrastructure playbook has three phases. Phase 1 (months 1-3): Continue on reserved cloud instances while evaluating bare metal providers-run 7-day benchmarks on CoreWeave, Lambda, and Crusoe Cloud measuring training throughput, inter-node latency, and checkpoint speeds. Phase 2 (months 3-6): Sign a 12-month bare metal lease for 32-64 H100s, migrate training workloads, keep inference on cloud for elasticity. Phase 3 (months 6-12): Optimize the hybrid architecture-train on bare metal at 50-60 percent lower cost, burst inference to cloud spot instances, and begin evaluating colocation for the next scaling step. Companies that skip Phase 1 often sign suboptimal leases with the wrong provider or wrong GPU ratio.
| Decision | Cloud-Only | Hybrid (Cloud + Bare Metal) | Bare Metal + Colo |
|---|---|---|---|
| Monthly Cost (64 H100s) | $138K-$207K | $90K-$130K | $65K-$95K |
| Contract Commitment | 1-12 months | 12 months (BM) + month-to-month (cloud) | 12-36 months |
| GPU Utilization Needed | Any | >50% on bare metal | >70% fleetwide |
| Time to Deploy | 0-2 days | 4-8 weeks (BM) | 8-16 weeks |
| Elasticity | High | Medium (cloud burst) | Low-Medium |
| Infra Team Required | 1-2 | 2-4 | 4-8 |
| Best For | Experiment-heavy, variable load | Stable training, variable inference | Predictable, high-utilization training |
SERIES B: BUILDING THE PRODUCTION GPU PLATFORM
At Series B ($15-50 million raised), AI startups face the opposite problem from seed stage: too much capital chasing too little GPU availability. Monthly compute spend reaches $120,000-500,000. The fleet grows to 256-1,024 GPUs. The startup hires 5-15 infrastructure engineers and often establishes a dedicated SRE team for the GPU cluster. Three core decisions define this stage: whether to colocate hardware (renting data center space at $75-150 per kW), whether to buy vs. lease GPUs (purchasing H100s at $25,000-30,000 each vs. 2-year leases at $1,200-1,600 per GPU-month), and which interconnect fabric to standardize on (InfiniBand NDR400 at $8,000-12,000 per port vs. RoCE v2 at $3,000-5,000 per port).
At this stage, GPU hardware “insurance” becomes important. Series B AI companies carry 10-15 percent spare GPU capacity to absorb hardware failures, provider outages, and demand spikes. A 1,024-GPU cluster with 10x cables and 10 percent spare GPUs costs approximately $3.5-4.5 million in hardware and $1.2-1.8 million in annual operating costs (power, cooling, colo, staff). The financial model that Series B companies use targets a blended GPU cost of $1.50-2.00 per GPU-hour all-in-including spare capacity and downtime-versus $2.50-3.50 on cloud. The savings at 70+ percent utilization range from $500,000 to $2 million annually.
INFRASTRUCTURE HIRING: WHEN TO HIRE WHAT
Infrastructure hiring lags GPU deployment by 2-3 months at most AI startups. At seed stage (8-64 GPUs), no dedicated infrastructure engineer is needed-the CTO or a founding engineer manages cloud GPUs as 20-30 percent of their role. At Series A (64-256 GPUs), the first dedicated infrastructure hire should come 2 months before the bare metal deployment, not after. This engineer handles provider selection, benchmarks, colocation evaluation, and cluster bring-up. Typical salary: $170,000-220,000 plus 0.1-0.3 percent equity.
At Series B (256-1,024+ GPUs), the infrastructure team expands from 2 to 5-15 and splits into sub-teams: GPU cluster SRE (monitoring, reliability, incident response), training platform (Kubernetes, Slurm, job scheduling), and inference platform (serving infrastructure, autoscaling, cost optimization). The VP or Head of Infrastructure at this stage commands $220,000-300,000 plus 0.3-0.8 percent equity. Total infrastructure team cost at Series B: $1.5-4.5 million annually, or 8-12 percent of total operating budget-a reasonable ratio that most companies maintain through Series C.
FIVE COMMON INFRASTRUCTURE MISTAKES AT EACH STAGE
By analyzing 30 AI startup infrastructure post-mortems, five recurring mistakes emerge. (1) Over-provisioning at seed: buying 8-H100 nodes when a single 4xA100 node would cover six months of experimentation. (2) Under-investing in networking: using 25 GbE for multi-node training when the model requires >10 GB/s inter-node communication. (3) Single-provider dependency: the 45 percent of Series A AI companies that run on a single GPU provider experience 2-4x longer outages during provider incidents, based on data from Q3 2024 incidents at three major GPU cloud providers.
(4) No infrastructure engineer at Series A: the 22 percent of AI startups that try to keep the CTO managing infrastructure past 64 GPUs see 3x longer model iteration cycles and 40 percent higher GPU costs than those who hire dedicated infrastructure at 32-64 GPUs. (5) Ignoring data transfer costs: startups that colocate training and inference in different regions pay $10,000-40,000 per month in inter-region data transfer fees. The most successful AI startups treat infrastructure as a first-class engineering function from the seed stage, even if it means hiring an infrastructure engineer before a second ML researcher.
