All essays
GuideGUIDEFEB 2026

Case Study: AI Voice/Call Center Startup GPU Deployment and Infrastructure

GPU infrastructure requirements for AI voice agents and call center startups. Real-time speech transcription, voice synthesis latency budgets, model serving architectures for 100-10,000 concurrent calls.

01

THE 200MS REAL-TIME BARRIER: WHY VOICE AI INFRASTRUCTURE IS UNIQUE

Voice AI has a hard real-time constraint that text-based AI does not: the acceptable latency for a conversational turn is 200-400 milliseconds end-to-end. Beyond 500ms, humans perceive the conversation as awkward or robotic. This latency budget must accommodate three sequential GPU-bound operations: automatic speech recognition (ASR, typically Whisper or a distilled variant: 30-150ms), LLM inference for response generation (100-500ms depending on model size), and text-to-speech synthesis (TTS, 50-200ms for high-quality output). The sum of these components-before any network latency-consumes 180-850ms with standard models, leaving zero headroom for multi-provider queuing, load balancing, or retries.

The infrastructure solution is a three-stage pipeline architecture where each stage runs on dedicated GPU instances. ASR runs on A10G or L4 GPUs (preferred for their lower cost and sufficient VRAM for Whisper-large-v3 at 10 GB), LLM inference runs on H100s (needed for model sizes of 7B-70B parameters), and TTS runs on the same GPU as the LLM or dedicated A10G instances (depending on model: CosmosVoice, Bark, or Play.ht’s proprietary models). The pipeline introduces a 5-20ms buffering delay between stages, but the total pipeline latency is 200-600ms p50 and 800ms-1.2s p95-marginally acceptable but leaving little room for spikes.

Pipeline StageModelGPU RequiredLatency (p50)GPU Cost/hrScaling Unit
ASR (Speech-to-Text)Whisper-large-v3 distilA10G / L430-80ms$0.80-$1.101 GPU per 200 concurrent calls
LLM (Response Gen)Llama 3.1 8B or Mixtral 8x7BH100 / A100-80G100-350ms$2.50-$3.501 GPU per 50-80 concurrent calls
TTS (Text-to-Speech)Play.ht / CosmosVoiceA10G / H100 (shared)50-150ms$0.80-$3.50 (shared)1 GPU per 100-200 concurrent calls
Agent OrchestrationCustom event loopCPU (8-16 vCPU)5-15ms$0.20-$0.501 instance per 500-1K calls
End-to-End (Combined)Full pipeline2-3 GPUs per flow200-600ms p50$4.50-$8.50See above
02

MODEL DISTILLATION STRATEGIES FOR SUB-200MS VOICE INFERENCE

Hitting sub-200ms end-to-end latency-the gold standard for natural conversation-requires every pipeline component to be distilled and optimized. For ASR, the standard is distilled Whisper variants (tiny or base distilled by OpenAI, or third-party distills like Whisper-Medusa). The original Whisper-large-v3 (1.5B parameters) runs at 80-150ms on an H100. Distilled to 32x attention reduction, Whisper-tiny (39M parameters) achieves 10-15ms latency with a 3-8 percent word error rate increase. Companies like Retell and Play.ai use a hybrid: Whisper-medium (769M params, 30-60ms latency) for the first pass, with a confidence threshold triggering a re-run of Whisper-large only when uncertainty exceeds 0.3.

For LLM inference in voice, a 7B-8B model is the sweet spot. Companies testing 70B models for voice find that at any batch size, the 300-600ms latency violates the real-time budget for anything except the most complex intent classification. The 8B Llama 3.1 class hits 100-200ms prompt processing and 50-100ms per token generation on H100, fitting 80-120 concurrent calls per GPU with continuous batching. Bland AI’s published benchmarks show their distilled 8B model achieves 280ms p50 and 480ms p95 end-to-end latency on 8x H100, serving 400 concurrent calls with a 4-percent abandonment rate (calls dropped by the customer due to silence).

Whisper VariantParametersLatency (H100)WER (LibriSpeech Clean)Ideally Used For
large-v31,550M80-150ms1.9%Quality-first, async transcription
medium769M30-60ms3.7%Default voice agent ASR
small244M20-35ms5.1%High-throughput, budget-optimized
tiny39M8-15ms8.3%Real-time streaming, first pass
distil-large-v3756M25-45ms2.4%Best quality-speed tradeoff
03

SCALING FROM 100 TO 10,000 CONCURRENT CALLS: THE GPU MATH

Scaling voice AI from 100 to 10,000 concurrent calls requires a fleet of 50-500 H100s. At 100 concurrent calls, a single Nvidia H100 GPU running an 8B LLM with continuous batching handles the LLM inference stage for 50-80 calls. ASR requires 1 A10G per 200 calls. TTS adds another GPU per 200-300 calls. Total: 4-7 GPUs for 100 concurrent calls. At 10,000 concurrent calls, the fleet scales to 200-350 H100s for LLM, 50-80 A10Gs for ASR, and 35-50 A10Gs for TTS, totaling 285-480 GPUs.

The non-linearity comes from batching efficiency. At 100 concurrent calls, batching is inefficient (batch size 2-4) because calls arrive asynchronously and the batch window (typically 50-100ms for voice) limits accumulation. At 10,000 concurrent calls, batch sizes reach 50-100, improving throughput per GPU by 2-3x. This means 10,000 calls require roughly 15-20x the GPUs of 100 calls, not 100x. The scaling inflection point occurs around 500-1,000 concurrent calls, where batch efficiency stabilizes and GPU requirements become approximately linear with call volume.

04

INFRASTRUCTURE PROFILES: BLAND AI, RETELL, AND PLAY.AI

Bland AI, the leader in outbound AI calling, processes 5-15 million calls per month on a GPU fleet estimated at 500-800 H100s. Their infrastructure is optimized for call connection rates (60-70 percent on outbound, with 10-30 second dial time per number). Bland uses a custom voice-optimized 8B LLM fine-tuned on 5 million+ call transcripts, plus a proprietary streaming ASR engine that begins transcription after 150ms of speech (instead of waiting for end-of-utterance), reducing perceived latency by 300-600ms. Bland’s abandonment rate is 2-3 percent on outbound calls (versus industry average 8-15 percent), which they attribute to sub-350ms end-to-end response latency.

Retell AI focuses on inbound voice agents for customer support. Their GPU deployment uses 128 H100s as of early 2026, with a tiered routing system: simple FAQ calls route to a distilled 3B model on A10Gs (180ms end-to-end), while complex support calls escalate to a 34B model on H100s (450ms end-to-end). Retell reports that 73 percent of calls are handled entirely by the distilled tier, meaning 73 percent of their call volume uses only 20 percent of their GPU fleet. Play.ai, focused on entertainment and character-based voice, uses 64 H100s for TTS synthesis running CosmosVoice and a proprietary voice cloning model, and routes all LLM inference to partner GPU providers.

05

THE UNIT ECONOMICS OF AI VOICE GPU INFRASTRUCTURE

The most detailed published voice AI infrastructure economics come from Bland AI’s 2025 blog post on unit economics. Bland’s cost per call minute breaks down as: compute GPU $0.0128 (at 60 percent utilization on 500 H100s averaging $1.85/hr with reserved pricing), ASR GPU $0.0042, TTS GPU $0.0031, telephony infrastructure $0.0085 (Twilio or competing SIP trunk provider), and overhead (orchestration, logging, monitoring) $0.0024. Total: $0.031 per call minute. At $0.08-0.15 per minute pricing to customers, gross margins are 55-79 percent.

The sensitivity of margins to GPU utilization is extreme. At 40 percent utilization (typical for a company with peak/off-peak call patterns and no load balancing across time zones), cost per minute rises to $0.046 and margins drop to 43-69 percent. At 80 percent utilization (achievable with multi-time-zone call routing and off-peak batch processing for voicemail or async voice), cost drops to $0.023 and margins reach 71-85 percent. Voice AI companies that achieve high GPU utilization through multi-region routing and async workload mixing have a 2x margin advantage over those that don’t.

Cost ComponentCost per Call MinutePercentage of TotalOptimization Lever
LLM Inference (H100)$0.012841%Model distillation, continuous batching
ASR (A10G/L4)$0.004214%Confidence-thresholded re-run strategy
TTS (A10G/H100)$0.003110%Shared GPU with LLM staging
Telephony (SIP)$0.008527%Direct SIP peering, volume discounts
Orchestration + Observability$0.00248%CPU-only serving (don’t waste GPUs)
Total Cost per Minute$0.031100%Target: <$0.02 at scale
Revenue per Minute$0.08-$0.15N/AVariable by vertical and volume
Filed under
AI Voice AgentsCall Center GPUReal-Time SpeechWhisper ASRVoice SynthesisConversational AI Infrastructure