All essays
TechnicalDEEP DIVEFEB 2026

Case Study: AI Search Startup Infrastructure-Perplexity-Style Architecture

How AI search engines like Perplexity handle GPU infrastructure for real-time web search with LLM synthesis. Indexing pipelines, inference cluster design, RAG infrastructure, and the economics of answering queries at scale.

01

THE AI SEARCH PAINT: BEYOND SIMPLE RAG

AI search engines like Perplexity, Arc Search, You.com, and Komo operate a fundamentally different pipeline from chatbot LLMs. A query goes through five sequential stages: query rewriting and intent classification (CPU, 5-20ms), web search and crawl (2-10 API calls to Bing/Google indexes, typically 200-800ms), content extraction and chunking (CPU, 50-200ms for 5-20 web pages), retrieval-augmented generation (RAG) inference (LLM synthesizes answers from retrieved chunks, 300ms-2s for 7B-34B models with 16K-32K context), and citation-augmented response formatting (CPU, 20-50ms). Each stage is separately scalable, but the RAG inference stage is the GPU bottleneck.

The RAG inference architecture differs from chatbot inference in two critical ways. First, context windows are 2-4x larger: a typical AI search query concatenates 8-20 web page chunks into a 6K-16K token prompt, versus 1K-4K for a chatbot. This means prompt processing (prefill) dominates inference latency rather than token generation. With vLLM’s chunked prefill, a 16K-token prompt processes in 150-400ms on H100 (versus 50-100ms for a 2K prompt). Second, cache hit rates are lower: web search queries are unique and ephemeral (cache hit rates of 5-15 percent versus 40-60 percent for coding assistants), meaning prefix caching helps less and each query requires near-full inference.

Pipeline StageHardwareLatency BudgetCost per 1K QueriesScaling Unit
Query Rewriting (Classifier)CPU (8 vCPU)5-20ms$0.0011 instance per 5K QPS
Web Search API CallsExternal API200-800ms$0.07-$0.30 (Bing API)Rate-limited by search API
Content ExtractionCPU (16 vCPU)50-200ms$0.0051 instance per 1K QPS
RAG Inference (7B-34B)H100300ms-2s$0.50-$2.001 GPU per 10-30 QPS
Result FormattingCPU (4 vCPU)20-50ms$0.0011 instance per 3K QPS
Total per QueryMixed0.5-3.5s$0.06-$0.25See above
02

THE REAL-TIME INDEXING CHALLENGE: 100M+ PAGES WITH SUB-10 MINUTE FRESHNESS

Unlike traditional search engines that build offline indexes updated daily, AI search engines must index web pages in near real-time to answer queries about breaking news, product launches, and current events. Perplexity’s indexing infrastructure processes an estimated 50-200 million web pages daily, with a target index-to-search latency of 2-10 minutes for high-authority domains. The indexing pipeline uses GPU-accelerated embedding models (e.g., E5-mistral-7b-instruct or BGE-large-en-v1.5) running on 64-128 H100s, generating 768-1024-dimensional embeddings for each document chunk at a rate of 100-500 chunks per second per GPU.

The embedding model choice determines index quality and GPU cost. A 7B embedding model produces higher-quality retrievals (NDCG@10 scores 3-5 percent higher on standard IR benchmarks) but requires 2-3x more GPU compute per document than a 350M-parameter model. The practical compromise: use a 7B embedding model for authoritative sources (Wikipedia, news outlets, academic papers) and a 350M-1.5B model for long-tail web pages, reducing total embedding compute by 60-70 percent while maintaining top-k retrieval quality. Vector storage adds another cost dimension: a 100M-document index with 768-dimensional embeddings at FP32 requires 307 GB of GPU or CPU memory for the index alone, typically served from a vector database cluster (Pinecone, Weaviate, Qdrant) costing $5,000-20,000 per month for production-scale deployments.

03

PERPLEXITY-SCALE GPU REQUIREMENTS: ESTIMATES AND EXTRAPOLATIONS

Based on public statements, job postings, and infrastructure footprint analysis, Perplexity likely operates 256-512 H100-equivalent GPUs for inference serving and an additional 64-128 GPUs for indexing and embedding. At $2.50-3.00 per GPU-hour on reserved instances (Perplexity uses a mix of cloud and bare metal), monthly inference compute cost is approximately $460,000-1,100,000 for serving. Search API costs (Bing or Google) add $200,000-500,000 per month. Total monthly infrastructure cost: $800,000-1,800,000, serving an estimated 15-25 million queries per day.

The GPU-to-query ratio is the critical efficiency metric. Perplexity reportedly achieves approximately 30-50 queries per GPU-hour (implying 1-2 minutes of GPU time per query, factoring in caching and batching). This is 5-10x less efficient than a standard chatbot (which achieves 200-500 queries per GPU-hour for a 7B model) due to the longer context windows and the sequential prefill-dominant workload. The efficiency gap explains why AI search companies need significantly larger GPU fleets than LLM chatbot companies of equivalent user scale-and why infrastructure cost is the primary constraint on growth.

04

STREAMING INFERENCE AND CACHE ARCHITECTURE FOR AI SEARCH

AI search engines use streaming inference (Server-Sent Events) to deliver partial results within 200-500ms, masking the total generation time of 1-4 seconds. Users see the first tokens streamed within 300-800ms of query submission, even though the full answer takes 2-4 seconds. This requires careful GPU scheduling: inference servers (vLLM or TensorRT-LLM) must preempt ongoing prompt processing when a streaming slot opens. Perplexity’s engineering blog noted that their streaming architecture can serve first-token latency of 350ms p50 and 800ms p95 for 7B models-competitive with traditional search result latency.

The cache architecture for AI search is fundamentally different from chatbots. Chatbots cache KV states for repeated queries (40-60 percent hit rates). AI search engines cache at the web result level: if two users ask “What is the weather in Tokyo?” within 5 minutes, the crawled web results are identical, but the LLM synthesis may differ because of random sampling. The cache strategy is to cache search results (not generation) at the chunk level, reducing web crawling by 30-50 percent for hot queries, and to use semantic caching for query embeddings (identical or near-identical queries within a time window). A 24-hour cache of hot query embeddings (10,000-50,000 queries) requires 8-40 GB of GPU memory, typically allocated on a separate embedding cache server.

Cache LayerCache DurationHit RateMemory RequiredGPU Cost Saved
Web Search Result Cache (URL level)5-30 minutes30-50%50-200 GB (NVMe)Significant (spares API costs)
Embedding Cache (Query level)1-24 hours10-20%8-40 GB (GPU RAM)Moderate (spares embedding compute)
LLM KV Cache (Prefix level)Not effective (<5% hits)5-15%High (varies)Minimal for search use case
Full Response Cache (Exact match)Varies by domain5-10%100-500 GB (NVMe/DRAM)Low (unique queries dominate)
05

COST STRUCTURE AND INFRASTRUCTURE EVOLUTION: THE ROAD TO <$0.01 PER QUERY

The long-term viability of AI search depends on reducing GPU cost per query from the current $0.06-0.25 to $0.01-0.03-competitive with traditional search monetization ($0.02-0.05 in ad revenue per search). The path to this cost is: model distillation (smaller models for simpler queries: 3B-7B for quick factual answers vs. 34B-70B for complex synthesis), speculative decoding (2-3x throughput improvement), and aggressive quantization (FP8 KV cache reducing per-query memory by 50 percent). Perplexity’s reported $20-per-month Pro subscription implies at least $10-15 of monthly infrastructure cost per heavy user, meaning the company is likely investing in model efficiency to expand margins.

The infrastructure evolution path for AI search engines over the next 12-18 months includes: replacing 7B-34B single-model deployments with 3-5 model tiers (3B for simple factual queries, 8B for standard synthesis, 34B for complex reasoning), implementing agentic search (multiple rounds of retrieval and synthesis for a single query, 3-5x more GPU compute but higher quality), and deploying retrieval-specialized accelerators (Groq LPUs or custom ASICs for embedding computation at 10-100x lower cost per query). The winners in AI search will be those who optimize the GPU-to-quality ratio-delivering answers that are better than traditional search at a cost structure that supports ad-supported or low-margin subscription pricing.

Filed under
AI Search EnginePerplexity AIRAG InfrastructureSearch GPULLM RetrievalReal-Time Web Search