All essays
TechnicalDEEP DIVEFEB 2026

Multimodal AI Infrastructure: Deploying CLIP, LLaVA, and GPT-4V-Style Vision-Language Models

Production serving patterns for vision-language models: CLIP embedding, LLaVA inference, multimodal RAG, and GPT-4V-style architectures on GPU clusters. Latency benchmarks, cross-modal routing, and cost per query for VLM serving at scale.

01

THE VISION-LANGUAGE MODEL ARCHITECTURE SPECTRUM

Multimodal models divide into three serving profiles with fundamentally different GPU requirements. Dual-encoder models like CLIP and SigLIP compute separate image and text embeddings and compare them in a shared latent space; these are the cheapest to serve and dominate retrieval and classification use cases. Fusion-encoder models like FLAVA combine modalities at the encoder level. Decoder-based VLMs like LLaVA, Qwen-VL, and GPT-4V connect a vision encoder to a large language model, enabling open-ended visual question answering.

The compute cost scales dramatically: a CLIP ViT-L/14 embedding query costs 15-25 ms on an H100 with 3 GB VRAM. LLaVA-1.6 34B inference costs 800-1500 ms on 2-4 H100 GPUs with 140 GB VRAM total. This 100-1000x cost difference drives the architecture of production multimodal systems, which typically use CLIP as a first-pass filter before routing to a VLM for deep analysis.

02

CLIP AND SIGLIP: HIGH-THROUGHPUT EMBEDDING SERVING

CLIP and its successor SigLIP serve as the backbone for most multimodal retrieval systems. The models are dual ViT-based encoders: image embeddings from a ViT-L/14 (428M parameters) and text embeddings from a transformer (123M parameters). Serving requires 2.8 GB VRAM for the image encoder and 0.8 GB for the text encoder at FP16. Throughput on an H100 reaches 4,800 images per second for embedding generation at batch size 256.

CLIP-based multimodal retrieval at 100M+ image scale is now standard infrastructure. A typical deployment stores CLIP embeddings in a vector database (Milvus, Qdrant) with HNSW indexing. The retrieval pipeline: encode query text via CLIP text encoder, perform ANN search in under 50 ms, then re-rank top-100 results with a cross-encoder model on GPU.

SigLIP provides a 2-3 percent recall improvement over CLIP with identical compute cost. For cost-sensitive serving at very high throughput (>100K QPS), INT8 quantization reduces VRAM to 1.5 GB per model and increases throughput by 2.1x with less than 1.0 percent mAP degradation.

MetricCLIP ViT-B/32CLIP ViT-L/14SigLIP ViT-L/16EVA-CLIP-18B
Parameters151M428M + 123M430M + 130M18B (MoE)
Embed Dim5127687681024
H100 Throughput12,500 img/s4,800 img/s4,400 img/s240 img/s
VRAM per Model0.6 GB2.8 GB3.1 GB36 GB
INet Zero-shot63.2%75.5%77.8%82.1%
Cost/1M Embed$0.04$0.10$0.11$2.10
Rec. GPUAny 8 GB+L4 / A10L4 / A101-2 H100
03

LLAVA AND OPEN-SOURCE VLM DEPLOYMENT

LLaVA-1.6 uses a CLIP ViT-L vision encoder, a two-layer MLP projection, and a large language model backbone. The 34B variant requires 4 H100 GPUs with tensor parallelism for real-time inference (batch size 1, 1000-1500 ms first-token latency). The 7B variant fits on a single H100 with 350-500 ms first-token latency.

The production serving architecture uses a vision encoder pool separated from the LLM pool. The vision encoder runs on 1-2 L40S GPUs at 2,000-3,000 images per second. The LLM pool receives pre-computed visual embeddings, avoiding redundant encoder computation for multi-turn conversations. vLLM's prefix caching is particularly valuable for caching the visual token KV cache across conversation turns.

VLMLLaVA 7BLLaVA 34BQwen-VL 7BQVQ-72B
GPUs1 H1004 H1001 H1004-8 B200
Vision EncoderCLIP ViT-L/14CLIP ViT-L/14SigLIP L/16InternViT-6B
Visual Tokens576 (336px)576 (336px)256 (224px)1,444 (448px)
First Token Lat350-500 ms1,000-1,500 ms280-400 ms2,500-4,000 ms
Output Throughput85 tok/s40 tok/s95 tok/s18 tok/s
Cost per QA$0.001-0.003$0.008-0.015$0.001-0.002$0.04-0.10
04

MULTIMODAL RAG: CROSS-MODAL RETRIEVAL AND GENERATION

Multimodal RAG requires three components: a multimodal embedding model to index images with text descriptions in the same vector space, a dual-retrieval pipeline that fetches top-k images and top-k text chunks, and a VLM to synthesize from the retrieved multimodal context.

The GPU cost is dominated by the VLM generation stage. A typical query lifecycle: text embedding (1-2 ms, $0.00001), image similarity search (20-50 ms, $0.00005), top-4 re-ranking (80-200 ms, $0.0005-0.002), and VLM answer generation (800-4000 ms, $0.002-0.05). The VLM stage accounts for 80-95 percent of total latency and cost.

A pragmatic infrastructure uses a three-tier model cascade: CLIP + HNSW for initial retrieval (tier 1, 20 ms), a 7B VLM for straightforward answering (tier 2, 500 ms, 85 percent of queries), and a 34B VLM for complex reasoning (tier 3, 1500 ms, 15 percent). This cascade reduces GPU cost by 60-70 percent versus calling a 34B VLM on every query.

05

CROSS-MODAL ROUTING AND MODEL GATEWAYS

Production multimodal systems embed a routing layer that inspects each query and dispatches it to the appropriate model. The router is typically a lightweight classifier (DeBERTa-v3 200M) running on CPU or small GPU in 2-5 ms. It evaluates query characteristics: image resolution, question complexity, latency budget, and cost cap.

Model gateways built on Envoy with custom gRPC extensions manage request routing, authentication, rate limiting, and cost accounting. Each model variant registers its latency profile, cost per query, and supported resolutions. The gateway implements cost-aware least-loaded balancing, preferring cheaper models when requirements are satisfied.

06

FUTURE DIRECTIONS: B200 AND LARGER VISION MODELS

The B200's 192 GB VRAM enables full-parameter serving of 34B-class VLMs on a single GPU, eliminating the tensor parallelism overhead that adds 15-25 percent latency. Single-GPU 34B VLM serving on B200 is projected at 600-800 ms first-token latency, a 40-50 percent improvement over 4x H100.

Emerging native multimodal architectures (DeepSeek-VL2, CogVLM2) embed vision tokens directly into the LLM input space without a separate encoder. This simplifies GPU serving to standard LLM inference with a prefill phase accepting interleaved image and text tokens, eliminating the need for separate encoder and decoder GPU pools.

Filed under
Multimodal AICLIP Embedding ServingLLaVA InferenceVision Language Model GPUVLM Serving ArchitectureMultimodal RAGGPT-4V Infrastructure