All essays
TechnicalDEEP DIVEFEB 2026

Multi-Modal Model GPU Serving: Images, Text, and Audio Joint Inference Infrastructure

GPU serving infrastructure for multi-modal models like GPT-4o, Gemini 2.5, and open-source VLM/audio models. Modality-specific GPU compute profiles, joint inference latency budgets, batch scheduling across modalities, and cost per multi-modal query on H100 and B200 clusters.

01

MODALITY-SPECIFIC GPU COMPUTE PROFILES

Multi-modal models encode each modality through a specialized encoder before joint processing in a shared transformer backbone. The compute and memory profiles differ by 50-100x between modalities, creating asymmetric GPU requirements within a single request. Image encoding through SigLIP-SO400M consumes 1.5 ms on H100 for a 448px image with 0.8 GB VRAM. Audio encoding through Whisper-Large-v3 (1.5B params) processes a 30-second audio clip in 45 ms on H100 with 4.1 GB VRAM. Text tokenization and embedding is negligible at 0.1 ms for 1K tokens. The joint transformer (2-10B params) then processes all modalities together through self-attention, where the image tokens dominate the sequence length: a 224px ViT outputs 196 tokens, a 30-second Whisper output is 150 tokens, and a 1K-token text prompt is 1,000 tokens. The joint attention matrix grows quadratically in _all_ tokens, making the text input the dominant compute cost despite its cheaper encoding.

For GPT-4o class models with 8B to 100B total parameters, the VRAM profile cumulates: SigLIP encoder 0.8 GB + Whisper encoder 4.1 GB + joint transformer 16-200 GB (FP8) + KV cache. At 100B parameters and FP8, the joint transformer requires 100 GB + 16 GB KV cache at 8K for text tokens + image token cache (2 GB for 512 image tokens). Total: ~120 GB, requiring 2 H100 GPUs with TP=2. The bottleneck is the joint decoder's cross-attention between text and image tokens: a 128-image-token * 4,096-text-token cross-attention layer produces 500K attention scores per head, dominating the compute (62% of total FLOPs) and memory (48% of activation storage).

ModalityEncoder ModelEncoder ParamsEncode Time (H100)Tokens GeneratedVRAM for Encoder
TextEmbedding0.1B0.1 ms / 1K tok1,000 (1K prompt)0.1 GB
Image 448pxSigLIP-SO400M0.4B1.5 ms1960.8 GB
Image 1024pxSigLIP-ViT-G/141.8B4.2 ms5123.6 GB
Audio 30sWhisper-Large-v31.5B45 ms1504.1 GB
Video 5s (8fps)VideoMAE-L3.0B120 ms1,02412.0 GB
02

JOINT INFERENCE LATENCY BREAKDOWN

A multi-modal query follows a pipeline: input encoding (variable by modality), joint prefill (process all tokens together), and autoregressive decode (generate text). For an image+text query with GPT-4o equivalent on 2x H100 TP=2, the end-to-end latency at BS=1 breaks down as: image encode 1.5 ms + audio encode (if present) 45 ms + joint prefill of 1,196 tokens (1K text + 196 image tokens) 380 ms + decode 24 ms per output token. TTFT is 426 ms for image-only, 471 ms with audio. TPOT is 24 ms/token. Total for a 512-token response with image input: 12.7 seconds.

The joint prefill dominates TTFT because the attention computation scales quadratically with total tokens. With 1,196 tokens, the prefill attention produces 1.4M attention scores per layer. At 32 layers on 2x H100 TP=2, this takes 380 ms with FlashAttention-3. For multi-image inputs (10 images, 1,960 image tokens + 1K text = 2,960 tokens), prefill rises to 2,300 ms. Audio adds 150 tokens for 30 seconds, raising prefill to 3,100 ms for text+image+audio queries. The decode phase is faster than standalone text because the KV cache includes image and audio token states: 11,520 tokens of KV cache (1K text + 196 images + 150 audio) at 32 layers = 176 GB, requiring the full 160 GB of 2x H100 and limiting batch size to 1.

03

BATCHING AND SCHEDULING ACROSS MODALITIES

Multi-modal batching is more complex than text-only because modalities have different sequence lengths and compute profiles. A naive batch of 4 requests, each with 1 image (196 image tokens each) and varying text prompts (500-2,000 tokens), creates total token counts of 696-2,196 tokens per request. FlashAttention handles variable-length sequences by padding to the longest sequence: 2,196 * 4 = 8,784 total tokens in one batch, requiring 8,784^2/2 attention scores = 38.6M scores per layer, using 72 GB of VRAM just for attention maps. This exceeds a single H100's 80 GB with the model loaded. The practical solution is dynamic batching with modality-aware scheduling: group requests by total token count (image tokens + text tokens), cap at 4,000 total tokens per batch.

A production scheduler uses three queues: image-first (fast encode, high prefill), audio-first (slow encode, moderate prefill), and text-only (no encode, fast prefill). On 8x H100 with TP=4, 2 replicas, each replica processes batches from the combined queue with modality-aware packing. Throughput varies by modality mix: text-only achieves 2,100 tok/s/replica, image+text achieves 680 tok/s/replica, image+text+audio achieves 320 tok/s/replica. The cost per query ranges from $0.001 for text-only to $0.008 for image+text to $0.015 for image+text+audio at H100 $2.50/hr rates. For a workload with 70% text-only, 20% image+text, and 10% full multi-modal, the blended cost is $0.003 per query.

Modality MixQueries/sec per 8xH100Avg Latency P50Cost per QueryVRAM per Replica
Text only8400.8s$0.00122 GB
Image + text (1 img)1402.8s$0.00436 GB
Image + text (5 img)427.5s$0.01252 GB
Audio + text (30s)854.2s$0.00642 GB
Image + audio + text329.5s$0.01564 GB
Video 5s + text822s$0.04274 GB
04

OPEN-SOURCE MULTI-MODAL MODEL GPU REQUIREMENTS

Open-source vision-language models (VLM) offer varying GPU footprints. LLaVA-NeXT 7B (LLaMA 2 7B + CLIP ViT-L): fits on a single L40S (6 GB weights + 4 GB image encoder + 4 GB KV cache at 8K = 14 GB), serving 24 tok/s with 0.5s TTFT for single-image + 256 token text. Qwen-VL 9B: fits on single L40S (18 GB weights + 0.8 GB encoder + 4 GB KV = 22.8 GB), serving 18 tok/s. InternVL2 40B MoE: requires 2 H100s (80 GB weights + 1.6 GB encoder + 8 GB KV = 89.6 GB), TP=2, serving 32 tok/s. For audio-visual multi-modal, the only open-source option is AudioLLaMA (12B total: 7B LLM + 1.5B Whisper + 3.5B image encoder), requiring 2x H100 with TP=2, serving 14 tok/s with audio+image+text inputs.

The cost to serve a production multi-modal pipeline depends on the model size and modality count. LLaVA-NeXT on L40S: $1.10/hr / 24 tok/s = $0.013 per 1K output tokens (excluding image encode). InternVL2 on 2x H100: $5.00/hr / 32 tok/s = $0.043 per 1K tokens. AudioLLaMA on 2x H100: $5.00/hr / 14 tok/s = $0.099 per 1K tokens. For context, GPT-4o API pricing is approximately $0.01 per image input + $0.015 per 1K output tokens. Open-source multi-modal serving on H100 clusters achieves comparable cost per query at 4-8x scale, while LLaVA-NeXT on L40S undercuts API pricing by 2-3x when latency requirements are relaxed.

Filed under
Multi-Modal GPU InferenceVLM GPU RequirementsGPT-4o GPU ServingGemini 2.5 GPU ClusterVision Language Model GPUAudio Model GPU InferenceJoint Modality Serving