All essays
TechnicalDEEP DIVEFEB 2026

AI Image Generation at Scale: Stable Diffusion, Flux, and DALL-E Inference Infrastructure

Production-scale image generation serving infrastructure for Stable Diffusion, Flux, DALL-E, and Midjourney. Latency benchmarks, batch serving patterns, LoRA adapters, VAE optimization, and total cost per image across GPU configurations.

01

THE IMAGE GENERATION INFRASTRUCTURE SPECTRUM

Image generation at scale comprises three model families: denoising diffusion models (SD 3.5, Flux, SDXL) requiring 4-8 GB VRAM and 20-50 steps; autoregressive models (DALL-E 3, Parti) generating image tokens sequentially; and consistency models (SD Turbo, LCM-LoRA) reducing inference to 1-4 steps with 5-10x throughput improvement.

The key metric is images per GPU-hour. SDXL at 1024x1024: 1,200 im/GPU-h on H100. Flux.1-dev: 150 im/GPU-h. SD Turbo: 15,000 im/GPU-h. A deployment at web scale processing 1-10 images per second requires 24-240 H100 GPUs for SDXL or 200-2000 for Flux, with $6,000/day GPU cost at 100 H100s.

02

FLUX AND SD3.5: THE NEW GENERATION OF TRANSFORMER MODELS

Flux.1-dev (12B, rectified flow transformer) requires 4x H100 with sequence parallelism. Each of 50 steps totals ~25 PFLOPS per image. At 4x H100 FP16: 18-24 seconds per image. FP8 improves to 10-14 seconds with no measurable FID degradation. Flux.1-schnell (4-step) produces acceptable results in 2-3 seconds on 4x H100 or 8-12 seconds on single H100.

SD3.5 is more GPU-efficient: a single H100 generates 1024x1024 in 4-6 seconds (28 steps). The MM-DiT architecture's shared parameterization reduces per-step compute by 30 percent versus Flux's dual-stream design. The T5-XXL encoder (11B) adds 12 GB VRAM overhead; in production it runs as a separate service with KV cache for frequent prompts.

ModelParamsResStepsim/GPU-h H100im/GPU-h FP8VRAM
SD 1.5860M512x5122012,00018,0002.5 GB
SDXL2.6B1024x1024281,2002,0005.5 GB
SD 3.5 Medium2.5B1024x1024289001,5008 GB
SD 3.5 Large8.7B1024x10242830050018 GB
Flux 1-schnell12B1024x102441803004x H100
Flux 1-dev12B1024x10245024454x H100
03

VAE, CONTROLNET, AND LORA: THE AUXILIARY COMPUTE STACK

The VAE decoder accounts for 10-15 percent of total inference latency: 80-150 ms for SDXL on H100. VAE optimization via TensorRT reduces this to 30-50 ms. ControlNet adds 100-300 ms per step (3-12 seconds per image total). The recommended architecture has dedicated ControlNet GPU workers sharing diffusion weights via unified memory.

Image LoRA adapters (typically 5-100 MB each) modify cross-attention key projections. With 10 popular style LoRAs pre-loaded in GPU memory, 80 percent of user requests are covered. Dynamic loading of long-tail LoRAs adds 200-500 ms cold-start latency, similar to FTaaS adapter routing.

04

BATCH SERVING AND SCHEDULING FOR DIFFUSION MODELS

Diffusion batching uses step-synchronous batching where all images advance through denoising steps together. Optimal batch size for SDXL on H100 is 8-16 at FP16 (75-85 percent utilization). Prompt-batching (same prompt, multiple seeds) achieves 3-4x throughput; resolution-batching achieves 2-3x.

Queue management: a request queuing layer (Redis) groups requests by class for 50-200 ms, forms the largest feasible batch, and sends to GPU. This adds 50-200 ms queue wait but improves throughput by 3-5x. An interactive priority queue handles 10-20 percent of requests with <500 ms queue latency.

05

TOTAL COST PER IMAGE: GPU, STORAGE, AND NETWORK

GPU compute cost per image at scale: SDXL 1024x1024 = $0.0015-0.0025, Flux.1-dev = $0.05-0.08, SD Turbo = $0.00015-0.00025. At 50 percent utilization (common for bursty traffic), costs increase by 60 percent. Storage and CDN add negligible cost compared to GPU compute.

Break-even for a $10/month subscription at 200 images/user/month: maximum GPU cost is $0.05/image. This makes SDXL and SD Turbo viable, while Flux.1-dev requires higher tiers ($30-50/month). Flux.1-schnell at $0.003-0.005/image leaves healthy margins at $10/month.

Cost/1M imagesSDXLFlux-schnellFlux-devSD Turbo
GPU (80% util)$2,000$4,200$62,500$200
GPU (50% util)$3,200$6,700$100,000$320
Storage 30d$35$35$35$15
CDN Egress$40$40$40$15
Serving Overhd$300$500$3,000$100
Total/img opt$0.0024$0.0048$0.066$0.00033
06

B200: THE IMAGE GENERATION GAME CHANGER

B200 enables batch sizes of 32-64 for SDXL (versus 8-16 on H100), translating to 2.5-3.5x im/GPU-h improvement. SDXL on B200 at FP8 is projected at 4,000-5,000 im/GPU-h versus 1,200 on H100 FP16.

For Flux, B200's 192 GB VRAM allows Flux.1-dev to fit on a single GPU at FP8, eliminating 4x H100 TP overhead that accounts for 30-40 percent of inference time. Flux.1-dev on B200 is projected at 80-120 seconds per image but at 2-3x lower cost per image due to single-GPU footprint.

Filed under
Stable Diffusion GPUFlux Model ServingDALL-E InfrastructureImage Generation ScaleSD LoRA ServingDiffusion Model GPUAI Image Production