THE IMAGE GENERATION INFRASTRUCTURE SPECTRUM
Image generation at scale comprises three model families: denoising diffusion models (SD 3.5, Flux, SDXL) requiring 4-8 GB VRAM and 20-50 steps; autoregressive models (DALL-E 3, Parti) generating image tokens sequentially; and consistency models (SD Turbo, LCM-LoRA) reducing inference to 1-4 steps with 5-10x throughput improvement.
The key metric is images per GPU-hour. SDXL at 1024x1024: 1,200 im/GPU-h on H100. Flux.1-dev: 150 im/GPU-h. SD Turbo: 15,000 im/GPU-h. A deployment at web scale processing 1-10 images per second requires 24-240 H100 GPUs for SDXL or 200-2000 for Flux, with $6,000/day GPU cost at 100 H100s.
FLUX AND SD3.5: THE NEW GENERATION OF TRANSFORMER MODELS
Flux.1-dev (12B, rectified flow transformer) requires 4x H100 with sequence parallelism. Each of 50 steps totals ~25 PFLOPS per image. At 4x H100 FP16: 18-24 seconds per image. FP8 improves to 10-14 seconds with no measurable FID degradation. Flux.1-schnell (4-step) produces acceptable results in 2-3 seconds on 4x H100 or 8-12 seconds on single H100.
SD3.5 is more GPU-efficient: a single H100 generates 1024x1024 in 4-6 seconds (28 steps). The MM-DiT architecture's shared parameterization reduces per-step compute by 30 percent versus Flux's dual-stream design. The T5-XXL encoder (11B) adds 12 GB VRAM overhead; in production it runs as a separate service with KV cache for frequent prompts.
| Model | Params | Res | Steps | im/GPU-h H100 | im/GPU-h FP8 | VRAM |
|---|---|---|---|---|---|---|
| SD 1.5 | 860M | 512x512 | 20 | 12,000 | 18,000 | 2.5 GB |
| SDXL | 2.6B | 1024x1024 | 28 | 1,200 | 2,000 | 5.5 GB |
| SD 3.5 Medium | 2.5B | 1024x1024 | 28 | 900 | 1,500 | 8 GB |
| SD 3.5 Large | 8.7B | 1024x1024 | 28 | 300 | 500 | 18 GB |
| Flux 1-schnell | 12B | 1024x1024 | 4 | 180 | 300 | 4x H100 |
| Flux 1-dev | 12B | 1024x1024 | 50 | 24 | 45 | 4x H100 |
VAE, CONTROLNET, AND LORA: THE AUXILIARY COMPUTE STACK
The VAE decoder accounts for 10-15 percent of total inference latency: 80-150 ms for SDXL on H100. VAE optimization via TensorRT reduces this to 30-50 ms. ControlNet adds 100-300 ms per step (3-12 seconds per image total). The recommended architecture has dedicated ControlNet GPU workers sharing diffusion weights via unified memory.
Image LoRA adapters (typically 5-100 MB each) modify cross-attention key projections. With 10 popular style LoRAs pre-loaded in GPU memory, 80 percent of user requests are covered. Dynamic loading of long-tail LoRAs adds 200-500 ms cold-start latency, similar to FTaaS adapter routing.
BATCH SERVING AND SCHEDULING FOR DIFFUSION MODELS
Diffusion batching uses step-synchronous batching where all images advance through denoising steps together. Optimal batch size for SDXL on H100 is 8-16 at FP16 (75-85 percent utilization). Prompt-batching (same prompt, multiple seeds) achieves 3-4x throughput; resolution-batching achieves 2-3x.
Queue management: a request queuing layer (Redis) groups requests by class for 50-200 ms, forms the largest feasible batch, and sends to GPU. This adds 50-200 ms queue wait but improves throughput by 3-5x. An interactive priority queue handles 10-20 percent of requests with <500 ms queue latency.
TOTAL COST PER IMAGE: GPU, STORAGE, AND NETWORK
GPU compute cost per image at scale: SDXL 1024x1024 = $0.0015-0.0025, Flux.1-dev = $0.05-0.08, SD Turbo = $0.00015-0.00025. At 50 percent utilization (common for bursty traffic), costs increase by 60 percent. Storage and CDN add negligible cost compared to GPU compute.
Break-even for a $10/month subscription at 200 images/user/month: maximum GPU cost is $0.05/image. This makes SDXL and SD Turbo viable, while Flux.1-dev requires higher tiers ($30-50/month). Flux.1-schnell at $0.003-0.005/image leaves healthy margins at $10/month.
| Cost/1M images | SDXL | Flux-schnell | Flux-dev | SD Turbo |
|---|---|---|---|---|
| GPU (80% util) | $2,000 | $4,200 | $62,500 | $200 |
| GPU (50% util) | $3,200 | $6,700 | $100,000 | $320 |
| Storage 30d | $35 | $35 | $35 | $15 |
| CDN Egress | $40 | $40 | $40 | $15 |
| Serving Overhd | $300 | $500 | $3,000 | $100 |
| Total/img opt | $0.0024 | $0.0048 | $0.066 | $0.00033 |
B200: THE IMAGE GENERATION GAME CHANGER
B200 enables batch sizes of 32-64 for SDXL (versus 8-16 on H100), translating to 2.5-3.5x im/GPU-h improvement. SDXL on B200 at FP8 is projected at 4,000-5,000 im/GPU-h versus 1,200 on H100 FP16.
For Flux, B200's 192 GB VRAM allows Flux.1-dev to fit on a single GPU at FP8, eliminating 4x H100 TP overhead that accounts for 30-40 percent of inference time. Flux.1-dev on B200 is projected at 80-120 seconds per image but at 2-3x lower cost per image due to single-GPU footprint.
