DIT VS U-NET: THE GPU INFERENCE SHIFT
Diffusion models are migrating from U-Net backbones to diffusion transformers (DiT) across all major image and video generators. Stable Diffusion 3 uses three DiT variants (800M, 2B, 8B) with MMDIT architecture that joint-attends text and image features. Flux augments this with dual-text encoders (T5-XXL and CLIP) before the transformer backbone, adding 4.7B text encoder parameters to the total model. The compute shift is material: SD1.5's U-Net runs 25 denoising steps in 2.1 seconds on L40S; SD3's 8B DiT runs 28 steps in 12.8 seconds. The DiT architecture replaces the efficient U-Net convolutional downsampling with full-scale attention across all spatial tokens, generating comparable images at 4-6x the compute cost.
The GPU implication is memory-dominated inference. Flux.1-dev at FP16 requires 14 GB for the transformer (3.8B params), 9.4 GB for T5-XXL text encoder (4.7B params), and 0.8 GB for the VAE decode, totaling 24.2 GB minimum VRAM before activations. At 1024x1024 resolution, the activation memory during the 28-step denoising reaches 4.2 GB for the attention maps and intermediate latents. Minimum GPU for Flux at full precision is an A100 40 GB or L40S 48 GB. On L40S, Flux generates one 1024px image in 14.5 seconds at BS=1. On H100, the same generation takes 5.8 seconds, 2.5x faster due to FP8 transformer support in the Flux architecture.
| Model | Parameters | Text Encoder | Min GPU | Latency BS=1 (1024px) | Cost/Image (H100 $2.50/hr) |
|---|---|---|---|---|---|
| SD3 Medium | 2.0B | CLIP-L + T5-XXL | L40S 48GB | 6.2s | $0.0043 |
| SD3 Large | 8.0B | CLIP-L + T5-XXL | A100 80GB | 12.8s | $0.0089 |
| Flux.1-dev | 3.8B + 4.7B T5 | CLIP-L + T5-XXL | L40S 48GB | 14.5s | $0.0101 |
| Flux.1-schnell | 3.8B + 4.7B T5 | CLIP-L + T5-XXL | L40S 48GB | 4.2s | $0.0029 |
| SDXL | 2.6B | CLIP-L + CLIP-G | A10 24GB | 4.8s | $0.0033 |
| SD1.5 | 0.9B | CLIP-L | T4 16GB | 2.1s | $0.0015 |
BATCH SERVING AND THROUGHPUT OPTIMIZATION
Diffusion model batch inference differs fundamentally from LLM batching. In LLMs, batching increases throughput linearly by sharing KV cache and attention computation. In diffusion models, batched inference runs the denoising steps in parallel for each image in the batch, scaling memory and compute linearly with batch size. A batch of 4 images on Flux at 1024px requires 4x the activation memory (16.8 GB) and produces one image every 14.5 seconds on L40S. There is no KV cache or attention reuse between batch items, making each image independent. This memory scaling means the maximum batch size is limited by VRAM rather than compute.
The practical optimization is step scheduling: Flux.1-schnell uses 4-step denoising via rectified flow distillation versus Flux.1-dev's 28 steps. This 7x step reduction translates to 7x lower compute and 7x lower VRAM for activations per image. On H100, Flux.1-schnell at BS=8 produces 64 images in 33 seconds (1.9 images/sec), costing $0.023 for 64 images or $0.00036 per image. This is the most cost-efficient configuration: batch size 8, 4-step generation on H100, achieving 1,900 images per hour at $0.69 per 1,000 images. For SD3 Medium on L40S, batch size 4 produces 2,260 images per hour at $0.49 per 1,000 images, making L40S the best cost-per-image GPU for medium-quality generation.
| Config | Images per min | Cost per 1K Images | Cost per Hour | Total VRAM |
|---|---|---|---|---|
| SD3 Medium BS=4 L40S | 38 | $0.49 | $0.67 | 28 GB |
| SD3 Large BS=4 H100 | 19 | $2.19 | $2.50 | 44 GB |
| Flux.1-dev BS=1 H100 | 10 | $4.12 | $2.50 | 32 GB |
| Flux.1-schnell BS=8 H100 | 116 | $0.69 | $2.50 | 36 GB |
| Flux.1-schnell BS=4 L40S | 57 | $0.64 | $1.10 | 24 GB |
| SDXL BS=8 A100 | 100 | $0.58 | $1.75 | 28 GB |
VIDEO DIFFUSION: SORA AND COGVIDEOX GPU REQUIREMENTS
Video diffusion models extend DiT with a temporal dimension, processing 3D latent tensors of shape (C, T, H, W) where T is the number of frames. Sora's rumored architecture uses a 10-12B parameter DiT operating on spacetime patches of 4x4x4 voxels. A 5-second 1080p clip at 24 FPS produces 120 frames, compressed through a 3D-VAE to 30 latent frames at 8x spatial compression. The DiT processes approximately 15,360 tokens per clip (30 * 32 * 16 latent patches). Inference requires 8-16 H100s in a single NVLink domain, with tensor parallelism across 8 GPUs handling the 12B parameter memory load. Sora generates one 1080p 5-second clip in 40-55 seconds at BS=1.
CogVideoX (open-source) uses a 5B parameter DiT with a 3D causal attention mask. On 8x H100 with TP=8, a 6-second 720p clip at 8 FPS (48 frames) generates in 65 seconds. The memory profile: 10 GB for weights (FP8), 12 GB for temporal KV cache, 6 GB for 3D-VAE, 18 GB for intermediate activations, totaling 46 GB per GPU before batch overhead. BS=1 fits; BS=2 would require 72 GB per GPU, exceeding H100's 80 GB with no headroom. The temporal attention mechanism is the bottleneck, consuming 58% of total step time versus 28% for spatial attention and 14% for feedforward. Temporal attention's quadratic scaling in frame count means 48-frame clips cost 9x more compute than 16-frame clips of the same per-frame resolution.
COST MODELING FOR DIFFUSION INFERENCE CLUSTERS
Diffusion inference clusters have a different capacity profile than LLM serving. Each generation occupies a GPU for 5-60 seconds with no opportunity for inter-generation batching. A single H100 at $2.50/hr running Flux.1-schnell at BS=8 produces 1,900 images/hr, costing $0.0013 per image. The same H100 running SD3 Medium at BS=4 produces 2,260 images/hr at $0.0011 per image. For a service serving 1M image generations daily, the GPU-hour requirement at SD3 Medium BS=4 L40S is 26 GPU-hours daily (cost $28.60), achievable with a single L40S running 24/7. At 10M images daily, 260 GPU-hours daily ($286) requires roughly 11 L40S GPUs.
The economics favor L40S for high-volume, latency-tolerant workloads due to its $1.10/hr cost versus H100 at $2.50/hr. For image generation services, the cluster sizing rule of thumb: each GPU serves 1,500-2,500 images per hour at SD-medium quality, or 500-1,000 at Flux-quality. For video, each 8-GPU Sora node serves 15-22 clips per hour at $20/hr total GPU cost, or $0.90-1.30 per clip. As DiT models continue adopting rectified flow distillation (4-8 steps instead of 28-50), effective throughput per GPU will increase 3-5x, making GPU supply constraints driven by diffusion workloads likely to ease by late 2026.
