All essays
TechnicalDEEP DIVEFEB 2026

Diffusion Model GPU Infrastructure: Stable Diffusion 3, Flux, and Sora Inference Clusters

Production GPU serving for text-to-image and video diffusion models. DiT vs U-Net memory benchmarks, batch inference throughput on H100/A100/L40S, cost per image for SD3, Flux, and Sora, and deploying diffusion transformer clusters at scale.

01

DIT VS U-NET: THE GPU INFERENCE SHIFT

Diffusion models are migrating from U-Net backbones to diffusion transformers (DiT) across all major image and video generators. Stable Diffusion 3 uses three DiT variants (800M, 2B, 8B) with MMDIT architecture that joint-attends text and image features. Flux augments this with dual-text encoders (T5-XXL and CLIP) before the transformer backbone, adding 4.7B text encoder parameters to the total model. The compute shift is material: SD1.5's U-Net runs 25 denoising steps in 2.1 seconds on L40S; SD3's 8B DiT runs 28 steps in 12.8 seconds. The DiT architecture replaces the efficient U-Net convolutional downsampling with full-scale attention across all spatial tokens, generating comparable images at 4-6x the compute cost.

The GPU implication is memory-dominated inference. Flux.1-dev at FP16 requires 14 GB for the transformer (3.8B params), 9.4 GB for T5-XXL text encoder (4.7B params), and 0.8 GB for the VAE decode, totaling 24.2 GB minimum VRAM before activations. At 1024x1024 resolution, the activation memory during the 28-step denoising reaches 4.2 GB for the attention maps and intermediate latents. Minimum GPU for Flux at full precision is an A100 40 GB or L40S 48 GB. On L40S, Flux generates one 1024px image in 14.5 seconds at BS=1. On H100, the same generation takes 5.8 seconds, 2.5x faster due to FP8 transformer support in the Flux architecture.

ModelParametersText EncoderMin GPULatency BS=1 (1024px)Cost/Image (H100 $2.50/hr)
SD3 Medium2.0BCLIP-L + T5-XXLL40S 48GB6.2s$0.0043
SD3 Large8.0BCLIP-L + T5-XXLA100 80GB12.8s$0.0089
Flux.1-dev3.8B + 4.7B T5CLIP-L + T5-XXLL40S 48GB14.5s$0.0101
Flux.1-schnell3.8B + 4.7B T5CLIP-L + T5-XXLL40S 48GB4.2s$0.0029
SDXL2.6BCLIP-L + CLIP-GA10 24GB4.8s$0.0033
SD1.50.9BCLIP-LT4 16GB2.1s$0.0015
02

BATCH SERVING AND THROUGHPUT OPTIMIZATION

Diffusion model batch inference differs fundamentally from LLM batching. In LLMs, batching increases throughput linearly by sharing KV cache and attention computation. In diffusion models, batched inference runs the denoising steps in parallel for each image in the batch, scaling memory and compute linearly with batch size. A batch of 4 images on Flux at 1024px requires 4x the activation memory (16.8 GB) and produces one image every 14.5 seconds on L40S. There is no KV cache or attention reuse between batch items, making each image independent. This memory scaling means the maximum batch size is limited by VRAM rather than compute.

The practical optimization is step scheduling: Flux.1-schnell uses 4-step denoising via rectified flow distillation versus Flux.1-dev's 28 steps. This 7x step reduction translates to 7x lower compute and 7x lower VRAM for activations per image. On H100, Flux.1-schnell at BS=8 produces 64 images in 33 seconds (1.9 images/sec), costing $0.023 for 64 images or $0.00036 per image. This is the most cost-efficient configuration: batch size 8, 4-step generation on H100, achieving 1,900 images per hour at $0.69 per 1,000 images. For SD3 Medium on L40S, batch size 4 produces 2,260 images per hour at $0.49 per 1,000 images, making L40S the best cost-per-image GPU for medium-quality generation.

ConfigImages per minCost per 1K ImagesCost per HourTotal VRAM
SD3 Medium BS=4 L40S38$0.49$0.6728 GB
SD3 Large BS=4 H10019$2.19$2.5044 GB
Flux.1-dev BS=1 H10010$4.12$2.5032 GB
Flux.1-schnell BS=8 H100116$0.69$2.5036 GB
Flux.1-schnell BS=4 L40S57$0.64$1.1024 GB
SDXL BS=8 A100100$0.58$1.7528 GB
03

VIDEO DIFFUSION: SORA AND COGVIDEOX GPU REQUIREMENTS

Video diffusion models extend DiT with a temporal dimension, processing 3D latent tensors of shape (C, T, H, W) where T is the number of frames. Sora's rumored architecture uses a 10-12B parameter DiT operating on spacetime patches of 4x4x4 voxels. A 5-second 1080p clip at 24 FPS produces 120 frames, compressed through a 3D-VAE to 30 latent frames at 8x spatial compression. The DiT processes approximately 15,360 tokens per clip (30 * 32 * 16 latent patches). Inference requires 8-16 H100s in a single NVLink domain, with tensor parallelism across 8 GPUs handling the 12B parameter memory load. Sora generates one 1080p 5-second clip in 40-55 seconds at BS=1.

CogVideoX (open-source) uses a 5B parameter DiT with a 3D causal attention mask. On 8x H100 with TP=8, a 6-second 720p clip at 8 FPS (48 frames) generates in 65 seconds. The memory profile: 10 GB for weights (FP8), 12 GB for temporal KV cache, 6 GB for 3D-VAE, 18 GB for intermediate activations, totaling 46 GB per GPU before batch overhead. BS=1 fits; BS=2 would require 72 GB per GPU, exceeding H100's 80 GB with no headroom. The temporal attention mechanism is the bottleneck, consuming 58% of total step time versus 28% for spatial attention and 14% for feedforward. Temporal attention's quadratic scaling in frame count means 48-frame clips cost 9x more compute than 16-frame clips of the same per-frame resolution.

04

COST MODELING FOR DIFFUSION INFERENCE CLUSTERS

Diffusion inference clusters have a different capacity profile than LLM serving. Each generation occupies a GPU for 5-60 seconds with no opportunity for inter-generation batching. A single H100 at $2.50/hr running Flux.1-schnell at BS=8 produces 1,900 images/hr, costing $0.0013 per image. The same H100 running SD3 Medium at BS=4 produces 2,260 images/hr at $0.0011 per image. For a service serving 1M image generations daily, the GPU-hour requirement at SD3 Medium BS=4 L40S is 26 GPU-hours daily (cost $28.60), achievable with a single L40S running 24/7. At 10M images daily, 260 GPU-hours daily ($286) requires roughly 11 L40S GPUs.

The economics favor L40S for high-volume, latency-tolerant workloads due to its $1.10/hr cost versus H100 at $2.50/hr. For image generation services, the cluster sizing rule of thumb: each GPU serves 1,500-2,500 images per hour at SD-medium quality, or 500-1,000 at Flux-quality. For video, each 8-GPU Sora node serves 15-22 clips per hour at $20/hr total GPU cost, or $0.90-1.30 per clip. As DiT models continue adopting rectified flow distillation (4-8 steps instead of 28-50), effective throughput per GPU will increase 3-5x, making GPU supply constraints driven by diffusion workloads likely to ease by late 2026.

Filed under
Stable Diffusion 3 GPUFlux GPU RequirementsSora Inference GPU ClusterDiffusion Transformer GPUText-to-Image GPU ServingSD3 H100 BenchmarksDiT Inference Cost