THE VIDEO GENERATION MODEL LANDSCAPE
Video generation models have converged on the diffusion transformer (DiT) architecture, replacing the earlier U-Net backbones that dominated image generation. OpenAI's Sora, Runway's Gen-3 Alpha, Pika 2.0, Kuaishou's Kling, and the open-source CogVideoX all use variants of the 3D-VAE + DiT design, where spatial and temporal attention blocks process a latent video tensor. The fundamental difference from image models is the additional temporal dimension: a 5-second clip at 24 FPS produces 120 frames, expanding the latent tensor by roughly 120x compared to a single image before compression through the 3D-VAE. This expansion drives the VRAM requirements 8-12x higher than equivalent-resolution image generation.
The inference compute profile differs markedly from LLMs. Video models process a single clip in 20-120 seconds on an H100, versus milliseconds per token for autoregressive text models. The workload is throughput-bound at the batch level rather than latency-bound: users expect results in 30-180 seconds, not sub-second. This makes video generation a natural fit for large-batch inference on H100 SXM clusters with NVLink interconnects, where the 3.35 TB/s HBM3 bandwidth directly accelerates the tensor operations in each attention layer.
SORA: OPENAI'S VIDEO DIT AT SCALE
Sora uses a denoising diffusion transformer with approximately 10-12 billion parameters, processing 1080p video in 20-second clips through what OpenAI describes as a spacetime latent patch representation. The model operates on a compressed latent grid where each patch encodes a 4x4x4 spacetime region, reducing the effective sequence length while preserving temporal coherence. Inference requires 8-16 H100 GPUs in a single NVLink domain for the full 1080p 20-second generation in under 60 seconds, with the 10B-level parameter count distributed across tensor parallelism.
The current deployment pattern at Microsoft Azure assigns Sora inference to a dedicated H100-NVL-32 partition (4x DGX H100 nodes, 32 GPUs). This allows simultaneous 1080p 20-second clips to be processed with a batch size of 4-8, achieving approximately 2-3 clips per minute per partition. The critical bottleneck is the temporal attention mechanism, which computes attention across all frames and consumes 60-70 percent of total FLOPs per step. OpenAI uses temporal attention splitting with T5-style relative position biases in the RoPE formulation to keep memory within H100's 80 GB HBM3.
| Configuration | Sora 1080p 20s | Runway Gen-3 1080p 10s | Pika 2.0 720p 5s | Kling 1080p 5s |
|---|---|---|---|---|
| Min GPUs | 8 H100 | 4 H100 | 1 H100 | 4 H100 |
| VRAM Required | 640 GB (8x80) | 320 GB (4x80) | 80 GB (1x80) | 320 GB (4x80) |
| Latency per Clip | 45-60s | 25-40s | 12-20s | 20-35s |
| Parallelism | TP=8, DP=1-4 | TP=4, DP=2-4 | Single GPU | TP=4, DP=1-2 |
| Peak Mem per GPU | 72 GB | 68 GB | 48 GB | 64 GB |
| Cost per Clip | $0.35-0.55 | $0.12-0.25 | $0.04-0.08 | $0.10-0.20 |
RUNWAY GEN-3 ALPHA: PRODUCTION-OPTIMIZED INFERENCE
Runway Gen-3 Alpha uses a mixture-of-experts DiT architecture with approximately 7 billion active parameters and 15 billion total parameters across 8 experts. The MoE design allows the model to activate only the relevant expert pathways per timestep, reducing effective FLOPs by 35-40 percent compared to a dense 15B model while maintaining quality. Runway deploys Gen-3 on a dedicated cluster of H100 GPUs with a custom inference engine built on TensorRT-LLM, achieving 25-40 seconds per 1080p 10-second clip versus 45-60 seconds for Sora on the same hardware.
The performance advantage comes from three optimizations: quantization of the VAE decoder to FP8 with minimal quality loss (0.03 FVD increase on UCF-101); temporal attention caching that reuses keys and values from the first half of the denoising trajectory for the second half, reducing attention FLOPs by 30 percent; and a CFG parallelization scheme that computes conditional and unconditional scores in a single forward pass with dual output heads.
PIKA 2.0 AND KLING: SINGLE-GPU AND SCALABLE APPROACHES
Pika 2.0 targets short-form 3-5 second clips at 720p resolution with a 3B parameter DiT that fits on a single H100 GPU with FP16 weights and full attention. The model uses a 3D-VAE with 8x spatial and 4x temporal compression, producing a latent tensor of shape [4, 128, 128, 20] for a 5-second clip at 24 FPS. At 20 denoising steps with DDIM sampling, inference completes in 12-20 seconds on one H100 and 8-12 seconds with FP8 optimization. The single-GPU deployment model makes Pika the most cost-efficient option for high-volume, short-form video generation at approximately $0.03-0.06 per clip.
Kling uses a 5B parameter diffusion model combining a 3D-VAE with a latent transformer for 1080p generation up to 5 seconds. Kling requires 4 H100 GPUs with tensor parallelism but achieves the lowest latency-to-quality ratio in the 5-second category. Kuaishou's production deployment runs on a 512-GPU H800 cluster using a custom scheduling layer that batches requests by duration and resolution to optimize GPU utilization.
B200 AND THE NEXT INFRASTRUCTURE WAVE
NVIDIA's B200 GPU with 192 GB HBM3e memory and 8 TB/s memory bandwidth fundamentally changes video model serving economics. The primary constraint on current H100-based video inference is the 80 GB VRAM ceiling, which forces multi-GPU tensor parallelism even for moderate-sized models. The B200's 192 GB capacity allows the full Sora-class models to fit on 2 GPUs instead of 8, reducing inference latency by 40-50 percent through lower communication overhead and enabling DP-parallel serving of multiple clips on a single node.
For Pika-class models (3B and below), B200 enables a batch size of 64-128 on a single GPU, approximately 4x the H100's throughput. At equivalent total cost of ownership, early projections from hyperscaler partners suggest B200 will deliver 3.5-4.2x more video clips per dollar than H100 for models under 7B parameters.
| Model Class | H100 (80 GB) | H200 (141 GB) | B200 (192 GB) | B200 vs H100 Lift |
|---|---|---|---|---|
| 3B (Pika class) | 1 GPU / 16-32 b | 1 GPU / 32-64 b | 1 GPU / 64-128 b | 3.5-4.2x thru |
| 7B (Runway class) | 4 GPUs / 8-16 b | 2 GPUs / 16-32 b | 2 GPUs / 32-64 b | 3.0-3.8x thru |
| 10B+ (Sora class) | 8 GPUs / 4-8 b | 4 GPUs / 8-16 b | 2-4 GPUs / 16-32 b | 3.5-4.5x thru |
| Reserved Cost/Clip | $0.18-0.35 | $0.12-0.22 | $0.06-0.12 | 3x reduction |
| P99 Lat 10s clip | 35-40s | 25-30s | 15-20s | 50% reduction |
PRODUCTION SERVING STRATEGIES FOR VIDEO MODELS
Video model serving in production requires a fundamentally different architecture from LLM serving. The key insight is that video generation is a persistent GPU workload (20-120 seconds per request) rather than a token-generation loop with variable length. GPU allocation must be reservation-based rather than demand-based: pre-allocate GPU partitions for video generation and queue requests against them, rather than using the dynamic batching typical of text inference.
A recommended deployment pattern uses a two-tier architecture: a CPU-based request queue with priority classes (real-time <30s, batch <120s, bulk >120s) fronting a GPU worker pool where each worker holds a 1-8 GPU partition loaded with a specific model variant. This architecture achieves 70-85 percent GPU utilization on the worker pool versus 30-50 percent with ad hoc assignment. For teams building on open-source models like CogVideoX, vLLM's experimental diffusion model support and the Hugging Face Diffusers library with pipeline parallelism provide a starting point.
