All essays
TechnicalDEEP DIVEFEB 2026

AI Video Generation Infrastructure: GPU Requirements for Sora, Runway, and Pika

Deployment guide for Sora, Runway Gen-3, Pika 2.0, and Kling on H100 and B200 GPU clusters. VRAM benchmarks, latency targets, batch inference strategies, and cost modeling for production video AI workloads at scale.

01

THE VIDEO GENERATION MODEL LANDSCAPE

Video generation models have converged on the diffusion transformer (DiT) architecture, replacing the earlier U-Net backbones that dominated image generation. OpenAI's Sora, Runway's Gen-3 Alpha, Pika 2.0, Kuaishou's Kling, and the open-source CogVideoX all use variants of the 3D-VAE + DiT design, where spatial and temporal attention blocks process a latent video tensor. The fundamental difference from image models is the additional temporal dimension: a 5-second clip at 24 FPS produces 120 frames, expanding the latent tensor by roughly 120x compared to a single image before compression through the 3D-VAE. This expansion drives the VRAM requirements 8-12x higher than equivalent-resolution image generation.

The inference compute profile differs markedly from LLMs. Video models process a single clip in 20-120 seconds on an H100, versus milliseconds per token for autoregressive text models. The workload is throughput-bound at the batch level rather than latency-bound: users expect results in 30-180 seconds, not sub-second. This makes video generation a natural fit for large-batch inference on H100 SXM clusters with NVLink interconnects, where the 3.35 TB/s HBM3 bandwidth directly accelerates the tensor operations in each attention layer.

02

SORA: OPENAI'S VIDEO DIT AT SCALE

Sora uses a denoising diffusion transformer with approximately 10-12 billion parameters, processing 1080p video in 20-second clips through what OpenAI describes as a spacetime latent patch representation. The model operates on a compressed latent grid where each patch encodes a 4x4x4 spacetime region, reducing the effective sequence length while preserving temporal coherence. Inference requires 8-16 H100 GPUs in a single NVLink domain for the full 1080p 20-second generation in under 60 seconds, with the 10B-level parameter count distributed across tensor parallelism.

The current deployment pattern at Microsoft Azure assigns Sora inference to a dedicated H100-NVL-32 partition (4x DGX H100 nodes, 32 GPUs). This allows simultaneous 1080p 20-second clips to be processed with a batch size of 4-8, achieving approximately 2-3 clips per minute per partition. The critical bottleneck is the temporal attention mechanism, which computes attention across all frames and consumes 60-70 percent of total FLOPs per step. OpenAI uses temporal attention splitting with T5-style relative position biases in the RoPE formulation to keep memory within H100's 80 GB HBM3.

ConfigurationSora 1080p 20sRunway Gen-3 1080p 10sPika 2.0 720p 5sKling 1080p 5s
Min GPUs8 H1004 H1001 H1004 H100
VRAM Required640 GB (8x80)320 GB (4x80)80 GB (1x80)320 GB (4x80)
Latency per Clip45-60s25-40s12-20s20-35s
ParallelismTP=8, DP=1-4TP=4, DP=2-4Single GPUTP=4, DP=1-2
Peak Mem per GPU72 GB68 GB48 GB64 GB
Cost per Clip$0.35-0.55$0.12-0.25$0.04-0.08$0.10-0.20
03

RUNWAY GEN-3 ALPHA: PRODUCTION-OPTIMIZED INFERENCE

Runway Gen-3 Alpha uses a mixture-of-experts DiT architecture with approximately 7 billion active parameters and 15 billion total parameters across 8 experts. The MoE design allows the model to activate only the relevant expert pathways per timestep, reducing effective FLOPs by 35-40 percent compared to a dense 15B model while maintaining quality. Runway deploys Gen-3 on a dedicated cluster of H100 GPUs with a custom inference engine built on TensorRT-LLM, achieving 25-40 seconds per 1080p 10-second clip versus 45-60 seconds for Sora on the same hardware.

The performance advantage comes from three optimizations: quantization of the VAE decoder to FP8 with minimal quality loss (0.03 FVD increase on UCF-101); temporal attention caching that reuses keys and values from the first half of the denoising trajectory for the second half, reducing attention FLOPs by 30 percent; and a CFG parallelization scheme that computes conditional and unconditional scores in a single forward pass with dual output heads.

04

PIKA 2.0 AND KLING: SINGLE-GPU AND SCALABLE APPROACHES

Pika 2.0 targets short-form 3-5 second clips at 720p resolution with a 3B parameter DiT that fits on a single H100 GPU with FP16 weights and full attention. The model uses a 3D-VAE with 8x spatial and 4x temporal compression, producing a latent tensor of shape [4, 128, 128, 20] for a 5-second clip at 24 FPS. At 20 denoising steps with DDIM sampling, inference completes in 12-20 seconds on one H100 and 8-12 seconds with FP8 optimization. The single-GPU deployment model makes Pika the most cost-efficient option for high-volume, short-form video generation at approximately $0.03-0.06 per clip.

Kling uses a 5B parameter diffusion model combining a 3D-VAE with a latent transformer for 1080p generation up to 5 seconds. Kling requires 4 H100 GPUs with tensor parallelism but achieves the lowest latency-to-quality ratio in the 5-second category. Kuaishou's production deployment runs on a 512-GPU H800 cluster using a custom scheduling layer that batches requests by duration and resolution to optimize GPU utilization.

05

B200 AND THE NEXT INFRASTRUCTURE WAVE

NVIDIA's B200 GPU with 192 GB HBM3e memory and 8 TB/s memory bandwidth fundamentally changes video model serving economics. The primary constraint on current H100-based video inference is the 80 GB VRAM ceiling, which forces multi-GPU tensor parallelism even for moderate-sized models. The B200's 192 GB capacity allows the full Sora-class models to fit on 2 GPUs instead of 8, reducing inference latency by 40-50 percent through lower communication overhead and enabling DP-parallel serving of multiple clips on a single node.

For Pika-class models (3B and below), B200 enables a batch size of 64-128 on a single GPU, approximately 4x the H100's throughput. At equivalent total cost of ownership, early projections from hyperscaler partners suggest B200 will deliver 3.5-4.2x more video clips per dollar than H100 for models under 7B parameters.

Model ClassH100 (80 GB)H200 (141 GB)B200 (192 GB)B200 vs H100 Lift
3B (Pika class)1 GPU / 16-32 b1 GPU / 32-64 b1 GPU / 64-128 b3.5-4.2x thru
7B (Runway class)4 GPUs / 8-16 b2 GPUs / 16-32 b2 GPUs / 32-64 b3.0-3.8x thru
10B+ (Sora class)8 GPUs / 4-8 b4 GPUs / 8-16 b2-4 GPUs / 16-32 b3.5-4.5x thru
Reserved Cost/Clip$0.18-0.35$0.12-0.22$0.06-0.123x reduction
P99 Lat 10s clip35-40s25-30s15-20s50% reduction
06

PRODUCTION SERVING STRATEGIES FOR VIDEO MODELS

Video model serving in production requires a fundamentally different architecture from LLM serving. The key insight is that video generation is a persistent GPU workload (20-120 seconds per request) rather than a token-generation loop with variable length. GPU allocation must be reservation-based rather than demand-based: pre-allocate GPU partitions for video generation and queue requests against them, rather than using the dynamic batching typical of text inference.

A recommended deployment pattern uses a two-tier architecture: a CPU-based request queue with priority classes (real-time <30s, batch <120s, bulk >120s) fronting a GPU worker pool where each worker holds a 1-8 GPU partition loaded with a specific model variant. This architecture achieves 70-85 percent GPU utilization on the worker pool versus 30-50 percent with ad hoc assignment. For teams building on open-source models like CogVideoX, vLLM's experimental diffusion model support and the Hugging Face Diffusers library with pipeline parallelism provide a starting point.

Filed under
AI Video GenerationSora GPU RequirementsRunway Gen-3 InferenceH100 Video AIB200 GPU VideoVideo Diffusion ModelsGPU Video Serving