HISTORICAL TRAJECTORY
Inference cost for GPT-3 class models dropped from approximately $20 per million tokens in 2022 to under $0.40 in 2026, a 50x reduction. When including quality-adjusted improvements from GPT-4 to GPT-5 class models, the effective cost-per-unit-intelligence fell over 1,000x. This collapse was driven by three compounding forces: GPU hardware efficiency (3-5x), model architecture improvements (4-8x), and inference optimization techniques (3-5x). The combined effect outstrips any single factor.
WHAT COLLAPSED IN PRICE
Standard LLM text generation experienced the most dramatic price collapse. OpenAI reduced GPT-4o pricing from $15/M input tokens to $0.50/M over three years. Open-source self-hosting costs dropped even more dramatically, from approximately $8/M tokens on A100 in 2023 to under $0.10/M on B200 with FP4 and speculative decoding in 2026. Image generation costs fell 20-30x as diffusion models matured and moved from 50-step to 4-step samplers like LCM and SDXL Turbo.
WHAT STAYED EXPENSIVE
Video generation costs remain stubbornly high, dropping only 3-5x since 2024. A 30-second 1080p video clip still costs $3-15 in GPU compute depending on model and provider. Real-time voice and audio generation maintained higher margins due to streaming latency requirements that prevent large-batch efficiency. Fine-tuning costs dropped only 2-3x because the training process is fundamentally memory-bandwidth-bound rather than compute-bound, limiting the impact of hardware improvements.
PROCUREMENT IMPLICATIONS
The cost collapse changes long-term GPU procurement strategy. Locking into multi-year GPU reservations at today's prices risks overpaying as costs continue to fall 40-60% annually. Teams should prefer 3-6 month renewable contracts with price renegotiation clauses tied to hardware generation milestones. The accelerating pace of cost reduction makes spot and on-demand GPU strategies more attractive relative to reserved commitments.
PROVIDER STRATEGY AND MARGINS
Inference providers maintain 70-90% gross margins despite the 50x price collapse, driven by continuous optimization. Together AI, Fireworks, and Groq all demonstrated in 2026 that model-optimized inference stacks can undercut hyperscalers by 3-5x while maintaining profitability. The key insight is that inference is a technology-driven business where margins depend on optimization deployment speed, not hardware cost advantage.
2027 OUTLOOK AND STRATEGY
Inference costs will likely continue falling 40-60% annually through 2027 as Vera Rubin and AMD MI400 enter production. The next wave of optimization will come from inference-specialized hardware and model-optimized architectures rather than general GPU improvements. Teams should plan GPU budgets assuming 50% annual cost decline for text inference and 30% for multimodal. Avoid fixed-price multi-year inference commitments and maintain flexibility to adopt new hardware generations within 6 months of availability.
