FROM DEEP LEARNING WORKSTATIONS TO GPU CLOUD
Lambda began in 2012 as a deep learning hardware company, building and selling GPU workstations pre-configured for deep learning frameworks. By 2019, Lambda was shipping thousands of workstations annually to universities and research labs, with the Lambda Blade workstation becoming a standard tool in academic AI labs. This hardware expertise gave Lambda deep knowledge of GPU thermal management, power delivery, and system reliability that most cloud providers lack.
The pivot to GPU cloud began in 2020, launched as a natural extension: if customers wanted GPU access without upfront hardware purchases, Lambda could provision the same hardware it sold as workstations in data centers. The cloud launched with A100-80G GPUs at $1.10 per GPU-hour, a price that was 50-60 percent below AWS and GCP at the time and remains the company's flagship pricing strategy: transparent, simple, and aggressively competitive.
BARE METAL ARCHITECTURE AND GPU CONFIGURATIONS
Lambda's cloud is built on bare metal servers with no virtualization layer. Each server is a standard Lambda-built system with 8x H100 SXM or H200 SXM GPUs connected via NVLink, running Ubuntu with NVIDIA drivers and Docker. Customers receive SSH root access to a fully dedicated server, not a VM or container. This architecture eliminates the hypervisor overhead and GPU memory contention that affects virtualized cloud instances, delivering 95-100 percent of theoretical GPU performance on training workloads.
The bare metal approach has implications for customer experience. Provisioning takes 2-5 minutes for a single node, versus seconds for cloud VMs, because the system must boot a full OS. But the performance is consistent and predictable: no noisy neighbor issues because there is no sharing. Lambda offers single-node rentals (8 GPUs) and cluster rentals (2 to 128+ nodes connected via InfiniBand). Cluster deployments include a dedicated Slurm or Kubernetes control plane and a shared NFS filesystem.
| Configuration | GPUs | VRAM Total | Interconnect | Price/hr | Best For |
|---|---|---|---|---|---|
| 1x H100 SXM | 8 H100 | 640 GB | NVLink + InfiniBand | $16.40-$18.40 | Training, finetuning |
| 1x H200 SXM | 8 H200 | 1,128 GB | NVLink + InfiniBand | $22.40-$26.40 | Large model training |
| 1x A100-80G SXM | 8 A100 | 640 GB | NVLink + InfiniBand | $8.80-$10.40 | Budget training |
| 1x A100-PCIe | 8 A100 | 320 GB | InfiniBand | $6.40-$7.60 | Inference, batch jobs |
| Cluster 8x H100 | 64 H100 | 5,120 GB | NDR InfiniBand | $131-$147 | Multi-node FSDP |
TRANSPARENT PRICING: THE $1.10 A100 LEGACY
Lambda's pricing strategy is built on transparency. H100 pricing starts at $2.05 per GPU-hour on a 1-month commitment and drops to $1.60 on 12-month reservations. A100-80G starts at $1.10 per GPU-hour, a price that has become Lambda's signature offering and remains among the lowest in the market for reserved instances. There are no charges for data transfer, no egress fees, and no separate storage costs: the quoted price includes a 1 TB NVMe SSD per node and 20 TB NFS storage per cluster.
The simplicity of Lambda's pricing is a deliberate competitive moat. While AWS, GCP, and CoreWeave have complex pricing with instance types, burstable credits, data transfer tiers, and storage classes, Lambda quotes a single per-GPU-hour price that includes everything except additional storage. This has made Lambda particularly popular with academic labs and budget-conscious startups that need predictable costs. Lambda's bet is that the operational simplicity of a single price outweighs the optimization opportunities of complex pricing structures for most customers.
| GPU | On-Demand/hr | 1-Month/hr | 3-Month/hr | 12-Month/hr | Includes |
|---|---|---|---|---|---|
| A100-80G (SXM) | $1.50 | $1.10 | $1.00 | $0.90 | 1 TB NVMe, 20 TB NFS |
| A100-PCIe | $1.10 | $0.80 | $0.72 | $0.65 | 1 TB NVMe, 20 TB NFS |
| H100 SXM | $2.80 | $2.30 | $2.05 | $1.60 | 1 TB NVMe, 20 TB NFS |
| H200 SXM | $3.80 | $3.30 | $2.80 | $2.40 | 1 TB NVMe, 20 TB NFS |
| RTX 6000 Ada | $1.00 | $0.75 | $0.65 | $0.55 | 500 GB NVMe, 10 TB NFS |
GPU AVAILABILITY AND DATA CENTER FOOTPRINT
Lambda operates data centers in the US (Santa Clara, CA; Dallas, TX; Ashburn, VA; Chicago, IL; and Phoenix, AZ opened in Q1 2026) and Europe (Amsterdam, NL; Frankfurt, DE; London, UK). Total GPU count is approximately 24,000 H100s and 8,000 A100s, making them roughly half the size of CoreWeave's fleet but significantly larger than mid-tier providers like RunPod or Vast.ai.
Lambda's GPU availability is generally better than hyperscalers but more constrained than CoreWeave's. As of 2026, H100 availability has normalized: on-demand access for 1-8 GPUs is nearly immediate, 8-64 GPUs require 1-3 days approval, and 64+ GPUs require 2-4 weeks for cluster provisioning. H200 and B200 availability remains tighter, with B200 expected in late 2026.
TARGET AUDIENCE: ACADEMIA, RESEARCH, AND COMPUTE-AWARE STARTUPS
Lambda's core demographic is academic researchers and computationally sophisticated startups. The combination of transparent pricing, SSH root access, and a hardware-first ethos resonates with users who know how to manage their own infrastructure and resent paying premium prices for managed services they do not need. Lambda also runs a grant program that has provided over $10 million in free GPU credits to academic researchers since 2022, with individual grants of $5,000-50,000 in compute credits.
Lambda's weaknesses are the inverse of its strengths. There are no managed inference endpoints, no serverless GPU options, no integrated MLOps platform, and limited multi-region deployment options. The console is functional but basic compared to hyperscaler portals. Teams that want a turnkey AI platform will find Lambda frustrating. Teams that want raw GPU access at transparent prices with minimal overhead will find Lambda ideal.
COMPETITIVE POSITION IN 2026
Lambda's position is unique: they are the only major provider that started as a hardware company, giving them a cost structure advantage in GPU procurement and system design. Their H100 12-month reserved price of $1.60/GPU-hour is competitive with CoreWeave ($2.00) and well below hyperscaler 1-year reserved rates ($2.80-3.20 on AWS/GCP). For academic budgets and seed-stage startups, Lambda is often the most accessible option outside of cloud credits.
However, Lambda faces increasing pressure from both above and below. CoreWeave and hyperscalers are expanding their bare metal options, reducing differentiation. Smaller providers like RunPod offer serverless GPU at lower entry prices. Lambda's response has been to add cluster-level capabilities (Slurm integration, shared filesystem, InfiniBand fabric) while maintaining pricing discipline. Lambda reported revenue of approximately $250 million in 2025, with 80 percent from cloud and 20 percent from hardware sales.
