THE MIGRATION DECISION
On-premise at 85% utilization: $0.85-1.20/GPU-hr fully loaded vs cloud on-demand $2.50-3.80/hr. At 50% utilization, on-premise rises to $1.70-2.40/hr, erasing the advantage. Tipping point: utilization below 60-70%.
Additional cloud advantages: faster GPU upgrades, global availability, reduced ops headcount.
| Workload Profile | On-Prem $/GPU-hr | Cloud $/GPU-hr | Recommended | Priority |
|---|---|---|---|---|
| 85%+ utilization 24/7 | $0.85-1.20 | $2.50-3.80 | On-Prem | Low |
| 50-70% utilization | $1.20-1.70 | $2.50-3.80 | Hybrid | Medium |
| <50% utilization | $1.70-2.40 | $2.50-3.80 | Cloud | High |
| Variable burst | N/A | $3.50-4.20 | Cloud | Immediate |
DATA TRANSFER AND STORAGE MIGRATION
Transferring 1 PB over 10 Gbps dedicated link takes 12 days. Over 1 Gbps internet: 120 days. Options: AWS Snowball Edge (80 TB/device, $300 + shipping) or direct connect ($2-10K/month).
Staged strategy: replicate hot datasets to cloud, validate pipelines, cut over training, maintain on-prem cold storage 3-6 months. Data transfer costs for 1 PB: $30-60K. Data pipeline re-engineering: $50-150K.
Replace on-prem NFS/Lustre with cloud FSx for Lustre, Filestore, or managed NFS.
NETWORKING RE-ARCHITECTURE
Multi-node training requires tuning for cloud interconnect: AWS EFA, GCP GPUDirect-TCPX, Azure InfiniBand. Benchmark all-reduce before migrating (target within 15% of on-prem). NCCL tuning: 2-4 weeks.
Cloud networking is usage-based: data transfer between nodes (same AZ) is free, but to object storage costs ~$1,800/month per 100 TB read.
OPERATIONAL READINESS
On-premise 128-GPU cluster requires 3-5 FTEs ($300-500K/year). Cloud reduces to 1-2 FTEs ($150-250K/year). Combined with no hardware depreciation ($500-800K/year), total operational savings: $650K-1.05M/year.
Phased migration: weeks 1-4 setup, weeks 5-8 migrate lower-priority jobs, weeks 9-16 production migration, weeks 17-24 decommission on-prem. Budget 15-20% overhead for unexpected issues.
