The Migration Assessment Framework for AI Workloads
AI workload migration is different from standard application migration because the bottleneck is not CPU or memory but data throughput and GPU interconnect topology. A standard cloud migration checklist (lift the VM, replicate the database, redirect DNS) misses the critical constraints: data gravity, network topology sensitivity, and storage IOPS requirements that change by orders of magnitude between GPU generations.
The assessment framework has five dimensions. Workload characteristics: is this training, inference, or data preprocessing? Training requires sustained high-bandwidth GPU-to-GPU communication and tolerates almost no latency jitter. Inference requires low-latency response and can be more geographically distributed. Data preprocessing is IO-bound and cares more about storage throughput than GPU proximity. Data gravity: how much data does the workload generate or consume, where is it currently stored, and what does it cost to move it? A 10TB dataset is trivial to migrate. A 5PB training corpus changes the entire cost analysis. Interconnect requirements: does the workload need NVLink, InfiniBand, or is Ethernet sufficient? If your on-premise cluster uses NVLink for tensor parallelism and the cloud provider cannot offer the same topology, the workload will not run the same way. Compliance and latency: do data residency requirements force a specific geographic location? Does the inference workload have a P99 latency target that rules out certain regions? Cost structure: what is the per-unit cost of compute, storage, and data transfer at the source and destination, and how does the total cost of ownership compare at different utilization levels?
Most teams rush to the cost comparison before completing the workload characteristic analysis. This produces migration plans that technically work but perform 30-60% worse than expected because the topology does not match the training framework's assumptions. The assessment should take 2-4 weeks and produce a scored compatibility matrix, not just a cost spreadsheet.
Data Transfer Costs: What Cloud Providers Charge
Data egress is the hidden cost that can double your cloud bill during a migration. Every major provider charges for data leaving their network. Ingress is usually free. Egress rates have been stable since 2024 and show no signs of dropping in mid-2026. The table below shows current published rates for transferring data out of each provider's cloud.
The critical insight is that egress costs dominate for data-intensive AI workloads. Moving 500TB of training data out of AWS to a colocation facility costs $45,000-$54,000 at standard rates. That is not insignificant, but it is a one-time cost. The ongoing egress cost for a hybrid setup that passes training checkpoints or inference results back to a cloud service can accumulate quickly if not architected correctly.
| Provider | First 10TB/mo | Next 40TB/mo | Next 100TB/mo | 500TB+/mo |
|---|---|---|---|---|
| AWS | $0.09/GB | $0.085/GB | $0.07/GB | $0.05/GB |
| Azure | $0.087/GB | $0.083/GB | $0.07/GB | $0.05/GB |
| GCP | $0.08/GB | $0.04/GB | $0.02/GB | Custom |
| Oracle Cloud | $0.0085/GB | $0.0085/GB | $0.0085/GB | $0.0085/GB |
| CoreWeave / Lambda | $0.05/GB | $0.05/GB | $0.05/GB | $0.04/GB |
| Direct Connect / ExpressRoute | $0.02-$0.04/GB | $0.02-$0.04/GB | Negotiated | Negotiated |
Network Latency: When It Breaks Your Training Pipeline
Network latency is the most underappreciated migration risk for AI training workloads. Distributed training frameworks like PyTorch DDP, FSDP, and DeepSpeed assume intra-node and inter-node latencies within specific bounds. When those bounds change, the training job does not fail; it just runs 20-60% slower, and the team wastes weeks debugging performance before realizing the root cause is network topology.
The relevant thresholds for training: intra-node NVLink communication should be under 1 microsecond (it is on the order of hundreds of nanoseconds on physical NVLink). InfiniBand inter-node latency should be under 2-3 microseconds. Cloud provider inter-node latency across Kubernetes pods or VMs in the same availability zone typically runs 3-10 microseconds. Across availability zones or regions, it jumps to 50-200 microseconds. The effect on all-reduce operations is measurable: a 32-GPU all-reduce that completes in 2ms on InfiniBand can take 15-25ms across cloud VMs without topology awareness.
For inference workloads, latency sensitivity depends on the use case. Real-time chatbots require P99 response times under 500ms, which typically restricts inference to a single region or even a single availability zone. Batch inference can tolerate cross-region latency because throughput costs are the binding constraint. The migration plan for inference workloads should include latency budget analysis before provider selection, not after.
Lift-and-Shift vs Re-Architecture: The Real Cost Difference
Lift-and-shift for AI workloads means moving the same training scripts, the same storage layout, and the same cluster scheduling setup to the new environment with minimal changes. It sounds simpler than re-architecture, but the practical cost difference is smaller than most teams expect because AI infrastructure tooling has standardized significantly since 2024.
A real lift-and-shift still requires: reconfiguring your Slurm or K8s GPU operator for the new network topology, reconfiguring your storage mount paths and permissions, updating your container registry references and secrets management, adjusting your data pipeline to point to new storage buckets, and re-validating your training framework's NCCL communication settings for the new interconnect. This is not zero work. Most teams underestimate the re-validation effort by 2-3x.
Re-architecture means changing the workload to take advantage of the new environment. Examples: refactoring from FSDP to DeepSpeed Ulysses if the cloud provider has better inter-node bandwidth, switching to a different inference serving framework that supports the provider's GPU topology, or restructuring the data pipeline to use cloud-native object storage instead of NFS. Re-architecture typically adds 4-8 weeks to the migration timeline but produces 15-30% better performance in the target environment.
Our recommendation based on 30+ observed migrations: do lift-and-shift for the initial migration to unblock the team and validate the target environment, then schedule re-architecture as a Phase 2 project. The cost of doing both simultaneously is high rework risk if the lift-and-shift reveals unexpected constraints in the target environment.
Hybrid Strategies: Keeping the Best of Both Worlds
Hybrid AI infrastructure separates workloads by characteristic. Training goes to whichever environment offers the best cost-per-FLOP for sustained compute. Inference runs close to users, which usually means a geographically distributed deployment across multiple cloud regions or edge locations. Data preprocessing and storage sits where the data already lives to avoid egress costs.
There are three viable hybrid patterns in mid-2026. Cloud-burst: run baseline training on reserved on-premise GPU capacity, and burst to cloud spot instances during training spikes or when testing new model architectures. This requires your training framework to support elastic compute (fault-tolerant checkpointing with dynamic node addition). Data-local: keep training data on-premise or in a colocation facility with high-bandwidth storage, and stream data to cloud GPUs during training jobs. This avoids egress costs for permanent data storage but creates a data pipeline dependency on network bandwidth between locations. Split-workload: run training in the cloud where GPU availability is highest, host inference on-premise where latency control is tighter, and use a streaming replication layer to keep data synchronized.
The hybrid approach adds operational complexity that every team under-estimates. You now need two sets of GPU cluster management tools, two storage systems, two network configurations, and twice the monitoring surface area. The operational overhead of hybrid typically adds 0.5-1.0 FTE in infrastructure engineering time. It is only worth it when the cost savings from workload placement exceed $200,000-$300,000 per year, which means at least 100+ GPUs in sustained use.
Migration Timeline: What Realistic Phasing Looks Like
A complete AI workload migration from on-premise to cloud (or vice versa) takes 12-20 weeks for a team with dedicated infrastructure engineering resources. The phasing below assumes a team of 2-3 infrastructure engineers working full-time on the migration alongside their normal responsibilities.
Weeks 1-3 (Assessment and Planning): complete the five-dimension assessment framework, build a cost model for the target environment, identify workload compatibility issues, and create the migration runbook. This phase produces a go/no-go decision. We have seen two teams reach week 3 and discover that their InfiniBand-dependent training framework would lose 50% throughput on a cloud fabric that could not match the topology. Both teams correctly halted the migration and re-architected before proceeding.
Weeks 4-8 (Pilot Migration): select one representative workload and migrate it completely. Validate throughput, latency, cost, and reliability. Run the workload in production for 2-4 weeks to surface any issues. This phase typically reveals 3-5 unexpected problems that require configuration changes or minor code modifications. The pilot workload should be non-critical but representative enough that you can extrapolate to the full migration.
Weeks 9-16 (Phased Migration): migrate workload groups in order of decreasing criticality, with 1-2 weeks of validation between each group. The first group should be the workload that provides the most cost savings or performance improvement. Leave the most complex or highest-risk workloads for last, when the team has accumulated migration experience.
Weeks 17-20 (Optimization and Decommission): decommission the source environment, finalize the monitoring and cost management tooling in the target environment, and run a post-migration optimization pass to adjust resource allocations, refine auto-scaling parameters, and address any performance regressions discovered during the phased migration.
Risk Factors That Kill AI Workload Migrations
The most common migration failure mode is NCCL timeout increases. When a training job that ran reliably on-premise with InfiniBand starts experiencing random NCCL timeouts in the cloud, teams spend weeks debugging before realizing the cloud's virtualized network has different time-out characteristics. The fix is usually adjusting the NCCL timeout parameter and aligning the cloud scheduler's GPU placement policy to keep training jobs on the same network switch, but if the cloud provider cannot guarantee topology-aware placement for multi-node jobs, the workload may never run reliably.
The second most common failure is storage performance mismatch. On-premise training jobs often use NFS or parallel filesystem storage (Lustre, WekaFS) configured for the specific GPU cluster topology. Cloud object storage (S3, GCS, Blob) has fundamentally different latency and throughput profiles. Training frameworks that read and write checkpoints frequently can see significant performance degradation. The fix is usually increasing checkpoint intervals, adding local NVMe caching, or using a cloud-managed parallel filesystem (Amazon FSx for Lustre, Azure Managed Lustre) at an additional cost that was not in the original budget.
The third risk is cost forecasting error. The cost model built during assessment inaccurately projects GPU utilization or underestimates data transfer costs. The most common error is assuming 100% GPU utilization in the target environment when the source environment ran at 60-70%. Add networking time, data loading delays, and rescheduling overhead, and the effective utilization in the new environment often drops to 40-50% for the first 2-3 months. The budget must account for this utilization gap or the migration gets canceled mid-way.
