Migration Strategies

Topics Covered

The Six Rs of Migration

Rehost (Lift and Shift)

Replatform (Lift and Reshape)

Refactor (Re-architect)

Repurchase (Replace with SaaS)

Retire (Decommission)

Retain (Keep On-Prem)

Migration Assessment: Inventory and Prioritize

Strangler Fig Pattern

How It Works

Why Incremental Beats Big-Bang

Feature Parity Milestones

Routing Strategies

Common Strangler Fig Pitfalls

Data Migration

Bulk Transfer (Offline Migration)

CDC-Based Sync (Continuous Migration)

Dual Writes

Data Validation

Choosing the Right Data Migration Strategy

Cutover Planning

Cutover Rehearsal

Go/No-Go Checklist

Parallel Running

DNS Cutover vs Proxy-Based Routing

Rollback Plan

Migration Velocity Metrics

Every migration begins with a question: what should we do with each workload? Moving everything to the cloud the same way is a recipe for wasted effort and blown budgets. Some applications benefit from a full re-architecture. Others just need to be moved as-is. Some should be shut down entirely. The Six Rs framework gives you a structured way to categorize every workload in your portfolio so each one gets the right treatment.

The six strategies form a spectrum from least effort to most transformation. Understanding each strategy and when to apply it is one of the most practical skills for any engineer involved in cloud adoption, platform modernization, or datacenter consolidation. The framework originated at Gartner and was popularized by AWS as a standard vocabulary for migration planning.

Six dispositions for one application portfolio, from rehost through to retire, with what each one leaves you.

Rehost (Lift and Shift)

Rehosting means moving a workload to the cloud without changing its code, architecture, or configuration. You take the virtual machine image, replicate it in the cloud, point DNS at the new address, and you are done. This is the fastest migration strategy and the one with the lowest risk.

When does rehosting make sense? When the application works fine but the datacenter lease is expiring. When the team does not have capacity for a rewrite. When you need to migrate 200 servers in 6 months and there is no time to re-architect each one. AWS reports that organizations often rehost 80% of their portfolio just to get out of the datacenter, then optimize later.

The tradeoff is clear: you gain cloud infrastructure (elastic scaling, managed networking, pay-per-use billing) but you do not gain cloud-native benefits (auto-scaling, managed databases, serverless compute). A VM running in the cloud is still a VM. You are paying cloud prices for on-prem architecture.

Rehosting is also the natural first step for organizations with limited cloud expertise. The team learns cloud operations (networking, IAM, monitoring) without simultaneously learning new application architectures. Once the team is comfortable operating in the cloud, they can selectively optimize individual workloads through replatforming or refactoring.

One often-overlooked benefit of rehosting is that it provides a baseline for future optimization. Once the application runs in the cloud, you can measure its resource usage with cloud-native monitoring tools. A rehosted VM that uses 10% of its allocated CPU reveals an over-provisioned instance that can be downsized immediately, saving cost before any code changes are made. This "rehost, then right-size" approach delivers quick wins that help justify the broader migration program to finance teams.

Replatform (Lift and Reshape)

Replatforming makes targeted optimizations during the move without changing the core architecture. You swap the self-managed MySQL instance for Amazon RDS. You replace the on-prem message queue with a managed service like SQS. You containerize the application for ECS or EKS without rewriting its internals.

This is the pragmatic middle ground. You get meaningful cloud benefits (managed patching, automated backups, scaling policies) without the cost and risk of a full re-architecture. The application logic stays the same. The infrastructure underneath it improves.

The key decision in replatforming is which components to swap. Focus on the components that cause the most operational pain: the database that requires manual patching every month, the message queue that crashes under load, the logging infrastructure that fills up disks. Each swap should deliver a clear operational improvement that justifies the migration effort.

A useful mental model for replatforming: ask "what wakes us up at 3 AM?" The components that cause on-call pages are the best candidates for replacement with managed services. Managed RDS eliminates database patching emergencies. Managed Elasticsearch eliminates cluster scaling incidents. Managed Kafka eliminates broker rebalancing failures. Each swap reduces operational burden by eliminating an entire category of incidents.

Refactor (Re-architect)

Refactoring means redesigning the application to be cloud-native. A monolith becomes microservices. Synchronous batch processing becomes event-driven pipelines. A relational database gets decomposed into purpose-built stores (DynamoDB for key-value, ElasticSearch for search, S3 for objects). This is the most expensive and risky migration strategy, but it unlocks the full value of cloud computing.

Refactor when the application cannot scale to meet demand in its current form. Refactor when the business needs features (multi-region deployment, sub-second scaling) that the current architecture cannot support. Do not refactor a stable internal tool with 50 users just because microservices are trendy.

The cost of refactoring is not just engineering time. It includes the risk of regression bugs, the learning curve for new architectures, the operational complexity of running more services, and the testing effort required to validate that the new system matches the old one's behavior. A refactoring project that takes 6 months to build and 3 months to stabilize is a 9-month investment. Make sure the business value justifies that timeline.

When refactoring does make sense, break it into phases. Do not attempt to refactor an entire monolith into microservices in one project. Extract one bounded context at a time using the strangler fig pattern (covered in the next section). Each extraction is a self-contained project with its own testing and cutover. This phased approach limits risk and delivers incremental value rather than betting everything on a single release.

A common trap with refactoring is scope creep. The team starts by extracting the user service, then decides to also modernize the authentication system, then adds a new API gateway, then redesigns the event bus. Each addition is individually reasonable but collectively turns a 3-month project into a 12-month project. Set clear boundaries at the start: "We are extracting the user service. Authentication, API gateway, and event bus are out of scope for this phase." Enforce these boundaries rigorously.

Refactoring also requires investment in testing infrastructure that most legacy systems lack. Before you can safely refactor a monolith into microservices, you need comprehensive integration tests that define the current behavior. Without these tests, you have no way to verify that the refactored system behaves identically to the original. Many refactoring projects spend the first 2-3 months just building the test suite before any architectural changes begin. This feels slow but it is the foundation that makes everything else safe.

Repurchase (Replace with SaaS)

Repurchasing means replacing a self-managed application with a commercial SaaS product. Your custom-built CRM becomes Salesforce. Your self-hosted email server becomes Google Workspace. Your internal analytics platform becomes Snowflake.

This strategy makes sense when the application is not a competitive differentiator. Running your own email server consumes engineering time that could be spent on your core product. The SaaS vendor has a larger team, better security posture, and handles upgrades for you.

The hidden cost of repurchasing is data migration and integration. Your custom CRM stores data in a schema that Salesforce does not match. Your workflows depend on custom fields that the SaaS product handles differently. Budget time for data mapping, import scripting, and user retraining. The ongoing SaaS subscription cost must also be compared against the fully-loaded cost of running the self-managed alternative (engineering time, infrastructure, on-call rotations, security patching).

Another risk with repurchasing is vendor lock-in. Once your data and workflows live in a SaaS product, switching to a competitor or building in-house becomes expensive. Evaluate the SaaS vendor's data export capabilities, API coverage, and contract flexibility before committing. A vendor that makes it easy to extract your data gives you negotiating leverage at renewal time.

Consider the integration burden as well. Your existing systems likely integrate with the self-managed application through internal APIs, shared databases, or file drops. The SaaS replacement will have different integration patterns: webhooks, REST APIs, or proprietary connectors. Every integration must be rebuilt. For an application with 10 integrations, the integration rewrite alone can take longer than the migration of the application itself.

Run a proof-of-concept with the SaaS product before committing. Import a subset of your data, configure the key workflows, and have actual users test it for 2-4 weeks. The proof-of-concept reveals gaps between the SaaS product's capabilities and your actual requirements: workflow limitations, missing fields, reporting gaps, and performance issues with large datasets. Discovering these gaps during a proof-of-concept costs weeks. Discovering them after a full migration and contract signing costs months and significant budget.

Retire (Decommission)

Migration is the perfect time to audit what you actually need. Organizations typically discover that 10-20% of their application portfolio is unused, redundant, or replaceable. That internal dashboard nobody has opened in 18 months. The reporting tool that was replaced by a newer system but never shut down. The staging environment for a product that was discontinued.

Retiring these workloads saves migration effort, reduces cloud costs, and shrinks your attack surface. Every application you do not migrate is one you do not have to maintain, monitor, or secure in the new environment.

The retirement process requires diligence. Before shutting down an application, verify with all potential stakeholders that it is truly unused. Check access logs for the past 90 days. Search for references to its API endpoints in other codebases. Some applications look unused because their users access them quarterly (compliance reports, annual audits). Archive the application's data and code before decommissioning, so it can be restored if someone discovers a dependency after shutdown.

A useful technique is "dark decommissioning": instead of shutting the application down immediately, stop advertising it and monitor access logs for 30 more days. If no one accesses it during that period, you can be confident it is truly unused. If someone does access it, you have identified a stakeholder who was not captured in the initial audit.

Retirement also applies to data, not just applications. Databases often contain tables that no process reads, logs that no one queries, and backups of systems that were decommissioned years ago. Migrating this dead data wastes time and cloud storage costs. Audit each table's access patterns as part of the migration assessment. Archive or delete data that has no active consumers before migrating the remaining data.

Retain (Keep On-Prem)

Some workloads should not move. Mainframe applications with decades of embedded business logic. Systems with regulatory requirements that mandate on-premises hosting. Applications with hardware dependencies (specialized GPUs, FPGA boards) that cloud instances cannot replicate cost-effectively.

Retaining is not failure. It is an acknowledgment that cloud migration is a business decision, not a technical mandate. A hybrid architecture with some workloads on-prem and others in the cloud is a valid long-term state.

When retaining a workload, document the decision and the reasons behind it. Include the conditions under which the decision should be revisited: "Retain until the regulatory landscape changes," or "Retain until the mainframe vendor ends support in 2028." This prevents the retained workload from becoming a forgotten liability. It also prevents well-meaning engineers from periodically re-proposing migration without new information.

For retained workloads that need to communicate with cloud-hosted services, plan the network connectivity carefully. A VPN or Direct Connect link between the on-premises datacenter and the cloud provider is typically required. This link must have sufficient bandwidth for the traffic volume and low enough latency for real-time interactions. Monitor this cross-environment link as carefully as any other production dependency: if it goes down, every cloud service that depends on the retained on-premises workload fails.

Security and compliance requirements for retained workloads may also affect the cloud architecture. If the retained workload handles PII or financial data, any cloud service that communicates with it must meet the same compliance standards. This can restrict which cloud regions, instance types, and networking configurations are allowed. Involve your security and compliance teams early in the migration planning to identify these constraints before they become blockers.

Interview Tip

In interviews, when asked about migration strategy, always start with the assessment phase. List the Six Rs, explain how you would categorize each workload, and then describe the migration plan for the highest-priority category. This demonstrates that you think in terms of portfolio-level decisions, not just individual application moves.

Migration Assessment: Inventory and Prioritize

Before migrating anything, you need a complete picture of what you have. The assessment phase has three steps.

Inventory: Catalog every application, database, and service. Record its dependencies, traffic volume, data sensitivity, and the team that owns it. Tools like AWS Application Discovery Service or manual spreadsheets both work. The goal is a complete list with no surprises during migration.

Categorize by 6R: For each workload, assign one of the six strategies. This decision depends on business value, technical complexity, and team capacity. A decision matrix helps: high business value plus high technical debt points toward refactor. Low business value plus low usage points toward retire. Everything else starts as rehost or replatform.

Prioritize by value and difficulty: Migrate low-risk, high-value workloads first. These early wins build team confidence and establish migration patterns that later workloads can follow. Save the complex refactors for when the team has migration experience and the tooling is mature.

A common mistake is starting with the hardest, most critical application because it delivers the most business value. This is backwards. The first migration is where the team learns its processes, discovers gaps in tooling, and builds muscle memory. Start with a low-risk application where a failed cutover has minimal business impact. The patterns, scripts, and runbooks developed during this first migration become templates for every subsequent one. By the time the team reaches the critical applications, they have migrated ten workloads and the process is well-understood.

Dependencies between applications complicate prioritization. If Application A depends on Application B's database, migrating A to the cloud while B stays on-premises creates a cross-environment dependency with added latency. Map all dependencies before setting the migration order. When possible, migrate clusters of tightly-coupled applications together to avoid cross-environment calls. When that is not possible, ensure the network path between cloud and on-premises has sufficient bandwidth and acceptable latency for the dependency traffic.

Create a dependency graph for your entire application portfolio. Nodes are applications. Edges are dependencies (API calls, shared databases, file shares, message queues). The graph reveals migration clusters: groups of applications that must migrate together because they are too tightly coupled to separate. It also reveals the natural migration order: applications with no outgoing dependencies (leaf nodes) can migrate first because they have no cross-environment concerns. Applications at the center of the graph with many dependencies should migrate last, once their dependencies are already in the cloud.

The dependency graph also exposes hidden coupling. Two applications that share a database are implicitly coupled even if they do not call each other's APIs. Migrating one without the other means the shared database must be accessible from both environments, which adds network complexity and latency. Identifying these hidden couplings early prevents mid-migration surprises.

Shared databases are the most common form of hidden coupling and the most dangerous. When two applications read from and write to the same tables, migrating one application to a new database requires either splitting the database (complex data migration) or maintaining cross-environment database access (latency and security concerns). Neither option is cheap. The earlier you identify shared databases in the assessment phase, the better you can plan their decomposition.