Disaster Recovery

Topics Covered

RTO and RPO Fundamentals

Why These Numbers Matter

How to Calculate RTO and RPO

DR Strategies Spectrum

Choosing the Right Strategy

The Cost-Availability Trade-off

Common RTO/RPO Mistakes

Types of Disasters

Backup Strategies

Automated Snapshots

Point-in-Time Recovery

Cross-Region Backup Copies

Backup Retention and Lifecycle

Infrastructure as Code for Recovery

Backup Security

Cross-Region Replication

Asynchronous Replication

Multi-Layer Replication

Object Storage Replication

Replication Monitoring

Data Consistency During Failover

Split-Brain Scenarios

DNS and Traffic Switching

Application-Level DR Considerations

DR Testing and Runbooks

Types of DR Testing

Why Teams Avoid DR Testing (and Why That Is Dangerous)

Building Effective Runbooks

Communication During Disasters

Post-Incident Reviews

Testing Cadence

Automated DR Validation

The DR Maturity Model

Every production system will eventually face a failure that takes it offline. Hardware dies, regions go down, engineers make mistakes, and attackers exploit vulnerabilities. Disaster recovery (DR) is the set of strategies and processes that bring your system back to a functional state after such an event. The question is not whether a disaster will happen, but how quickly and completely you can recover when it does.

Disaster recovery starts with two numbers that every engineering team must define before building anything. These two numbers determine your architecture, your budget, and your operational complexity. Without them, you are guessing.

Recovery Time Objective (RTO) is the maximum acceptable downtime after a disaster. If your RTO is 4 hours, your system must be fully operational within 4 hours of a failure. If your RTO is 30 seconds, you need infrastructure that can failover almost instantly. RTO answers the question: "How long can we be down before the business suffers unacceptable damage?" RTO is measured from the moment the disaster occurs (not from when you detect it) to the moment the system is fully functional again.

Recovery Point Objective (RPO) is the maximum acceptable data loss measured in time. If your RPO is 1 hour, you can lose up to 1 hour of data. If your RPO is zero, you cannot lose any committed transaction. RPO answers the question: "How much data can we afford to lose?" RPO is measured backward from the disaster: an RPO of 1 hour means you need a recovery point (backup, snapshot, or replicated state) that is no more than 1 hour old at the time of failure.

Two clocks pointing in opposite directions from the same moment of failure, with what each one is measuring.

Why These Numbers Matter

RTO and RPO are not technical metrics. They are business decisions. A payment processing system might have an RTO of 30 seconds and an RPO of zero because every minute of downtime costs revenue and every lost transaction creates legal liability. A batch analytics pipeline might tolerate an RTO of 24 hours and an RPO of 6 hours because yesterday's data is still useful and the pipeline runs daily anyway.

The relationship between RTO, RPO, and cost is direct: lower numbers cost more. An RTO of 4 hours might require only nightly backups stored in another region. An RTO of 30 seconds requires a fully running standby environment with real-time replication, costing 2x your infrastructure.

How to Calculate RTO and RPO

Start with the business, not the technology. Talk to stakeholders and ask concrete questions. "If our checkout system goes down, how long can we survive before customers switch to competitors?" That answer is your RTO. "If we lose the last hour of orders, can we reconstruct them from payment processor records, or are they gone forever?" That answer informs your RPO.

For revenue-generating systems, the calculation is quantitative. If your system generates $10,000 per hour in revenue and your DR strategy costs $5,000 per month, you break even if the strategy prevents just 30 minutes of downtime per year. For systems where the cost of downtime is reputational rather than directly financial (a social media platform, a healthcare portal), the calculation requires judgment about brand damage, user trust, and regulatory consequences.

Document these decisions explicitly. Write down: "We accept an RTO of 4 hours for the admin dashboard because the cost of downtime is limited to 3 internal users being unable to generate reports. We accept an RPO of 24 hours because the dashboard pulls from the analytics warehouse, which is rebuilt nightly." When a disaster occurs, this documentation prevents arguments about priorities. The team executes the plan rather than debating which system to recover first.

Interview Tip

In interviews, always ask what the RTO and RPO requirements are before proposing a DR strategy. Jumping straight to active-active replication for a system that tolerates 4 hours of downtime shows you are optimizing for technology, not business needs. The best engineers pick the cheapest strategy that meets the requirements.

DR Strategies Spectrum

There are four standard DR strategies, each trading cost for speed of recovery. Think of them as a spectrum from cheapest and slowest to most expensive and fastest.

Five DR strategies priced against how fast each comes back, how much it loses, and what it costs sitting idle.

Backup and Restore is the simplest strategy. You take periodic backups (database snapshots, file system copies) and store them in a separate region. When disaster strikes, you provision new infrastructure, restore the backups, and redirect traffic. RTO is typically 4-24 hours depending on data volume and infrastructure complexity. RPO equals the time since the last backup. Cost is minimal because you only pay for backup storage, not running infrastructure.

Pilot Light keeps the absolute minimum infrastructure running in a DR region at all times. Typically this means a replicated database and perhaps a few core services. When disaster strikes, you scale up the remaining compute, deploy application servers, and switch traffic. RTO is 15-60 minutes. RPO depends on replication lag, usually minutes. Cost is low because you run only the data layer.

The half of an architecture that stays switched on, and what the other half costs in minutes to start.

The name "pilot light" comes from gas furnaces: a small flame that stays lit so the furnace can ignite quickly. In DR terms, the database replication is the pilot light. It is always burning (running), so when you need the full system, you are not starting from cold. The key insight is that data is the bottleneck in recovery. Compute is fast to provision (minutes with containers, seconds with pre-built AMIs). Data is slow to restore (hours for large databases). Pilot light keeps the slow part warm.

Warm Standby runs a scaled-down but complete copy of your production environment in the DR region. All services are running but at reduced capacity (perhaps 20-30% of production scale). When disaster strikes, you scale up to full capacity and switch traffic. RTO is 1-15 minutes. RPO is near-zero if replication is synchronous, seconds to minutes if asynchronous. Cost is moderate because you run the full stack at reduced scale. The advantage over pilot light is that scaling up existing instances is faster than launching new ones, and you can serve a portion of traffic during the scale-up window rather than serving nothing.

Active-Active runs full production environments in multiple regions simultaneously. Traffic is distributed across all regions using global load balancing. When one region fails, the others absorb the traffic with no failover needed because every region is already handling production traffic. RTO is near-zero. RPO is near-zero. Cost is the highest because you are running 2x or more infrastructure at full scale, plus the engineering complexity of handling multi-region writes, conflict resolution, and data consistency. Active-active is the gold standard for availability but the most complex to implement correctly. Most systems that claim to be active-active are actually active-passive with fast failover.

Choosing the Right Strategy

The right strategy depends on three factors: the business impact of downtime, the cost of the strategy, and the operational complexity your team can sustain. A startup with a 5-person engineering team should not attempt active-active multi-region deployments. The operational burden will consume more engineering time than the downtime it prevents. A payment processor that loses $100,000 per minute of downtime cannot afford a 4-hour RTO from backup-restore.

Map each of your systems to its RTO/RPO requirements individually. Your payment system might need warm standby while your admin dashboard is fine with backup-restore. Not every system needs the same level of protection.

The Cost-Availability Trade-off

DR is an insurance policy, and like all insurance, you are paying a known cost to protect against an uncertain loss. The challenge is right-sizing the premium.

A rough cost model for each strategy on a system spending $10,000/month on primary infrastructure:

Backup-restore adds approximately 5-10% ($500-1000/month) for backup storage and cross-region copies. Pilot light adds approximately 15-25% ($1500-2500/month) for a replicated database running continuously. Warm standby adds approximately 40-60% ($4000-6000/month) for a scaled-down full stack. Active-active adds approximately 100-150% ($10000-15000/month) for full duplicate infrastructure plus the engineering overhead of multi-region coordination.

These numbers help frame the conversation with stakeholders. "We can achieve an RTO of 30 minutes for $2000/month, or an RTO of 30 seconds for $12000/month. Which does the business need?" is a more productive discussion than debating DR architecture in abstract technical terms.

Common RTO/RPO Mistakes

The most common mistake is setting RTO and RPO without understanding the cost. A team writes "RPO: 0, RTO: 5 minutes" in a design document because it sounds professional, then implements daily backups because active-active is too expensive. The document says one thing, the implementation does another, and nobody discovers the gap until a disaster reveals it.

The second mistake is treating all data equally. A user's profile photo and a user's financial transaction have very different RPO requirements. The photo can be re-uploaded. The transaction cannot be recreated. Design your DR strategy to protect critical data paths more aggressively than non-critical ones.

The third mistake is measuring RTO from the start of recovery rather than from the start of the outage. RTO includes detection time (how long until you know there is a problem), decision time (how long until someone initiates failover), and recovery time (how long the failover takes). If your monitoring takes 10 minutes to alert, your decision process takes 15 minutes, and your failover takes 5 minutes, your actual RTO is 30 minutes even though the technical recovery is only 5 minutes. Automated failover reduces all three components but introduces the risk of false positives triggering unnecessary failovers.

Types of Disasters

DR planning must account for different categories of failure, each with different characteristics.

Single-component failures (one server, one disk, one process) are the most common and are typically handled by redundancy within a region: multiple instances behind a load balancer, RAID arrays, database replicas. These are not usually classified as "disasters" because the system self-heals.

Regional failures (entire AWS/GCP/Azure region becomes unavailable) are rare but have occurred multiple times for every major cloud provider. These require cross-region DR capabilities. The cause is typically a cascading failure in the control plane, a network partition, or physical infrastructure damage (power grid failure, cooling system failure).

Logical failures (data corruption from bugs, accidental deletion, schema migration gone wrong) are the most insidious because they are not detected by health checks. The system appears healthy but is serving wrong data. Replication faithfully copies the corruption everywhere. Only backups with sufficient retention protect against logical failures.

Security incidents (ransomware encrypting databases, attacker deleting resources, compromised credentials used to destroy infrastructure) require backups stored in separate accounts with separate credentials. If the attacker has root access to your primary AWS account, and your backups are in the same account, the attacker can destroy both.

Dependency failures (a critical third-party service goes down: your DNS provider, your payment processor, your CDN) are outside your control. You cannot failover your DNS provider. You can, however, reduce dependency risk by using multiple providers, implementing circuit breakers, and having fallback behavior (queue payments for retry if the processor is down, serve stale content if the CDN is unreachable).

Each category requires a different response. Component failures are automated. Regional failures require cross-region failover. Logical failures require point-in-time recovery from backups. Security incidents require immutable, isolated backups and a clean infrastructure rebuild. Dependency failures require redundancy at the service level and graceful degradation in your application.

Understanding which type of disaster you are protecting against determines which DR strategy is appropriate. A team that builds active-active replication to protect against logical errors has wasted their budget: replication faithfully copies the error to all regions. They needed better backups, not more replicas.