0%
Cloud Architecture Patterns
Cloud Foundations
Compute Patterns
Storage and Databases
Application Patterns
Infrastructure as Code
Advanced Patterns
High Availability Patterns
Availability answers a simple question: when a user sends a request, does the system respond correctly? It is expressed as a ratio of successful responses to total requests, or equivalently, uptime divided by total time. A system that is available 99.9% of the time sounds almost perfect, but the gap between 99.9% and 99.99% is enormous in practice.
The Nines of Availability
Availability is measured in "nines." Each additional nine cuts allowed downtime by a factor of ten:
- 99% (two nines): 3.65 days of downtime per year. Acceptable for internal tools that are only used during business hours.
- 99.9% (three nines): 8.7 hours per year. This is the baseline for most production web applications.
- 99.99% (four nines): 52 minutes per year. Required for payment systems, authentication services, and other critical infrastructure.
- 99.999% (five nines): 5.2 minutes per year. Reserved for core infrastructure like DNS, load balancers, and databases that everything else depends on.
The cost of each additional nine is not linear. Going from 99.9% to 99.99% might require adding a standby database, geographic redundancy, and automated failover. Going from 99.99% to 99.999% might require cell-based architecture, active-active deployments across multiple regions, and a dedicated reliability engineering team. Each nine roughly doubles your infrastructure and operational cost.
In interviews, do not promise five nines unless the problem specifically demands it. State the availability target explicitly, justify it with business requirements, and explain the tradeoffs you are making to achieve it. Saying 99.99% and explaining why is far more impressive than saying 99.999% without understanding the cost.
Serial and Parallel Component Availability
Real systems are composed of multiple components. How you combine them determines overall availability.
Serial components (all must work for the system to work): multiply their individual availabilities. A system with a web server at 99.9% and a database at 99.9% has combined availability of 99.9% x 99.9% = 99.8%. Each component you add to the serial chain makes the system less available. A chain of five components each at 99.9% drops to 99.5%.
Parallel components (any one working is sufficient): use the formula 1 - (1 - A1)(1 - A2). Two databases in parallel, each at 99.9%, give you 1 - (0.001 x 0.001) = 99.9999%. Redundancy is how you beat the multiplication penalty of serial chains.
This is why high availability architecture is fundamentally about redundancy. Every critical component needs at least one backup that can take over without human intervention. The art is in deciding which components get redundancy, what kind of redundancy they get, and how fast the switchover happens.
The critical assumption in the parallel availability formula is independence: both components must fail for independent reasons. If two database replicas share the same power supply and that power supply fails, both go down together. The formula 1 - (1-A1)(1-A2) no longer applies because the failures are correlated. True redundancy means eliminating shared failure modes: separate power supplies, separate network switches, separate racks, and for the highest availability requirements, separate data centers.
Measuring Availability in Practice
Production availability is not measured by whether the server process is running. It is measured by whether users are getting correct responses within acceptable latency. A server that is running but returning errors is not available. A server that is running but responding in 30 seconds instead of 300 milliseconds is arguably not available either.
Modern systems track availability through success rate: the percentage of requests that return a non-error response within the latency SLO. This is more accurate than uptime monitoring because it captures partial failures, performance degradation, and error spikes that traditional "is the process alive" checks would miss entirely.
Why Availability Is Not Free
Every unit of availability requires investment. Two nines (99%) means you can tolerate the occasional multi-hour outage, so a single server with manual restart might suffice. Three nines (99.9%) means you need automated restarts and probably a standby instance. Four nines (99.99%) means you need redundancy at every layer, automated failover, and zero-downtime deployments. Five nines (99.999%) means every component is redundant, every failover is automatic and tested regularly, and you have a dedicated team monitoring everything around the clock.
The cost-benefit calculation is straightforward: how much does one minute of downtime cost your business? An e-commerce site processing $10 million per day loses roughly $7,000 per minute of downtime. For that business, spending $500,000 per year to move from 99.9% to 99.99% (saving approximately 7.8 hours of downtime) is clearly justified. An internal wiki used by 20 engineers might cause $50 per minute in lost productivity. Spending the same $500,000 to protect it makes no sense. Match your availability investment to your downtime cost.
Planned vs Unplanned Downtime
Availability budgets must account for both planned and unplanned downtime. Planned downtime includes maintenance windows, database migrations, and infrastructure upgrades. Unplanned downtime includes crashes, network partitions, and software bugs.
Many teams mistakenly exclude planned maintenance from their availability calculations. From the user's perspective, the system is unavailable regardless of whether the cause is a planned migration or an unexpected crash. A 99.9% SLO with monthly 2-hour maintenance windows is effectively a 99.6% SLO. If maintenance windows are necessary, schedule them during lowest-traffic periods and use techniques like rolling restarts and blue-green deployments to minimize user impact. The gold standard is zero-downtime maintenance, which is achievable for most web applications but requires investment in deployment tooling and database migration strategies that work without locking tables.
Common Availability Anti-Patterns
Several patterns look like they improve availability but actually undermine it:
Single points of failure hiding behind redundancy: Two database replicas in parallel look redundant, but if they share a single network switch, losing that switch takes both down. True redundancy requires eliminating correlated failure modes at every level: power, network, rack, and data center.
Over-relying on monitoring instead of prevention: Monitoring tells you when something broke. It does not prevent breakage. Teams that invest heavily in dashboards but not in redundancy, circuit breakers, and graceful degradation are optimizing for fast detection rather than fast recovery. Detection without automated recovery means a human must respond, which adds minutes to hours of downtime.
Treating all components as equally critical: If every service is marked as "critical" and gets the same availability investment, you spread resources thinly across everything. In practice, 20% of services handle 80% of revenue-impacting requests. Focus your highest availability investments on those services and accept lower availability for everything else.
Ignoring correlated failures: Running two replicas in the same availability zone feels redundant, but a zone-wide failure (power outage, network partition, cooling failure) takes both down simultaneously. Distributing replicas across availability zones eliminates this correlation. For the highest availability, distribute across regions, accepting the latency cost of cross-region replication.
Confusing uptime with availability: A process running with 100% uptime but returning 50% errors has 50% availability. Uptime is a necessary but insufficient condition for availability. Always measure availability from the user's perspective: did the request succeed within the latency target? This reframes availability as a user experience metric, not an infrastructure metric.
Understanding these anti-patterns is as important as understanding the patterns themselves. In interviews, pointing out what not to do demonstrates deeper understanding than simply listing what to do. It shows you have seen real-world systems fail and understand why common assumptions about availability break down under production conditions.