Designing for Failure

Topics Covered

Partial Failure in Distributed Systems

Byzantine vs Non-Byzantine Faults

Graceful Degradation

The Failure Dependency Map

Timeouts and Retries

Timeout Tuning

Retry Strategies

Why Jitter Matters

Retry Budgets

Distinguishing Retryable from Non-Retryable Errors

Idempotency Keys for APIs

Why the Client Generates the Key

Server-Side Implementation

Edge Cases

Real-World Idempotency Key Implementations

The Saga Pattern

How It Works

Choreography vs Orchestration

When Sagas Get Complex

Saga State Persistence

Chaos Engineering

Netflix's Chaos Monkey

Game Days

The Experiment Loop

Building Confidence Through Controlled Experiments

Common Chaos Engineering Tools

In a single-process application, failure is binary: the process either works or it crashes. Distributed systems break this assumption. A request from Service A to Service B might succeed at B, but the response never arrives back at A. Service C might be healthy while Service D is down. The database accepts the write but the cache update fails. This is partial failure: some components fail while others continue operating, and the failing components may not even know they have failed.

Partial failure is not an edge case. It is the normal operating mode of any system running across multiple machines. Networks drop packets, disks fill up, garbage collection pauses make services unresponsive for seconds, and cloud providers occasionally lose entire availability zones.

It is also arithmetically unavoidable once a request touches more than a couple of services. If a request needs every one of nn dependencies to succeed, their availabilities multiply:

Asystem=∏i=1nAiA_{\text{system}} = \prod_{i=1}^{n} A_i

Ten dependencies at a respectable 99.9 percent each give a system availability of 99.0 percent, which is 87 hours of downtime a year rather than the 8.8 hours any single component promises. Thirty of them give 97.0 percent, or 259 hours. Nobody wrote a bug to produce that number; it is the direct consequence of a fan-out where every call is required. The escape is not more reliable components, which buys a factor at growing cost, but removing the requirement that every call succeed: make the non-essential ones optional, cache their last good answer, or move them off the request path entirely.

Why is partial failure harder than total failure? Because it creates ambiguity. When everything is down, you know there is a problem. When one service out of twenty is degraded, you might not notice until a customer reports inconsistent data. The request might have partially succeeded: the order was created but the notification was not sent, the payment was charged but the inventory was not reserved. This ambiguity is what makes designing for failure fundamentally different from designing for correctness.

Designing for failure means accepting that partial failures will happen and building systems that continue serving users when they do.

One slow dependency filling thread pools upstream until four healthy services go down, and where the chain could be cut.

Byzantine vs Non-Byzantine Faults

Distributed systems literature distinguishes two categories of failure. Non-Byzantine (crash-stop) faults are when a component stops responding entirely: it crashes, loses power, or gets disconnected from the network. The component is either working correctly or completely silent. Most failures in well-managed data centers fall into this category: servers crash, processes get killed by the OS, network cables get unplugged.

Byzantine faults are far worse: a component continues operating but produces incorrect or contradictory results. A corrupted disk returns wrong data. A buggy service sends malformed responses that look valid. A compromised node deliberately lies. Byzantine fault tolerance is expensive to implement (it requires 3f+1 nodes to tolerate f faulty nodes) and is rarely needed outside of blockchain systems, aerospace, or nuclear safety. For most web-scale systems, designing for crash-stop faults is sufficient.

The practical implication: when a service stops responding, you can safely assume it has crashed or lost connectivity. You do not need to worry about it actively sending you wrong data. This simplification lets you use much simpler failure detection (heartbeats, timeouts) instead of full Byzantine consensus.

Gray failures sit between these two categories. A service is technically running and responding, but its responses are degraded: elevated latency, increased error rates, or intermittently correct results. Gray failures are harder to detect than crash-stop faults because health checks may still pass while user-facing traffic suffers. A database with a failing disk might serve reads from cache (healthy) while silently dropping writes (broken). Monitoring that only checks "is the process alive?" misses gray failures entirely. Effective detection requires checking response quality, not just response presence.

Graceful Degradation

The goal of failure-tolerant design is not to prevent all failures. It is to ensure that when components fail, the system degrades gracefully instead of collapsing entirely.

Features sorted into degradation tiers ahead of time, with the fallback, the user-visible effect, and the shedding order of each.

Serve partial results. When a product page depends on the product catalog, the recommendation engine, and the reviews service, a failure in recommendations should not prevent the page from loading. Return the product details and reviews, and replace the recommendation section with a generic "Popular items" fallback. The user still gets a useful page.

Queue for later. If the email service is down when a user places an order, do not fail the entire checkout. Write the email job to a persistent queue and process it when the service recovers. The user gets their order confirmation 5 minutes late instead of getting a checkout error.

Show cached data. If the pricing service is slow, serve the last known prices from cache with a "prices may be slightly outdated" notice. Stale data is better than no data for most read-heavy use cases.

Shed non-critical load. Under extreme pressure, disable non-essential features first. Turn off personalization, analytics tracking, and A/B test evaluation before touching core functionality like search and checkout. Feature flags make this possible: a single configuration change disables a feature without a code deployment.

Level Expectations

Mid-level engineers handle individual failure cases reactively. Senior engineers design systems with failure budgets and degradation tiers defined upfront. Staff engineers establish organization-wide resilience patterns, define SLOs that account for partial failure, and ensure every service has a documented degradation strategy before it ships.

The Failure Dependency Map

A practical tool for designing degradation strategies is a failure dependency map: for each external dependency your service calls, document three things. First, what is the expected failure mode (timeout, error response, data corruption)? Second, what is the user-visible impact if this dependency is unavailable? Third, what is the fallback behavior?

For example, a checkout service might depend on the pricing API, the tax calculator, the inventory service, and the fraud detection engine. If the fraud detection engine is down for 30 seconds, do you block all purchases (safe but costly) or temporarily allow purchases below a dollar threshold with a flag for post-hoc review (risky but maintains revenue)? There is no universal right answer. The right answer depends on your business context, and the point of the failure dependency map is to force that conversation before the outage happens, not during it.

Services without a documented failure dependency map default to the worst possible degradation strategy: total failure on any dependency outage. The key mindset shift is moving from "how do I prevent failure?" to "what does the user experience look like when this component is down?" Every external dependency in your system should have an answer to that question before it reaches production.