Incident Response

Topics Covered

Incident Classification

Severity Levels

Detection Methods

Declaring an Incident

Severity Re-Classification During an Incident

Communication During Incidents

Real-World Severity Classification Examples

On-Call and Escalation

On-Call Rotations

On-Call Handoff

Incident Roles

The Response Lifecycle

Escalation Paths

On-Call Health and Burnout Prevention

Runbook Design

Postmortems and Learning

Blameless Postmortems

Timeline

Root Cause Analysis

The Five Whys

Contributing Factors

Action Items

Running the Postmortem Meeting

Postmortem Anti-Patterns

Incident Metrics and Trending

Chaos Engineering

Principles of Chaos Engineering

Common Chaos Experiments

Game Days

Learning from Incidents at Scale

Chaos Engineering Maturity Levels

Building a Chaos Engineering Culture

Chaos Engineering Tools

Connecting Chaos Engineering to Incident Response

Every production system breaks. The difference between a well-run organization and a chaotic one is not whether incidents happen, but how quickly the team classifies the problem, mobilizes the right people, and starts working toward resolution. Incident classification is the triage step that determines everything downstream: who gets paged, how fast they respond, and what communication channels activate.

Without a clear severity framework, teams waste critical minutes debating how bad things are instead of fixing them. A well-defined classification system turns that debate into a lookup table. Think of it like triage in an emergency room: a nurse does not debate each patient's condition from scratch. They check vital signs against predefined criteria and assign a priority level. The same principle applies to production incidents.

The classification system also determines how the organization tracks reliability over time. When every incident is tagged with a severity level, the team can measure trends: are SEV1 incidents decreasing quarter over quarter? Are SEV2 incidents concentrating in a specific service? Without consistent classification, this data is meaningless.

Some organizations add a fourth level, SEV0, for existential threats: events that could threaten the company's survival, such as a massive data breach, a compliance violation that could result in regulatory shutdown, or complete loss of customer data. SEV0 activates executive leadership in addition to engineering teams and may involve legal, PR, and compliance departments. The distinction between SEV0 and SEV1 is that SEV0 has consequences beyond the immediate technical failure.

Four severity levels with a real example on each, and the line below which nothing pages anybody at night.

Severity Levels

Most organizations use a three-tier or four-tier severity model. The most common framework uses SEV1 through SEV3, where each level defines the blast radius, the response urgency, and the team mobilization pattern.

SEV1 (Critical): Total service outage or data loss affecting all users. Revenue is actively being lost. Examples include the primary database going down, a payment processing pipeline failing, or a security breach exposing user data. Response expectation is all hands on deck within 5 minutes. The incident commander role activates immediately. External communication goes out within 15 minutes. Everything else stops until this is resolved.

SEV2 (Major): Significant degradation affecting a large subset of users, but the service is not completely down. Examples include one region's servers failing while others remain healthy, search functionality broken but browsing works, or latency spiked to 10x normal. Response expectation is on-call engineer plus escalation to the team lead within 15 minutes. If not resolved within 30 minutes, escalate to SEV1 procedures.

SEV3 (Minor): Limited impact affecting a small percentage of users or a non-critical feature. Examples include a dashboard rendering incorrectly, a batch job failing that can be rerun, or elevated error rates on a single endpoint that has retry logic. Response expectation is on-call engineer investigates during business hours. No immediate escalation needed unless the problem worsens.

Interview Tip

In interviews, when discussing incident classification, always tie severity to business impact rather than technical symptoms. A database failover that completes in 2 seconds is not a SEV1 even though the database went down. A CSS bug that hides the checkout button on mobile is a SEV1 even though no server is failing. Impact on users and revenue determines severity, not the sophistication of the technical failure.

Detection Methods

Incidents surface through three channels, and mature organizations invest in all three because no single channel catches everything.

Automated alerting is the first line of defense. Monitoring systems (Datadog, PagerDuty, Prometheus with Alertmanager) watch metrics like error rates, latency percentiles, CPU utilization, and queue depths. When a metric crosses a threshold, an alert fires and pages the on-call engineer. The key design challenge is alert quality: too many false positives cause alert fatigue, and engineers start ignoring pages. Too few alerts mean real incidents go undetected. Good alert design uses multi-signal confirmation (error rate AND latency both elevated) and suppresses known-noisy alerts during deployments.

User reports catch problems that metrics miss. A subtle data corruption bug might not spike error rates, but users notice their data is wrong and file support tickets. Organizations that pipe support ticket trends into their incident detection systems catch these issues hours earlier than those that treat support and engineering as separate silos.

Dashboard monitoring fills the gap between automated alerts (which fire on thresholds) and user reports (which arrive slowly). Engineers watching real-time dashboards during deployments or high-traffic events can spot anomalies before they cross alerting thresholds. The traffic pattern looks unusual, conversion rates dipped slightly, or a new error type appeared at low volume. This manual monitoring is most valuable during known risk windows.

Declaring an Incident

The moment someone suspects a SEV1 or SEV2 issue, they declare an incident. This is a deliberate, low-friction action: create a dedicated communication channel (Slack channel, Zoom bridge), page the incident commander, and start the incident clock. The threshold for declaration should be low. Declaring an incident that turns out to be a false alarm costs 30 minutes of overhead. Not declaring an incident that turns out to be real costs hours of uncoordinated debugging.

Many organizations use an incident bot that automates this process: a single /incident command creates the channel, pages the on-call, starts the timeline log, and posts the initial status update. Reducing the friction of declaration to a single command removes the "is this really worth an incident?" hesitation that delays response.

Severity Re-Classification During an Incident

Severity is not static. An incident can escalate or de-escalate as new information emerges. A SEV3 that started as "one endpoint returning occasional 500s" might escalate to SEV2 when the team discovers the endpoint serves the checkout flow and the error rate is climbing. Conversely, a SEV1 declared during a full outage might be downgraded to SEV2 once the team confirms that only one region is affected and traffic has been rerouted.

The re-classification triggers should be explicit. Escalation triggers include: the blast radius is expanding, the estimated time to resolution exceeds the target for the current severity level, or a new dimension of impact is discovered (data loss in addition to availability loss). De-escalation triggers include: a workaround is found that restores most user functionality, or the affected population is smaller than initially estimated.

Teams that do not practice re-classification tend to either leave everything at SEV1 (wasting resources on minor issues) or leave everything at SEV3 (under-responding to growing problems). Dynamic classification keeps the response proportional to the actual impact at every point during the incident.

The person who declares the re-classification should state the reason explicitly in the incident channel: "Escalating from SEV2 to SEV1 because the error rate is now affecting all regions, not just us-east-1." This creates an audit trail that is invaluable during the postmortem. It also prevents confusion: without an explicit announcement, some team members may still be operating under the original severity assumptions, leading to mismatched urgency levels within the response team.

Communication During Incidents

Effective incident communication follows a cadence, not ad-hoc updates. For SEV1 incidents, status updates go out every 15 minutes, even if the update is "still investigating, no new information." Silence during an incident is worse than repetitive updates because stakeholders interpret silence as "nobody is working on this."

Status page updates should be honest about impact without creating panic. Good messaging acknowledges the problem, states what users might experience, and provides an expected time for the next update. Avoid technical jargon in customer-facing communication. "We are experiencing elevated error rates on our payment processing service" is better than "The payment-service pods are OOMKilling due to a memory leak in the gRPC handler."

Internal communication follows a different pattern. Engineering teams need technical detail: which services are affected, what hypotheses are being investigated, which dashboards to watch. Product and support teams need impact summaries: how many users are affected, what workarounds exist, and when the fix is expected. The communications lead routes the right level of detail to each audience.

Real-World Severity Classification Examples

Understanding severity classification becomes concrete with real-world examples from well-known incidents.

In 2017, Amazon S3 experienced a four-hour outage that took down thousands of websites and services. This was a clear SEV1: a single engineer's typo during routine maintenance removed more servers than intended, cascading into a regional outage. The blast radius was massive because so many services depended on S3. The key lesson was that severity is determined not just by the failing service but by its position in the dependency graph. A seemingly minor service that sits at the bottom of the stack can cause catastrophic cascading failures.

Contrast this with a typical SEV3: a batch processing job that generates weekly analytics reports fails at 2 AM on a Tuesday. No users notice. No revenue is affected. The data is not lost, just not processed. The on-call engineer re-runs the job during business hours. This is a SEV3 because the impact is limited to an internal reporting delay with no user-facing consequences.

The gray area is where classification gets interesting. Consider a mobile push notification system that stops sending notifications. No errors are visible to users, no pages load slowly, and the app works perfectly. But engagement metrics will drop over the next 24 hours as users stop being reminded to check the app. Is this a SEV2 (significant business impact over time) or a SEV3 (no immediate user-facing issue)? The answer depends on how the organization values engagement-driven revenue versus direct transaction revenue. These edge cases are where a well-documented severity framework earns its value.