Every complex system fails sometimes. What matters is how quickly the team detects and recovers, and whether each incident leaves the system safer. Blameless postmortems and service level objectives are the two most widely used tools for that.
This guide covers the incident lifecycle and its measures, how to set an SLO and compute an error budget, how to run a postmortem that looks at contributing factors instead of blame, and a worked example in which one long incident consumes more than four times a month's budget.
Before You Start
Why Incident Management and Postmortems Matter
Outages Will Happen
Complex systems fail. What separates teams is how fast they detect and recover, and whether they learn.
Blame Hides Causes
If people fear punishment, they hide details, and the real causes stay in place. A blameless approach gets the facts.
Repeat Incidents Are Waste
The same failure twice usually means the first postmortem produced no lasting change.
Reliability Is a Trade-Off
Service level objectives make the trade-off explicit, so teams can balance new features against stability.
The Incident Lifecycle and Its Measures
| Stage | Measure | Meaning |
|---|---|---|
| Detection | Time to detect (MTTD) | From the start of the problem to when it is noticed, ideally by monitoring, not by customers. |
| Response | Time to acknowledge (MTTA) | From the alert to when someone owns it. |
| Mitigation | Time to mitigate | From detection to when customer impact is stopped, even if the root cause is not fixed. |
| Recovery | Time to restore (MTTR) | Total time to restore normal service. |
| Learning | Postmortem completed, actions closed | Whether the organization improved. |
Service Level Objectives and Error Budgets
A service level indicator (SLI) measures some aspect of service, such as the share of requests that succeed. A service level objective (SLO) sets a target for it, for example 99.9% availability over 30 days. The error budget is the allowed unreliability: 100% minus the SLO. The approach is described in Google's Site Reliability Engineering.
When the budget is spent, the team slows feature releases and invests in reliability until it recovers. When plenty of budget remains, it can take more risk. Choose an SLO that reflects what users need; a target higher than users notice is expensive without benefit.
The Blameless Postmortem
A postmortem is a written review of an incident that focuses on how the system and process allowed it, not on who made a mistake. John Allspaw's Blameless PostMortems and a Just Culture (2012) describes the idea: people usually act sensibly given what they knew at the time, so ask what made the action seem reasonable, and what would make the failure harder next time.
- Timeline. Build it from logs, chat, and alerts. Include when the problem started, was detected, was acknowledged, was mitigated, and was resolved.
- Contributing factors. List several, across tooling, process, communication, design, and monitoring. Avoid a single "root cause" that is really a person.
- What went well. Note what worked, such as an alert or a runbook, so it is repeated.
- Action items. Each has an owner, a due date, and a type: prevent, detect, mitigate, or learn. Track them to completion.
Record the analysis in the Blameless Postmortem Template.
Worked Example: A Month Against an SLO
A service has a 99.9% availability SLO over a 30-day month. It had five incidents of 12, 45, 8, 90, and 25 minutes. The numbers are illustrative.
| Measure | Calculation | Result |
|---|---|---|
| Total downtime | 12 + 45 + 8 + 90 + 25 | 180 minutes |
| Availability | 1 − 180 / 43,200 | 99.58% |
| Error budget | 43,200 × 0.001 | 43.2 minutes |
| Budget consumed | 180 / 43.2 | about 4.2 times the budget |
| Mean time to restore | 180 / 5 | 36 minutes |
What the team does. The 90-minute incident accounts for half of the downtime, so it gets a full postmortem. The timeline shows that detection took 22 minutes because the alert fired on CPU load, not on failed requests. Contributing factors also include a deployment without a canary stage and a runbook that was out of date. Actions: add an alert on request success rate (detect), introduce canary releases for this service (mitigate), and update and rehearse the runbook (prevent recurrence). Each has an owner and a date.
Because the budget is spent, the team also pauses non-essential feature releases for two weeks and puts reliability work first. The lesson is not that 99.9% was impossible, but that one long incident can dominate a month, and shortening detection and recovery time is often more valuable than preventing every failure.
Self-Assessment Questions
- Do we detect incidents with monitoring before customers report them?
- Do we have SLOs that reflect what users need, and do we track the error budget?
- Do postmortems focus on system and process causes, not on blaming people?
- Does every action item have an owner and a date, and do we close them?
- Do we look for patterns across incidents, not only within one?
Common Mistakes
Blaming an Individual
"Human error" is where the investigation should start, not end. Ask what made the error easy to make.
Postmortems Without Actions
A document nobody acts on teaches nothing. Track actions like any other work.
SLOs Nobody Uses
If the error budget does not change decisions, it is decoration. Agree in advance what happens when it is spent.
Measuring Only MTTR
Mean time to restore hides the spread. Look at each stage of the timeline, and at the longest incidents.
Incident Management and Blameless Postmortems: Frequently Asked Questions
What is a blameless postmortem?
A blameless postmortem is a written review of an incident that focuses on how the system, tools, and processes allowed it, rather than on who made a mistake. It assumes people acted reasonably given what they knew, and asks what would make the failure harder or the recovery faster next time, ending with tracked action items.
What is an error budget?
An error budget is the amount of unreliability a service is allowed under its service level objective, equal to 100% minus the SLO. For a 99.9% SLO over 30 days, it is 43.2 minutes. When the budget is spent, the team prioritizes reliability work over new features until service recovers.
What is the difference between MTTD, MTTA, and MTTR?
MTTD is the mean time to detect a problem from when it starts, MTTA is the mean time to acknowledge an alert, and MTTR is the mean time to restore normal service, though some teams use R for repair or resolve. Definitions vary, so state yours, and look at each stage of the timeline because averages can hide long incidents.
Sources and Further Reading
- Betsy Beyer, Chris Jones, Jennifer Petoff, and Niall Richard Murphy (eds.), Site Reliability Engineering, Google, 2016.
- John Allspaw, "Blameless PostMortems and a Just Culture," Etsy Code as Craft, 2012.
- Sidney Dekker, The Field Guide to Understanding 'Human Error'.
- Nicole Forsgren, Jez Humble, and Gene Kim, Accelerate, on time to restore service.