Written by David Rodgers

Quality and Operations Perspective

Written by David Rodgers, Lean Six Sigma Black Belt and ASQ-certified quality leader. This guide applies quality and process-improvement methods to software and IT operations from a quality and operations perspective. The author is not a software engineer, site reliability engineer, or security professional.

Last editorial review: September 24, 2026. Educational content only: not medical, legal, or regulatory advice. Follow your organization's policies and the requirements that apply to you, and have subject-matter experts review any change to a live process.

  • Lean Six Sigma Black Belt
  • ASQ CQE
  • ASQ CMQ/OE
  • Quality systems and process improvement

Every complex system fails sometimes. What matters is how quickly the team detects and recovers, and whether each incident leaves the system safer. Blameless postmortems and service level objectives are the two most widely used tools for that.

This guide covers the incident lifecycle and its measures, how to set an SLO and compute an error budget, how to run a postmortem that looks at contributing factors instead of blame, and a worked example in which one long incident consumes more than four times a month's budget.

Get the Postmortem Template Open the Reliability Calculator

Before You Start

Educational content. This guide applies quality and process-improvement methods to software delivery and IT operations. It is not security, legal, compliance, or engineering advice. Practices, tools, and risks differ between teams and systems, so have qualified engineers and security professionals review changes to production systems and controls.

Why Incident Management and Postmortems Matter

Outages Will Happen

Complex systems fail. What separates teams is how fast they detect and recover, and whether they learn.

Blame Hides Causes

If people fear punishment, they hide details, and the real causes stay in place. A blameless approach gets the facts.

Repeat Incidents Are Waste

The same failure twice usually means the first postmortem produced no lasting change.

Reliability Is a Trade-Off

Service level objectives make the trade-off explicit, so teams can balance new features against stability.

The Incident Lifecycle and Its Measures

StageMeasureMeaning
DetectionTime to detect (MTTD)From the start of the problem to when it is noticed, ideally by monitoring, not by customers.
ResponseTime to acknowledge (MTTA)From the alert to when someone owns it.
MitigationTime to mitigateFrom detection to when customer impact is stopped, even if the root cause is not fixed.
RecoveryTime to restore (MTTR)Total time to restore normal service.
LearningPostmortem completed, actions closedWhether the organization improved.

Service Level Objectives and Error Budgets

A service level indicator (SLI) measures some aspect of service, such as the share of requests that succeed. A service level objective (SLO) sets a target for it, for example 99.9% availability over 30 days. The error budget is the allowed unreliability: 100% minus the SLO. The approach is described in Google's Site Reliability Engineering.

Error budget (minutes) = period minutes × (1 − SLO). For 30 days and 99.9%: 43,200 × 0.001 = 43.2 minutes

When the budget is spent, the team slows feature releases and invests in reliability until it recovers. When plenty of budget remains, it can take more risk. Choose an SLO that reflects what users need; a target higher than users notice is expensive without benefit.

The Blameless Postmortem

A postmortem is a written review of an incident that focuses on how the system and process allowed it, not on who made a mistake. John Allspaw's Blameless PostMortems and a Just Culture (2012) describes the idea: people usually act sensibly given what they knew at the time, so ask what made the action seem reasonable, and what would make the failure harder next time.

  • Timeline. Build it from logs, chat, and alerts. Include when the problem started, was detected, was acknowledged, was mitigated, and was resolved.
  • Contributing factors. List several, across tooling, process, communication, design, and monitoring. Avoid a single "root cause" that is really a person.
  • What went well. Note what worked, such as an alert or a runbook, so it is repeated.
  • Action items. Each has an owner, a due date, and a type: prevent, detect, mitigate, or learn. Track them to completion.

Record the analysis in the Blameless Postmortem Template.

Worked Example: A Month Against an SLO

A service has a 99.9% availability SLO over a 30-day month. It had five incidents of 12, 45, 8, 90, and 25 minutes. The numbers are illustrative.

MeasureCalculationResult
Total downtime12 + 45 + 8 + 90 + 25180 minutes
Availability1 − 180 / 43,20099.58%
Error budget43,200 × 0.00143.2 minutes
Budget consumed180 / 43.2about 4.2 times the budget
Mean time to restore180 / 536 minutes
0 50 100 150 200 Incident 1 12 min Incident 2 45 min Incident 3 8 min Incident 4 90 min Incident 5 25 min 12 57 65 155 180 Error budget 43.2 min Incident minutes in the month (bars) and cumulative total (line)
The budget is exhausted during the second incident, and by month's end the service has used more than four times what the SLO allowed.

What the team does. The 90-minute incident accounts for half of the downtime, so it gets a full postmortem. The timeline shows that detection took 22 minutes because the alert fired on CPU load, not on failed requests. Contributing factors also include a deployment without a canary stage and a runbook that was out of date. Actions: add an alert on request success rate (detect), introduce canary releases for this service (mitigate), and update and rehearse the runbook (prevent recurrence). Each has an owner and a date.

Because the budget is spent, the team also pauses non-essential feature releases for two weeks and puts reliability work first. The lesson is not that 99.9% was impossible, but that one long incident can dominate a month, and shortening detection and recovery time is often more valuable than preventing every failure.

Self-Assessment Questions

  • Do we detect incidents with monitoring before customers report them?
  • Do we have SLOs that reflect what users need, and do we track the error budget?
  • Do postmortems focus on system and process causes, not on blaming people?
  • Does every action item have an owner and a date, and do we close them?
  • Do we look for patterns across incidents, not only within one?

Common Mistakes

Blaming an Individual

"Human error" is where the investigation should start, not end. Ask what made the error easy to make.

Postmortems Without Actions

A document nobody acts on teaches nothing. Track actions like any other work.

SLOs Nobody Uses

If the error budget does not change decisions, it is decoration. Agree in advance what happens when it is spent.

Measuring Only MTTR

Mean time to restore hides the spread. Look at each stage of the timeline, and at the longest incidents.

Incident Management and Blameless Postmortems: Frequently Asked Questions

What is a blameless postmortem?

A blameless postmortem is a written review of an incident that focuses on how the system, tools, and processes allowed it, rather than on who made a mistake. It assumes people acted reasonably given what they knew, and asks what would make the failure harder or the recovery faster next time, ending with tracked action items.

What is an error budget?

An error budget is the amount of unreliability a service is allowed under its service level objective, equal to 100% minus the SLO. For a 99.9% SLO over 30 days, it is 43.2 minutes. When the budget is spent, the team prioritizes reliability work over new features until service recovers.

What is the difference between MTTD, MTTA, and MTTR?

MTTD is the mean time to detect a problem from when it starts, MTTA is the mean time to acknowledge an alert, and MTTR is the mean time to restore normal service, though some teams use R for repair or resolve. Definitions vary, so state yours, and look at each stage of the timeline because averages can hide long incidents.

Sources and Further Reading

  • Betsy Beyer, Chris Jones, Jennifer Petoff, and Niall Richard Murphy (eds.), Site Reliability Engineering, Google, 2016.
  • John Allspaw, "Blameless PostMortems and a Just Culture," Etsy Code as Craft, 2012.
  • Sidney Dekker, The Field Guide to Understanding 'Human Error'.
  • Nicole Forsgren, Jez Humble, and Gene Kim, Accelerate, on time to restore service.