Skip to content
Business · 6 min read

Write the incident review to fix the system, not the person

The claim After an outage, the most tempting explanation is that someone made a mistake, and the most useless response is to make sure that person is more careful next time. People...

A Written by Administrator
Write the incident review to fix the system, not the person

The claim

After an outage, the most tempting explanation is that someone made a mistake, and the most useless response is to make sure that person is more careful next time. People are careful; they still make mistakes, because the capacity to make that particular mistake was built into the system and left there. An incident review that concludes with "we told the engineer to be more careful" has learned nothing and fixed nothing, and it guarantees the same failure will recur with a different name attached. The purpose of the review is to change the system so the mistake becomes impossible or harmless, not to identify who to blame.

Why blame destroys the information you need

The practical case against blame is not that it is unkind, though it is. It is that blame makes the review useless by destroying its inputs. The moment people understand that the outcome of an incident review is finding fault, they stop telling you what actually happened. The engineer who made the change stops volunteering the detail that would explain why it seemed correct at the time; the person who noticed the early warning sign stops mentioning that they were unsure whether to escalate. You end up with a sanitised account that protects people and teaches nobody, and the real chain of events — the one you needed to break — stays hidden because surfacing it is now dangerous.

A review conducted without blame gets the opposite: people explain their reasoning candidly, including the reasoning that turned out to be wrong, because there is no penalty for having been wrong in a way the system made easy. That candour is the entire value of the exercise. You cannot fix a failure you do not understand, and you cannot understand it if the people who were there are managing their exposure instead of telling you what happened.

The question that redirects the review

The discipline is to keep asking not "who did this" but "how was this possible". Every human error is an invitation to find the system that permitted it:

  • An engineer deployed a breaking change on a Friday afternoon. Why was it possible to deploy a change that breaks production at all? Where was the check?
  • Someone ran a destructive command against the wrong environment. Why do the production and staging environments look identical at the command line? Why does that command have no confirmation?
  • A configuration typo took the site down. Why did a syntactically valid but wrong config reach production without being caught? Where was the validation?

In each case, the human action is real, but it is the last link in a chain, and every earlier link is a place the system could have stopped the outcome and did not. Those earlier links are what the review exists to find, because they are what you can actually change.

What a useful review produces

The output of the review is not a narrative of who did what. It is a small number of specific, assigned, dated changes to the system, each of which would have prevented or reduced the incident regardless of who was at the keyboard:

Incident: production outage, 47 minutes
Contributing factors (not causes-of-blame):
- deploy tool allowed release with a failing health check
- staging and prod prompts visually identical
- no automated rollback; recovery was manual

Actions:
- [ ] deploy tool blocks release if health check fails   @alex  by Oct 20
- [ ] prod shell prompt shows red "PRODUCTION" banner     @sam   by Oct 15
- [ ] one-command rollback script, tested                 @alex  by Oct 24

Notice that none of the actions is "be more careful" and none names a fault. Each is a change to the system that makes the next occurrence of the same human action harmless. That is the test of a good action item: it works even if the person makes the exact same mistake again, because the point was never the person.

The distinction that keeps it honest

Removing blame is not removing accountability, and conflating the two is how teams either descend into finger-pointing or drift into consequence-free carelessness. Accountability in a healthy review is collective and forward-looking: the team owns the commitment to make the specific changes that prevent recurrence, and is answerable for whether those changes actually get made. What is removed is the backward-looking search for an individual to punish, which produces neither safety nor learning. The engineer who caused the outage is frequently the best person to help design the fix, because they understand the failure most intimately — and they will only do that freely in a review that is not trying to convict them.

Make the actions real

The most common way incident reviews fail is not blame but evaporation: the review happens, the actions are listed, and nothing is built because the actions have no owner and no date and compete with feature work that feels more urgent. An action item without a named owner and a deadline is a wish. Treat the review's actions as real work, tracked like any other commitment, with the same visibility, and revisit them — a review whose actions are still open when the next similar incident occurs has told you exactly what to prioritise, and also told you that your process for closing the loop is itself a system that needs fixing.

The standard to hold

Judge every incident review by one question: after this review, is the system different in a way that makes this failure less likely or less severe, regardless of who is operating it? If yes, the review did its job. If the only thing that changed is that someone feels bad and has resolved to be more careful, then nothing has changed, because carefulness was never the missing ingredient — the missing ingredient was a system that did not depend on a tired human getting every step right at the worst possible moment. Build that system, one incident at a time, and the reviews compound into genuine reliability instead of a recurring search for someone to blame.

#incident response #culture #reliability #process

Keep reading