Most engineering organizations write postmortems. Far fewer can say what happened to the action items from last quarter's postmortems. The document gets written, the meeting happens, a handful of tickets are filed, and then the next incident arrives and the tickets slide down the backlog. Months later a similar failure recurs, and someone finds the old postmortem that predicted it.
This article covers the two halves of a postmortem process that actually reduces repeat incidents: writing it blamelessly, so people tell you what really happened, and closing the loop on action items, so what you learned changes the system. It draws on the postmortem chapters of Google's SRE book and SRE Workbook and on PagerDuty's publicly documented incident response process.
The SRE book defines a postmortem as "a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root cause(s), and the follow-up actions to prevent the incident from recurring" (Google SRE book, Postmortem Culture). It names the primary goals as documenting the incident, understanding all contributing causes and, "especially", putting effective preventive actions in place.
That last word matters. A postmortem that produces an accurate narrative but no change in the system is a history lesson. The SRE Workbook quotes Ben Treynor Sloss on exactly this point: "To our users, a postmortem without subsequent action is indistinguishable from no postmortem" (Google SRE Workbook, Postmortem Culture).
If the decision to write a postmortem is made after each incident, it becomes a negotiation, and uncomfortable incidents tend to be the ones that escape review. The SRE book recommends defining postmortem criteria before an incident occurs, and lists common triggers:
It also notes that any stakeholder may request a postmortem for an event. If you already run SLOs, an error-budget threshold is a natural trigger: an incident that consumed a meaningful share of the budget gets a postmortem automatically. Our guide to SLO burn-rate alerting covers how to measure that consumption.
The SRE book is direct about what blameless means: the postmortem "must focus on identifying the contributing causes of the incident without indicting any individual or team for bad or inappropriate behavior", and it "assumes that everyone involved in an incident had good intentions and did the right thing with the information they had." The reason is practical rather than kind. Where people are shamed for mistakes, the book warns, they "will not bring issues to light for fear of punishment", and incidents get swept under the rug.
Blameless does not mean vague. The same chapter says a postmortem should still call out where and how services can be improved. The shift is in the question being asked: not "who did the wrong thing?" but "why did the system allow, or even encourage, a reasonable person to do this?"
A useful test for any contributing cause: if the fix is "be more careful", you have not finished the analysis. The Workbook puts it plainly: "trying to change human behavior is less reliable than changing automated systems and processes." Keep asking why until you reach something you can change in tooling, defaults, guardrails, alerting or process. A command that can wipe a production cluster with no confirmation is a system problem, whoever typed it.
PagerDuty's published postmortem process is a good concrete template because it assigns clear responsibilities and deadlines:
The SRE book adds a review step that many teams skip: "An unreviewed postmortem might as well never have existed." Its suggested review criteria include whether the impact assessment is complete, whether the root cause is sufficiently deep, and whether the action plan is appropriate with fixes at the right priority. A senior engineer outside the affected team makes a good reviewer, because they will ask the questions insiders have stopped noticing.
If you use distributed tracing, pull traces for the incident window into the timeline while they are still retained. Our OpenTelemetry production guide covers the instrumentation side.
The SRE Workbook's contrast between a poor and a good postmortem is most useful in its treatment of action items. The poor example fails in recognizable ways: items that are mostly mitigative with little prevention, one "preventative" item asking to make humans less error-prone, every item tagged with equal priority, vague verbs like "improve" and "make better", and only one item with a tracking bug. The Workbook's conclusion is blunt: "Without a formal tracking process, action items from postmortems are often forgotten, resulting in outages."
The good example's action items share five characteristics, which make a practical checklist:
The Workbook's example also tags each item by type: investigate, prevent, mitigate, repair. That labelling is worth copying. When a postmortem's items are all "repair", the team has cleaned up but not reduced the chance of recurrence.
PagerDuty's guidance warns against creating too many tickets and generally limits follow-ups to P0 and P1 work that absolutely should be dealt with. That discipline matters. A postmortem with twenty action items tells the organization that nothing in particular is important, and the list decays together. Record lower-priority ideas in the document itself as observations, and file tickets only for the items you intend to finish.
Everything above can be done well and the loop can still stay open. Closing it takes a small amount of recurring process, owned by someone with the standing to push back on roadmap pressure.
Many preventative action items are guardrails: confirmation steps for destructive operations, input validation on admin tools, automatic rollback on failed health checks. These are good candidates for automation, provided the automation is itself tested and its actions are logged. Our article on self-healing infrastructure discusses where automated remediation is appropriate and where a human decision should stay in the loop.
The SRE book's goal is to share postmortems "to the widest possible audience that would benefit", and PagerDuty's process ends by communicating results and key learnings internally. Other teams often run the same patterns and carry the same latent risks. A short internal summary, covering what happened, what the contributing causes were and what changed, spreads the lesson further than the full document ever will. Keep anything that could identify end users out of shared documents.
If you are formalizing incident response alongside wider platform work, our automation practice covers runbooks and remediation tooling.
No. Blameless means the analysis looks for contributing causes in systems and processes rather than punishing individuals for reasonable decisions. Accountability moves to the action items, each of which has a named owner and a verifiable end state.
Soon enough that memories and logs are fresh. PagerDuty's documented process schedules the meeting within 3 calendar days for a SEV-1 and 5 business days for a SEV-2. Whatever you choose, set the deadline in advance rather than per incident.
As few as will genuinely reduce recurrence. PagerDuty generally limits tickets to P0 and P1 work. A short list of owned, measurable items is far more likely to be completed than a long list of good intentions.