Reliability

Blameless Postmortems and Action Items

ARQQ · September 28, 2026 · 9 min read

Most engineering organizations write postmortems. Far fewer can say what happened to the action items from last quarter's postmortems. The document gets written, the meeting happens, a handful of tickets are filed, and then the next incident arrives and the tickets slide down the backlog. Months later a similar failure recurs, and someone finds the old postmortem that predicted it.

This article covers the two halves of a postmortem process that actually reduces repeat incidents: writing it blamelessly, so people tell you what really happened, and closing the loop on action items, so what you learned changes the system. It draws on the postmortem chapters of Google's SRE book and SRE Workbook and on PagerDuty's publicly documented incident response process.

What a postmortem is for

The SRE book defines a postmortem as "a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root cause(s), and the follow-up actions to prevent the incident from recurring" (Google SRE book, Postmortem Culture). It names the primary goals as documenting the incident, understanding all contributing causes and, "especially", putting effective preventive actions in place.

That last word matters. A postmortem that produces an accurate narrative but no change in the system is a history lesson. The SRE Workbook quotes Ben Treynor Sloss on exactly this point: "To our users, a postmortem without subsequent action is indistinguishable from no postmortem" (Google SRE Workbook, Postmortem Culture).

Decide in advance when a postmortem is required

If the decision to write a postmortem is made after each incident, it becomes a negotiation, and uncomfortable incidents tend to be the ones that escape review. The SRE book recommends defining postmortem criteria before an incident occurs, and lists common triggers:

  • User-visible downtime or degradation beyond a set threshold.
  • Data loss of any kind.
  • On-call intervention such as a release rollback or rerouting traffic.
  • Resolution time above a set threshold.
  • A monitoring failure, which usually means the incident was discovered manually.

It also notes that any stakeholder may request a postmortem for an event. If you already run SLOs, an error-budget threshold is a natural trigger: an incident that consumed a meaningful share of the budget gets a postmortem automatically. Our guide to SLO burn-rate alerting covers how to measure that consumption.

Writing it blamelessly

The SRE book is direct about what blameless means: the postmortem "must focus on identifying the contributing causes of the incident without indicting any individual or team for bad or inappropriate behavior", and it "assumes that everyone involved in an incident had good intentions and did the right thing with the information they had." The reason is practical rather than kind. Where people are shamed for mistakes, the book warns, they "will not bring issues to light for fear of punishment", and incidents get swept under the rug.

Blameless does not mean vague. The same chapter says a postmortem should still call out where and how services can be improved. The shift is in the question being asked: not "who did the wrong thing?" but "why did the system allow, or even encourage, a reasonable person to do this?"

Language that helps

  • Describe actions, not character. "The deploy was run against the production cluster" rather than "the engineer carelessly deployed to production".
  • Explain the information available at the time. What did the dashboard show? What did the runbook say? Was the warning visible, or buried in a log?
  • Use roles, not names, in the analysis. The SRE Workbook's worked example of a poor postmortem singles out named individuals in its causes and action items, and the Workbook explains that this leads people to become risk-averse for fear of being publicly shamed.
  • Replace drama with data. The Workbook also cautions against animated language and exclamations about luck, and asks for verifiable data to justify statements of severity.

Stop at systems, not people

A useful test for any contributing cause: if the fix is "be more careful", you have not finished the analysis. The Workbook puts it plainly: "trying to change human behavior is less reliable than changing automated systems and processes." Keep asking why until you reach something you can change in tooling, defaults, guardrails, alerting or process. A command that can wipe a production cluster with no confirmation is a system problem, whoever typed it.

Running the process: owner, timeline, review

PagerDuty's published postmortem process is a good concrete template because it assigns clear responsibilities and deadlines:

  • A single owner is designated by the incident commander at the end of the incident call or shortly after. The owner populates the document, gathers logs, manages the follow-up investigation and keeps interested parties informed.
  • A fixed meeting deadline: PagerDuty schedules the postmortem meeting within 3 calendar days for a SEV-1 and 5 business days for a SEV-2, and books it before the document is filled in so it is on the calendar.
  • Timeline first. The timeline is the initial focus, with each entry linked to the metric, graph or log search that supports it, and the commands used to gather data recorded so others can see how it was obtained.
  • A short meeting. PagerDuty's meetings generally last 15 to 30 minutes and are a wrap-up: the goal is to surface disagreement on facts, analysis or actions, not to write the document live.

The SRE book adds a review step that many teams skip: "An unreviewed postmortem might as well never have existed." Its suggested review criteria include whether the impact assessment is complete, whether the root cause is sufficiently deep, and whether the action plan is appropriate with fixes at the right priority. A senior engineer outside the affected team makes a good reviewer, because they will ask the questions insiders have stopped noticing.

If you use distributed tracing, pull traces for the incident window into the timeline while they are still retained. Our OpenTelemetry production guide covers the instrumentation side.

Writing action items that get done

The SRE Workbook's contrast between a poor and a good postmortem is most useful in its treatment of action items. The poor example fails in recognizable ways: items that are mostly mitigative with little prevention, one "preventative" item asking to make humans less error-prone, every item tagged with equal priority, vague verbs like "improve" and "make better", and only one item with a tracking bug. The Workbook's conclusion is blunt: "Without a formal tracking process, action items from postmortems are often forgotten, resulting in outages."

The good example's action items share five characteristics, which make a practical checklist:

  1. Ownership. Every item has a single named owner and a tracking number. The Workbook notes that items without clear owners are less likely to be resolved, and that one owner with several collaborators beats shared ownership.
  2. Prioritization. Every item has a priority, so the team knows what to do first.
  3. Measurability. Every item has a verifiable end state. "Add an alert when more than X% of machines are removed from service" can be checked; "improve monitoring" cannot.
  4. Preventative action. Each theme includes items that prevent or mitigate recurrence, not only items that repair the immediate damage.
  5. Blamelessness. Items target systems and processes, not individuals.

The Workbook's example also tags each item by type: investigate, prevent, mitigate, repair. That labelling is worth copying. When a postmortem's items are all "repair", the team has cleaned up but not reduced the chance of recurrence.

Fewer, better items

PagerDuty's guidance warns against creating too many tickets and generally limits follow-ups to P0 and P1 work that absolutely should be dealt with. That discipline matters. A postmortem with twenty action items tells the organization that nothing in particular is important, and the list decays together. Record lower-priority ideas in the document itself as observations, and file tickets only for the items you intend to finish.

Closing the loop

Everything above can be done well and the loop can still stay open. Closing it takes a small amount of recurring process, owned by someone with the standing to push back on roadmap pressure.

  • Track postmortem items separately. Label every postmortem ticket so they can be listed together, with incident, severity and due date. A single view of open postmortem work is the most important artefact in the whole process.
  • Review open items on a fixed cadence. A short weekly or fortnightly review of overdue items, attended by engineering managers, turns forgotten tickets into visible decisions: finish it, re-plan it, or explicitly accept the risk.
  • Make "won't do" explicit. Sometimes an action item is no longer worth doing. That is fine, but the decision should be recorded with a reason and an approver, rather than the ticket quietly ageing out.
  • Verify, don't just close. For preventative items, check that the fix works: re-run the failure in a test environment, or confirm the new alert fires as designed. A closed ticket is not the same as a closed risk.
  • Link recurrences. When a new incident resembles an old one, link the postmortems and check the status of the old action items. This is the clearest signal of whether the process is working.
  • Connect to reliability policy. If your SLOs have an error-budget policy, overdue P0 postmortem items are a strong candidate for the reliability work that takes priority when the budget is exhausted.

Where automation helps

Many preventative action items are guardrails: confirmation steps for destructive operations, input validation on admin tools, automatic rollback on failed health checks. These are good candidates for automation, provided the automation is itself tested and its actions are logged. Our article on self-healing infrastructure discusses where automated remediation is appropriate and where a human decision should stay in the loop.

Sharing what you learn

The SRE book's goal is to share postmortems "to the widest possible audience that would benefit", and PagerDuty's process ends by communicating results and key learnings internally. Other teams often run the same patterns and carry the same latent risks. A short internal summary, covering what happened, what the contributing causes were and what changed, spreads the lesson further than the full document ever will. Keep anything that could identify end users out of shared documents.

A minimal process to start with

  1. Write down your postmortem triggers and publish them.
  2. Name one owner per postmortem, and book the review meeting within a fixed number of days.
  3. Build the timeline first, with links to evidence.
  4. Analyse to systems, not people; reject "be more careful" as a fix.
  5. File a small number of owned, prioritized, measurable action items, each typed as investigate, prevent, mitigate or repair.
  6. Have someone outside the team review the document before it is published.
  7. Review open postmortem items on a fixed cadence until each is verified done or explicitly declined.

If you are formalizing incident response alongside wider platform work, our automation practice covers runbooks and remediation tooling.

FAQ

Does blameless mean nobody is accountable?

No. Blameless means the analysis looks for contributing causes in systems and processes rather than punishing individuals for reasonable decisions. Accountability moves to the action items, each of which has a named owner and a verifiable end state.

How soon after an incident should the postmortem happen?

Soon enough that memories and logs are fresh. PagerDuty's documented process schedules the meeting within 3 calendar days for a SEV-1 and 5 business days for a SEV-2. Whatever you choose, set the deadline in advance rather than per incident.

How many action items should a postmortem have?

As few as will genuinely reduce recurrence. PagerDuty generally limits tickets to P0 and P1 work. A short list of owned, measurable items is far more likely to be completed than a long list of good intentions.

← Back to Insights