postmortem

/post-MOR-tum/ · noun · Development · Origin: 2004

Definitions

  1. A structured document written after an incident that analyzes what happened, why, and how to prevent recurrence. Good postmortems are blameless (focusing on systems, not individuals), thorough (establishing a detailed timeline), and actionable (producing tracked follow-up items). The practice originated in medicine and was adapted for software by Google's SRE team.

    In plain English: A document written after something goes wrong that explains what happened and how to prevent it next time — the key rule is to blame systems, not people.

    Example: The postmortem revealed that the outage had five contributing factors and the 'root cause' everyone assumed was actually just the trigger for a deeper issue.

Origin Story

Learning from Failure Without Pointing Fingers

A postmortem document (also called an incident review or retrospective) is a structured analysis of a system failure or outage, written to understand what happened, why it happened, and how to prevent recurrence. The practice of blameless postmortems was pioneered at Google in the early 2000s and formalized in the company's Site Reliability Engineering (SRE) methodology. Ben Treynor Sloss, who founded Google's SRE team around 2003, championed the idea that incidents are learning opportunities, not occasions for punishment. The key insight was that blaming individuals discourages honest reporting, while focusing on systemic factors reveals actionable improvements. A good postmortem typically includes a timeline of events, root cause analysis, impact assessment, what went well, what went poorly, and concrete action items with owners and deadlines. John Allspaw, then CTO of Etsy, further developed the practice with his influential 2012 writing on 'blameless postmortems,' arguing that human error is a symptom of system design flaws, not a root cause. The practice has since spread across the tech industry, with companies like Amazon, Microsoft, and Meta all maintaining formal postmortem processes. Many organizations publish public postmortems for major incidents, contributing to a shared body of operational knowledge.

Coined by: Google SRE team (popularized by Ben Treynor Sloss and John Allspaw)

Context: Formalized at Google in the early 2000s as part of Site Reliability Engineering practices.

Fun fact: Google maintains an internal postmortem archive with thousands of entries. Former Googlers who joined other companies often cite the postmortem culture as one of the practices they most wanted to replicate at their new organizations.

Related Terms