Document coach
Incident Postmortem
Blameless root-cause analysis that your future self will thank you for.
Turn a painful incident into institutional knowledge. This guided postmortem walks you through the timeline, quantified customer impact, a genuine five-whys root cause analysis, and action items with owners and deadlines — the way experienced SRE teams do it. Blameless by design: the questions target systems and safeguards, never people, so the write-up you end with is one your whole team can learn from.
What this expert will cover
- 2 questions
Audience
working notesWho this postmortem must serve and what it should enable them to do — a working note that steers tone and depth, never part of the write-up. Complete means: the readers named (the future on-call, the wider team, leadership), and the decision or change the document should enable for them.
- 3 questions
Incident Summary
A reader who sees only this section knows what broke, when, for how long, how severe it was, and the one-line cause. Complete means: severity, dates/times, blast radius, and current status — in under a paragraph.
- 3 questions
Timeline
A precise chronology from first trigger to full resolution. Complete means: specific timestamps with timezone for the triggering change, first customer impact, detection, key response decisions, mitigation, and all-clear.
- 3 questions
Customer Impact
The quantified cost of the incident. Complete means: how many users or requests were affected, which capabilities degraded, error rates or latency numbers, and any SLA, revenue, or support-ticket consequences.
- 3 questions
Root Cause Analysis
The causal chain, not just the trigger. Complete means: the five-whys goes at least three levels past the immediate cause, distinguishes trigger from root cause, and names the missing safeguard that would have prevented or contained the incident.
- 2 questions
Detection & Response
How you found out and how the response went. Complete means: what detected the incident (alert vs. human vs. customer report), the gap between impact and detection, and an honest read on what went well and what dragged during response.
- 3 questions
Action Items
Committed remediation, not a wish list. Complete means: each item is specific and verifiable, has exactly one named owner and a due date, and the highest-priority item directly addresses the missing safeguard from the root cause analysis.
- 2 questions
Lessons Learned
What the organization now knows that it did not before. Complete means: at least one concrete, transferable lesson (not a platitude), and a note on whether this class of incident could recur elsewhere in the system.
The kind of questions it asks
- Who needs to read this postmortem — the future on-call, the wider engineering org, leadership? Name the readers, because the depth and tone of every section follows from who they are.
- In one or two sentences: what broke, and what could customers not do while it was broken? Plain language — imagine explaining it to a stakeholder outside engineering.
- When did the incident actually begin — the timestamp of the triggering event (deploy, config change, traffic shift), not when you noticed? Include the timezone.
Ready when you are.
Every answer inks the page in. Skip anything; return anytime.