loading…
A postmortem is a learning and prioritization tool, not a courtroom transcript. It should make the next operator faster and the next failure less damaging.
Include customer impact, detection, precise timeline, contributing conditions, root cause, mitigation, recovery, what worked, what failed, and action items with owners and dates.
Avoid “human error” as a root cause. Ask why one action could create that impact and why controls did not catch it.
Write a postmortem for the double-refund incident. Separate causal facts from hypotheses. Create one action that prevents recurrence, one that improves detection, and one that reduces recovery time.
Every action must change a system, control, test, runbook, or ownership boundary. “Be more careful” is not an action item.