insomnia.club back to site
[ playbook · engineering ]

Incident postmortems drafted from the record, not from memory

After an incident closes, an agent rebuilds what happened from the actual record: alert times, chat messages in the incident channel, deploys, config changes, and status page updates. It drafts the timeline and a blameless postmortem, so the review meeting starts from facts instead of reconstruction.

who owns it

Incident commander for the specific incident, with the SRE or platform lead owning the process

what starts it

An incident is marked resolved in the incident tool

01the problem and who owns it

Postmortems are due within a few days and written a week late, from a scrolled-through Slack channel and fading memory. Timestamps are approximate, the first alert that everyone ignored is left out, and action items get vague because the writer is tired.

The incident commander owns the write-up while also catching up on the work the incident displaced. The platform lead owns the process and sees the same contributing causes appear in different postmortems, with nobody connecting them.

02what the AI does, step by step

  1. Gather the recordThe agent pulls alerts and acknowledgements from the paging tool, messages from the incident channel, deploy and config change events, relevant dashboards, and status page posts for the incident window.
  2. Build a precise timelineEvents are merged into one timeline with real timestamps and sources: first signal, detection, escalation, diagnosis milestones, mitigation, and resolution. Gaps where nothing was recorded are marked as gaps.
  3. Compute the key intervalsTime to detect, acknowledge, mitigate, and resolve are calculated from the timeline, along with customer impact windows from the status page and error rates, rather than estimated.
  4. Draft the narrativeFollowing your postmortem template, the model writes the summary, impact, contributing factors, and what went well, in blameless language that describes systems and decisions rather than people.
  5. Propose action itemsEach proposed action ties to a specific contributing factor and is phrased as a concrete change with a suggested owning team. Duplicates of open actions from earlier incidents are linked instead of re-created.
  6. Cross-reference historyThe agent searches past postmortems for similar contributing factors and lists them, so the review can ask why a known weakness bit again.

03systems it connects to

04human checkpoints

05what to measure

06risks and guardrails

07build vs buy

Modern incident management platforms include timeline capture and AI summaries, and if you already run incidents through one of them, turn those features on before building anything.

Custom work makes sense when the evidence is spread across tools those platforms do not read, when your postmortem format is specific, or when you want analysis across years of past incidents.

Browse every engineering playbook or the full library.

want this running in your business?

We can connect your paging tool, chat, and deploy history, draft postmortems for your last few incidents, and let your team judge the drafts against the ones they wrote by hand.

See how we deliver it: ai agent development.

book a call drop your number

info@insomnia.club