Incident postmortems drafted from the record, not from memory
After an incident closes, an agent rebuilds what happened from the actual record: alert times, chat messages in the incident channel, deploys, config changes, and status page updates. It drafts the timeline and a blameless postmortem, so the review meeting starts from facts instead of reconstruction.
Incident commander for the specific incident, with the SRE or platform lead owning the process
An incident is marked resolved in the incident tool
01the problem and who owns it
Postmortems are due within a few days and written a week late, from a scrolled-through Slack channel and fading memory. Timestamps are approximate, the first alert that everyone ignored is left out, and action items get vague because the writer is tired.
The incident commander owns the write-up while also catching up on the work the incident displaced. The platform lead owns the process and sees the same contributing causes appear in different postmortems, with nobody connecting them.
02what the AI does, step by step
- Gather the recordThe agent pulls alerts and acknowledgements from the paging tool, messages from the incident channel, deploy and config change events, relevant dashboards, and status page posts for the incident window.
- Build a precise timelineEvents are merged into one timeline with real timestamps and sources: first signal, detection, escalation, diagnosis milestones, mitigation, and resolution. Gaps where nothing was recorded are marked as gaps.
- Compute the key intervalsTime to detect, acknowledge, mitigate, and resolve are calculated from the timeline, along with customer impact windows from the status page and error rates, rather than estimated.
- Draft the narrativeFollowing your postmortem template, the model writes the summary, impact, contributing factors, and what went well, in blameless language that describes systems and decisions rather than people.
- Propose action itemsEach proposed action ties to a specific contributing factor and is phrased as a concrete change with a suggested owning team. Duplicates of open actions from earlier incidents are linked instead of re-created.
- Cross-reference historyThe agent searches past postmortems for similar contributing factors and lists them, so the review can ask why a known weakness bit again.
03systems it connects to
- Paging and incident tools. PagerDuty, Opsgenie, incident.io, or similar for alerts and incident state.
- Team chat. Slack or Microsoft Teams incident channels.
- Deploy and change history. CI/CD logs, infrastructure as code changes, and feature flag audit logs.
- Docs and tracker. Confluence, Notion, or Google Docs for the postmortem, and Jira or Linear for action items.
04human checkpoints
- Incident commander edits and owns. The draft is a starting point. The commander corrects it, adds context only participants know, and is the named author.
- Review meeting. The team reviews the finished postmortem together. Action items are created in the tracker only after the meeting agrees them.
- External communication. Any customer-facing incident report is approved by the owner of customer communication and, where contracts require it, legal.
05what to measure
- Time from resolution to published postmortem. Tracked per incident.
- Timeline corrections. How many timeline entries the commander changes, as a check on accuracy.
- Action item completion. Share of postmortem actions closed by their due date.
- Repeat contributing factors. How often the same factor appears across incidents over a quarter.
06risks and guardrails
- Blame creeping in. Chat logs name people. Instruct and check the draft for blameless phrasing, and let the commander strike anything that reads as fault-finding.
- Sensitive data in channels. Incident channels often contain customer identifiers, tokens pasted in a hurry, or security details. Redact before drafting, and restrict security incident postmortems to the appropriate audience.
- Fabricated causes. A model may fill gaps with plausible explanations. Every causal claim in the draft must cite a timeline event, and unexplained gaps stay labeled as unknown.
07build vs buy
Modern incident management platforms include timeline capture and AI summaries, and if you already run incidents through one of them, turn those features on before building anything.
Custom work makes sense when the evidence is spread across tools those platforms do not read, when your postmortem format is specific, or when you want analysis across years of past incidents.
08related playbooks
Browse every engineering playbook or the full library.
want this running in your business?
We can connect your paging tool, chat, and deploy history, draft postmortems for your last few incidents, and let your team judge the drafts against the ones they wrote by hand.
See how we deliver it: ai agent development.
book a call drop your number