Incident postmortem template

Copy the blameless template below into a markdown file within 48 hours of the incident, fill the timeline from your alert timestamps in UTC, and end it with at most three action items that each have an owner and a date.

A postmortem exists so the same outage does not happen twice. Blameless means you ask which part of the system let a reasonable person make the mistake, not who made it. Name a person and the next incident gets hidden. Name the missing lock timeout and you fix it once. The template below is short on purpose: a postmortem nobody finishes teaches nothing.

Blameless postmortem template

terminal
cat > "postmortem-$(date +%F).md" <<'EOF'
# Postmortem: <what users saw, in one line>

Date: YYYY-MM-DD | Author: <name> | Status: draft

## Summary
<Two sentences: what broke, for whom, for how long.>

## Impact
- Duration: HH:MM to HH:MM UTC (<n> minutes)
- Users affected: <number>
- Failed requests: <number, from logs>
- Cost: <refunds, credits, missed SLA>

## Timeline (UTC)
- HH:MM  Change that triggered it
- HH:MM  First failed check, alert fired
- HH:MM  Someone acknowledged
- HH:MM  Cause identified
- HH:MM  Fix shipped
- HH:MM  Monitor back up

## Root cause
<The mechanism, not the person.>

## What went well
## What went wrong
## Where we got lucky

## Action items
| Action | Owner | Due | Ticket |
|--------|-------|-----|--------|
EOF

How to fill in the postmortem template

  • Write it within 48 hours. After a week the timeline is a guess and the chat scrollback is gone.
  • Every timestamp in UTC. Mixed time zones turn a 20-minute outage into a 2-hour one on paper.
  • Impact in numbers: minutes, users, failed requests, money. Some users is not a number.
  • The root cause is a mechanism. If the answer is that someone forgot, ask why the system let them.
  • Where we got lucky is the section that finds the next outage. Write down what would have made this one worse.
  • Three action items at most, each with an owner, a date and a ticket. Fifteen items with no owner is zero items.

Post incident report example

Checkout API down for 23 minutes. At 14:02 UTC a deploy ran a migration that locked the orders table, and checkout requests queued until the connection pool ran dry. At 14:03 the health check could not get a connection, the monitor flipped to down and a Telegram alert fired. 14:09 acknowledged, 14:17 migration identified, 14:25 rolled back, up alert the same minute. Impact: 412 failed checkouts, 9 refunds. Root cause: migrations run inside the deploy with no lock timeout. Actions: a 5-second lock timeout on migrations, owner Ana, due Friday; migrations as a separate step before the deploy, owner Tom, due next sprint.

Where the timeline comes from

The two timestamps everyone argues about are the start and the end. A Logdash monitor sends a Telegram message saying the service is down on the first failed check and another saying it is up on the first healthy one, so the chat history brackets the incident to within one check interval: 5 minutes on the free plan, 1 minute on Builder, 15 seconds on Pro. For check-level detail, a published status page returns the last 100 checks of each monitor. Run this right after you resolve, because 100 checks at one a minute is under two hours of history.

failed checks, oldest first
curl -s https://api.logdash.io/v1/status_pages/your-status-page-id \
  | jq -r '.monitors[] | .name as $m | .pings[]
      | select(.statusCode < 200 or .statusCode >= 400)
      | "\(.createdAt)  \($m)  \(.statusCode)  \(.responseTimeMs)ms"'
  1. Add the monitor Point a Logdash HTTP monitor at your health endpoint. Each check stores the status code, the response time and the time it ran.
  2. Connect Telegram Attach a Telegram channel to the monitor. Every down and up message carries its own timestamp, which is the timeline you will paste later.
  3. Rehearse the first line Make the endpoint return 503. The down alert reaches Telegram on the next check. Restore it and the up alert follows. Those two messages are the first and last line of your next postmortem.

This template plus Logdash vs incident.io

FeatureLogdashincident.io
PriceFree template, monitoring free for 5 servicesBasic free, Team $19 per user a month
TimelineYou assemble it from alerts and chatRecorded from the Slack channel as it happens
First draftYou write itAI draft from the incident data on Pro
Action itemsA table you copy into your trackerFollow-ups exported to Jira or Linear
Where it livesA markdown file in your repo, reviewed in a PRInside incident.io
Setup for a team of oneNoneSlack workspace, roles, workflows

When incident.io is the better pick

  • More than a handful of people respond to incidents, and they already do it in Slack.
  • You run several incidents a month and rebuilding each timeline by hand costs hours.
  • An auditor wants a review workflow and a trail, not a markdown file.
What is a blameless postmortem template?
A postmortem template whose questions point at the system, not at people: what failed, why the system allowed it, what made it worse. It has no field for who caused it. The one above is blameless by structure, with root cause defined as a mechanism.
What goes in an incident postmortem template?
A one-line title, a two-sentence summary, impact in numbers, a UTC timeline from trigger to recovery, the root cause, what went well and badly, where you got lucky, and at most three owned action items. Anything longer rarely gets finished.
Is there a post incident report example I can copy?
Yes, the checkout example on this page: a 23-minute outage from a migration that locked the orders table, with timeline, impact, root cause and two owned actions. Replace the times and numbers with your own.
Is there a postmortem template for a small team?
This one. It fits on one page and needs no tool: one markdown file per incident in the repo. A dedicated incident platform starts paying off once several people respond to incidents every month.

Point it at your own URL and watch it for real.