Incident postmortem template
Copy the blameless template below into a markdown file within 48 hours of the incident, fill the timeline from your alert timestamps in UTC, and end it with at most three action items that each have an owner and a date.
A postmortem exists so the same outage does not happen twice. Blameless means you ask which part of the system let a reasonable person make the mistake, not who made it. Name a person and the next incident gets hidden. Name the missing lock timeout and you fix it once. The template below is short on purpose: a postmortem nobody finishes teaches nothing.
Blameless postmortem template
cat > "postmortem-$(date +%F).md" <<'EOF'
# Postmortem: <what users saw, in one line>
Date: YYYY-MM-DD | Author: <name> | Status: draft
## Summary
<Two sentences: what broke, for whom, for how long.>
## Impact
- Duration: HH:MM to HH:MM UTC (<n> minutes)
- Users affected: <number>
- Failed requests: <number, from logs>
- Cost: <refunds, credits, missed SLA>
## Timeline (UTC)
- HH:MM Change that triggered it
- HH:MM First failed check, alert fired
- HH:MM Someone acknowledged
- HH:MM Cause identified
- HH:MM Fix shipped
- HH:MM Monitor back up
## Root cause
<The mechanism, not the person.>
## What went well
## What went wrong
## Where we got lucky
## Action items
| Action | Owner | Due | Ticket |
|--------|-------|-----|--------|
EOFHow to fill in the postmortem template
- Write it within 48 hours. After a week the timeline is a guess and the chat scrollback is gone.
- Every timestamp in UTC. Mixed time zones turn a 20-minute outage into a 2-hour one on paper.
- Impact in numbers: minutes, users, failed requests, money. Some users is not a number.
- The root cause is a mechanism. If the answer is that someone forgot, ask why the system let them.
- Where we got lucky is the section that finds the next outage. Write down what would have made this one worse.
- Three action items at most, each with an owner, a date and a ticket. Fifteen items with no owner is zero items.
Post incident report example
Checkout API down for 23 minutes. At 14:02 UTC a deploy ran a migration that locked the orders table, and checkout requests queued until the connection pool ran dry. At 14:03 the health check could not get a connection, the monitor flipped to down and a Telegram alert fired. 14:09 acknowledged, 14:17 migration identified, 14:25 rolled back, up alert the same minute. Impact: 412 failed checkouts, 9 refunds. Root cause: migrations run inside the deploy with no lock timeout. Actions: a 5-second lock timeout on migrations, owner Ana, due Friday; migrations as a separate step before the deploy, owner Tom, due next sprint.
Where the timeline comes from
The two timestamps everyone argues about are the start and the end. A Logdash monitor sends a Telegram message saying the service is down on the first failed check and another saying it is up on the first healthy one, so the chat history brackets the incident to within one check interval: 5 minutes on the free plan, 1 minute on Builder, 15 seconds on Pro. For check-level detail, a published status page returns the last 100 checks of each monitor. Run this right after you resolve, because 100 checks at one a minute is under two hours of history.
curl -s https://api.logdash.io/v1/status_pages/your-status-page-id \
| jq -r '.monitors[] | .name as $m | .pings[]
| select(.statusCode < 200 or .statusCode >= 400)
| "\(.createdAt) \($m) \(.statusCode) \(.responseTimeMs)ms"'- Add the monitor Point a Logdash HTTP monitor at your health endpoint. Each check stores the status code, the response time and the time it ran.
- Connect Telegram Attach a Telegram channel to the monitor. Every down and up message carries its own timestamp, which is the timeline you will paste later.
- Rehearse the first line Make the endpoint return 503. The down alert reaches Telegram on the next check. Restore it and the up alert follows. Those two messages are the first and last line of your next postmortem.
This template plus Logdash vs incident.io
| Feature | Logdash | incident.io |
|---|---|---|
| Price | Free template, monitoring free for 5 services | Basic free, Team $19 per user a month |
| Timeline | You assemble it from alerts and chat | Recorded from the Slack channel as it happens |
| First draft | You write it | AI draft from the incident data on Pro |
| Action items | A table you copy into your tracker | Follow-ups exported to Jira or Linear |
| Where it lives | A markdown file in your repo, reviewed in a PR | Inside incident.io |
| Setup for a team of one | None | Slack workspace, roles, workflows |
When incident.io is the better pick
- More than a handful of people respond to incidents, and they already do it in Slack.
- You run several incidents a month and rebuilding each timeline by hand costs hours.
- An auditor wants a review workflow and a trail, not a markdown file.