MTTR
MTTR is the average time it takes to get a broken service working again: add up the downtime from every incident and divide by the number of incidents, so four outages of 12, 47, 6 and 31 minutes give an MTTR of 24 minutes.
MTTR is the number you quote after an outage. The arithmetic is one division; the hard part is where the clock starts. Start it when someone opens a laptop and every minute nobody knew about the outage drops out of the report, and those are usually the longest minutes.
MTTR meaning: four metrics, one acronym
The R means four different things depending on who wrote the runbook. Check which one before you compare two numbers.
- Mean time to repair: from the start of the repair work to the fix. It leaves out the time it took to notice.
- Mean time to recovery, or restore: from the moment the service broke to the moment users could use it again. This matches what your customers felt, and it is the one this page means.
- Mean time to respond: from the first alert to a human acknowledging it and starting work. Some teams call this MTTA, mean time to acknowledge.
- Mean time to resolve: from the break to the root cause being fixed for good. Often days, because it includes the follow-up after service is back.
MTTR formula
MTTR = total downtime / number of incidents. Take a month with four outages of 12, 47, 6 and 31 minutes. That is 96 minutes over 4 incidents, an MTTR of 24 minutes. One bad night dominates the average: drop the 47 and it falls to 16 minutes. Once you have more than a handful of incidents, report the median next to the mean.
# minutes from each down alert to its up alert
printf '%s\n' 12 47 6 31 |
awk '{ total += $1; n++ } END { printf "MTTR %.1f min over %d incidents\n", total / n, n }'
# MTTR 24.0 min over 4 incidentsHow to calculate MTTR from uptime alerts
You need two timestamps per incident: when it went down and when it came back. An external uptime monitor records both for you. Logdash sends one alert when a monitor flips to down and another when it flips back to up, so the gap between the red and the green message is your recovery time, measured from outside your stack.
One caveat: the clock starts on the first failed check, not the first failed request, so a 5-minute interval can hide up to 5 minutes of every incident. At 15 seconds that error nearly disappears. The receiver below turns the alerts into a running MTTR. It needs a webhook channel set to POST, which is a paid-plan option; the free plan sends a GET with no body.
// mttr.mjs - run with: node mttr.mjs
// Point a Logdash webhook channel (method POST) at this server.
import { createServer } from 'node:http';
const downSince = new Map();
const outages = [];
createServer((req, res) => {
let body = '';
req.on('data', (chunk) => (body += chunk));
req.on('end', () => {
res.end('ok');
let event;
try {
event = JSON.parse(body);
} catch {
return; // not a Logdash alert
}
const { httpMonitorId, newStatus, name } = event;
if (newStatus === 'down') downSince.set(httpMonitorId, Date.now());
if (newStatus !== 'up' || !downSince.has(httpMonitorId)) return;
outages.push(Date.now() - downSince.get(httpMonitorId));
downSince.delete(httpMonitorId);
const mttr = outages.reduce((sum, ms) => sum + ms, 0) / outages.length;
console.log(`${name} is back. MTTR over ${outages.length} outages: ${(mttr / 60000).toFixed(1)} min`);
});
}).listen(3000);MTTR vs MTBF
MTBF, mean time between failures, is total uptime divided by the number of failures. MTTR says how fast you recover, MTBF says how often you have to. Together they give availability: MTBF / (MTBF + MTTR). A service that breaks once a month, an MTBF of about 43,200 minutes, and takes 43 minutes to recover sits at 99.9%. Halve the MTTR and it is at 99.95% without preventing a single failure.
For a small team, MTTR is the lever you can actually pull. Failures come from deploys, hosts and third parties you only partly control. Recovery time is mostly detection plus finding the cause, and both shrink with a shorter check interval and the error logs one click from the monitor.
- Point a monitor at the service Add the health URL as an HTTP monitor. Every check stores the status code and response time, every 5 minutes on the free plan, every minute on Builder and every 15 seconds on Pro.
- Connect Telegram Add a Telegram channel for whoever fixes things. If you want the MTTR computed for you, add a POST webhook pointing at the receiver above as a second channel.
- Break it and time the recovery Stop the app, wait a few minutes, start it again. Your first MTTR data point is the gap between the down alert and the up alert that reach you on Telegram.