MTTR

MTTR is the average time it takes to get a broken service working again: add up the downtime from every incident and divide by the number of incidents, so four outages of 12, 47, 6 and 31 minutes give an MTTR of 24 minutes.

MTTR is the number you quote after an outage. The arithmetic is one division; the hard part is where the clock starts. Start it when someone opens a laptop and every minute nobody knew about the outage drops out of the report, and those are usually the longest minutes.

MTTR meaning: four metrics, one acronym

The R means four different things depending on who wrote the runbook. Check which one before you compare two numbers.

  • Mean time to repair: from the start of the repair work to the fix. It leaves out the time it took to notice.
  • Mean time to recovery, or restore: from the moment the service broke to the moment users could use it again. This matches what your customers felt, and it is the one this page means.
  • Mean time to respond: from the first alert to a human acknowledging it and starting work. Some teams call this MTTA, mean time to acknowledge.
  • Mean time to resolve: from the break to the root cause being fixed for good. Often days, because it includes the follow-up after service is back.

MTTR formula

MTTR = total downtime / number of incidents. Take a month with four outages of 12, 47, 6 and 31 minutes. That is 96 minutes over 4 incidents, an MTTR of 24 minutes. One bad night dominates the average: drop the 47 and it falls to 16 minutes. Once you have more than a handful of incidents, report the median next to the mean.

terminal
# minutes from each down alert to its up alert
printf '%s\n' 12 47 6 31 |
  awk '{ total += $1; n++ } END { printf "MTTR %.1f min over %d incidents\n", total / n, n }'
# MTTR 24.0 min over 4 incidents

How to calculate MTTR from uptime alerts

You need two timestamps per incident: when it went down and when it came back. An external uptime monitor records both for you. Logdash sends one alert when a monitor flips to down and another when it flips back to up, so the gap between the red and the green message is your recovery time, measured from outside your stack.

One caveat: the clock starts on the first failed check, not the first failed request, so a 5-minute interval can hide up to 5 minutes of every incident. At 15 seconds that error nearly disappears. The receiver below turns the alerts into a running MTTR. It needs a webhook channel set to POST, which is a paid-plan option; the free plan sends a GET with no body.

mttr.mjs
// mttr.mjs - run with: node mttr.mjs
// Point a Logdash webhook channel (method POST) at this server.
import { createServer } from 'node:http';

const downSince = new Map();
const outages = [];

createServer((req, res) => {
  let body = '';
  req.on('data', (chunk) => (body += chunk));
  req.on('end', () => {
    res.end('ok');
    let event;
    try {
      event = JSON.parse(body);
    } catch {
      return; // not a Logdash alert
    }
    const { httpMonitorId, newStatus, name } = event;

    if (newStatus === 'down') downSince.set(httpMonitorId, Date.now());
    if (newStatus !== 'up' || !downSince.has(httpMonitorId)) return;

    outages.push(Date.now() - downSince.get(httpMonitorId));
    downSince.delete(httpMonitorId);

    const mttr = outages.reduce((sum, ms) => sum + ms, 0) / outages.length;
    console.log(`${name} is back. MTTR over ${outages.length} outages: ${(mttr / 60000).toFixed(1)} min`);
  });
}).listen(3000);

MTTR vs MTBF

MTBF, mean time between failures, is total uptime divided by the number of failures. MTTR says how fast you recover, MTBF says how often you have to. Together they give availability: MTBF / (MTBF + MTTR). A service that breaks once a month, an MTBF of about 43,200 minutes, and takes 43 minutes to recover sits at 99.9%. Halve the MTTR and it is at 99.95% without preventing a single failure.

For a small team, MTTR is the lever you can actually pull. Failures come from deploys, hosts and third parties you only partly control. Recovery time is mostly detection plus finding the cause, and both shrink with a shorter check interval and the error logs one click from the monitor.

  1. Point a monitor at the service Add the health URL as an HTTP monitor. Every check stores the status code and response time, every 5 minutes on the free plan, every minute on Builder and every 15 seconds on Pro.
  2. Connect Telegram Add a Telegram channel for whoever fixes things. If you want the MTTR computed for you, add a POST webhook pointing at the receiver above as a second channel.
  3. Break it and time the recovery Stop the app, wait a few minutes, start it again. Your first MTTR data point is the gap between the down alert and the up alert that reach you on Telegram.
What is the MTTR formula?
Total downtime divided by the number of incidents. 96 minutes of downtime across 4 incidents is an MTTR of 24 minutes. Use the same start and end points for every incident, or the average means nothing.
MTTR vs MTBF: what is the difference?
MTTR measures how long a failure lasts, MTBF measures how long the service runs between failures. Availability is MTBF / (MTBF + MTTR), so you raise uptime either by failing less often or by recovering faster.
What is the MTTR meaning in incident management?
Usually mean time to recovery: from the service breaking to users being able to use it again. Repair, respond and resolve are the other three readings, and each starts or stops the clock at a different moment.
How to calculate MTTR if nobody logged the incidents?
Use your uptime alerts. Every incident has a down alert and an up alert; subtract and average. The result is accurate to one check interval, so 5 minutes on a 5-minute check and 15 seconds on a 15-second one.

Point it at your own URL and watch it for real.