Uptime vs downtime
Uptime is the share of a period in which your service answered correctly and downtime is the share in which it did not, so the two always add up to 100%: 99.9% uptime is 0.1% downtime, or 43.8 minutes in an average month.
Uptime is usually quoted as a percentage and downtime as a duration, but they are the same measurement read from opposite ends. A month of 43,830 minutes at 99.9% uptime has 43.8 minutes of downtime. Say the second number out loud and the first one stops sounding like a rounding error.
Uptime and downtime meaning
Both need a definition of up before they mean anything. For an HTTP monitor, up is a status code from 200 to 399 within the timeout. Down is everything else: a 404 on the health path, a 500 from a crashed handler, a 503 from a health check that could not reach the database, or no answer at all. Logdash gives a request 10 seconds and records a timeout as status code 0. A page that returns 200 with an error message in the body is up, as far as a status-code check can tell.
How to calculate uptime and downtime
There are two formulas. Counting minutes: uptime is minutes up divided by minutes in the period, which is how every SLA is written. Counting checks: uptime is successful checks divided by all checks, which is what a check-based monitor computes, Logdash included. Downtime is 100% minus uptime in both. You can run the second one yourself with nothing but curl and awk.
# one check, appended as "timestamp status" (000 = no answer in 10 s)
echo "$(date -u +%FT%TZ) $(curl -s -o /dev/null -m 10 -w '%{http_code}' https://example.com/health)" >> checks.log
# uptime and downtime from every check so far
awk '{ n++; if ($2 >= 200 && $2 < 400) up++ }
END { printf "uptime %.3f%% downtime %.3f%% failed %d of %d checks\n", 100*up/n, 100*(n-up)/n, n-up, n }' checks.log
# uptime 99.909% downtime 0.091% failed 8 of 8766 checksLogdash counts checks, SLAs count minutes
| Feature | Logdash | SLA math |
|---|---|---|
| Formula | Successful checks over all checks | Minutes up over minutes in the period |
| Data you need | None extra, the monitor counts checks as it runs | A start and end time for every outage |
| A 3-minute outage with 5-minute checks | 0 or 1 failed checks, so 0 or 5 minutes | 3 minutes |
| Precision | One interval: 5 minutes free, 1 minute on Builder, 15 seconds on Pro | To the minute, if someone wrote the times down |
| Works at 3am with nobody watching | Yes | Only if something recorded the outage |
When SLA math is the better pick
- You are checking a provider against its SLA, or writing one. Contracts count minutes, so convert to minutes before you compare.
- You have exact start and end times from logs or an incident timeline. That beats any check interval.
- Your checks run every 5 minutes and the outages are short. Check counts will round them to 0 or 5 minutes each.
Why downtime is the number to watch
99.9% and 99.5% look 0.4 points apart. In downtime it is 43.8 minutes a month against 3 hours 39 minutes, five times more. Uptime percentages compress exactly the part you care about, so set targets and write postmortems in minutes. Count the time to notice too: on a 5-minute interval an outage can be 5 minutes old before anyone knows, and those minutes are downtime whether or not a check saw them. A single one-hour outage takes that month to 99.86%, below three nines on its own.
- Add the URL Create a service in Logdash and paste the address of your health endpoint, or your homepage if the site is static. The first check runs straight away, so a wrong path shows up in seconds rather than during an incident.
- Pick the interval Every 5 minutes on the free plan, every minute on Builder, every 15 seconds on Pro. The shorter the interval, the closer the check count gets to the minute count.
- Break it on purpose Connect a Telegram channel, then stop the app or make the endpoint return 503. An alert you never tested is an alert you cannot trust. On the next check the monitor flips to down and a Telegram alert lands on your phone with the monitor name, the status code and the error.