Uptime SLA
An uptime SLA is a provider's written promise that a service will be available for a set share of each month, such as 99.9% or 99.99%, with a service credit off that month's bill when it falls short.
An SLA reads like an uptime guarantee. It is not one. It is a refund policy with a number on top. The provider promises a share of the month, measures it themselves, and if they miss, you get a percentage of that service's bill back, usually only after you ask for it. Knowing that changes how much weight a 99.99 on a pricing page should carry.
Uptime SLA meaning: the three parts
- The target. The number of nines, almost always per calendar month. 99.9% is 43.2 minutes in a 30-day month, 99.99% is 4.32 minutes.
- The measurement. How downtime is defined and who measures it. Usually minutes in which the service was unavailable, by the provider's own definition and from the provider's own data.
- The remedy. A service credit, a percentage of that month's bill for the affected service, in tiers that grow as uptime drops.
A real one: AWS EC2
AWS commits to 99.99% monthly uptime for EC2 at the region level and 99.5% for a single instance. Monthly uptime is 100% minus the percentage of minutes the service was unavailable. The region-level credits are 10% of the bill below 99.99%, 30% below 99.0% and 100% below 95.0%. You get nothing unless you open a support case by the end of the second billing cycle after the incident, and outages caused by your own software are excluded. A 4-hour outage in a 30-day month is 99.44% uptime, which buys a 10% credit on the EC2 bill, not on the launch day it cost you.
Uptime SLA 99.99
Four nines allow 4.38 minutes of downtime in an average month and 52.6 minutes in a year. That is tight enough that a 5-minute check cannot verify it: one failed check already counts as up to 5 minutes. At 15-second checks a 30-day month is 172,800 checks, and 17 failures fit inside 99.99%. If you promise four nines to your own customers, you need sub-minute checks just to know whether you kept the promise.
Uptime SLA calculator
def monthly_uptime(down_minutes, days_in_month=30):
total = days_in_month * 24 * 60
return 100 * (total - down_minutes) / total
def credit(uptime):
# AWS EC2 region-level tiers, as an example
if uptime >= 99.99:
return 0
if uptime >= 99.0:
return 10
if uptime >= 95.0:
return 30
return 100
uptime = monthly_uptime(52) # 52 minutes down in a 30-day month
print(f"{uptime:.3f}% uptime, {credit(uptime)}% credit")
# 99.880% uptime, 10% creditSwap the tiers for the ones in your provider's SLA and the days for the month you are claiming. The input that matters is down minutes, and the provider will use their number, not yours, so keep your own record with timestamps.
What an SLA will never cover
- Your code. A deploy that breaks login is 100% your downtime and 0% theirs.
- The chain. Three dependencies at 99.9% each give your app a ceiling of 99.7%, about 2 hours 11 minutes a month.
- Your losses. Credits are a share of the provider bill, and AWS issues none under one dollar.
- Add the URL Create a service in Logdash and paste the address of your health endpoint, or your homepage if the site is static. The first check runs straight away, so a wrong path shows up in seconds rather than during an incident.
- Pick the interval Match the interval to the target you care about. For 99.9% a 1-minute check on Builder is enough to see a breach. For 99.99% use 15 seconds on Pro. The down and up alerts carry timestamps, which is the record you take to a claim.
- Break it on purpose Connect a Telegram channel, then stop the app or make the endpoint return 503. An alert you never tested is an alert you cannot trust. On the next check the monitor flips to down and a Telegram alert lands on your phone with the monitor name, the status code and the error.