Backup monitoring

Backup monitoring means hearing about a backup that did not finish, and AWS Backup and Azure Backup already alert on their own jobs, so the gap is the pg_dump, restic and rsync scripts nothing watches.

A backup that fails loudly is the easy case. The hard case is the one that stops: the cron entry lost in a server rebuild, the full disk that left pg_dump writing 0 bytes, the SSH key the NAS stopped accepting. Each of them leaves the last good backup getting older, quietly, until the day you need it. Backup monitoring is the alarm on that age.

AWS backup monitoring is already built in

If AWS Backup runs the job, use AWS. An EventBridge rule on Backup Job State Change with state FAILED, pointed at an SNS topic, is the setup the AWS docs describe, and AWS Backup sends those events on a best-effort basis every 5 minutes. For the backup that never ran, Backup Audit Manager has a Last recovery point was created control that flags resources with no recovery point in the last 1 to 744 hours. It is evaluated once every 24 hours and is not switched on for existing frameworks.

Azure backup monitoring is already built in

Azure Backup raises built-in Azure Monitor alerts for failed backup and restore jobs, on by default, within about 20 minutes of the failure. They notify nobody until you add an action group and an alert processing rule, which can route to email or a webhook. Do that once and the vaults are covered.

The backups no cloud console sees

Neither one watches pg_dump on a Hetzner box, restic on a home server or rsync to a NAS. That is where Logdash fits, with a push monitor on the Pro plan. There is no per-job schedule: on Pro, Logdash expects a ping in every 15-second window and alerts on the first empty one. So the script touches a stamp only when every step passed, and a heartbeat line pings while the stamp is fresh.

/usr/local/bin/pg-backup.sh
#!/bin/sh
# Credentials come from ~/.pgpass. Any failing step exits before the touch.
set -eu
FILE=/backups/app-$(date +%F).dump

pg_dump --format=custom --file="$FILE" app
pg_restore --list "$FILE" >/dev/null
rsync -a "$FILE" backup@nas:/srv/backups/
touch /var/lib/heartbeat/pg-backup
crontab -e
30 2 * * * /usr/local/bin/pg-backup.sh
* * * * * for i in $(seq 12); do find /var/lib/heartbeat/pg-backup -mmin -1500 2>/dev/null | grep -q . && curl -fsS -m 4 -o /dev/null -X POST https://api.logdash.io/ping/68b4c1f0e3a2d5c7b9f01234; sleep 5; done

The -mmin -1500 is 25 hours, a nightly run plus an hour of slack. Past it the pings stop and the alert goes out within 30 seconds. pg_restore --list proves the file is a readable archive, not that every row restores. A restore into a scratch database once a month is the only test that proves that.

  1. Create a push monitor On Pro, add a service per backup, set its monitor to push and copy the ping URL into the heartbeat line.
  2. Run the backup once by hand Create /var/lib/heartbeat and run the script. When the stamp appears the heartbeat starts and the monitor goes up.
  3. Fail it on purpose Connect Telegram to the monitor, then backdate the stamp with touch -d '2 days ago', which is what a missed night looks like a day later. Within 30 seconds Telegram shows the backup name, is down, and Did not receive call for this time range.

Picking a backup monitoring tool

Logdash vs native AWS and Azure alerts

FeatureLogdashAWS and Azure native
AWS Backup and Azure Backup jobsNot integratedJob events and alerts built in
pg_dump, restic and rsync scriptsOne push monitor per scriptOut of scope
Alert after a failed jobWhen the stamp passes your thresholdAbout 5 minutes on AWS, 20 on Azure
EmailNot built, Telegram or webhookThrough SNS or an action group
Setup per backupTwo crontab linesOne rule or one vault setting

When AWS and Azure native is the better pick

  • Every backup you have is an AWS Backup plan or an Azure Backup vault. Native alerts cover it, and Logdash would only be watching the watcher.
  • You need email. Both clouds deliver it and Logdash does not.
  • You want the alert minutes after a failed job, not after a staleness threshold. The cloud events fire on the failure itself.
What is backup monitoring?
Checking that backups finish and stay recent, and alerting when they stop. The useful signal is the age of the last good backup, because a backup that never ran produces no error to alert on.
How does Azure backup monitoring alert on a failed job?
Azure Backup raises a built-in Azure Monitor alert for failed backup and restore jobs, on by default. To be told about it, create an action group with email or a webhook and an alert processing rule that routes backup alerts to it.
Does AWS backup monitoring catch a backup that never ran?
An EventBridge rule on FAILED only fires for jobs that started. For jobs that never started, enable the Backup Audit Manager control Last recovery point was created. It runs every 24 hours and flags resources with no recovery point inside the window you set.
Which backup monitoring tool should I use for scripts?
One that alerts on silence, not only on errors: Healthchecks.io, Cronitor, Dead Man's Snitch or a Logdash push monitor. Logdash makes sense if you already watch the app there on Pro. For backups alone, the free tiers of Healthchecks.io or Dead Man's Snitch do the job.

Point it at your own URL and watch it for real.