Backup monitoring
Backup monitoring means hearing about a backup that did not finish, and AWS Backup and Azure Backup already alert on their own jobs, so the gap is the pg_dump, restic and rsync scripts nothing watches.
A backup that fails loudly is the easy case. The hard case is the one that stops: the cron entry lost in a server rebuild, the full disk that left pg_dump writing 0 bytes, the SSH key the NAS stopped accepting. Each of them leaves the last good backup getting older, quietly, until the day you need it. Backup monitoring is the alarm on that age.
AWS backup monitoring is already built in
If AWS Backup runs the job, use AWS. An EventBridge rule on Backup Job State Change with state FAILED, pointed at an SNS topic, is the setup the AWS docs describe, and AWS Backup sends those events on a best-effort basis every 5 minutes. For the backup that never ran, Backup Audit Manager has a Last recovery point was created control that flags resources with no recovery point in the last 1 to 744 hours. It is evaluated once every 24 hours and is not switched on for existing frameworks.
Azure backup monitoring is already built in
Azure Backup raises built-in Azure Monitor alerts for failed backup and restore jobs, on by default, within about 20 minutes of the failure. They notify nobody until you add an action group and an alert processing rule, which can route to email or a webhook. Do that once and the vaults are covered.
The backups no cloud console sees
Neither one watches pg_dump on a Hetzner box, restic on a home server or rsync to a NAS. That is where Logdash fits, with a push monitor on the Pro plan. There is no per-job schedule: on Pro, Logdash expects a ping in every 15-second window and alerts on the first empty one. So the script touches a stamp only when every step passed, and a heartbeat line pings while the stamp is fresh.
#!/bin/sh
# Credentials come from ~/.pgpass. Any failing step exits before the touch.
set -eu
FILE=/backups/app-$(date +%F).dump
pg_dump --format=custom --file="$FILE" app
pg_restore --list "$FILE" >/dev/null
rsync -a "$FILE" backup@nas:/srv/backups/
touch /var/lib/heartbeat/pg-backup30 2 * * * /usr/local/bin/pg-backup.sh
* * * * * for i in $(seq 12); do find /var/lib/heartbeat/pg-backup -mmin -1500 2>/dev/null | grep -q . && curl -fsS -m 4 -o /dev/null -X POST https://api.logdash.io/ping/68b4c1f0e3a2d5c7b9f01234; sleep 5; doneThe -mmin -1500 is 25 hours, a nightly run plus an hour of slack. Past it the pings stop and the alert goes out within 30 seconds. pg_restore --list proves the file is a readable archive, not that every row restores. A restore into a scratch database once a month is the only test that proves that.
- Create a push monitor On Pro, add a service per backup, set its monitor to push and copy the ping URL into the heartbeat line.
- Run the backup once by hand Create /var/lib/heartbeat and run the script. When the stamp appears the heartbeat starts and the monitor goes up.
- Fail it on purpose Connect Telegram to the monitor, then backdate the stamp with touch -d '2 days ago', which is what a missed night looks like a day later. Within 30 seconds Telegram shows the backup name, is down, and Did not receive call for this time range.
Picking a backup monitoring tool
Logdash vs native AWS and Azure alerts
| Feature | Logdash | AWS and Azure native |
|---|---|---|
| AWS Backup and Azure Backup jobs | Not integrated | Job events and alerts built in |
| pg_dump, restic and rsync scripts | One push monitor per script | Out of scope |
| Alert after a failed job | When the stamp passes your threshold | About 5 minutes on AWS, 20 on Azure |
| Not built, Telegram or webhook | Through SNS or an action group | |
| Setup per backup | Two crontab lines | One rule or one vault setting |
When AWS and Azure native is the better pick
- Every backup you have is an AWS Backup plan or an Azure Backup vault. Native alerts cover it, and Logdash would only be watching the watcher.
- You need email. Both clouds deliver it and Logdash does not.
- You want the alert minutes after a failed job, not after a staleness threshold. The cloud events fire on the failure itself.