Restic backup monitoring
Chain restic backup, restic check and a timestamp so the stamp only moves when both commands exit 0, then let a heartbeat ping Logdash while the stamp is under 26 hours old, and a failed, partial or skipped backup becomes a Telegram alert.
Restic sends no notifications. It exits with a code, writes to whatever log you pointed it at, and goes quiet. That is fine on the nights it works. On the night the NAS stops accepting the SSH key, restic backup exits non-zero and saves nothing, and it does the same the next night and every night after, while the log grows and nobody reads it. Most people find out on the day they need to restore.
The restic backup check script
#!/bin/sh
export RESTIC_REPOSITORY=sftp:backup@nas:/srv/restic
export RESTIC_PASSWORD_FILE=/etc/restic/password
restic backup /srv /etc --exclude-caches \
&& restic check --read-data-subset=5% \
&& touch /var/lib/heartbeat/resticThree details carry the weight. The && chain means the stamp moves only if both commands exit 0. Exit code 3 matters most: restic made a snapshot but could not read some files, and counting that as a failure is how you hear about the database file that was locked every night. And --read-data-subset=5% downloads and verifies about 5% of the pack data per run, around 10 GB a night on a 200 GB repository, so budget the egress if the repository lives in a cloud bucket.
Why restic monitoring needs a heartbeat
Logdash watches this with a push monitor, a Pro plan feature. It has no schedule or grace setting: on Pro it expects a ping in every 15-second window, alerts on the first empty one and recovers on the next ping. A single curl at the end of a nightly backup would read as down nearly all day. So the script touches a stamp, and a second crontab line turns the age of that stamp into a steady heartbeat.
0 3 * * * /usr/local/bin/restic-backup.sh >>/var/log/restic-backup.log 2>&1
* * * * * for i in $(seq 12); do find /var/lib/heartbeat/restic -mmin -1560 2>/dev/null | grep -q . && curl -fsS -m 4 -o /dev/null -X POST https://api.logdash.io/ping/68b4c1f0e3a2d5c7b9f01234; sleep 5; doneThe -mmin -1560 is 26 hours: one nightly run plus two hours for a slow upload. Past that the pings stop and the alert goes out within 30 seconds. If the machine running restic dies outright, cron dies with it and the alert comes just as fast, without waiting for the threshold.
Restic backup notification in three steps
- Create a push monitor On Pro, add a service named after the repository, set its monitor to push and copy the ping URL into the heartbeat line.
- Run the script once by hand Create /var/lib/heartbeat, then run the script. The first check takes the longest. When it finishes the stamp exists, the heartbeat starts and the monitor goes up.
- Break it on purpose Connect Telegram to the monitor, then backdate the stamp with touch -d '2 days ago'. Within 30 seconds Telegram shows the repository name, is down, and Did not receive call for this time range. In real life the same message arrives about two hours after a missed night.
Logdash vs Healthchecks.io for restic
| Feature | Logdash | Healthchecks.io |
|---|---|---|
| Nightly schedule | Encoded in the stamp threshold | Period or cron expression, plus grace, per check |
| Restic output in the alert | Not stored | Ping body up to 100 kB, so the log rides along |
| Exit code reporting | Silence only | Ping with the exit status, non-zero alerts at once |
| Backup host dies | Alert within 30 seconds | Alert after the period plus grace |
| Uptime checks and logs alongside | HTTP monitors and eight SDKs | Heartbeats only |
| Price for one nightly backup | Push monitors need Pro | Free, 20 checks |
When Healthchecks.io is the better pick
- You run restic on many machines with different schedules. A cron expression per check is less to maintain than a threshold in every crontab.
- You want the restic output attached to the alert, so the failing line is the first thing you read.
- You want the alert by email. Logdash sends Telegram and webhooks only.
- Backups are the only thing you monitor. Healthchecks.io does it for free, and Logdash push monitors start at Pro.