Dead man's switch monitoring
A dead man's switch alerts on the absence of a signal: the job checks in after every successful run, and when the check-ins stop, for any reason at all, the switch fires.
Normal monitoring waits for something to go wrong and say so: a crash, a 500, an error line in the log. That model is blind to failures that make no noise. The server that was terminated, the crontab a deploy overwrote, the script hung on a network mount since Tuesday. None of them report anything, because the thing that would report is the thing that stopped.
How a dead man's switch works
The name comes from trains. The driver holds a handle down, and if the driver collapses the handle rises and the brakes engage. Nobody has to notice the collapse. In software your job is the hand. It checks in after every good run, a service outside your infrastructure holds the timer, and the alert fires when the check-ins stop. Silence is the alert. That is also why the switch has to live somewhere else: a switch on the same server dies with it.
Dead man's switch cron setup
In Logdash the switch is a push monitor, on the Pro plan. It has no per-job interval. On Pro it checks every 15 seconds, goes down on the first window without a ping and comes back up on the next one. That suits a process that is always running. For a cron job that runs once a night, the job leaves proof on disk and a heartbeat line keeps the handle held while the proof is fresh.
# The job leaves proof only when it exits 0.
0 2 * * * /srv/app/bin/nightly-export && touch /var/lib/heartbeat/nightly-export
# The hand on the handle: ping every 5 seconds while the proof is under 25 hours old.
* * * * * for i in $(seq 12); do find /var/lib/heartbeat/nightly-export -mmin -1500 2>/dev/null | grep -q . && curl -fsS -m 4 -o /dev/null -X POST https://api.logdash.io/ping/68b4c1f0e3a2d5c7b9f01234; sleep 5; doneRead the second line as the hand. Every 5 seconds it checks that the export finished in the last 25 hours and pings if it did. If the export fails, the stamp ages out and the pings stop. If the server dies, cron stops and the pings stop. Either way Logdash hears nothing and fires, within 30 seconds of the last ping.
Dead man's switch email, Telegram or webhook
Logdash does not send email. Alerts go to Telegram or to a webhook, and that is the full list. Set the webhook method to POST and it receives the body below. GET, the default, carries no body. If the switch has to end in an inbox, the webhook is where you attach a mailer you run yourself.
{
"httpMonitorId": "68b4c1f0e3a2d5c7b9f01234",
"newStatus": "down",
"name": "nightly-export",
"url": "push monitor",
"errorMessage": "Did not receive call for this time range",
"statusCode": "0"
}- Create a push monitor On Pro, add a service, set its monitor to push and copy the ping URL into the heartbeat line.
- Hold the handle Paste both lines into the crontab and run the job once by hand so the stamp exists. The monitor goes up on the first ping.
- Let go Connect Telegram to the monitor, then comment out the heartbeat line. Within 90 seconds, once the loop already running finishes its minute, Telegram reads nightly-export is down, with Did not receive call for this time range. Put the line back and the is up message follows.
Logdash vs Dead Man's Snitch
| Feature | Logdash | Dead Man's Snitch |
|---|---|---|
| Interval per job | None, set it as the stamp threshold | Pick an interval per snitch |
| Email alerts | Not built, Telegram or webhook | Email by default, repeated each failed period |
| Check in by email | Not built | Every snitch has an address that counts as a check-in |
| Host goes dark | Alert within 30 seconds | Alert at the end of the interval |
| Uptime checks for the API | HTTP monitors on the same account | Not built |
| Free plan | Push monitors need Pro | One snitch free, three for $5 a month |
When Dead Man's Snitch is the better pick
- You want the alert in your inbox. Dead Man's Snitch emails by default and sends one more for every failed period until you act.
- Your job can send email but not run curl. Email check-ins cover that, and Logdash has nothing like it.
- You watch one nightly job and nothing else. One snitch is free and you skip the stamp file entirely.