Celery beat monitoring
Have beat send a canary task every 10 seconds that posts to a Logdash push monitor, and the first 15 seconds without a ping tells you beat, the broker or every worker has stopped.
Celery beat fails without raising anything. When the beat process dies, workers sit idle, the queue stays empty, Flower shows healthy workers, and the error tracker stays quiet because no task ran. You find out when a customer asks why the weekly report never arrived.
Celery beat health check: a canary task
The cheapest proof that beat is scheduling and a worker is consuming is a task that does nothing except say so. Beat sends it every 10 seconds, a worker runs it, the task posts to Logdash. If beat, the broker or every worker goes down, the posts stop. The expires option matters after an outage: returning workers discard canaries older than 10 seconds instead of chewing through hundreds of stale ones before real work.
from urllib.request import Request, urlopen
from celery import Celery
app = Celery("tasks", broker="redis://localhost:6379/0")
PING_URL = "https://api.logdash.io/ping/68b4c1f0e3a2d5c7b9f01234"
app.conf.beat_schedule = {
"logdash-heartbeat": {
"task": "tasks.heartbeat",
"schedule": 10.0,
# A canary that waited in the queue is stale. Drop it.
"options": {"expires": 10},
},
}
@app.task(ignore_result=True)
def heartbeat():
urlopen(Request(PING_URL, data=b"", method="POST"), timeout=5)Run beat as its own process with celery -A tasks beat. Celery's docs call the embedded -B flag not recommended for production, and only one beat may run per schedule or every task fires twice. With django-celery-beat the entry still works: the database scheduler copies beat_schedule into its table on start.
What the 15 seconds mean
Push monitors are a Pro feature. On Pro, Logdash looks for a ping every 15 seconds and marks the monitor down on the first window without one. There is no grace setting, which is why the canary runs every 10 seconds. It also means a worker pool stuck on long tasks for 15 seconds pages you. If that is the outage you want to hear about, keep the canary on the main queue. If it is noise, route it to its own queue with one dedicated worker and accept that you then watch beat and the broker, not the main pool.
Celery periodic task monitoring for nightly jobs
The canary proves the machinery runs. It does not prove the 3am invoice task succeeded, and a task that runs once a day cannot feed a monitor that wants a call every 15 seconds. For those, have the task set a Redis key with a 26-hour expiry when it finishes, serve a route that returns 503 once the key is gone, and point an ordinary HTTP monitor at it. Or use a tool built around schedules.
- Create a push monitor Add a service on Pro, set the monitor to push and paste the id from its ping URL into PING_URL.
- Deploy beat and a worker Start celery -A tasks worker and celery -A tasks beat. The monitor turns up on the first check after the first canary.
- Kill beat Connect a Telegram channel, then stop the beat process. Within 30 seconds a Telegram alert says the monitor is down with "Did not receive call for this time range". Start beat and the up message follows.
Logdash vs Cronitor for Celery
| Feature | Logdash | Cronitor |
|---|---|---|
| Setup | One canary task and a push monitor | cronitor.celery.initialize discovers every beat task |
| Per-task schedules | One canary for the whole system | A monitor per periodic task, schedule read from beat |
| Time to alert when beat dies | 15 to 30 seconds | When the next scheduled task is late |
| django-celery-beat | Works, the canary is a normal entry | Auto-discovery does not support it yet |
| Free plan | Push monitors start at Pro, $15 a month | 5 monitors |
When Cronitor is the better pick
- You have a dozen periodic tasks on different schedules and want each one watched with its own late alert.
- Your important tasks run hourly or nightly. Cronitor knows the schedule; Logdash needs the canary plus a freshness route per task.
- You want job monitoring on a free plan, or alerts by email and Slack rather than Telegram and webhooks.
- You want to see that a task started, ran too long or failed, not only that the canary arrived.