pg_cron monitoring

pg_cron writes every run to cron.job_run_details and never alerts on it, so monitor it with a second pg_cron job that pings an external monitor every 10 seconds, but only while the last finished run of your job succeeded inside the window you expect.

pg_cron runs SQL on a schedule and says nothing when it fails. Each run becomes a row in cron.job_run_details, and that is the end of it. No email, no webhook, no retry. A nightly rollup failing on a permission error since a migration three weeks ago looks exactly like one that works, until someone reads the table.

pg_cron job status in job_run_details

The table has ten columns: jobid, runid, job_pid, database, username, command, status, return_message, start_time and end_time. Join it to cron.job on jobid to get the job name. A run passes through starting and running, and finishes as succeeded or failed. For a failure, return_message holds the Postgres error. For a success it holds the command tag, such as 1 row. This query lists every pg_cron failed job from the last day.

Failed runs in the last 24 hours
psql "$DATABASE_URL" -c "
  select j.jobname, d.status, d.return_message, d.start_time
  from cron.job_run_details d
  join cron.job j using (jobid)
  where d.status = 'failed'
    and d.start_time > now() - interval '1 day'
  order by d.start_time desc;"

pg_cron monitoring with a 10-second heartbeat

Logdash push monitors are on the Pro plan, $15 a month. On Pro, Logdash checks each push monitor every 15 seconds, marks it down when no ping arrived in that window, and alerts on the change. A nightly job cannot ping that often, so a second pg_cron job pings for it every 10 seconds, but only while the newest finished run of your job succeeded within a window you pick. That window is your grace period, written in SQL.

Schedule the heartbeat and the cleanup
psql "$DATABASE_URL" <<'SQL'
select cron.schedule('logdash-heartbeat', '10 seconds', $$
  select net.http_post('https://api.logdash.io/ping/68b4c1f0e3a2d5c7b9f01234')
  from (
    select d.status, d.end_time
    from cron.job_run_details d
    join cron.job j using (jobid)
    where j.jobname = 'nightly-rollup'
      and d.status in ('succeeded', 'failed')
    order by d.start_time desc
    limit 1
  ) last_run
  where last_run.status = 'succeeded'
    and last_run.end_time > now() - interval '25 hours'
$$);

select cron.schedule('prune-run-details', '0 * * * *',
  $$delete from cron.job_run_details where end_time < now() - interval '2 days'$$);
SQL
  • A failed run stops the pings within 10 seconds, because the newest finished row is no longer succeeded.
  • A job that stops being scheduled, unscheduled by a migration or set inactive, stops the pings 25 hours after its last good run.
  • A run in progress does not count against you. The filter judges the last finished run.
  • Postgres down, or the pg_cron worker dead: the heartbeat dies with them and the alert lands within 30 seconds. No check inside the database can report that.

It needs pg_cron 1.5 or newer for second-based schedules and pg_net for the HTTP call. Supabase ships both. Without pg_net, run the same query from a 10-second loop with psql and curl. The heartbeat adds 8,640 rows a day, hence the prune job.

  1. Create a push monitor Add a service in Logdash on Pro, set the monitor to push and copy the monitor id from the ping URL. Name it after the job, since the alert shows the name.
  2. Schedule the heartbeat Run the psql block with your job name, monitor id and window. Use the job interval plus the lateness you can live with: 25 hours for a nightly job, 70 minutes for an hourly one. The monitor goes green within 15 seconds.
  3. Break it on purpose Run select cron.unschedule('logdash-heartbeat'). Within 30 seconds a Telegram message lands saying the monitor is down, status code 0, did not receive call for this time range. Schedule it again and the up message follows.

Logdash vs Healthchecks.io for pg_cron

FeatureLogdashHealthchecks.io
Schedule model15-second heartbeat, window written in your SQLCron schedule and grace time per check
Extra work in PostgresA 10-second job and a prune jobOne ping appended to the job command
Notices Postgres itself is downWithin 30 secondsAt the next missed run plus grace
HTTP uptime checks and logs in the same toolSame service, same timelineJobs only
Free plan for job monitoringNone, push monitors need Pro20 jobs free
Source codeAGPL-3.0 on GitHubOpen source on GitHub

When Healthchecks.io is the better pick

  • You want one ping per run instead of a heartbeat job. Append a net.http_get to the end of the job command and they apply the schedule and grace time for you.
  • You do not want 8,640 extra rows a day in job_run_details, prune job or not.
  • You are not on Pro. Their free plan covers 20 jobs, and Logdash has no free push monitors.
  • You have many pg_cron jobs. One check per job with its own schedule beats one hand-written window per job.
How do I check pg_cron job status?
Query cron.job_run_details joined to cron.job on jobid and order by start_time descending. The status column reads running while a job runs, then succeeded or failed. return_message carries the error text for failures.
How do I get alerted on pg_cron failed jobs?
pg_cron has no alerting of its own. Schedule a 10-second heartbeat that pings an external monitor only while the last finished run succeeded. A failure stops the pings and Logdash sends a Telegram alert within 30 seconds.
Does pg_cron job_run_details grow forever?
Yes. pg_cron never cleans it, and a job running every 10 seconds adds 8,640 rows a day. Schedule a delete for rows older than a few days. Setting cron.log_run to off stops the logging, but then there is nothing left to monitor.
Does pg_cron monitoring work on Supabase?
Yes. Supabase offers both pg_cron and pg_net as extensions, so the heartbeat runs inside your project with no extra server. Enable pg_net first, then schedule the job.

Point it at your own URL and watch it for real.