Kubernetes CronJob monitoring
Kubernetes records when each CronJob last succeeded in status.lastSuccessfulTime and never alerts on it, so either alert on that timestamp in Prometheus through kube-state-metrics, or run a small watcher that pings an external monitor every 10 seconds while it is fresh.
A CronJob can fail quietly in more ways than a crontab line. The pod exits non-zero until the Job hits its backoffLimit. The image tag was deleted and the pod sits in ImagePullBackOff. Someone set suspend to true during an incident and forgot. The controller was down past startingDeadlineSeconds and skipped the run. kubectl get cronjob shows a LAST SCHEDULE column, and nothing pages anyone.
One field covers all of these: status.lastSuccessfulTime, which the CronJob controller sets when a Job completes. Every failure above stops it moving. So monitoring a CronJob means alerting when that timestamp is older than the schedule plus some slack.
Kubernetes CronJob monitoring with Prometheus
If kube-state-metrics already runs in the cluster, it is two rules. time() - kube_cronjob_status_last_successful_time > 90000 fires when a job has not succeeded in 25 hours, and kube_job_status_failed > 0 catches each failed Job, with a reason label. Alertmanager routes both. For a team that already operates that stack, this is the answer. For a team that does not, it is three components to run, and they go down with the cluster they are meant to watch.
Kubernetes CronJob failed alert without Prometheus
The Logdash version is one Deployment that reads lastSuccessfulTime every 10 seconds and posts to a push monitor while it is fresh. Push monitors are on Pro, $15 a month, and Pro checks each one every 15 seconds: no ping in that window, the monitor goes down. That is why the CronJob does not call Logdash itself. A curl at the end of a nightly job is one ping a day against a 15-second window.
kubectl create serviceaccount cron-watch
kubectl create role cron-watch --verb=get --resource=cronjobs
kubectl create rolebinding cron-watch --role=cron-watch --serviceaccount=default:cron-watchapiVersion: apps/v1
kind: Deployment
metadata:
name: cron-watch
spec:
replicas: 1
selector:
matchLabels: { app: cron-watch }
template:
metadata:
labels: { app: cron-watch }
spec:
serviceAccountName: cron-watch
containers:
- name: watch
image: alpine/k8s:1.36.5
env:
- { name: JOB, value: nightly-report }
- { name: MAX_AGE, value: "90000" } # 25 hours in seconds
- { name: PING, value: https://api.logdash.io/ping/68b4c1f0e3a2d5c7b9f01234 }
command: ["/bin/sh", "-c"]
args:
- |
while true; do
kubectl get cronjob "$JOB" -o json \
| jq -e "now - (.status.lastSuccessfulTime | fromdateiso8601) < $MAX_AGE" >/dev/null \
&& curl -fsS -m 5 -X POST "$PING"
sleep 10
doneThe alpine/k8s image ships kubectl, jq and curl. jq -e exits non-zero when the timestamp is stale or missing, so a CronJob that has never succeeded sends nothing. Set MAX_AGE to the schedule interval plus the lateness you accept. Because the watcher runs in the cluster, a dead cluster or a broken network goes quiet too, and that reaches you as well.
- Create a push monitor Add a service in Logdash on Pro, set the monitor to push and copy the monitor id from the ping URL. Name it after the CronJob, since the alert shows the name.
- Apply the watcher Run the three kubectl lines, put the CronJob name, window and monitor id into the env block and apply the file. If the job succeeded inside the window, the monitor goes green within 15 seconds.
- Break it on purpose Scale the watcher to zero replicas. Within 30 seconds a Telegram message lands saying the monitor is down, status code 0, did not receive call for this time range. Scale it back and the up message follows.
Logdash vs Prometheus for CronJobs
| Feature | Logdash | Prometheus and kube-state-metrics |
|---|---|---|
| Setup if you already run it | RBAC, a Deployment and a monitor per job | Two alert rules |
| Setup from zero | One Deployment, alerting lives outside | kube-state-metrics, Prometheus and Alertmanager |
| New CronJobs | One more watcher and monitor each | Covered by the same rule |
| Why it failed | Only that the job went stale | Failure reason as a label |
| Whole cluster down | Alert within 30 seconds | Silent unless you add an outside check |
| Cost | Pro, $15 a month | Open source, you run and store it |
When Prometheus is the better pick
- You already run Prometheus with kube-state-metrics. Two rules beat a new Deployment.
- You have dozens of CronJobs. One rule covers all of them, including the ones added next month.
- You want to know why a Job failed. The reason label separates BackoffLimitExceeded from DeadlineExceeded.
- Alerts have to reach Slack or PagerDuty. Logdash sends Telegram and webhooks only.