Hetzner server monitoring
Hetzner Cloud draws CPU, disk and network graphs for every server but sends no alert when one stops answering, so you add an outside HTTP check against a health path on the server or load balancer and route its alert to Telegram.
Hetzner servers are cheap, and the monitoring matches the price. The Cloud Console graphs CPU, disk IOPS and network traffic for every server, measured from the hypervisor. That is enough to spot a runaway process. It cannot see inside the guest, so a full disk, an OOM-killed Postgres or nginx answering 502 to every visitor all look normal on those graphs while CPU sits at 12%. And the graphs never message anyone: there is no threshold alert in the Console.
Hetzner cloud monitoring: what you get
- Cloud servers: CPU, disk and network graphs in the Console and through the metrics API. No memory graph, no disk usage, no alerts.
- Dedicated servers in Robot: System Monitor (SysMon), free ping, port and HTTP checks. Email only, and it alerts on the second consecutive failure, not the first. Its HTTP check accepts only 200, 301 and 302.
- Load balancers: active health checks over HTTP, HTTPS or TCP, every 15 seconds by default with 3 retries. An unhealthy target is pulled from rotation and nobody is told.
Hetzner load balancer health check
If you run a load balancer, its health check is already watching. The catch is the default: the root path, and any 2xx or 3xx counts as healthy. That passes as long as the web server answers, even when the app behind it cannot reach its database. Point it at a real health path, expect exactly 200, and keep the timeout well under the interval.
# Check /health on the app port instead of /
hcloud load-balancer update-service my-lb \
--listen-port 443 \
--health-check-protocol http \
--health-check-port 3000 \
--health-check-http-path /health \
--health-check-http-status-codes 200 \
--health-check-interval 15s \
--health-check-timeout 5s \
--health-check-retries 3
# Which targets does it call healthy right now?
hcloud load-balancer describe my-lb
# What a user gets through the load balancer
curl -sS -o /dev/null -w '%{http_code} in %{time_total}s\n' \
https://example.com/healthThat check protects users from one bad target. It does nothing when every target is bad, which on a one or two server setup is the usual outage: the load balancer has nowhere to send traffic, every request fails, and Hetzner sends no notification.
Hetzner VPS monitoring from outside
An outside HTTP monitor asks the question your users ask, through the same DNS, TLS and load balancer. Logdash sends a GET every 5 minutes on the free plan, every minute on Builder at $9 a month and every 15 seconds on Pro at $15, records the status code and response time, and alerts on the first failed check. Anything outside 200-399, a refused connection or no answer within 10 seconds counts as down. It only calls public addresses, so a private network IP like 10.0.0.2 is rejected: point it at the load balancer or the server's public hostname.
- Add the URL Create a service in Logdash and paste https://example.com/health, the same path the load balancer checks. The first check runs straight away, so a wrong path shows up now.
- Connect Telegram Add a Telegram channel once and attach it to the monitor. A webhook works the same way if you route alerts through your own handler.
- Break it on purpose Stop the app on every target. The next check fails, the monitor flips to down and a Telegram alert lands with the monitor name, the status code and the error. Start the app again and a second message says it is back up.
Logdash vs Hetzner built-in monitoring
| Feature | Logdash | Hetzner |
|---|---|---|
| Alert when a Cloud server stops answering | Telegram or webhook on the first failed check | None, graphs only |
| Dedicated server checks | HTTP only | SysMon: ping, port, HTTP, DNS and mail protocols |
| Alert channels | Telegram and webhook | Email, from SysMon only |
| Removes a broken server from traffic | No, it only watches | Load balancer drops failing targets |
| Cost | Free for 5 services at 5-minute checks | Included with the server |
When Hetzner built-in tools is the better pick
- You run dedicated servers and email is enough. SysMon is free, already in Robot, and checks ping and ports, which Logdash does not.
- You only need traffic to avoid one broken server. The load balancer health check does that every 15 seconds with no outside tool.
- You need memory and disk usage alerts. Neither Logdash nor the Console sees inside the guest, so run node_exporter with Prometheus for those.