Fly.io monitoring
Fly.io gives you health checks that steer traffic between machines and Grafana dashboards of your metrics, but nothing that tells you when the app goes down, so pair the fly.toml check with an outside HTTP monitor that sends you a Telegram message.
Fly.io ships two monitoring pieces, and both are good at their job. Service checks in fly.toml stop the Fly Proxy routing to a machine that fails them, and hold back a deploy that never comes up healthy. Metrics land in a managed Prometheus with Grafana dashboards at fly-metrics.net and about 15 days of retention. Neither piece tells you anything. Fly's docs say it plainly: no built-in alerting on metrics, and a failing check does not restart or stop the machine.
What Fly.io health checks do
A service check under [[http_service.checks]] runs at the proxy every 30 seconds by default. Fail it and the proxy marks that machine unhealthy and stops sending it traffic until it passes. That is the whole effect. The machine keeps running, nobody gets a message, and fly checks list is the only place the red shows up. The design assumes two or more machines. A side project on one machine has nothing to route to instead, and you hear about it from a user.
The check to paste
# Add a service check to fly.toml, then deploy
cat >> fly.toml <<'TOML'
[[http_service.checks]]
grace_period = "10s"
interval = "30s"
method = "GET"
timeout = "5s"
path = "/health"
TOML
fly deploy
fly checks list
# The request an outside monitor makes, through DNS and the Fly edge
curl -s -o /dev/null -w '%{http_code} in %{time_total}s\n' \
https://your-app.fly.dev/healthPoint the path at a route that runs one cheap database query and returns 503 when it fails, or it passes every check while the database is gone. The curl line takes the path your users take, through public DNS and the Fly edge. The service check never leaves the platform.
What Fly.io monitoring leaves out
- Alerts. A message means writing Grafana alert rules, or running your own Prometheus and Alertmanager. The old Slack and PagerDuty check handlers are gone.
- Restarts. A machine failing its check stays up until the app exits with a non-zero code and the restart policy brings it back.
- The outside view. A check inside Fly cannot see a custom domain pointing at the wrong place or a problem between your users and the edge.
- Fly's own incidents. status.flyio.net reports the platform, not your app.
Fly.io uptime from outside, with machines that auto-stop
An outside check is a request like any other. With auto_start_machines on, it starts a stopped machine, and the proxy's stop loop only runs every few minutes, so a 15-second check keeps the app running for good. Logdash gives up after 10 seconds, so a cold boot slower than that reads as down. Either set min_machines_running = 1 and pay for one machine that stays up, or check every 5 minutes and accept that some checks measure boot time.
- Add the URL Create a service in Logdash and point its HTTP monitor at https://your-app.fly.dev/health, or the custom domain your users type. Every check stores the status code and response time.
- Pick the interval Every 5 minutes on the free plan, every minute on Builder, every 15 seconds on Pro. 200 to 399 is up. Any other code, or no answer in 10 seconds, is down.
- Break it on purpose Stop the database or deploy a /health that returns 503. On the next check the monitor flips to down and a Telegram message arrives with the monitor name, status code and error, then another on recovery.
Logdash vs Fly.io built-in checks
| Feature | Logdash | Fly.io |
|---|---|---|
| Routes traffic away from a bad machine | No, it only watches | Yes, that is what service checks are for |
| Holds back a broken deploy | No | Yes, deploys wait for checks to pass |
| Tells you the app is down | Telegram or webhook on every down and up | Nothing built in, Grafana alert rules you write |
| Check interval | 5 minutes free, 1 minute Builder, 15 seconds Pro | You set it, 30 seconds by default |
| CPU, memory and OOM data | Only what your app sends through an SDK | Built in, about 15 days in Grafana |
| Price | Free for 5 services | Included with the app |
When Fly.io built-in checks is the better pick
- You run two or more machines and only need traffic to avoid a broken one. Fly does that, Logdash cannot.
- You already have Grafana alert rules sending to Slack or email. A second tool adds the outside view and little else.
- You need machine CPU, memory and out-of-memory data. That is Fly metrics, not an uptime monitor.