Fly.io monitoring

Fly.io gives you health checks that steer traffic between machines and Grafana dashboards of your metrics, but nothing that tells you when the app goes down, so pair the fly.toml check with an outside HTTP monitor that sends you a Telegram message.

Fly.io ships two monitoring pieces, and both are good at their job. Service checks in fly.toml stop the Fly Proxy routing to a machine that fails them, and hold back a deploy that never comes up healthy. Metrics land in a managed Prometheus with Grafana dashboards at fly-metrics.net and about 15 days of retention. Neither piece tells you anything. Fly's docs say it plainly: no built-in alerting on metrics, and a failing check does not restart or stop the machine.

What Fly.io health checks do

A service check under [[http_service.checks]] runs at the proxy every 30 seconds by default. Fail it and the proxy marks that machine unhealthy and stops sending it traffic until it passes. That is the whole effect. The machine keeps running, nobody gets a message, and fly checks list is the only place the red shows up. The design assumes two or more machines. A side project on one machine has nothing to route to instead, and you hear about it from a user.

The check to paste

terminal
# Add a service check to fly.toml, then deploy
cat >> fly.toml <<'TOML'

[[http_service.checks]]
  grace_period = "10s"
  interval = "30s"
  method = "GET"
  timeout = "5s"
  path = "/health"
TOML

fly deploy
fly checks list

# The request an outside monitor makes, through DNS and the Fly edge
curl -s -o /dev/null -w '%{http_code} in %{time_total}s\n' \
  https://your-app.fly.dev/health

Point the path at a route that runs one cheap database query and returns 503 when it fails, or it passes every check while the database is gone. The curl line takes the path your users take, through public DNS and the Fly edge. The service check never leaves the platform.

What Fly.io monitoring leaves out

  • Alerts. A message means writing Grafana alert rules, or running your own Prometheus and Alertmanager. The old Slack and PagerDuty check handlers are gone.
  • Restarts. A machine failing its check stays up until the app exits with a non-zero code and the restart policy brings it back.
  • The outside view. A check inside Fly cannot see a custom domain pointing at the wrong place or a problem between your users and the edge.
  • Fly's own incidents. status.flyio.net reports the platform, not your app.

Fly.io uptime from outside, with machines that auto-stop

An outside check is a request like any other. With auto_start_machines on, it starts a stopped machine, and the proxy's stop loop only runs every few minutes, so a 15-second check keeps the app running for good. Logdash gives up after 10 seconds, so a cold boot slower than that reads as down. Either set min_machines_running = 1 and pay for one machine that stays up, or check every 5 minutes and accept that some checks measure boot time.

  1. Add the URL Create a service in Logdash and point its HTTP monitor at https://your-app.fly.dev/health, or the custom domain your users type. Every check stores the status code and response time.
  2. Pick the interval Every 5 minutes on the free plan, every minute on Builder, every 15 seconds on Pro. 200 to 399 is up. Any other code, or no answer in 10 seconds, is down.
  3. Break it on purpose Stop the database or deploy a /health that returns 503. On the next check the monitor flips to down and a Telegram message arrives with the monitor name, status code and error, then another on recovery.

Logdash vs Fly.io built-in checks

FeatureLogdashFly.io
Routes traffic away from a bad machineNo, it only watchesYes, that is what service checks are for
Holds back a broken deployNoYes, deploys wait for checks to pass
Tells you the app is downTelegram or webhook on every down and upNothing built in, Grafana alert rules you write
Check interval5 minutes free, 1 minute Builder, 15 seconds ProYou set it, 30 seconds by default
CPU, memory and OOM dataOnly what your app sends through an SDKBuilt in, about 15 days in Grafana
PriceFree for 5 servicesIncluded with the app

When Fly.io built-in checks is the better pick

  • You run two or more machines and only need traffic to avoid a broken one. Fly does that, Logdash cannot.
  • You already have Grafana alert rules sending to Slack or email. A second tool adds the outside view and little else.
  • You need machine CPU, memory and out-of-memory data. That is Fly metrics, not an uptime monitor.
What does Fly.io monitoring include?
Service checks that steer routing and deploys, a managed Prometheus with a Grafana instance at fly-metrics.net, and about 15 days of metrics. There is no built-in alerting, so a down app stays silent until you add Grafana alert rules or an outside monitor.
Do Fly.io health checks restart a machine?
No. A machine that fails its service check is marked unhealthy and taken out of routing, but it keeps running until someone restarts it. If you want a restart, make the app exit with a non-zero code when it is stuck and the default restart policy brings it back.
Do Fly.io health checks send alerts?
No. fly checks list shows the current state and the old Slack and PagerDuty handlers have been removed. For a message you need Grafana alerting on Fly metrics, or an outside monitor such as Logdash, which sends Telegram or webhook alerts.
How do I track Fly.io uptime for my app?
Point an outside HTTP monitor at a health route on your public hostname. status.flyio.net only reports the Fly platform, not whether your app answers. Logdash checks every 5 minutes on the free plan and every 15 seconds on Pro, and keeps the uptime history and response times.

Point it at your own URL and watch it for real.