GraphQL API monitoring
Monitor a GraphQL API through a GET readiness route next to /graphql that runs one database query and returns 503 when it fails, because uptime monitors read the status code and a GraphQL server answers 200 even when a resolver throws.
GraphQL breaks the usual uptime check twice. First the request: queries go out as a POST with a JSON body, and Logdash monitors send a plain GET with no body and no custom headers. Then the answer: when a resolver throws, the server still returns 200 with an errors array next to the data. A monitor that reads the status line, which is all Logdash reads, would stay green through a dead database.
So do not monitor /graphql. Monitor a route next to it that fails for the same reason the resolvers would fail, usually the database, with a status code a monitor can read. At 15-second intervals that route runs 5,760 times a day, so keep it to one select 1.
GraphQL health check
The usual advice is GET /graphql?query={__typename}. Two problems. Apollo Server 4 and later ship with CSRF prevention switched on, which answers a GET with 400 unless it carries a JSON Content-Type, an x-apollo-operation-name header or an apollo-require-preflight header. Logdash cannot add that header, so the monitor would be red from its first check. And __typename never reaches a resolver or the database, so it proves the process runs and nothing more.
GraphQL Yoga does most of this for you. It answers /health with a 200 as a liveness check, and the useReadinessCheck plugin serves /ready, which returns 503 when your check returns false or throws.
import { createServer } from 'node:http';
import { createSchema, createYoga, useReadinessCheck } from 'graphql-yoga';
import pg from 'pg';
const pool = new pg.Pool({ connectionString: process.env.DATABASE_URL });
const yoga = createYoga({
schema: createSchema({
typeDefs: 'type Query { hello: String }',
resolvers: { Query: { hello: () => 'world' } },
}),
plugins: [
useReadinessCheck({
endpoint: '/ready',
check: async () => {
try {
await pool.query('select 1');
return true; // 200
} catch {
return false; // 503, no body
}
},
}),
],
});
// /graphql for clients, /ready for the monitor.
createServer(yoga).listen(4000);On Apollo Server with Express, the same idea is an app.get('/ready') handler registered before the GraphQL middleware, running the same select 1 and answering 503 when it throws. Apollo Server 4 and later have no built-in health route, so this one is yours to write.
GraphQL uptime monitoring
- Ship the readiness route Deploy the server and open https://api.yourapp.com/ready. A 200 with an empty body is correct.
- Point a monitor at it Add a monitor with that URL. It checks every 5 minutes on the free plan, every minute on Builder and every 15 seconds on Pro, and stores the status code and response time of each check.
- Break it on purpose Connect Telegram and stop the database. /ready returns 503, the monitor flips to down on the next check, and a Telegram alert names the API and shows the 503.
What this does not catch: a resolver bug that only one query hits, or an expired token for the third-party API one field depends on. The server is up, the database answers, and that field returns an error inside a 200. Those are errors, not downtime. Log them from your error handler, or use a tool that sends real queries.
GraphQL monitoring with real queries
Logdash vs Checkly
| Feature | Logdash | Checkly |
|---|---|---|
| Send a GraphQL query as a POST | No, a GET with no body or headers | Yes, with a GraphQL body type and custom headers |
| Assert on the response body | No, status code and response time only | JSON body assertions, so an errors array can fail the check |
| Fastest check interval | 15 seconds on Pro | 1 minute on Hobby and Starter, 30 seconds on Team |
| Alert channels | Telegram and webhook | Email, Slack and webhook on the free plan |
| Free plan | 5 monitors, checked every 5 minutes | 10,000 API check runs a month |
When Checkly is the better pick
- You need to know that a real query returns real data, not only that the server and the database answer.
- Your API sits behind auth and the check has to send a token.
- A 200 carrying an errors array is the outage you actually worry about.