Alerts
What opens an alert, how it clears itself, and what acknowledging actually does.
Everything else in the dashboard waits for you to look at it. Alerts are the part that does the looking: a job re-checks every active school on a schedule and records anything wrong as an incident you can see, claim, and watch close itself.
Nothing here is configured from the dashboard. The rules are fixed, the thresholds come from
each school's own sync interval, and the schedule is ALERT_CRON, a server environment
variable that defaults to every minute.
Detection runs separately from the API
Alerts are opened and closed by the same process that runs the retention sweep, not by the API. If only the API process is running, no alert is ever created — the bell reads Everything is healthy no matter how badly a school is failing.
What opens an alert
Each pass evaluates every active school and records the conditions firing right now.
| Alert | Opens when | Severity |
|---|---|---|
| Agent offline | A school stops checking in | Warning, then critical |
| Queue lag | Changes sit unsent in either direction | Warning, then critical |
| Run failed | The most recent sync run recorded an error | Always critical |
| Conflicts open | A school has unresolved conflicts | Warning, critical past 100 |
| Backfill failed | A backfill request ended in failure | Warning |
| Backfill stalled | An approved backfill stops releasing rows for 15 minutes | Warning, critical past an hour |
| Retention failed | The last completed sweep recorded an error | Warning |
| Capture out of date | A database's capture triggers have not matched its sync tables for 30 minutes | Warning, critical past a day |
Queue lag is tracked per direction, so a school can have a separate alert for changes waiting online and changes waiting locally.
Retention failures belong to the deployment rather than any one school, so they appear without a school attached.
Capture drift is tracked per database. The 30 minutes is room to re-run the setup SQL after changing a table, which briefly puts the triggers out of date by design. A trigger left over from an operation since turned off never raises this alert, because what it captures is discarded; it shows only on the school's overview.
A stalled backfill matters more than its row count suggests: a request stuck part-way through its release blocks every later backfill for the same table, and nothing else reports it.
A school that has never checked in is not offline
Detection skips schools with no recorded check-in at all. That is a school whose agent has not been installed yet — a setup step, not a fault.
Thresholds follow each school's own clock
A school syncing every five minutes and one syncing every hour cannot be judged against the same stopwatch. Both the offline and lag rules are multiples of that school's configured interval rather than fixed durations.
| Elapsed | State |
|---|---|
| Up to 3 intervals | Healthy |
| Past 3 intervals | Warning |
| Past 10 intervals | Critical |
So the hourly school is not called offline until three hours of silence, while the five-minute school is flagged in fifteen minutes. Severity is re-evaluated on every pass — a warning becomes critical on its own if the condition keeps getting worse, without opening a second alert.
One condition, one alert
The same problem detected a hundred times in a row is one row in the list, not a hundred. Each condition has a key, and while an alert for that key is live it is updated rather than duplicated.
That is also what closes them. Every pass compares what is firing against what is live, and anything that has stopped firing is resolved automatically:
The move to Resolved is never a human decision. An agent-offline alert closes when the agent checks in, queue lag closes when the queue drains, a failed run closes when the next run succeeds, and conflicts close when the last disputed row is settled.
Acknowledging
Acknowledge means someone is on this. It does not mean fixed.
Acknowledging stops the alert counting toward the unread badge and stops critical alerts re-announcing themselves, while leaving the incident live and recording who claimed it and when. A school that has genuinely been offline for three days should stay visible; it just should not keep interrupting the room about it.
Three states are easy to confuse:
| Scope | Meaning | |
|---|---|---|
| Read | Just you | You have seen it |
| Acknowledged | Everyone | Someone is handling it |
| Resolved | Everyone | The condition is gone |
Marking everything read only clears your own badge. Acknowledging clears it for the whole team.
Where alerts show up
The bell in the header carries the unread count and lists what is live, with the same count mirrored on the sidebar. The count turns red when any unread alert is critical and amber otherwise. New critical alerts also raise a toast, once each — the backlog present when you open the app is never announced.
The Alerts page is the full view, split into Live and Resolved, with search and filters by severity, kind, school, and date.
Sending alerts to Discord
Set DISCORD_WEBHOOK_URL on the server to an incoming webhook and every alert is posted to
that channel as it opens, colour-coded by severity, with the school, type, and a link back to
the dashboard. When the condition clears, a green all-clear follows.
Leave the variable unset and alerts stay inside the app. There is nothing to configure in the dashboard either way.
A few details that matter in practice:
One message per incident, not per pass. Delivery follows the same deduplication as the list, so a school that has been offline for a week produced one message, not ten thousand.
A failed post is retried, not dropped. An alert is only marked as delivered once Discord accepts it, so a webhook outage delays notifications rather than losing them.
Recoveries are only announced for incidents that were announced. If the channel was never told about a problem, it is not told that the problem went away.
Delivery rides on the same process
The webhook is posted by the cron process, so it inherits the same dependency: no cron, no detection, and therefore nothing to send.
Limits worth knowing
There is no email or push delivery. Discord is the only outbound channel, and there is no per-user or per-school routing — every alert goes to the one webhook.
There is no manual resolve. For most kinds that is correct, since the detector can see whether the condition still holds. The exception is a failed backfill: it clears only when the request is retried, so one that is abandoned stays live indefinitely. Acknowledging it hides the badge but does not close it.
Acknowledging is not reflected in Discord. Claiming an alert in the dashboard quiets the badge, but the channel has no way to show that someone picked it up.
Nothing escalates. An acknowledged alert that is still live a week later looks exactly like one acknowledged a minute ago.
What happens to old alerts
Resolved alerts stay as history and are removed by the retention sweep once they pass the history window, measured from when they resolved. Unresolved alerts are never removed at any age — an open incident is still your problem however long it has been sitting there.