Sync Monitoring

Alerts

What opens an alert, how it clears itself, and what acknowledging actually does.

Everything else in the dashboard waits for you to look at it. Alerts are the part that does the looking: a job re-checks every active school on a schedule and records anything wrong as an incident you can see, claim, and watch close itself.

Nothing here is configured from the dashboard. The rules are fixed, the thresholds come from each school's own sync interval, and the schedule is ALERT_CRON, a server environment variable that defaults to every minute.

Detection runs separately from the API

Alerts are opened and closed by the same process that runs the retention sweep, not by the API. If only the API process is running, no alert is ever created — the bell reads Everything is healthy no matter how badly a school is failing.

What opens an alert

Each pass evaluates every active school and records the conditions firing right now.

AlertOpens whenSeverity
Agent offlineA school stops checking inWarning, then critical
Queue lagChanges sit unsent in either directionWarning, then critical
Run failedThe most recent sync run recorded an errorAlways critical
Conflicts openA school has unresolved conflictsWarning, critical past 100
Backfill failedA backfill request ended in failureWarning
Backfill stalledAn approved backfill stops releasing rows for 15 minutesWarning, critical past an hour
Retention failedThe last completed sweep recorded an errorWarning
Capture out of dateA database's capture triggers have not matched its sync tables for 30 minutesWarning, critical past a day

Queue lag is tracked per direction, so a school can have a separate alert for changes waiting online and changes waiting locally.

Retention failures belong to the deployment rather than any one school, so they appear without a school attached.

Capture drift is tracked per database. The 30 minutes is room to re-run the setup SQL after changing a table, which briefly puts the triggers out of date by design. A trigger left over from an operation since turned off never raises this alert, because what it captures is discarded; it shows only on the school's overview.

A stalled backfill matters more than its row count suggests: a request stuck part-way through its release blocks every later backfill for the same table, and nothing else reports it.

A school that has never checked in is not offline

Detection skips schools with no recorded check-in at all. That is a school whose agent has not been installed yet — a setup step, not a fault.

Thresholds follow each school's own clock

A school syncing every five minutes and one syncing every hour cannot be judged against the same stopwatch. Both the offline and lag rules are multiples of that school's configured interval rather than fixed durations.

ElapsedState
Up to 3 intervalsHealthy
Past 3 intervalsWarning
Past 10 intervalsCritical

So the hourly school is not called offline until three hours of silence, while the five-minute school is flagged in fifteen minutes. Severity is re-evaluated on every pass — a warning becomes critical on its own if the condition keeps getting worse, without opening a second alert.

One condition, one alert

The same problem detected a hundred times in a row is one row in the list, not a hundred. Each condition has a key, and while an alert for that key is live it is updated rather than duplicated.

That is also what closes them. Every pass compares what is firing against what is live, and anything that has stopped firing is resolved automatically:

The move to Resolved is never a human decision. An agent-offline alert closes when the agent checks in, queue lag closes when the queue drains, a failed run closes when the next run succeeds, and conflicts close when the last disputed row is settled.

Acknowledging

Acknowledge means someone is on this. It does not mean fixed.

Acknowledging stops the alert counting toward the unread badge and stops critical alerts re-announcing themselves, while leaving the incident live and recording who claimed it and when. A school that has genuinely been offline for three days should stay visible; it just should not keep interrupting the room about it.

Three states are easy to confuse:

ScopeMeaning
ReadJust youYou have seen it
AcknowledgedEveryoneSomeone is handling it
ResolvedEveryoneThe condition is gone

Marking everything read only clears your own badge. Acknowledging clears it for the whole team.

Where alerts show up

The bell in the header carries the unread count and lists what is live, with the same count mirrored on the sidebar. The count turns red when any unread alert is critical and amber otherwise. New critical alerts also raise a toast, once each — the backlog present when you open the app is never announced.

The Alerts page is the full view, split into Live and Resolved, with search and filters by severity, kind, school, and date.

Sending alerts to Discord

Set DISCORD_WEBHOOK_URL on the server to an incoming webhook and every alert is posted to that channel as it opens, colour-coded by severity, with the school, type, and a link back to the dashboard. When the condition clears, a green all-clear follows.

Leave the variable unset and alerts stay inside the app. There is nothing to configure in the dashboard either way.

A few details that matter in practice:

One message per incident, not per pass. Delivery follows the same deduplication as the list, so a school that has been offline for a week produced one message, not ten thousand.

A failed post is retried, not dropped. An alert is only marked as delivered once Discord accepts it, so a webhook outage delays notifications rather than losing them.

Recoveries are only announced for incidents that were announced. If the channel was never told about a problem, it is not told that the problem went away.

Delivery rides on the same process

The webhook is posted by the cron process, so it inherits the same dependency: no cron, no detection, and therefore nothing to send.

Limits worth knowing

There is no email or push delivery. Discord is the only outbound channel, and there is no per-user or per-school routing — every alert goes to the one webhook.

There is no manual resolve. For most kinds that is correct, since the detector can see whether the condition still holds. The exception is a failed backfill: it clears only when the request is retried, so one that is abandoned stays live indefinitely. Acknowledging it hides the badge but does not close it.

Acknowledging is not reflected in Discord. Claiming an alert in the dashboard quiets the badge, but the channel has no way to show that someone picked it up.

Nothing escalates. An acknowledged alert that is still live a week later looks exactly like one acknowledged a minute ago.

What happens to old alerts

Resolved alerts stay as history and are removed by the retention sweep once they pass the history window, measured from when they resolved. Unresolved alerts are never removed at any age — an open incident is still your problem however long it has been sitting there.

On this page