Retention
What the nightly sweep removes, what it keeps, and why some rows are never pruned.
Every table the monitoring app writes grows with sync volume, and nothing in the normal flow deletes from them. Retention is what keeps the database bounded.
It is not something you configure from the dashboard. The schedule and the day counts are server environment variables, and the sweep runs in its own process. The outbox in each school's database is pruned by the agents instead.
Retention runs separately from the API
The server ships as two processes. If only the API process is running, nothing is ever pruned \u2014 there is no error and no signal in the dashboard, only a database that keeps growing.
What a sweep removes
| Records | Kept for | Measured from |
|---|---|---|
| Run history | RETENTION_HISTORY_DAYS | When the run was created |
| Applied rows | RETENTION_HISTORY_DAYS | When the row was applied |
| Agent logs | RETENTION_HISTORY_DAYS | When the log was written |
| Finished backfill requests | RETENTION_HISTORY_DAYS | When the request was created |
| Resolved conflicts | RETENTION_HISTORY_DAYS | When the conflict was resolved |
| Acknowledged queue rows | RETENTION_QUEUE_DAYS | When the row was queued |
RETENTION_HISTORY_DAYS defaults to 30 and RETENTION_QUEUE_DAYS to 7. The sweep itself runs
on RETENTION_CRON, which defaults to three in the morning.
Deleting a run takes its per-table detail with it, so a run and the changes it carried always disappear together rather than leaving orphaned detail behind.
What is never removed
Three rules matter more than the day counts, because they are what stop retention from breaking sync.
Queue rows are only removed below the cursor. A row at or above the other side's acknowledged position may still be pulled, however old it is. A school whose online agent has been down for a month keeps its entire backlog — age alone never discards an unsent change.
Unresolved conflicts are kept at any age. The record is what holds the disputed row: an agent refuses that key for as long as a conflict exists for it. Pruning an old open conflict would release the hold and let the row be overwritten, so only resolved records are eligible and their clock starts when they were resolved.
Backfill requests still in flight are kept. Only completed, rejected and failed requests age out.
Queue rows are transport, not history
Once the far side has acknowledged past a queue row it can never be pulled again, which is why it is safe to prune far sooner than the history describing it. The run and applied-row records are what you read afterwards.
How a sweep behaves
Deletes are batched, with a ceiling on how many batches one table takes per sweep. That keeps a large first sweep from holding a long transaction open, at the cost of needing a few nights to catch up if the backlog is very large.
A sweep still running when the next tick arrives is skipped rather than queued behind itself. A sweep that throws is logged and abandoned, and the next tick starts fresh — nothing is left half-done, because each batch is its own delete.
The outbox in each school's database
The sweep only covers the monitoring server's own tables. Each school database has its own
sync_outbox, which gains a row for every captured change and which the server cannot reach.
Agents clean that up themselves: after each push they delete outbox rows that were pushed
longer ago than RETENTION_OUTBOX_DAYS, which defaults to 7.
- Only pushed rows are deleted. A change not yet sent is never touched, however old it is.
- It is batched. At most 50,000 rows a cycle, in batches of 5,000, so an outbox with years of history clears over a few cycles without stalling the school's application.
- The window comes from the server, sent to agents with their settings each cycle, so
changing it needs no change at any school. Set it to
0to keep pushed rows forever. Unlike the other two, the API process reads it, not the retention process. - It needs agent 1.1.0 or later. Older agents never prune.
A week of pushed rows is enough to answer "was this change captured?" for a recent problem. Nothing reads pushed rows otherwise.
If pruning fails — usually because the agent's database user may not DELETE from
sync_outbox — syncing carries on, the agent says so in its output, and a warning appears in
the school's agent logs.
Deleting rows does not shrink the file
MySQL reuses the space freed by pruning for new rows but does not return it to the disk. To
reclaim it after clearing a large backlog, run OPTIMIZE TABLE sync_outbox; once, at a quiet
time: it rebuilds the table and briefly blocks writes to it, including the triggers'.
Choosing the numbers
RETENTION_HISTORY_DAYS decides how far back you can investigate. Thirty days covers the
usual case of someone reporting a problem well after it happened; a week is reasonable on a
busy deployment where the tables are large and questions arrive quickly.
RETENTION_QUEUE_DAYS has little to do with investigation, since acknowledged queue rows are
not what you read when diagnosing anything. Its default is mostly a safety margin — those rows
are already spent, and the run history describes them.
Both have a minimum of one day. There is no way to switch retention off through configuration;
if you need it off, do not run the process. RETENTION_OUTBOX_DAYS is the exception: 0
turns outbox pruning off.