Sync Monitoring

Retention

What the nightly sweep removes, what it keeps, and why some rows are never pruned.

Every table the monitoring app writes grows with sync volume, and nothing in the normal flow deletes from them. Retention is what keeps the database bounded.

It is not something you configure from the dashboard. The schedule and the day counts are server environment variables, and the sweep runs in its own process. The outbox in each school's database is pruned by the agents instead.

Retention runs separately from the API

The server ships as two processes. If only the API process is running, nothing is ever pruned \u2014 there is no error and no signal in the dashboard, only a database that keeps growing.

What a sweep removes

RecordsKept forMeasured from
Run historyRETENTION_HISTORY_DAYSWhen the run was created
Applied rowsRETENTION_HISTORY_DAYSWhen the row was applied
Agent logsRETENTION_HISTORY_DAYSWhen the log was written
Finished backfill requestsRETENTION_HISTORY_DAYSWhen the request was created
Resolved conflictsRETENTION_HISTORY_DAYSWhen the conflict was resolved
Acknowledged queue rowsRETENTION_QUEUE_DAYSWhen the row was queued

RETENTION_HISTORY_DAYS defaults to 30 and RETENTION_QUEUE_DAYS to 7. The sweep itself runs on RETENTION_CRON, which defaults to three in the morning.

Deleting a run takes its per-table detail with it, so a run and the changes it carried always disappear together rather than leaving orphaned detail behind.

What is never removed

Three rules matter more than the day counts, because they are what stop retention from breaking sync.

Queue rows are only removed below the cursor. A row at or above the other side's acknowledged position may still be pulled, however old it is. A school whose online agent has been down for a month keeps its entire backlog — age alone never discards an unsent change.

Unresolved conflicts are kept at any age. The record is what holds the disputed row: an agent refuses that key for as long as a conflict exists for it. Pruning an old open conflict would release the hold and let the row be overwritten, so only resolved records are eligible and their clock starts when they were resolved.

Backfill requests still in flight are kept. Only completed, rejected and failed requests age out.

Queue rows are transport, not history

Once the far side has acknowledged past a queue row it can never be pulled again, which is why it is safe to prune far sooner than the history describing it. The run and applied-row records are what you read afterwards.

How a sweep behaves

Deletes are batched, with a ceiling on how many batches one table takes per sweep. That keeps a large first sweep from holding a long transaction open, at the cost of needing a few nights to catch up if the backlog is very large.

A sweep still running when the next tick arrives is skipped rather than queued behind itself. A sweep that throws is logged and abandoned, and the next tick starts fresh — nothing is left half-done, because each batch is its own delete.

The outbox in each school's database

The sweep only covers the monitoring server's own tables. Each school database has its own sync_outbox, which gains a row for every captured change and which the server cannot reach. Agents clean that up themselves: after each push they delete outbox rows that were pushed longer ago than RETENTION_OUTBOX_DAYS, which defaults to 7.

  • Only pushed rows are deleted. A change not yet sent is never touched, however old it is.
  • It is batched. At most 50,000 rows a cycle, in batches of 5,000, so an outbox with years of history clears over a few cycles without stalling the school's application.
  • The window comes from the server, sent to agents with their settings each cycle, so changing it needs no change at any school. Set it to 0 to keep pushed rows forever. Unlike the other two, the API process reads it, not the retention process.
  • It needs agent 1.1.0 or later. Older agents never prune.

A week of pushed rows is enough to answer "was this change captured?" for a recent problem. Nothing reads pushed rows otherwise.

If pruning fails — usually because the agent's database user may not DELETE from sync_outbox — syncing carries on, the agent says so in its output, and a warning appears in the school's agent logs.

Deleting rows does not shrink the file

MySQL reuses the space freed by pruning for new rows but does not return it to the disk. To reclaim it after clearing a large backlog, run OPTIMIZE TABLE sync_outbox; once, at a quiet time: it rebuilds the table and briefly blocks writes to it, including the triggers'.

Choosing the numbers

RETENTION_HISTORY_DAYS decides how far back you can investigate. Thirty days covers the usual case of someone reporting a problem well after it happened; a week is reasonable on a busy deployment where the tables are large and questions arrive quickly.

RETENTION_QUEUE_DAYS has little to do with investigation, since acknowledged queue rows are not what you read when diagnosing anything. Its default is mostly a safety margin — those rows are already spent, and the run history describes them.

Both have a minimum of one day. There is no way to switch retention off through configuration; if you need it off, do not run the process. RETENTION_OUTBOX_DAYS is the exception: 0 turns outbox pruning off.

On this page