Issue 108: the artifact store has never collected anything. Fifty-three repositories on the machine that serves everything else, and the only outcome of leaving it is a full disk reported as somebody else's failure. The mesh decides what may go — from its own build records, so it never names a digest it did not put there — and the store reclaims the bytes in a nightly window with its server held still. Deletion on the one door takes nothing a push did not already have. Designs 18 and 20 amended; issue 108 resolved. Also issue 202, found running the controller's suite: a module whose required setting nobody set is left out of the machine in silence, and dnsmasq became that module this morning.
8.4 KiB
topic, status, date, deciders, reconstructed, extends
| topic | status | date | deciders | reconstructed | extends |
|---|---|---|---|---|---|
| the mesh | accepted | 2026-10-02 | jochen | false | 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md |
189. The store keeps what the records name, and a maintenance step holds its writers still
Context
The mesh's artifact store has never collected anything (issue 108). Every build pushes another layer set; nothing has ever removed one. The predecessor ran a routine on a timer — stop the registry, collect, start it — and the conversion carried the settings that routine depends on without the routine, because the routine was a script beside the module and not a resource in it. The store now holds fifty-three repositories on the machine that serves everything else, and the only outcome of leaving it is a full disk reported as somebody else's failure.
Three things stood in the way, and the issue names all three.
Nothing in the mesh's vocabulary expresses a maintenance window. The collector requires every
writer stopped while it runs. A run-once step runs beside containers, not instead of them, and
a scheduled step is the same container on a cadence. There is no way for a module to say hold this
container of mine still while this runs.
Deletion is not enabled, and the door it would be enabled on has no accounts. The store is internal, reached by name over the overlay, trusted because being on that network is the permission (ADR 0082). The predecessor kept deletion behind an authenticated door, which it could, having one.
Nothing says what may be removed. The registry's own answer — collect everything no tag names — is wrong here. The mesh pushes each artifact under one moving tag and pins machines by digest, so every build but the newest is untagged and some machine may still be running it.
Decision
1. Deletion is enabled on the store's one door, and the overlay stays the permission. The objection dissolves on inspection: that door already accepts a push, and a writer who can push can replace any tag in the store with anything it likes. Delete takes nothing a push did not already have, and the machines that can reach the door are the ones the mesh's own filter admits (ADR 0168). Putting an authenticated door in front of deletion while leaving push open would be a lock on the window beside an open door, and it would cost the thing ADR 0082 bought: a store every machine can reach without a credential to distribute first.
2. The mesh deletes what it made and no longer keeps; the store reclaims the bytes. Two halves, each doing what only it can.
The mesh decides. It does not need to enumerate the store to do it — it has never put anything there it did not record, so every digest it could remove is already in its own build records. It deletes those manifests through the store's door, by digest, and remembers that it did.
The store reclaims. A deleted manifest frees no bytes until the registry's own collector walks
the storage with nothing writing to it, so the module declares that collector as a scheduled step
with the server held still for its duration. Plain collection, not --delete-untagged: what the
mesh keeps is still a manifest in the store, so it is still referenced, so its blobs stay — the
dangerous flag is not needed at all once the mesh is the one deciding.
3. What the mesh keeps, stated as three reasons rather than a number. A digest is kept because:
- a definition names it — every artifact reference in any module's current recorded manifest, which is what the mesh would hand a machine now. No age limit: this is the floor;
- the mesh can still go back to it — every artifact of the five most recent successful builds of each module, so a release that turns out wrong has somewhere to return to;
- nothing else. An artifact older than that, which no definition names, is what the store is carrying for no stated reason.
A digest the mesh did not record making is never touched. That is not a safety margin, it is the whole rule restated: the mesh removes what it put there and can account for, and the images genesis pushed before any record existed are exactly what this must not reach (04-ISSUES/102, F4).
4. A scheduled step may hold its module's own containers still while it runs —
while-stopped, naming resource ids in the same module. The host stops each, runs the step, and
starts them again whatever the step did, including when it failed or the host was interrupted.
Three boundaries:
- Its own module's containers only. A module that could quiesce a neighbour could stop the mesh; a maintenance window is a statement about one service's own insides.
- Scheduled steps only, not
run-once. At apply time the host already has a window: the declaration is applied in order and a step gates what follows, so a one-time offline migration says before rather than instead of. A recurring window is the case order cannot express. - Restoring is not conditional. A step that fails must leave the service running; the whole risk of this field is a window that never closes.
5. The sweep runs where the records change — after a build the mesh recorded. That is the moment new bytes landed and the moment the keep set moved, and it needs no new timer. The store's collection runs nightly, because reclaiming is slow and the thing it reclaims is already unreferenced.
Consequences
- Disk stops growing without bound on the machine that serves the mesh. That is the whole point and it has no other way to be true.
- A machine behind by more than five builds of a module, which recreates a container, cannot pull what it was running. It is already a machine the mesh reports as behind, and the answer is the one the mesh already gives it: the current declaration. Stated here rather than discovered.
- The store is a little less of a museum. A digest in an old build record may no longer be fetchable, and the record still says what that build made — the record is history, not an index of what is on disk. The collected mark is kept beside it so the two can be told apart.
while-stoppedis a second thing the host does to a container it did not start this pass. It is deliberately the narrowest form: the module's own, by id, restored unconditionally.- The store is briefly unavailable each night, for as long as collection takes. Everything that pulls from it retries; nothing in the mesh treats a momentary store as a failure (ADR 0185).
How this is checked
- The host: a scheduled step with
while-stoppedstops the named containers before the run and starts them after; it starts them again when the step fails; it refuses an id that is not a container of the same module, its own id, andwhile-stoppedon arun-oncestep. Each refusal is tested for what it says, not only that it says something. - The controller: given build records and current manifests, the keep set holds every reference a manifest names and every reference of the five most recent builds per module, and nothing else; a reference the mesh never recorded is never in the delete set; a delete that answers 404 is recorded as collected rather than retried forever.
- The sweep is tested against a fake store that records what it was asked to delete, so what is asserted is the decision and not the registry's behaviour.
- Live: the store's size before and after the first nightly collection, read from the machine.