Files
hq/04-ISSUES/108-the-registry-has-no-garbage-collection-once-it-has-two-doors/00-report.md
T
jschoubben b69ae663bc ADR 0189: the store keeps what the records name, and a maintenance step holds its writers still
Issue 108: the artifact store has never collected anything. Fifty-three
repositories on the machine that serves everything else, and the only outcome
of leaving it is a full disk reported as somebody else's failure.

The mesh decides what may go — from its own build records, so it never names
a digest it did not put there — and the store reclaims the bytes in a nightly
window with its server held still. Deletion on the one door takes nothing a
push did not already have.

Designs 18 and 20 amended; issue 108 resolved.

Also issue 202, found running the controller's suite: a module whose required
setting nobody set is left out of the machine in silence, and dnsmasq became
that module this morning.
2026-10-02 21:48:48 +02:00

107 lines
6.7 KiB
Markdown

---
status: resolved
opened: 2026-09-23
located-in: [mesh-controller internal/inventory, mesh-controller internal/artifacts, mesh-host internal/apply, mesh-catalog modules/distribution]
fixed-by: 02-DECISIONS/0189-the-store-keeps-what-the-records-name.md
amended-design: 03-DESIGN/01-to-be/18-building-a-module.md
---
# 108 — The registry has no garbage collection, and two doors make it harder to add
## 2026-09-26 — the second door was attached to the wrong thing
*Left in the title and in the text below rather than rewritten, because the reasoning that assumed
two doors is what a reader needs to see retracted.*
Two separate things were treated as one. **The mesh's own artifact store holds the store seat**: it
is internal, reached by name over the overlay, with no accounts, because being on that network is the
permission ([ADR 0082](../../02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)).
**Serving a registry publicly is a service the mesh can host** — a module with its own name, its own
accounts and its own storage, like any other thing it runs for somebody. The conversion this report
was written beside gave the seat-holding store a second, public, authenticated door over the same
filesystem, which is neither of those: it is the internal store with an external face.
That second door is not being built. A publicly served registry, if one is wanted, is a module beside
the store rather than another way in to it, and it brings its own storage with it.
What that leaves here:
- **The complication in the title is gone.** One door means one registry process on the store's
filesystem, so the shared blob-descriptor cache and the deletion-cached-by-the-other-door problem
this report worried about do not arise at all.
- **The original issue is untouched, and is the whole of it.** The mesh's registry has no garbage
collection, never had, and the settings a collection routine depends on are not enabled.
- **One thing is sharpened rather than removed.** Enabling deletion on the only door enables it on a
door with no accounts, reachable by everything on the overlay. The predecessor kept deletion behind
its authenticated door — which it could, having one. So adding garbage collection now includes
deciding whether deletion is exposed on that door at all, or only ever performed by a routine the
mesh runs against its own store.
## What was observed
Reviewing the conversion that gives the mesh's image registry its public, authenticated name,
2026-09-23. The predecessor's registry module ran a maintenance routine on a timer: stop the
registry, `garbage-collect --delete-untagged`, restart, with a tag-retention step in front. The
mesh's registry has no such routine — it never did — and the conversion carried the settings the
routine depends on but not the routine.
The conversion also settled the public door as a **second registry process** on the same
filesystem, because the store's own door must stay account-free inside the mesh
([ADR 0082](../../02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)).
The registry's garbage collector requires every writer stopped while it runs. There are now two.
Nothing in the mesh's vocabulary says *stop these two containers, run a one-shot against their
shared volume, start them again*; a `run-once` step runs beside containers, not instead of them.
Meanwhile the merged store holds 53 repositories and every build the mesh will ever make pushes
another layer set into it. Disk grows without bound.
## Why it matters beyond this instance
A registry without garbage collection is a slow leak that presents as a full disk on a node that
serves everything else. The predecessor knew this and scheduled for it; the mesh forgot it in the
conversion because the routine was a script beside the module, not a resource in it — the shape
[issue 098](../098-taking-a-module-replaces-a-configuration-nobody-compared/00-report.md) describes
for configuration, here for behaviour.
It also asks something of the module system: a maintenance window over several containers is a
real thing services need, and the mesh cannot express one.
## Open questions
- Should the module system gain a maintenance step — a one-shot that quiesces named containers,
runs, and restores them — or is this better solved by a registry that does not need its writers
stopped (a storage backend the mesh does not run today)?
- Tag retention before GC: the predecessor kept the last N tags per repository; the mesh names
images by digest and moves by version — is retention "the digests no recorded build names"?
- Who owns the routine when the store and its public door are two modules — the store, since the
volume is its?
## Answered, 2026-10-02 — [ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md)
The three open questions, answered:
- **A maintenance step, or a backend that does not need its writers stopped?** The step. A
scheduled container may name `while-stopped` — resource ids of **its own module's** containers,
which the host stops before the run and starts again after it whatever the step did. A storage
backend the mesh does not run would be a bigger thing to own than the mechanism it avoids, and
the mechanism is wanted anyway: a service that cannot have work done underneath it is a real
shape and the mesh could not express it at all.
- **Is retention "the digests no recorded build names"?** Nearly. An artifact stays because a
definition the mesh holds names it (no age limit), or because it belongs to one of the five most
recent successful builds of its module. Last-N-tags was the predecessor's rule for a registry
that knew nothing else; this mesh knows what each digest is for.
- **Who owns the routine now the second door is gone?** Both halves, each where it can be. The
**mesh** decides what may go — only it holds the records — and asks the store to drop it. The
**store** reclaims the bytes, because only it can stop its own server. Neither half can be done
by the other.
And the sharpened point — enabling deletion on a door with no accounts — dissolved on inspection:
**that door already accepts a push**, so a writer who can reach it can already replace any tag.
Delete takes nothing a push did not have. What it does not do is undo ADR 0082's bargain, which
putting an authenticated door in front of deletion would have.
The second registry process is not built, as the 2026-09-26 note says, so the shared blob cache
and the deletion-cached-by-the-other-door problem never arise. Plain `garbage-collect` is enough:
what the mesh keeps is still a manifest in the store, so `--delete-untagged` — the flag that would
delete images machines are running — is not needed at all.