ADR 0189: the store keeps what the records name, and a maintenance step holds its writers still

Issue 108: the artifact store has never collected anything. Fifty-three
repositories on the machine that serves everything else, and the only outcome
of leaving it is a full disk reported as somebody else's failure.

The mesh decides what may go — from its own build records, so it never names
a digest it did not put there — and the store reclaims the bytes in a nightly
window with its server held still. Deletion on the one door takes nothing a
push did not already have.

Designs 18 and 20 amended; issue 108 resolved.

Also issue 202, found running the controller's suite: a module whose required
setting nobody set is left out of the machine in silence, and dnsmasq became
that module this morning.
This commit is contained in:
2026-10-02 21:48:48 +02:00
parent 0b9fc90885
commit b69ae663bc
6 changed files with 298 additions and 6 deletions
+26 -1
View File
@@ -5,11 +5,12 @@ code:
- mesh-catalog modules/showcase
- mesh-controller internal/builder
- mesh-sdk src
updated: 2026-09-30
updated: 2026-10-02
decisions:
- 02-DECISIONS/0150-a-modules-own-code-runs-as-supervised-processes-under-one-account.md
- 02-DECISIONS/0099-a-step-that-runs-once-names-what-it-reads.md
- 02-DECISIONS/0053-a-step-that-runs-on-a-schedule.md
- 02-DECISIONS/0189-the-store-keeps-what-the-records-name.md
- 02-DECISIONS/0074-the-wire-is-specified-not-the-types.md
- 02-DECISIONS/0040-what-a-module-is.md
- 02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md
@@ -208,3 +209,27 @@ is recreated with the new fact
([ADR 0099](../../02-DECISIONS/0099-a-step-that-runs-once-names-what-it-reads.md)). *How it is
checked:* the host's unit tests run a step again when its named file changed and not otherwise,
and recreate a container naming a step after the step ran.
## A recurring step may hold its own module's containers still
*Written 2026-10-02, from [ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md)
and [issue 108](../../04-ISSUES/108-the-registry-has-no-garbage-collection-once-it-has-two-doors/00-report.md).*
Some work cannot be done underneath a running service: an artifact store's collector walks the
storage and requires every writer stopped. A `run-once` step runs *beside* containers and a
scheduled one is the same container again, so until this a module had no way to say it — and the
mesh inherited a store that has never collected anything, because the predecessor said it with a
shell script and a script beside a module is not a resource in it.
A scheduled step may name `while-stopped`: resource ids of **its own module's** containers, which
the host stops before the run and starts again after it, in the reverse order, **whatever the step
did**. Three boundaries, each refused where it can be seen earliest — its own module's containers
only, because a module that could quiesce a neighbour could stop the mesh; scheduled steps only,
because at apply the declaration is applied in order and a step already gates what follows, so a
one-time offline job says *before* rather than *instead of*; and restoring that is not conditional
on anything, because the only real risk of the field is a window that never closes.
*How it is checked:* the host's unit tests assert stop–run–start in that order, the restart after a
step that **failed**, the reverse order for several containers, and a service left down said
loudly. The controller refuses, from the definition alone, a window with no schedule, one on a
run-once step, one naming a container the module does not declare, and one naming itself.