Review before merge: the multi-holder boundary, the window's open race, the sweep's bounds

ADR 0201 gains the boundary found reading it back: a consumer keeping several
holders of a deriving provider is refused, because the two ends have no way
to agree. ADR 0189 gains two consequences — the sweep is bounded because it
runs inside a build, and an apply arriving mid-window reopens it.

That last one is issue 224, recorded rather than fixed: the host's rule for a
stopped container is to replace it, and while-stopped is the first thing that
makes a stopped container intentional. Both candidate fixes are decisions with
their own cost. Nothing is worse than it was; the store has never collected.
This commit is contained in:
2026-10-04 03:27:32 +02:00
parent 0231974226
commit a3523617d3
3 changed files with 89 additions and 0 deletions
@@ -0,0 +1,64 @@
---
status: open
opened: 2026-10-04
located-in: [mesh-host internal/apply/apply.go, mesh-host internal/apply/schedule.go]
fixed-by:
amended-design:
---
# 224 — An apply arriving during a maintenance window reopens it, by recreating the container the window is holding still
## What was observed
Reviewing [ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md)'s
`while-stopped` before merging it, 2026-10-04. Found by reading, not by running.
A scheduled step may hold its module's containers still while it runs. The host stops them, runs
the step, starts them again. Nothing tells the **apply** that a window is open, and the apply's
rule for a container it finds stopped is to replace it:
```
case existed && (before.Spec == want || legacy) && before.Running && len(reasons) == 0:
out.Action = "unchanged"
case existed:
rm -f
```
`before.Running` is false for a container a window is holding, so the second branch takes it:
the container is removed and recreated, **running**, in the middle of the step that required it
to be still.
## Why it matters
For the store, which is what the field was built for, the chain is: a push lands at 03:30 → the
apply recreates the registry → the registry accepts an upload from a build running at the same
time → `garbage-collect`, already past its mark phase, sweeps the blob that upload just wrote.
The image is then in the store with a layer missing, and the build that made it reported success.
Two things have to coincide, so it is not likely. It is also not rare enough to leave unsaid: the
mesh pushes on every merge, at any hour, and a collection over a store this size is minutes rather
than seconds.
**The general shape is the one that matters.** `while-stopped` is the first thing in the mesh that
makes a container's stopped state *intentional*. Everything else in the host reads "stopped" as
"broken, fix it", which is right everywhere else and wrong here. Any future use of the field
inherits this.
## What this is not
Not a regression. The store has never collected anything, so nothing is worse than it was; this
is a hole in something new rather than something that broke.
## Open questions
- **Should a window take the apply lock?** The daemon already serialises applies with `applying`
and, across processes, with `store.Lock`. A window that held it would make the race impossible.
The cost is that a push arriving mid-window waits for minutes, and a push that waits is what
[issue 185](../185-a-refused-membership-publish-stops-the-controller/00-report.md)'s
family of outages looked like from outside.
- **Or should the apply learn that a container is held?** Narrower: the scheduler says which
containers a window currently holds, and `applyContainer` reports those unchanged instead of
recreating them. Nothing blocks, and the apply tells the truth for the minutes it matters —
at the cost of a second source for "is this container meant to be running".
- Either way: should the *report* say a window is open, so a machine that looks half-stopped at
03:31 reads as working rather than broken?