ADR 0201 gains the boundary found reading it back: a consumer keeping several holders of a deriving provider is refused, because the two ends have no way to agree. ADR 0189 gains two consequences — the sweep is bounded because it runs inside a build, and an apply arriving mid-window reopens it. That last one is issue 224, recorded rather than fixed: the host's rule for a stopped container is to replace it, and while-stopped is the first thing that makes a stopped container intentional. Both candidate fixes are decisions with their own cost. Nothing is worse than it was; the store has never collected.
3.1 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design | ||
|---|---|---|---|---|---|---|
| open | 2026-10-04 |
|
224 — An apply arriving during a maintenance window reopens it, by recreating the container the window is holding still
What was observed
Reviewing ADR 0189's
while-stopped before merging it, 2026-10-04. Found by reading, not by running.
A scheduled step may hold its module's containers still while it runs. The host stops them, runs the step, starts them again. Nothing tells the apply that a window is open, and the apply's rule for a container it finds stopped is to replace it:
case existed && (before.Spec == want || legacy) && before.Running && len(reasons) == 0:
out.Action = "unchanged"
case existed:
rm -f
before.Running is false for a container a window is holding, so the second branch takes it:
the container is removed and recreated, running, in the middle of the step that required it
to be still.
Why it matters
For the store, which is what the field was built for, the chain is: a push lands at 03:30 → the
apply recreates the registry → the registry accepts an upload from a build running at the same
time → garbage-collect, already past its mark phase, sweeps the blob that upload just wrote.
The image is then in the store with a layer missing, and the build that made it reported success.
Two things have to coincide, so it is not likely. It is also not rare enough to leave unsaid: the mesh pushes on every merge, at any hour, and a collection over a store this size is minutes rather than seconds.
The general shape is the one that matters. while-stopped is the first thing in the mesh that
makes a container's stopped state intentional. Everything else in the host reads "stopped" as
"broken, fix it", which is right everywhere else and wrong here. Any future use of the field
inherits this.
What this is not
Not a regression. The store has never collected anything, so nothing is worse than it was; this is a hole in something new rather than something that broke.
Open questions
- Should a window take the apply lock? The daemon already serialises applies with
applyingand, across processes, withstore.Lock. A window that held it would make the race impossible. The cost is that a push arriving mid-window waits for minutes, and a push that waits is what issue 185's family of outages looked like from outside. - Or should the apply learn that a container is held? Narrower: the scheduler says which
containers a window currently holds, and
applyContainerreports those unchanged instead of recreating them. Nothing blocks, and the apply tells the truth for the minutes it matters — at the cost of a second source for "is this container meant to be running". - Either way: should the report say a window is open, so a machine that looks half-stopped at 03:31 reads as working rather than broken?