Files
jschoubben a3523617d3 Review before merge: the multi-holder boundary, the window's open race, the sweep's bounds
ADR 0201 gains the boundary found reading it back: a consumer keeping several
holders of a deriving provider is refused, because the two ends have no way
to agree. ADR 0189 gains two consequences — the sweep is bounded because it
runs inside a build, and an apply arriving mid-window reopens it.

That last one is issue 224, recorded rather than fixed: the host's rule for a
stopped container is to replace it, and while-stopped is the first thing that
makes a stopped container intentional. Both candidate fixes are decisions with
their own cost. Nothing is worse than it was; the store has never collected.
2026-10-04 03:27:32 +02:00

3.1 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
open 2026-10-04
mesh-host internal/apply/apply.go
mesh-host internal/apply/schedule.go

224 — An apply arriving during a maintenance window reopens it, by recreating the container the window is holding still

What was observed

Reviewing ADR 0189's while-stopped before merging it, 2026-10-04. Found by reading, not by running.

A scheduled step may hold its module's containers still while it runs. The host stops them, runs the step, starts them again. Nothing tells the apply that a window is open, and the apply's rule for a container it finds stopped is to replace it:

case existed && (before.Spec == want || legacy) && before.Running && len(reasons) == 0:
        out.Action = "unchanged"
case existed:
        rm -f

before.Running is false for a container a window is holding, so the second branch takes it: the container is removed and recreated, running, in the middle of the step that required it to be still.

Why it matters

For the store, which is what the field was built for, the chain is: a push lands at 03:30 → the apply recreates the registry → the registry accepts an upload from a build running at the same time → garbage-collect, already past its mark phase, sweeps the blob that upload just wrote. The image is then in the store with a layer missing, and the build that made it reported success.

Two things have to coincide, so it is not likely. It is also not rare enough to leave unsaid: the mesh pushes on every merge, at any hour, and a collection over a store this size is minutes rather than seconds.

The general shape is the one that matters. while-stopped is the first thing in the mesh that makes a container's stopped state intentional. Everything else in the host reads "stopped" as "broken, fix it", which is right everywhere else and wrong here. Any future use of the field inherits this.

What this is not

Not a regression. The store has never collected anything, so nothing is worse than it was; this is a hole in something new rather than something that broke.

Open questions

  • Should a window take the apply lock? The daemon already serialises applies with applying and, across processes, with store.Lock. A window that held it would make the race impossible. The cost is that a push arriving mid-window waits for minutes, and a push that waits is what issue 185's family of outages looked like from outside.
  • Or should the apply learn that a container is held? Narrower: the scheduler says which containers a window currently holds, and applyContainer reports those unchanged instead of recreating them. Nothing blocks, and the apply tells the truth for the minutes it matters — at the cost of a second source for "is this container meant to be running".
  • Either way: should the report say a window is open, so a machine that looks half-stopped at 03:31 reads as working rather than broken?