--- status: open opened: 2026-10-04 located-in: [mesh-host internal/apply/apply.go, mesh-host internal/apply/schedule.go] fixed-by: amended-design: --- # 224 — An apply arriving during a maintenance window reopens it, by recreating the container the window is holding still ## What was observed Reviewing [ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md)'s `while-stopped` before merging it, 2026-10-04. Found by reading, not by running. A scheduled step may hold its module's containers still while it runs. The host stops them, runs the step, starts them again. Nothing tells the **apply** that a window is open, and the apply's rule for a container it finds stopped is to replace it: ``` case existed && (before.Spec == want || legacy) && before.Running && len(reasons) == 0: out.Action = "unchanged" case existed: rm -f ``` `before.Running` is false for a container a window is holding, so the second branch takes it: the container is removed and recreated, **running**, in the middle of the step that required it to be still. ## Why it matters For the store, which is what the field was built for, the chain is: a push lands at 03:30 → the apply recreates the registry → the registry accepts an upload from a build running at the same time → `garbage-collect`, already past its mark phase, sweeps the blob that upload just wrote. The image is then in the store with a layer missing, and the build that made it reported success. Two things have to coincide, so it is not likely. It is also not rare enough to leave unsaid: the mesh pushes on every merge, at any hour, and a collection over a store this size is minutes rather than seconds. **The general shape is the one that matters.** `while-stopped` is the first thing in the mesh that makes a container's stopped state *intentional*. Everything else in the host reads "stopped" as "broken, fix it", which is right everywhere else and wrong here. Any future use of the field inherits this. ## What this is not Not a regression. The store has never collected anything, so nothing is worse than it was; this is a hole in something new rather than something that broke. ## Open questions - **Should a window take the apply lock?** The daemon already serialises applies with `applying` and, across processes, with `store.Lock`. A window that held it would make the race impossible. The cost is that a push arriving mid-window waits for minutes, and a push that waits is what [issue 185](../185-a-refused-membership-publish-stops-the-controller/00-report.md)'s family of outages looked like from outside. - **Or should the apply learn that a container is held?** Narrower: the scheduler says which containers a window currently holds, and `applyContainer` reports those unchanged instead of recreating them. Nothing blocks, and the apply tells the truth for the minutes it matters — at the cost of a second source for "is this container meant to be running". - Either way: should the *report* say a window is open, so a machine that looks half-stopped at 03:31 reads as working rather than broken?