Merge pull request 'Issue 295 and ADR 0242 (proposed): a recorded build moves only by a person's push, and a send says what it recreates' (#171) from issues/295-a-recorded-build-was-carried-by-a-plans-send into main

This commit was merged in pull request #171.
This commit is contained in:
2026-10-07 18:33:14 +00:00
3 changed files with 224 additions and 0 deletions
@@ -0,0 +1,131 @@
---
topic: what runs on it
status: proposed
date: 2026-10-07
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md
---
# 242. A recorded build moves only by a person's push, and a send says what it recreates
## Context
[ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)
§4 keeps `record` for three kinds of module: those whose new build a rollback could not undo, the
network path, and the providers whose restart costs. A person takes each of their builds. Its §4a
says that "a gated send carries everything waiting on its machine, and its gate judges all of it".
A send carries the machine's whole declaration
([ADR 0221](0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md)),
and that declaration is composed from the builds the mesh holds. A recorded module's build is
registered at its merge, so every send to its machine composes the new build. "Waiting" and
"judged" were only ever computed for modules that roll out.
On 2026-10-07 ([issue 295](../04-ISSUES/295-a-recorded-build-was-carried-by-a-plans-send/00-report.md)),
a catalogue merge adopted the images' own health checks in twelve modules. Its plan's gated sends to
its two first machines carried the new builds of the database (on both machines) and of the
document store (on the control node). Both modules are `record`. No gate judged them, and nobody
pushed. The same send recreated nine of mail's eleven containers at once for a change that brought
no new image, and the operator's mail was down until they were up again. The change plan had said
only that mail "receives" a build.
Measured that evening:
- Fourteen modules record: the bus, two holding irreplaceable data, five on the network path, five
providers and keycloak.
- Each of them, waiting on a machine, was carried by the next send there for any other module that
rolled out. That happened twice that evening.
- One was still waiting afterwards: the media library, rebuilt at 19:01.
## Considered Options
1. **Refuse every send but a person's to a machine where a recorded build waits**, as the bus is
held today. Rejected. One recorded build would stop every plan on its machine until a person
pushed it. On the control node that is the controller's own path, and [issue 280](../04-ISSUES/280-a-rebuild-of-an-unchanged-source-was-read-as-a-new-bus/00-report.md)
is the evening it cost.
2. **Do not register a recorded build until a person pushes.** Rejected. The mesh would no longer
hold what was built. `status` could not say a machine is behind, and a person's push would have
nothing to send.
3. **Leave the recorded module out of the declaration** ([ADR 0163](0163-taking-a-module-over-is-a-comparison.md) rule 6: its containers
untouched). Rejected. A left-out module contributes nothing, so its consumers on the machine would
lose what it provides to them in the same send.
4. **Compose the recorded module at the build its machine runs.** Chosen. The machine runs what it
ran, everything else in the send moves, and the person's push remains the one way the new build
arrives.
For the interruption:
5. **The node-engine recreates a module's containers one at a time where the module allows it.**
Not decided here. It is the node-engine's apply order, and mail cannot tolerate even one of its
front, smtp and imap containers going down unseen. It stays open as the better answer for
modules with replicas.
6. **Schedule a recreation-only change into a quiet window a module declares.** Not decided here.
It needs a window per module and a clock in the walk, and a person's choice of moment does the
same today.
7. **A module that people use directly declares `record` and says why, and every send says what it
recreates.** Chosen. It uses the policy that exists, which 1–4 make hold, and it marks the
interruption where people read the send.
## Decision
1. **A recorded build moves only by a person's push.** A person's push names a machine. Every other
send composes each recorded module at the build that machine was last sent: the manifest of that
build, from the build records, standing for the one the mesh holds. It also records that it still
carries that build. This covers a plan's send to its first machine and to the rest, a release
plan's, a rollback's, a healer's and a rotation's.
- `status` still says the machine is behind, and `push <node>` sends the new build.
- **The bus step is a person's word for the bus alone.** It moves the bus, and every other
recorded module on that machine stays.
- A module the machine was never sent is composed as the mesh holds it, since there is nothing
running to keep.
- If the build to keep is no longer in the records, the send is refused and the refusal names
`push <node>`. Composing the new build there would be the very move this rule stops.
§4a's "a gated send carries everything waiting on its machine" now reads: everything *that rolls
out*.
2. **A send says what it recreates.** For every module a gated send moves, it compares the build
the machine ran with the build it is sent, container by container. It says how many containers
it recreates and of how many, which ones, and whether any image is new or only the declaration
changed. It warns when it recreates more than one container at once, because the module's service
is interrupted until they are up again. This goes in the walk's record (`plans <id>`, the
delivery's walk), the plan's note and the controller's log.
3. **A module people use directly declares `record` and says why.** Mail is the first. A change to
how its containers are declared recreates them together, so a person chooses the moment. Its
image updates wait for a person too. A module that can be interrupted unseen keeps rolling.
## Consequences
- **Plans no longer restart providers behind a person's back.** In exchange, a recorded module's
machines stay behind until somebody pushes, as `record` always promised. `status` and `upgrade
backlog` already say so.
- **A declaration can now hold a build older than the one the mesh holds.** The resolution sees the
kept manifest for that machine only. Other machines' consumers still resolve against the
provider's registered manifest.
- **Pruning the build records can make a kept build disappear.** [ADR 0189](0189-the-store-keeps-what-the-records-name.md)
keeps five builds per module. If a recorded module falls more than five builds behind, sends to its
machine are refused, said with the remedy, until a person pushes. That is loud rather than silent.
- **The marking comes from the build, not before it.** The change plan at merge time cannot yet
know which containers a build changes. The walk says it at the send. Saying it before the merge
needs the planner to read the merged manifests, which is a later step.
- **The time interrupted is not yet a number.** The mesh measures a send-to-report time per
machine (`durations`), not how long one container takes to restart. The send says
"interrupted until they are up again". A per-container start time in the node-engine's report
would let it name a number.
How it is checked: the controller's `TestARecordedBuildIsCarriedOnlyByAPersonsPush` replays
issue 295 (a recorded provider and a rolled module on one machine, both rebuilt, the gated send). It
fails on the commit before the fix. A mesh-lab replay is owed with the issue.
## References
- [Issue 295](../04-ISSUES/295-a-recorded-build-was-carried-by-a-plans-send/00-report.md), the
evidence.
- mesh-controller PR #115: `recordedKept`, `keepRecorded`, `sendKeeps`, `inventory.ManifestAt`,
`catalogue.Recreates`.
- mesh-catalog PR #106: mail's policy.
- [ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)
§4, §4a; [ADR 0221](0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md);
[ADR 0240](0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md), whose first
declarations were the change that recreated mail.
+1
View File
@@ -339,6 +339,7 @@ python3 00-META/checks/index.py fail if stale
- **0233** — [A module declares the data it holds, and the mesh protects and watches it from that declaration](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)
- **0235** — [The bus is backed up by its own snapshot of each stream, taken under the bus module's account](0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md)
- **0240** — [A module says how it is healthy, and the node-engine judges it](0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md)
- **0242** — [A recorded build moves only by a person's push, and a send says what it recreates](0242-a-recorded-build-moves-only-by-a-persons-push-and-a-send-says-what-it-recreates.md) *(proposed)*
- **0243** — [The agent module removes a home item it did not place only on the person's word, and keeps a copy](0243-the-agent-module-removes-a-home-item-it-did-not-place-only-on-the-persons-word-and-keeps-a-copy.md)
### How it is built
@@ -0,0 +1,92 @@
---
status: located
opened: 2026-10-07
located-in: [mesh-controller (cmd/mesh-controller release.go, plan.go, push.go), mesh-catalog modules/mailu]
fixed-by: mesh-controller PR #115; mesh-catalog PR #106
amended-design:
---
# 295. A recorded build was carried by a plan's send, and mail went down with a health check
## Symptom
At 18:54 on 2026-10-07 a catalogue merge was delivered. It added the first health declarations of
[ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md)
to twelve modules: the images' own health checks, adopted by name. Nothing else changed in them. Its
plan rebuilt all twelve. Two minutes later two things happened:
- **On the control node, every mail container but two was recreated in the same minute.** Nine
were recreated; the antivirus and the cache were not, because neither was given a check. The
operator's phone reported mail errors, and a `mailu … smtp not running` condition was raised
and then cleared.
- **The database and the document store were recreated on the control node, and the database on
the anchor as well.** Their upgrade policy is `record` ([ADR 0236](../../02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)
§4): they are providers whose restart drops every consumer on the machine, so a person takes
each build. Nobody pushed.
The mesh's own record shows it:
- **The plan.** Ten of its modules carry a gate line. postgres and mongodb read only "built from"
the merge's commit: no gate judged them, because the gate judges only what rolls out.
- **The gate verdicts.** Ten were kept for the merge's commit: seven on the anchor, three on the
control node. There are none for postgres or mongodb.
- **What each machine was last sent.** After the sends, both machines are recorded as carrying
postgres at the merge's commit, and the control node mongodb too. Before the merge they carried
the previous commit.
- **The controller's journal**, read through the service manager seat's `journal` verb:
1. 18:55:52: postgres registered.
2. 18:56:05: "sent 7 module(s) to [the anchor] first in one send (baserow, grafana, matrix,
mosquitto, nodered, redis, supabase)".
3. 18:56:17: "sent [the control node] 480 resource(s)", then "sent 3 module(s) to [the control
node] first in one send (mailu, step-ca, website)".
4. 18:57:26: the control node applied.
- **What each machine reports.** `postgres.server`, `mongodb.server` and nine `mailu.*` containers
have run only since 18:57 on the control node, and `postgres.server` since 18:56 on the anchor.
No walk held back for later was involved. The sends were the plan's own sends to its first machines.
## Why it happened
1. **A send carries the machine's whole declaration, composed from the builds the mesh holds**
([ADR 0221](../../02-DECISIONS/0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md)). The controller registers a recorded module's new build at
the merge. From that moment, any send to a machine running the module composes the new build,
whoever the send is for.
2. **ADR 0236 §4a guards that only for modules that roll out.** "A gated send carries everything
waiting on its machine, and its gate judges all of it." The controller's list of moves waiting
for a gate leaves out a module whose policy records, because "a recorded one is a person's
push". So the gate did not judge the recorded builds, the refusal of ungated sends did not see
them, and the gated send carried them. `record` held only against the end of a person's push to
another machine (issue 259's `heldBack`) and against the bus (the bus step). It did not hold
against a plan's walk, a rollback, a healer or a rotation.
3. **Nothing tells a person, before or during a send, that it will recreate every container of a
module at once.** A change that only alters how containers are declared reads as harmless. A
health check, an environment variable or a label each recreates the container all the same,
and the node-engine applies a module's containers together. The delivery plan said mail
"receives" a build. Nothing said its service would be interrupted, and mail's policy let a plan
decide when.
## What is at stake
`record` is how the mesh keeps the providers, the network path and the irreplaceable data out of
unattended walks. Tonight showed that it holds only until some other module on the same machine
rolls out. This was still true after the incident: plex (`record`, irreplaceable data) was rebuilt
by a media-catalogue merge at 19:01 and sent nowhere. The next gated send to the anchor, for any
module, would recreate it.
## Fix
Located and fixed in pull requests, not merged. Decision: [ADR 0242](../../02-DECISIONS/0242-a-recorded-build-moves-only-by-a-persons-push-and-a-send-says-what-it-recreates.md).
- **mesh-controller #115.** Every send except a person's push composes a recorded module from the
manifest of the build the machine was last sent, and records that it still carries that build.
The bus step lets the bus alone move. A kept build that is gone from the records refuses the send
and names `push <node>` as the remedy. Every move a gated send carries now says how many of the
module's containers it recreates, and whether with a new image or only a changed declaration, in
the plan's note, the controller's log and `plans <id>`. The test
`TestARecordedBuildIsCarriedOnlyByAPersonsPush` replays this case and fails without the fix.
- **mesh-catalog #106.** Mail declares `record`: a module people use directly moves at a moment a
person chooses. It holds only once #115 is live, so it merges second.
A replay in mesh-lab's register is still owed before this resolves (ADR 0237). The controller test
above shows the replay's shape: a recorded provider and a rolled module on one machine, both
rebuilt, and the plan's gated send.