Merge pull request 'ADR 0236, to-be 45 Phase 4: a build is judged on its first machine and put back by something other than itself' (#146) from feat/core-upgrades-that-roll-back into main
mesh/delivery held for a person: merged without a passing check: only a person decides that it goes on

This commit was merged in pull request #146.
This commit is contained in:
2026-10-06 17:15:09 +00:00
5 changed files with 346 additions and 3 deletions
@@ -61,6 +61,13 @@ is tried again. It is never reported as done "until the next push".
The plan records which machine went first, so a controller replaced mid-rollout resumes from there. A
policy of *together* keeps today's behaviour.
> **The mechanism changed — 2026-10-06, by [ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md).** What still stands: one machine
> first, the first by name, the rest only after it, a failed first machine stopping the module and the
> plan, *together* as it was. What moved: "reported applied and current" is no longer enough — the first
> machine is judged at a gate by the component's health, three times over at least two minutes within ten,
> before the rest are sent; and a first machine that fails, by its report or at the gate, is not left on
> the build that failed it: the previous build is registered again and sent to it, once per build.
**3. A newer merge takes over an older plan.** When a merge into a repository's branch makes a plan,
every open plan for the same repository and branch made before it is superseded, ordered by when each
plan was made, never by commit:
@@ -84,6 +84,13 @@ behind, held or not. It is the remedy `record` names when an upgrade is announce
when you want them"), and the one command for taking a recorded upgrade everywhere. Narrowing it would
leave `record` with no way to say "now".
> **The mechanism changed — 2026-10-06, by [ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md).** What still stands: a send records
> the builds it carried, a held machine is left and named, the named machine is sent everything, and
> `push --behind` takes a recorded upgrade everywhere. What moved: `record` is no longer the default — a
> build rolls out one machine first, gated, unless a person, the module, its irreplaceable data or the bus
> says otherwise — and one case of "the machine a push names is sent everything" is refused: a machine
> whose bus would move, which is replaced only as the planned step `bus upgrade`.
## Consequences
- `record` and "one machine first" hold a change back from every push that does not name the machine.
@@ -0,0 +1,272 @@
---
topic: the mesh
status: accepted
date: 2026-10-06
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md
---
# 236. A build is judged on its first machine and put back by something other than itself, and so it rolls out unattended
## Context
**The operator said "start phase 4"** of [to-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md)
on 2026-10-06: a health definition per core component, the gate on a plan's first machine, rollback by
a witness that is not the new build, the bus as a planned step ([ADR 0227](0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md)
rule 8). The same day's hand-act log says what the missing phase costs:
- **47 pushes by hand in one day**, the cause the log records most and the one that raised
`healer-wanted` (S15). By what they did: **21 walked a new node-engine build through the four machines
one at a time**, each pushed only after the one before was looked at; **8 carried a new controller's
grant into the bus's user list**, which the controller's own rollout does not send; **2 were `push
--behind`** to deliver catalogue builds that their policy held back; 16 were steps of other work a person
was walking through.
- **Every module but five had the upgrade policy `record`** — built on a merge, sent to no machine
until a person pushed. `record` was the column's default since [ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md)
§3, and nothing told a `record` somebody chose from the default it always was. The five that rolled out
were set by hand: the controller, the catalogue, the build agent, the intrusion filter and the records
module.
- **The one-machine-first rollout ([ADR 0218](0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md)
§2) judged the first machine by "reported applied".** A build that applies and then serves no tools,
crashes after the report, or makes the machine's own conditions fire passes that test. Nothing put a
failed build back: a first machine that failed stopped the plan, and the machine stayed on the build
that failed it.
- **The witness for the core is being built beside this.** mesh-host's node-engine keeps the previous
controller and node tools bundles beside the new ones, judges the new controller by the lease and the
node tools by their answer to the services protocol's ping, puts the previous build back when either is
not healthy in its bound, and says so in every report while it stands (mesh-host PR #40). The controller
must grant it what it reads, raise what it says, and never send what it put back again.
- **A merge that deleted a module's directory made the controller ask the build seat to build it.**
The merge of 2026-10-06 that folded `public-acme` into the proxy failed its plan: "has no module.json at
modules/public-acme"; the four other modules of its tier were built and never sent.
- **The bus's planned step had a snapshot to take since [ADR 0235](0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md)**:
the bus machine's backup holder snapshots every stream through the bus module's dump.
**Checked against GENESIS.** *Anything requiring a human to notice it will be noticed late*: a person
walking a build through four machines is that sentence, every day. *Failure must be loud*: a build put
back without saying it, or never put back, is the silent failure the core exists to end. Nothing here
conflicts with GENESIS.
## Considered Options
1. **Keep `record` as the default and heal the pushes** — a healer that pushes what is behind. Rejected:
it is a push with no judgement, the cascade of [ADR 0221](0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md)
with the hand taken out; a broken build would reach every machine as fast as before, and nothing would
put it back.
2. **Roll everything out, gate nothing, and rely on the conditions to say what broke.** Rejected: rule 8
says the core is judged by health, not by "applied"; a condition after a build is everywhere says what
broke after it broke everywhere.
3. **Gate only the core, and leave modules to `record`.** Rejected: the pushes are mostly the core's, but
the catalogue's are the ones that stay behind the longest, and a module has health the controller can
read as well as the core has.
4. **A health definition per core component, a gate on every plan's first machine that judges a module by
its own health, a rollback there by the ordinary path once per build, a witness on the machine for what
cannot put itself back, and then rolling out as the default.** Chosen.
## Decision
**1. A core component's health is a definition, written as probes.** In the self-check's registry
(to-be 45 §4), run every five minutes on every machine and by the gate on a first machine:
| Probe | Component | Healthy when |
|---|---|---|
| H-controller | controller | the lease is held, renewed within its age, by a controller that says it is ready: its self-check ran, and `status` answered in full within ten seconds in that run (D9) |
| H-engine | node-engine | every machine heard from has reported its current declaration, under a node-engine build it names; on its first machine, under the new build |
| H-tools | node tools | every machine heard from that runs them has them answering the bus's discovery |
| H-bus | bus | every stream and durable consumer the mesh defines is on the bus, and a request crosses it to the machines' node tools and back |
**2. Every release plan's first machine passes a gate before the rest are sent.** ADR 0218's first
machine is judged from the moment it was sent the build: it reported the build applied; no witness on it
put the build back; no condition was raised since about the machine, or about the module on it; and the
component's health holds — a core component's definition above, or a module's own: its tools are served
on that machine where it has tools and the machine runs the node tools that serve them. **Healthy three
times, at least forty seconds apart and two minutes after the send, within ten minutes of it.** A module
on one machine is judged the same way; only then does the plan go on. A policy of *together* is not
gated: it is the module saying it must change everywhere at once.
**3. A build that fails its gate is put back there, once, and never sent again on its own.** The plan
stops; the build is marked failed at its gate; the module's registered build goes back to the build the
first machine ran before — the newest successful build of that commit still kept ([ADR 0189](0189-the-store-keeps-what-the-records-name.md)
keeps five) — and that machine is sent it, by the ordinary send. The verdict is written before the send
and is never written over, so a controller replaced between the two does not send it twice. A build
marked failed is refused registration if its outcome is heard again, and `plans retry` refuses the plan;
a newer merge or a `rebuild` makes a new build, judged again. It is said as the condition
`build.<module>.<machine>.rolled-back` (a warning: the machine runs what it ran before) or
`core.<component>.<machine>.rolled-back` (urgent), `rollback-failed` (urgent, the operator's) when
nothing could be put back — no earlier build kept, the machine had never had the module, the send refused
— and as the event `rolled-back`. The probe DG keeps each until a newer build of the module passes.
**4. A build rolls out unattended by default.** With the gate and the rollback in place, a module's
build is sent one machine first, judged, then the rest, unless something says otherwise:
- **a person**, through `upgrade <module> roll-out|record|default`, `record` with why; kept, said with
who and why, and taken back by `default`;
- **the module**, in its manifest: `upgrade` with policy `roll`, `together` or `record`, the last two
with why;
- **its data**: a module that declares irreplaceable data ([ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md)),
its own or kept with a provider, records — a rollback cannot undo what a new build does to data that
cannot be had again;
- **the bus** records whatever anyone says (rule 6 below).
The store's `record` rows were the old default and become no choice; a person's `roll-out` is kept.
The catalogue keeps `record` where a module says why — the network path a rollback could not cross, and
the providers every consumer on a machine drops with, among them the two holding the photos. The
resulting policy of every module the mesh held on 2026-10-06:
| Policy | Modules | Why |
|---|---|---|
| record — the bus | nats | a planned step a person starts (rule 6) |
| record — irreplaceable data | plex, photos | the media library; the photos and their albums (the operator's ranking, ADR 0233) |
| record — the network path | dnsmasq, nftables, networkmanager, systemd-networkd, sshd | a build that cuts a machine off from the bus cannot be put back from outside it; sshd is the way in when the mesh cannot reach it, which no gate sees |
| record — a provider whose restart costs | postgres, mssql, mongodb, minio, keycloak | every consumer on the machine drops with it, a new version may change its data's format in place, and minio and mongodb hold the photos |
| roll — a person's choice, kept | mesh-controller, mesh-catalog, build-agent, fail2ban, records | — |
| roll — the default | every other module: 111 of the 129, ten of them on no machine, the node-engine and the node tools among them | gated on the first machine; the node-engine and node tools also witnessed on the machine |
The private network itself is provided by the controller and moves only with it; the modules that run on
no machine move nothing.
**4a. No build reaches a machine without a gate — the backlog included.** A send carries a machine's
whole declaration, so a plan sending one module, a cascade, a healer's resend or a whole-mesh push would
carry every other build waiting there. Under the old default builds were registered and sent nowhere;
on the day the default becomes `roll`, the next send of anything would restart them all at once, on every
machine. So:
- **a gated send carries everything waiting on its machine, and its gate judges all of it** — a plan's
first machine, a release plan's machine; a pass is each build's verdict, a failure puts back what was
found wanting, on the machines that were sent it;
- a module a send exists for may move where it goes — a policy of *together*, a rollback; a build that
passed a gate on one machine may go to the others;
- a person's send naming a machine (`push <node>`, the bus step) carries what it carries;
- **every other send is refused, or leaves the machine, while a build no gate has seen waits there** —
a plan's "rest", a cascade, a healer's, a whole-mesh push, the bus's user list carried — said with what
waits and the remedy;
- **a rebuild that made the same artifacts from the same manifest is no move.**
**The release plan walks what waits.** Whenever builds no gate has seen wait on machines and no plan that
has started walks them, the mesh opens a release plan: every such machine heard from, **one at a time,
the control node last**, each sent everything waiting there and judged by the gate before the next is
sent. A machine not heard from when its turn comes is left. A build asked outside a plan (a `rebuild`)
waits for it too, instead of being sent one machine after another unjudged. **A release plan that fails
puts back what failed and stops; the next one opens only when a person says `upgrade release-backlog
--why`**, and until then `release-held` says what waits. `upgrade backlog` lists it, read-only.
Chosen over holding the whole backlog for a person's release (safer by one human glance, but every
merge after the switch would then stall behind it) and over waves of a few modules (a send cannot carry
part of a declaration — ADR 0221's second option — so the bound that can be kept is one machine at a
time, which is the one kept). Measured on 2026-10-06 at the switch: the catalogue merge that rebuilt
103 modules for a change to the build agent made **88 of them byte-identical** to the builds before
(no move) and 15 different; every machine had already been sent all of them by hand that evening, so
**no build waited on any of the four machines** when this was decided.
**5. The controller's rollback is the node-engine's on its machine; the contract is written once on
each side.** mesh-controller `internal/lease/witness.go` and mesh-host `internal/witness/contract.go`,
held field for field:
- **What the host reads.** A direct get of the lease bucket's one key, `holder`: the lease's holder as
the controller writes it (instance, host, epoch, taken, renewed; anything else ignored). The node
principal of a machine assigned the controller is granted that one subject; every machine's node
principal is granted the ping of its own node tools. Reads only: the writers table still refuses any
write.
- **When it is healthy.** The holder's host is the machine, it took the key at or after the moment the
host started the new build (two seconds of skew), and renewed it within the key's fifteen seconds:
asked every five seconds, within sixty of the start. The node tools: they answer the ping within five
seconds, within sixty of the start.
- **What it starts.** The previous bundle it kept, as the previous declaration ran it.
- **What it says.** `rollbacks` on every report while the verdict stands — component, from, to, outcome,
why, when — and `witness` with the contract's version. The controller raises
`core.<component>.<machine>.<outcome>`: urgent for rolled-back, not-reversible, restore-failed and
halted; a warning for nothing-to-restore and unwitnessed; cleared by the first report without it. On a
first machine, a verdict made since the send fails the gate, so the build is marked and the registered
build put back.
- **What the controller adds.** It writes in the lease's value whether it is ready (rule 1), which the
host does not read and the gate does: the host rolls back a controller that never holds the lease, the
gate one that holds it and never becomes ready, by sending the previous build, which the host applies
as any declaration. A process's `witness` and `not-reversible` are the host's to read; the controller
sends neither yet — the host's defaults by name are the two witnesses above — and sends them only to a
machine whose report carries `witness`.
- **And the grant that follows the controller.** Once a new controller passes its gate, it sends the
machine holding the bus the user list it composes, when that differs from the one last sent and nothing
held back would go with it: the eight pushes of the day that carried a controller's grant by hand.
**6. The bus is never rolled out; its upgrade is a step a person starts.** Its policy is `record`
whatever is said, `upgrade` refuses it a roll-out, a plan builds it and sends nothing, a cascade holds its
machine (ADR 0221), and **no send reaches its machine while a new bus build waits for it** — a push naming
it, a plan's send for another module there, a healer's, a rollback's — each refused with the remedy. On
2026-10-06 a plan's send of the catalogue to the control node carried the bus's rebuilt image with it,
and the bus restarted under every machine with nobody having asked (`record` held a cascade, not the
machine a send was for). A rebuild that made the same artifacts from the same manifest is no move. A
change of the bus's **user list** is not a restart: the bus's image reloads its server in place when the
list it is written changes, and that stays the ordinary path. The step is `bus upgrade --why …`: refused unless the person says whether the new version can
be reverted (`--reversible`, or `--irreversible` as their explicit word that it runs anyway); the bus
machine's backup holder snapshots the streams first (ADR 0235) — a person who took one by hand says where
— and only then is the machine sent; recorded as a hand act; `bus-maintenance` (the probe DB) is open
while it runs, and the step ends done when the machine reported the new bus applied and H-bus passes, or
failed after fifteen minutes, said urgent with its snapshot as the way back while the bus is not healthy.
**7. A module deleted at its source is not built.** The forge's announcer says which of a merge's files
it deleted; a module whose manifest is among them is forgotten where nothing holds it and said otherwise
— never asked to build. A build that finds no manifest at a module's path, from an announcer that does
not say which files went, marks the module deleted in its plan, which goes on.
## Consequences
- **A merge reaches every machine with no hand**, one machine first, judged for at least two minutes —
and a build that breaks its first machine is put back there and goes nowhere else. The 21 pushes that
walked a node-engine build, and the 2 that delivered held catalogue builds, are the plan's.
- **A rollout is slower by the gate**: two minutes at the least per tier that rolls out, ten at the most
before a verdict. A plan with a module on one machine waits on that machine's gate before its next
tier.
- **The controller's grant gains one verb that acts**, `node-backup now`, used only by the bus step, and
an event, `rolled-back`; the installer's first user list carries both (mesh-host).
- **A new bus build stalls every send to the bus's machine until a person runs the step.** A plan whose
first machine is that machine waits, said in the plan and, past its bound, as `stalled`. That is the
point: the bus is replaced when a person is there, not as a side effect.
- **A change to the build agent rebuilds most of the catalogue** (its modules are built on what it
builds), which is how the bus came to be rebuilt by a merge that did not touch it. Whether that tiering
is right is left open here; a rebuild that changes nothing is at least no longer a move of the bus.
- **Thirteen modules still wait for a person**, each saying why in the catalogue or by its data; `status`
lists the machines behind them and `push <machine>` walks them, as before. A person who wants another
held says `upgrade <module> record --why …`.
- **What the gate cannot see.** A container that crash-loops after its compose applied is seen only
through what it breaks — its tools, a provider's failing word, the machine's own conditions — until the
node-engine reports container state, which is mesh-host's to add. A build whose migration cannot be
undone is put back all the same; marking such a build `not-reversible` for the witness is not composed
yet.
- **A merge after a release plan failed may stall** on a machine where a build waits for the person's
release: its first send carries and judges what waits there, but its "rest" waits. Said in the plan
and by `release-held`.
- **Rollout order**: mesh-host first (its genesis user list, and the witness, which reads nothing the
controller does not yet grant until then); the controller second, whose migration turns the store's
default `record` rows into no choice; the catalogue third — its manifests carry `upgrade`, which a
builder older than the controller refuses.
## How it is checked
| Rule | Checked by |
|---|---|
| the gate holds back the rest | the controller's test against a real store: a build that fails its gate on the first machine is put back there, registered at the previous build, marked, said as a condition and an event, never sent to the second machine, never registered or rolled back again, and `plans retry` refuses it |
| a passing build rolls everywhere | the controller's test: three healthy judgings, then the rest sent and the plan done, its verdict kept, no condition |
| the bound | the controller's test: a first machine never healthy fails at the bound and is put back |
| the health definitions | the controller's test per component: tools not served, a condition since the send, a witness's verdict, node tools not answering, a lease held by an older controller or one not ready |
| the policy | the catalogue's test: default roll, a module's word, irreplaceable data, the bus whatever it says; a record without why refused; the store's test: the current builds read the derived policy |
| the bus is never rolled out | the controller's test: the bus records whatever it says, a person's roll-out is refused, a plan sends nothing, no send may reach its machine while a new bus build waits — a rebuild with the same artifacts and manifest excepted — the step refuses without its word on reversibility and without its snapshot, and starts with both |
| no build without a gate | the controller's test: the backlog released one machine at a time, the first judged before the second is sent, each pass kept, a rebuild with the same bytes no move, a send that judges nothing refused; a release that fails puts back what it carried on its first machine, goes no further, holds the next until a person releases it with why; a cascade does not carry a build no gate has seen, and does once one passed |
| a deleted module is not built | the controller's test: a merge deleting a module's manifest asks no build and forgets it; a build finding no manifest leaves the plan going |
| the witness's contract | the lease package's test of the host's rule; the broker's test of the grants (the ping on every machine, the lease's key where the controller runs, no write); the controller's test that a witness's verdict is its condition while reports carry it and cleared after |
| live | the next merge to the catalogue sends one machine first and the rest after its gate, with no push; `plans <id>` shows the gate's record; the next week's hand-act log has no push for a build that rolled out |
## References
- [To-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) §8 and Phase 4 — the design
this builds and amends.
- [ADR 0218](0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md),
[ADR 0221](0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md),
[ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md) — the rollout, the hold, the policy,
whose mechanisms move here.
- [ADR 0233](0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md),
[ADR 0232](0232-a-binding-to-a-consumers-data-moves-only-by-a-person.md) — the data a rollout never
moves; [ADR 0235](0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md) — the snapshot the
bus step takes.
- mesh-controller, mesh-host, mesh-catalog: the branches `feat/core-upgrades-that-roll-back`;
mesh-host's witness `feat/core-upgrades-roll-back`.
+1
View File
@@ -204,6 +204,7 @@ python3 00-META/checks/index.py fail if stale
- **0230** — [A consumer the mesh stops asking for is retired, and deleted only by a person](0230-a-consumer-the-mesh-stops-asking-for-is-retired-and-deleted-only-by-a-person.md)
- **0231** — [A healer acts on what observation raised, and only observation says it worked](0231-a-healer-acts-on-what-observation-raised-and-only-observation-says-it-worked.md)
- **0234** — [The mesh holds a conversation with its operator, over channels that are seats, and an answer that performs an action is authorised by the controller](0234-the-mesh-holds-a-conversation-with-its-operator.md)
- **0236** — [A build is judged on its first machine and put back by something other than itself, and so it rolls out unattended](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)
### Its tiers, from the bottom up
@@ -4,6 +4,7 @@ status: in-progress
code: [mesh-controller, mesh-host, mesh-tools, mesh-catalog, mesh-sdk, mesh-lab]
updated: 2026-10-06
decisions:
- 02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md
- 02-DECISIONS/0234-the-mesh-holds-a-conversation-with-its-operator.md
- 02-DECISIONS/0233-a-module-declares-the-data-it-holds-and-the-mesh-protects-and-watches-it-from-that.md
- 02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md
@@ -241,7 +242,9 @@ producer that wrote it.
| D12 | every consumer of a provision that keeps its data is bound where it was last sent, or moves by a pin (`binding-kept`, `binding-moving`, `binding-moved`) | ADR 0232 |
| D13 | every item of data a machine declares is measured, is there, holds what it held, is written where it says it is, is backed up within its bound or sits on healthy redundant storage, and is no empty replacement of a copy kept elsewhere; an irreplaceable or valuable item a machine no longer declares is retired, not forgotten. Each machine's `node-backup` holder is asked `backed-up`; every provider of kept consumer data `provisioner_retirement` (with each held consumer's size) — see §2 for the kinds | ADR 0233, issue 273 |
| DW | the watchdogs of §3 ran within three of their intervals: the watchers are watched, and raise S10 when the self-check stops | rule 6 |
| H-* | the health probes of §8, run for every core component on every machine | rule 8 |
| H-* | the health probes of §8, run for every core component on every machine: H-controller, H-engine, H-tools, H-bus (as built, [ADR 0236](../../02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)) | rule 8 |
| DG | every build that failed its gate, and every core build a witness put back, is said — `build.<module>.<machine>.rolled-back`, `core.<component>.<machine>.<outcome>`, `rollback-failed` — until a newer build passes or the reports stop saying it (ADR 0236) | rule 8 |
| DB | a bus upgrade a person started is said as `bus-maintenance` while it runs, and ends healthy within its bound or is said `bus-upgrade-failed`, with its snapshot (ADR 0236) | rule 8 |
- **`doctor`** answers the last run's verdict at once: per probe, pass, fail or failed-to-run, and age.
**`doctor run`** runs now under a call id. **`doctor probes`** lists the registry; **`doctor
@@ -420,12 +423,51 @@ the other machines follow. "Reported applied" is not enough.
verdict, time to verdict, rolled back or not. Read through `plans`.
- **A rolled-back build is not retried** by the same plan. A newer merge makes a new plan.
**As built** ([ADR 0236](../../02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)):
- **Every release plan is gated, not only the core's.** A module is judged by its own health: it reported
applied, no witness put it back, no condition was raised since its send about the machine or about the
module there, and its tools are served on that machine where it has tools. Three healthy judgings at
least forty seconds apart and two minutes after the send, within ten minutes of it. A policy of
*together* is not gated.
- **The rollback is the ordinary path once per build.** The build is marked failed at its gate (never
registered again, `plans retry` refuses it); the module's registered build goes back to the build the
first machine ran before, still kept by ADR 0189; that machine is sent it. The verdict is written before
the send. Said as `build.<module>.<machine>.rolled-back` (warning) or `core.<component>.<machine>.
rolled-back` (urgent), `rollback-failed` (urgent, the operator's) when nothing could be put back, and as
the event `rolled-back`.
- **The witness, as built on the host** (mesh-host `internal/witness/contract.go`, the controller's half
`internal/lease/witness.go`): the controller's is the lease alone — the holder on this machine, taken
since the start, renewed within fifteen seconds, within sixty seconds of the start — and the node tools'
is their answer to the services protocol's ping within five seconds, within sixty. The controller's
"status in bound, doctor ran once" is the controller's own word in the lease's value, which the gate
reads and the host does not: a controller that holds the lease and never becomes ready is put back by
the gate sending the previous build. A witness says its verdict in `rollbacks` on every report while it
stands; the controller raises `core.<component>.<machine>.<outcome>` from it — urgent for rolled-back,
not-reversible, restore-failed and halted — and clears it with the first report without it.
- **No build reaches a machine without a gate**: a gated send carries and judges everything waiting on
its machine; every other send — a plan's rest, a cascade, a healer's, a whole-mesh push — is refused
or leaves the machine while a build no gate has seen waits there; a rebuild with the same artifacts and
manifest is no move. What waits is walked by a **release plan**, one machine at a time, the control
node last, each judged; one that fails holds the next until `upgrade release-backlog --why`.
- **A new controller that passes its gate sends the bus's machine the user list it composes**, when that
changed and nothing held back would go with it.
- **The default policy is to roll** (ADR 0236 §4); `record` stays where a module says why, keeps
irreplaceable data, or is the bus.
**The bus is a planned step.** A bus upgrade is a maintenance step a person starts through the
controller: the streams are snapshotted; a `bus-maintenance` condition is open for the step's duration;
the bus is replaced; afterwards D6, D7 and the round trip must pass, or the step is reported failed and
the snapshot is the way back. A step whose new version cannot be reverted (the bus's 2.10 → 2.11 is one)
says so before it starts and runs only on a person's explicit word, recorded as a hand act. Whether the
bus becomes a cluster that can be upgraded live is left to its own effort. *2026-10-06:* the snapshot
bus becomes a cluster that can be upgraded live is left to its own effort.
As built (ADR 0236): the verb is `bus upgrade --why … --reversible|--irreversible`; the snapshot is the
bus machine's backup holder backing the bus module up now ([ADR 0235](../../02-DECISIONS/0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md)),
or one a person took and names with `--snapshot-taken`; the step's bound is fifteen minutes; **no send
reaches the bus's machine while a new bus build waits for it** — a push, a plan's send for another
module, a healer's — after a plan's send restarted the bus unasked on 2026-10-06; a change of the bus's
user list is reloaded in place, not a restart. *2026-10-06:* the snapshot
and the way back from it exist — the bus image's own snapshot program, the same one the night's backup
runs, and a restore that builds a new store beside the live one for a person to swap in
([ADR 0235](../../02-DECISIONS/0235-the-bus-is-backed-up-by-its-own-snapshot-of-each-stream.md),
@@ -564,7 +606,7 @@ warning, not compared), and SMART read under an array.
| Repository | Delivers |
|---|---|
| mesh-controller | the health probes of §8; the gate on a core component's first machine; the rollout record in the plan; the bus maintenance step as a verb |
| mesh-controller | the health probes of §8; the gate on a core component's first machine — as built, on every plan's first machine (ADR 0236); the rollout record in the plan; the bus maintenance step as a verb; the default upgrade policy `roll` (ADR 0236) |
| mesh-host | keeping the previous controller and node tools builds; restoring one when its health is not met in bound, watching the lease bucket for the controller |
| mesh-tools | answering the health ping |
| mesh-lab | R3, R7, R8 |
@@ -574,6 +616,20 @@ one that starts and does nothing, one that crashes, one that cannot reach the bu
back with no hand, the mesh ends on the previous build, and a condition and a message say so. Live: the
next three core rollouts each record a health verdict.
**Built 2026-10-06** ([ADR 0236](../../02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md),
the operator's "start phase 4"): in mesh-controller (migration 0073), the probes H-controller, H-engine,
H-tools, H-bus, DG and DB; the gate on every plan's first machine, its record in the plan (`plans <id>`)
and in the store; the rollback by the ordinary path, once per build; the witnesses' verdicts as
conditions; the grants the witness reads; `bus` and `bus upgrade`; the default policy `roll`, `upgrade`
listing every module's policy and where it comes from; the user list carried after a controller passes;
a module deleted at its source forgotten, not built. In mesh-host, the witness (keeping the previous
controller and node tools, restoring one not healthy in bound, the launcher's for the node-engine) and the
installer's grants. In mesh-catalog, the modules that keep `record` say why, and the forge's announcer
says which files a merge deleted. On branches, not yet merged. **Not yet:** R3, R7, R8 on a lab mesh, and
so the *done when*; container state in the node-engine's report, without which a container that
crash-loops after its compose applied is seen only through what it breaks; composing `not-reversible` for
a build whose migration cannot be undone; the next three core rollouts' verdicts.
### Phase 5 — Checks before merge, and the replays
| Repository | Delivers |