3 Commits
Author SHA1 Message Date
jschoubben 0b1ddaa7e9 to-be 44: in progress — the three guards in the controller 2026-10-05 15:32:21 +02:00
jschoubben e7126bc1e6 ADR 0217, to-be 44: no change to a machine takes effect unseen
Two incidents in two days had one shape: a change took effect that nobody saw first. Three guards
where the change is made — a settings layer read before it is replaced, a push that says what it
will change, a running module's data move acknowledged — each silent when nothing is at stake.
2026-10-05 15:14:28 +02:00
jschoubben d222091fe9 Issues 245, 246; 240 resolved, 242 located
245: a media server's previews were reached through a link its container never mounted — fixed by
mounting them. 246: settings set replaces the whole layer with nothing to read it first and no
history. 240 is fixed by mesh-controller#47; 242 has backups running and records what the rollout
taught.
2026-10-05 14:54:38 +02:00
28 changed files with 270 additions and 795 deletions
@@ -60,14 +60,6 @@ with the server held still for its duration. Plain collection, not `--delete-unt
mesh keeps is still a manifest in the store, so it is still referenced, so its blobs stay — the
dangerous flag is not needed at all once the mesh is the one deciding.
> **Progressive insight — 2026-10-05.** "What the mesh keeps is still a manifest in the store" was true
> of images and false of archives: the builder published every archive as a bare blob no manifest names,
> and the store's collector keeps only what a manifest names. Its first night would have deleted every
> archive the mesh keeps ([issue 253](../04-ISSUES/253-the-stores-collector-would-delete-every-archive-the-mesh-keeps/00-report.md)).
> The decision stands — the mesh decides, the store reclaims with plain collection. What changes is how an
> archive is published: with a manifest that holds it, so the sentence becomes true of archives too. Until
> every kept archive is held, the collector runs as a dry run.
**3. What the mesh keeps, stated as three reasons rather than a number.** A digest is kept because:
- **a definition names it** — every artifact reference in any module's current recorded manifest,
@@ -0,0 +1,81 @@
---
topic: the mesh
status: accepted
date: 2026-10-05
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md
---
# 217. No change to a machine takes effect unseen
## Context
Two incidents in two days had one shape: **a change took effect that nobody saw before it did.**
- Issue 241: a provisioner read an unreadable grants file as "nobody asks" and withdrew every
consumer; the postgres provider dropped seven databases. Nothing announced the withdrawal before it
happened.
- Issue 246: one placement was set for one module on one machine with `settings set`. The layer is
replaced whole, so the module's other settings on that machine — three directory placements, eight
media accesses, an exposure, four endpoints, the account's ids — were dropped, and the command
answered only "places". The machine's plan then ran the module on an empty configuration
directory. It was caught by reading the plan before a push, and the lost layer was read back from a
database backup a few hours old.
Since issue 241 the host no longer destroys what it did not make (ADR 0030 for directories, issue
241's fix for files) and no provider drops a store. What remains is the class above: the mesh doing
exactly what it was told, when what it was told was not what the person meant, and nothing showing
the difference until a machine had changed.
## Considered Options
1. **Confirm every push.** Rejected: every push would ask, and a confirmation asked every time is
answered without reading. The guard must speak only when something is at stake.
2. **Make `settings set` merge.** Rejected: removing a key would then need a second command, and the
layer stops being a statement of the whole (ADR 0046, ADR 0164). The fault is that the removal was silent,
not that it was possible.
3. **Three guards, each where the change is made, each silent when nothing is at stake.** Adopted.
## Decision
**A change that removes, moves or replaces something a machine is running is shown before it takes
effect; one that only adds is not interrupted.**
1. **A settings layer is read before it is replaced, and what a replacement removes is said.**
`settings show` (and the verb) reads a layer. `settings set` answers with the keys it adds, changes
and removes, and **refuses to remove a key unless told `--replace`**. Every change keeps the layer
it replaced, so the previous value is one command away, not in a backup.
2. **A push says what it will change.** `plan <node> --diff` lists what differs from what the machine
was last sent: resources added, changed and removed, and every container that will be recreated
with the reason. **`push` with no machine named is refused** unless `--all` says so.
3. **Moving a running module's data is acknowledged.** When a push would give a running container a
different host directory at a mount it already has — its data moving, or not following — the push
to that machine is held, naming the module, the mount and both directories; the other machines in
the same push go ahead. It is sent when the operator says so for that module (`--move <module>`).
None of the three adds a step to an ordinary change: a new module, a rebuilt image, a setting that
only adds keys and a push to a named machine behave exactly as before.
## Consequences
- **The controller keeps what it last sent each machine**, not only its digest, so a plan can be
compared with it. Migration 0014 declined to keep it — "a second account of what a machine should
be" — and that reason stands for what it was about: the copy is never read as what a machine
*should* be, only as what it *was told*, which is the one thing a comparison needs and the digest
cannot give.
- Settings carry a history; `settings show` reads it.
- A script that ran `push` with no machine must say `--all`.
- A changed placement for a running module needs one acknowledgement, once.
## How it is checked
Tests in the controller: `settings set` that drops a key without `--replace` is refused and changes
nothing; with it, the answer names the removed key and the history holds the previous layer. `plan
--diff` of an unchanged node is empty. A push changing a running container's mount source is held,
and goes ahead with `--move` for that module. `push` with no machine and no `--all` is refused.
## References
- [04-ISSUES/241](../04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md), [04-ISSUES/246](../04-ISSUES/246-setting-a-modules-settings-replaces-the-whole-layer-and-nothing-shows-it-first/00-report.md)
- ADR 0030 (data outlives its declaration), ADR 0046 (a module's configuration is its assignments), ADR 0164 (a setting is declared with its default and its cost), ADR 0214 (backups)
@@ -1,100 +0,0 @@
---
topic: the mesh
status: accepted
date: 2026-10-05
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md
---
# 218. A plan sends grants before code, rolls a module out one machine first, and a newer merge takes over an older plan
## Context
On 2026-10-05 the delivery path was watched through a day of merges, by several sessions at once. Three
things went wrong, each recorded as an issue with its evidence.
- **Code arrived before the right to use it** ([issue 249](../04-ISSUES/249-a-modules-new-state-is-refused-until-a-push-the-merge-did-not-make/00-report.md)).
A merge gave a module a new key-value state. The plan sent the new bundle to every machine, and only
then issued the memberships that grant the state. On three machines the module's new state was refused
for two minutes, until a push made by hand. The order is written into the code on purpose: memberships
"after the declaration, because the runtime it is for arrives with it". That reason holds only for a
first assignment, and even then a membership is kept on the bus for the runtime that connects later
([ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)).
- **No machine went first.** The module's upgrade policy sends one machine at a time, but a plan's rollout
ignores it and sends every machine running the module at once. One at a time also never waited for the
first machine to come up healthy: it stopped only if the publish itself failed. A change was therefore
everywhere before anything had seen it run.
- **Plans for successive merges ran over each other** ([issue 254](../04-ISSUES/254-plans-for-successive-merges-run-over-each-other-and-one-was-left-open/00-report.md)).
Three merges to the catalogue within four minutes made three plans. Each sent the build agent to every
machine and asked for the same builds. One was still "building" hours later, with nothing left for it to
wait on. [ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md) decides one plan per merge
and says nothing about the next merge arriving while one is open. [Issue 219](../04-ISSUES/219-an-older-build-that-finishes-later-replaces-a-newer-one/00-report.md)
settled only which build's output wins.
## Considered Options
1. **Debounce merges:** wait a window before planning, so close merges make one plan. Rejected: it only
delays the overlap, does nothing for merges further apart than the window, and makes every merge slower.
2. **Queue plans:** a new plan waits until the older one is done. Rejected: the older plan builds what the
newer merge is about to replace, then the newer one builds it again.
3. **A newer merge's plan takes over the older plan's unfinished work, a plan rolls a module out one
machine first, and grants travel before code.** Chosen.
## Decision
**1. Grants before code.** Every send — a plan's rollout and a push alike — issues the memberships for
the machines it is about to send to before it sends their declarations, after raising the buckets they
name. When the composed list of bus users changes, the machine that holds the bus is sent first, because
that list travels in its declaration. A membership that could not be issued fails the send, and the send
is tried again. It is never reported as done "until the next push".
**2. One machine first.** A plan rolls a module out according to the module's upgrade policy
([ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md) §3). Unless the policy says
*together*:
- the module is sent to one machine first, the first by name of the machines running it;
- the rest are sent only once that machine has reported the new declaration applied and current;
- a first machine that reports a failure, or does not report in time, stops the module's rollout there.
The plan names the machine and the reason, and the other machines keep what they ran.
The plan records which machine went first, so a controller replaced mid-rollout resumes from there. A
policy of *together* keeps today's behaviour.
**3. A newer merge takes over an older plan.** When a merge into a repository's branch makes a plan,
every open plan for the same repository and branch made before it is superseded, ordered by when each
plan was made, never by commit:
- the modules the older plan had not yet built join the newer plan's set, before its tiers are computed;
- the older plan ends in a state of its own, *superseded*, naming the plan that took it over.
Builds the older plan already asked for still finish and register; issue 219's ordering keeps the newer
one current. A person can also close a plan that waits on nothing, by its id. The plan is marked closed
by hand and never resumed.
## Consequences
- A module that gains a state, an event or a tool can use it from its first start on every machine.
- A change reaches one machine before the rest. A change that breaks its first machine stops there, with
the reason in the plan.
- Successive merges build each module once, for the newest commit. The build agent is sent to the
machines once per run of merges, not once per merge.
- **What got harder:** a rollout takes one machine's report longer than before. A module that must change
everywhere at once says *together* in its policy. A plan's record now has a superseded state that
readers of the plans must know.
## How it is checked
| Rule | Checked by |
|---|---|
| grants before code | the controller's test: a send records memberships issued before any declaration; the machine holding the bus is sent first when the user list changes; a failed membership fails the send |
| one machine first | the controller's test: with a one-at-a-time policy, one machine is sent, the rest only after its applied and current report; a failed first machine stops the module; *together* sends all at once |
| a newer merge takes over | the controller's test: an older open plan for the same repository and branch is superseded, its unbuilt modules folded in; a plan for another repository is left alone; a superseded plan is not open |
| live | the next merge to the catalogue that gives a module a new state: no refusal of that state on any machine, the first machine named in the plan, one plan open per repository |
## References
- [Issue 249](../04-ISSUES/249-a-modules-new-state-is-refused-until-a-push-the-merge-did-not-make/00-report.md), [issue 254](../04-ISSUES/254-plans-for-successive-merges-run-over-each-other-and-one-was-left-open/00-report.md)
- [ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md) — plans and tiers, extended here
- [ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md) — memberships, kept on the bus
- [to-be 30](../03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md) — the design this amends
+1 -1
View File
@@ -194,7 +194,7 @@ python3 00-META/checks/index.py fail if stale
- **0207** — [A module depends on the node seats that apply its resources](0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md)
- **0210** — [A tool's configuration is its seat holder's, and every other module extends it through the seat](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)
- **0212** — [A seat says what it receives, and the machine's hotkeys are a seat](0212-a-seat-says-what-it-receives-and-the-machines-hotkeys-are-a-seat.md)
- **0218** — [A plan sends grants before code, rolls a module out one machine first, and a newer merge takes over an older plan](0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md)
- **0217** — [No change to a machine takes effect unseen](0217-no-change-to-a-machine-takes-effect-unseen.md)
### Its tiers, from the bottom up
+1 -16
View File
@@ -5,7 +5,7 @@ code:
- mesh-controller cmd/mesh-builder
- mesh-controller internal/builder
- mesh-catalog modules/build-agent
updated: 2026-10-05
updated: 2026-10-04
decisions:
- 02-DECISIONS/0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md
- 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md
@@ -346,21 +346,6 @@ scheduled step with the server held still — which is what `while-stopped` exis
keeps is still a manifest and so still referenced, and the dangerous flag is not needed once the
mesh is the one deciding.
**An archive is held by a manifest of its own** (2026-10-05,
[issue 253](../../04-ISSUES/253-the-stores-collector-would-delete-every-archive-the-mesh-keeps/00-report.md)).
The sentence above held for images and not for archives: an archive was published as a bare blob that no
manifest names, and the store's collector keeps only what a manifest names, so a nightly collection
would have removed every archive the mesh keeps, the current ones included. So:
- an archive is published with a manifest that names it and nothing else, built from the archive's digest
and size alone so it can be computed again from the record;
- the sweep makes sure every archive it keeps is held that way before it lets anything go, and lets go of
an archive by removing its manifest first;
- the reference a machine fetches is unchanged.
The collector runs as a dry run until the controller reports no kept archive unheld; only then does it
collect for real.
A machine behind by more than five builds of a module, recreating a container, cannot pull what it
was running. It is already a machine the mesh reports as behind, and the answer is the current
declaration.
@@ -2,9 +2,8 @@
layer: to-be
status: proposed
code: []
updated: 2026-10-05
updated: 2026-10-01
decisions:
- 02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md
- 02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md
- 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md
- 02-DECISIONS/0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md
@@ -130,25 +129,6 @@ controller replaced mid-plan resumes from the store. `status` lists open plans a
has waited too long. The transition discipline for breaking changes in the list above is still
unwritten, and still the next thing.
## How a plan sends (2026-10-05)
Revision, [ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md). Three rules on how a plan delivers what it built.
- **Grants travel before code.** A send issues the memberships for the machines it is about to send to
before their declarations. The machine holding the bus is sent first when the list of bus users changes.
A membership that could not be issued fails the send.
- **One machine first.** Unless a module's upgrade policy says *together*, a plan sends it to one machine,
the first by name, and to the rest only once that machine reports the new declaration applied and
current. A first machine that fails stops the module's rollout there, with the reason in the plan.
- **A newer merge takes over.** A merge's plan supersedes every older open plan for the same repository
and branch, and takes in the modules they had not yet built. A plan that waits on nothing can be closed
by hand, by its id.
What a merge changed is read from the forge whole, page by page
([issue 252](../../04-ISSUES/252-a-merges-changed-modules-were-read-wrong/00-report.md)). A changed path in a
module directory the mesh does not hold yet is that module's own, not shared code, when its definition is
among the changed paths.
## Why now, and why not yet
**Why it matters:** self-update is the difference between a mesh a person maintains by typing
@@ -0,0 +1,60 @@
---
layer: to-be
status: in-progress
code:
- mesh-controller: cmd/mesh-controller/unseen.go, cmd/mesh-controller/push.go (--all, --move, the hold), cmd/mesh-controller/plan.go (--diff), cmd/mesh-controller/modules.go (settings show, --replace), internal/inventory/unseen.go, internal/inventory/migrations/0057-what-a-change-replaces.sql
updated: 2026-10-05
decisions:
- 02-DECISIONS/0217-no-change-to-a-machine-takes-effect-unseen.md
- 02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md
---
# 44 — No change to a machine takes effect unseen
**A change that removes, moves or replaces something a machine runs is shown before it takes effect;
a change that only adds goes through as before** (ADR 0217). Three guards, each at the place the
change is made.
## 1. A settings layer is read before it is replaced
A module's settings are layers — the whole mesh, and one per machine — each replaced whole when set.
That stays. Around it:
- **Reading.** `settings show <module> [--node]` prints the layers as they are, and the console verb
answers the same.
- **Saying what changed.** Setting a layer answers with the keys it adds, changes and removes,
compared with the layer it replaces.
- **Refusing a silent removal.** A set that would remove a key is refused, naming the keys, unless
`--replace` says the removal is meant. A set that only adds or changes keys needs nothing.
- **Keeping the previous layer.** Each set and each clear records the layer it replaced and when, so
the previous value is read back with `settings show --history`, not from a backup.
## 2. A push says what it will change
- **What was sent is kept.** Each send records the declaration it sent the machine, beside the digest
already kept. It is read only to compare; what a machine *should* be is still composed from the
mesh's records every time.
- **`plan <node> --diff`** compares what would be sent now with what was sent last: resources added,
removed and changed by id, and for a changed container the fields that differ — which is what
recreates it. A machine with nothing to change shows nothing.
- **`push` with no machine is refused** unless `--all` names that intent. A push to one machine, or
`--behind`, is unchanged; the console verb never pushed every machine and still does not.
## 3. Moving a running module's data is acknowledged
Before sending, each machine's declaration is compared with what it was last sent. **A container
that keeps a mount at the same place inside it, with a different directory on the machine behind
it**, is a module's data moving — or, as on 2026-10-05, a module about to run on an empty directory
because a placement was lost. That machine's push is held, saying the module, the mount and both
directories; the other machines in the same push go ahead. `push … --move <module>` sends it.
A first send to a machine, a new container, a mount added or removed, and a changed image are not
moves and are not held.
## How it is checked
Controller tests: a set dropping a key without `--replace` is refused and changes nothing, and with
it the answer names the removed key and the history holds the layer it replaced; `plan --diff` of an
unchanged machine is empty and names a changed container's changed field; a push changing a running
container's mount source is held for that machine only and goes with `--move`; `push` with no machine
and no `--all` is refused.
@@ -1,8 +1,8 @@
---
status: resolved
status: open
opened: 2026-10-04
located-in: [mesh-host internal/apply/apply.go, mesh-host internal/apply/schedule.go]
fixed-by: mesh-host PR #23
fixed-by:
amended-design:
---
@@ -62,9 +62,3 @@ is a hole in something new rather than something that broke.
at the cost of a second source for "is this container meant to be running".
- Either way: should the *report* say a window is open, so a machine that looks half-stopped at
03:31 reads as working rather than broken?
## Resolved — 2026-10-05
An apply that arrives during a window now leaves the containers the window holds alone, and reports them `held-still`, naming the step. The first apply after the window converges them. The window is recorded under the host's state directory, the one place every applier on a machine shares: the daemon, a hand-run apply and the installer. It is released on every path, and lapses after six hours or when its process is gone. The machine's report lists open windows. Live on all four machines the same evening. Of the issue's two options, the narrower was taken: nothing blocks, so a push is never held for the length of a window.
Accepted and said in the change: an apply that inspected a container as running in the instant a window opens can still recreate it.
@@ -1,8 +1,8 @@
---
status: open
status: resolved
opened: 2026-10-04
located-in: []
fixed-by:
located-in: [mesh-controller cmd/mesh-controller/main.go, mesh-controller cmd/mesh-builder, mesh-controller internal/link/build.go]
fixed-by: mesh-controller#47
amended-design:
---
@@ -35,3 +35,11 @@ acts on it, review becomes a formality: the unreviewed definition reaches a mach
1. Where does the dry run's outcome enter the record — the build machine's `built` event, consumed as
any other build's?
2. Did the controller roll the dry run out, or did a later push compose from it?
## Resolution (2026-10-05)
A dry run is marked on the request (`DryRun`), the builder echoes the mark on its outcome, and the
controller's daemon sets a marked outcome aside: no record, no registration, no plan, nothing a push
could send (mesh-controller#47, with a test that the daemon takes a dry run in with no store at all).
Left open as a follow-up: the catalogue module also hears build outcomes and records their edges; it
should skip a dry run too.
@@ -1,8 +1,8 @@
---
status: open
status: located
opened: 2026-10-05
located-in: []
fixed-by:
located-in: [mesh-controller internal/catalogue/seats.go, mesh-catalog modules/restic]
fixed-by: mesh-controller#49, mesh-catalog#49, mesh-catalog#54, mesh-catalog#55, mesh-catalog#56, mesh-media-catalog#1
amended-design: 03-DESIGN/01-to-be/43-backups-against-mistakes.md
---
@@ -47,3 +47,27 @@ research 030; proposed as ADR 0214.
Issue 241's recovery cost a night and lost the forge's records of twelve days. With a nightly backup
held on another machine, it would have been a ten-minute restore of yesterday.
## Where it stands (2026-10-05)
Backups run (ADR 0214, to-be 43): the control node keeps nightly restore points of its eight stores
and services, the home server of its databases and its media apps' libraries and cover art; each on
its own machine, on its larger filesystem. The first runs were tried one module, then two, then a
whole node, each proven by a restore beside the live data. Not yet built, which is why this stays
open: the weekly test restore into a throwaway instance, the 48-hour status line, and failures
reaching the operator rather than the holder's log.
What the rollout taught, for the next module that takes contributions:
- **A contribution makes the contributing module depend on the seat.** Merging the stores'
contributions before a holder was assigned left every node's plan unresolvable — the whole mesh,
not only the machines running a store — until the holder was assigned. Assign the holder in the
same step as the merge.
- **A brand-new module is not built by the push that adds it**; its first build is asked for by hand.
- **The holder must look as the account that can see.** Its first run called a store's dumps missing
because it checked as the runtime's account, which cannot see inside the store's own directory.
- **A kept single file is not a directory.** restic restores a snapshot's subfolder, not a file;
the first restore of a module keeping its settings file refused it.
- **An update kills a running backup** — the holder's restic is its child. A hand-off to the new
version, the run living on as its own unit under the machine's service manager, is the idea to
take forward.
@@ -1,8 +1,8 @@
---
status: resolved
status: located
opened: 2026-10-05
located-in: [mesh-catalog modules/claude-code, mesh-catalog modules/claude-licence-manager]
fixed-by: mesh-catalog PR #45
fixed-by:
amended-design:
---
@@ -31,9 +31,3 @@
**Ruled out.** The bus delivered every binding: each machine that took the fourth binding did so in the
same second it was published. The licences themselves were sound. The remaining licence refreshed
on every attempt, and the second account's login was adopted from its first report.
## Resolved — 2026-10-05
Live on all four machines and the manager the same morning. After the restart every machine reported
the generation it held, and the manager's bindings matched them. The manager logged no failure while
moving its sequence past them.
@@ -0,0 +1,39 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-media-catalog modules/plex]
fixed-by: mesh-media-catalog#2
amended-design:
---
# 245 — A media server's previews were reached through a link its container never mounted
## What was observed
On the home server the media server's scrub previews — 423 GB, about 57 000 preview files, one per
video — had not grown in seven months: the newest was from the day its preview folder was moved off
the server's own disk onto the large storage pool. The server's log repeated, live, that it could not
create a directory under its preview folder.
## Why
The move left the preview folder as a **symbolic link** inside the server's configuration directory,
pointing at the pool. The container mounts the configuration directory and the media libraries, and
not the pool's path, so inside the container the link pointed at nothing: the server could neither
show the previews it had nor make new ones, and said so only in its own log. Nothing the mesh reports
showed it — the container ran, and answered.
It is the failure the mesh's rule against symbolic links exists for: a link resolves differently in
every place that reads it, and a container is such a place.
## Resolution
The previews are a directory of the module, mounted at the server's preview path; where that
directory lives on a machine is the machine's placement setting, here the pool path the files were
already in. The link was removed at the cutover, the container recreated, the preview folders visible
inside it again and the log's errors gone. The previews are not backed up, by the operator's choice.
## How it is checked
The catalogue check that every mount is declared passes over the module. On the machine: the preview
folder seen from inside the container lists the same folders as the pool path.
@@ -1,50 +0,0 @@
---
status: open
opened: 2026-10-05
located-in: []
fixed-by:
amended-design:
---
# 245 — `status` calls a module behind when only its repository moved
## What was observed
2026-10-05. After three catalogue merges that changed 11 modules, `status` listed **69 modules
behind their source**, each as `holds <older commit>, source has <newer commit>`, with the advice
"`build --behind` builds them; `push --behind` sends them on".
Between the two commits, `git diff --name-only` shows changes under 11 module directories only. For
58 of the 69 modules listed, for example a Bluetooth module, the container runtime's, the forge's,
the package manager's and the bus's, the diff of the module's own path is empty. Their sources did
not change; only the repository's commit did.
The merges' plans were right: they rebuilt the changed modules and those that depend on them, by
tier. Only the report was wrong. An agent following the report's own advice ran `build --behind`,
which rebuilt all 69. The rebuilt bus module was rolled out, and its container was replaced on the
control node. Every node's runtime lost the bus for about a minute.
## Why it matters beyond this instance
"Behind" is the word a person and an agent act on, and `status` attaches a command to it. A module
is held to a commit of its repository, so after any merge almost every module of that repository
reads behind. The list then says nothing about what needs building: it hides the few modules that
really are behind among the many that are not, and it invites a rebuild of everything, which is not
a harmless act (above).
## The operator's direction (2026-10-05)
The only truth is the outcome of the build plan. A plan already decides, from a change, which modules
it affects: those whose sources changed and those that depend on them, tier by tier. "Behind" means
a module that a plan has decided to rebuild and has not yet rebuilt or rolled out, and nothing else.
No second comparison beside the plan is made, whether of commits, of folders or of files, because a
second answer to the same question is how the two came to disagree.
## Open questions
1. Where does `status` read "behind" from today, and what replaces it: the open plans' remaining
tiers?
2. What does a module's recorded commit mean once a plan that leaves it untouched has run? Does it
move forward, or does the record stop carrying a commit that only says when it was last built?
3. Should `build --behind` and `push --behind` take their lists from the same place, so that they can
never act on a module no plan named?
@@ -0,0 +1,45 @@
---
status: open
opened: 2026-10-05
located-in: [mesh-controller cmd/mesh-controller (settings), mesh-controller internal/inventory (SetSettings)]
fixed-by:
amended-design:
---
# 246 — Setting a module's settings replaces the whole layer, and nothing shows it first
## What was observed
An operator's agent set one placement — where a media server's preview folder lives on one machine —
with `settings set <module> {"places": {…}} --node <machine>`. The command answered that the setting
was recorded. The machine's plan then mounted the server's configuration from an empty default
directory: the node's layer had held the placements of three other directories, eight media
accesses, a public exposure, four endpoints and the account's ids, and every one of them was gone.
Caught before any push, by reading the plan. The previous layer was read back from the controller
database's nightly dump — the backups that issue 242 asked for, a few hours old.
## Why
The layer is a statement of the whole, by design (the inventory's `SetSettings`: "replacing rather
than merging … removing a key is done by leaving it out"). That design is sound; what is missing
around it is everything that makes it safe to use:
- **There is no way to read a layer.** `settings` has `set` and `clear`, no `show`; the console verb
likewise. To change one key, a person must already know every other key in the layer.
- **There is no history.** The row is updated in place; the previous values exist nowhere but a
database backup.
- **The answer does not say what was dropped.** "places" was reported as set; the six keys removed
were not mentioned.
## What would fix it
1. A way to read a layer — `settings show <module> [--node]` and the same on the verb.
2. `set` answers with what changed: keys added, changed and **removed**, so dropping one is never
silent. A removal could even require saying so.
3. The previous value kept: a settings history row per change, so an undo needs no backup.
## Status
Open. Until it is fixed: read the layer (from the store, read-only) before setting it, and compare
the node's plan before and after.
@@ -1,65 +0,0 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-tools]
fixed-by: mesh-tools pull request 13 — discovery waits for every runtime that answered PING, names the ones it missed, and mesh_runtimes says who answered
amended-design:
---
# 246 — The console says a module runs nowhere when a runtime answers late
## What was observed
2026-10-05. On the laptop, the console answered `mesh_machine` for the laptop with no modules and no
seats, and a call to one of the laptop's modules with "nothing in the mesh is called slack". The
laptop's own runtime said at the same moment that it served 275 tools for 48 modules. A minute
earlier, a call to another laptop module had been answered with "it does not run on the laptop; it
runs on the workstation", and a retry of the same call worked. Nothing was logged anywhere.
An agent worked around it by calling the runtime's local MCP port directly, which gives the right
answer and bypasses everything the console stands for: one way in, one account, one record of what
was called. The operator asked for a tool instead.
## What was measured
A read-only probe on the laptop's own runtime credential timed the discovery answers over 25 rounds
against the live bus, whose round trip from the laptop was about 40 ms:
- The laptop's runtime answer was the largest on the mesh, at about 164 kB with 341 endpoints. That
is far below the bus's message limit, and it was never shortened.
- It arrived last in every round: a median of about 365 ms, and once 813 ms. The other runtimes
answered within 180 to 275 ms, and the controller within 50 ms.
- The console gathered discovery answers for a fixed 750 ms. Inside the console, the gather runs
beside two controller calls and about 570 kB of answers on the same link, and it is slower still
while a runtime re-serves after a restart.
## Root cause
Discovery decided who was there by who answered in a fixed window. A late answer was not a failure
to anyone, so nobody said it. The index simply lacked that runtime, and every answer built on the
index then stated as fact that the runtime's modules did not exist, or ran only elsewhere.
Ruled out by measurement or by reading the code: an answer too large for the bus, subscriptions lost
when the bus reconnects, the console not counting its own machine's answer, and the merge of two
answers dropping a machine.
## Resolution
- Discovery asks who is there (PING, a hundred bytes, answered at once) beside what each serves
(INFO). It waits at least the old window, and then up to five seconds for every instance that said
it is there, so a large answer is waited for and a quiet mesh costs nothing extra.
- A runtime that said it is there and did not say what it serves in time, or that answered recently
and not now, is named. While one is unheard, the console never says an address is missing or runs
elsewhere: it says which runtime was not heard, and where the controller's records place the module.
- A new console tool, `mesh_runtimes`, says for every runtime how long its answer took, how large it
was, how many modules and tools it announced, whether it was shortened, and when it was last heard,
and which runtimes or machines were not heard.
- An announcement still too large after its descriptions are cut to their first line now leaves the
descriptions out, and says so.
## How it is checked
The fix ships with tests against a real bus: a runtime that answers after the old window is found and
called (the same test fails with the fixed window), a runtime that answers PING and never INFO is
named, and a restarted runtime, which answers under a new instance, is not reported as missed. Live,
`mesh_runtimes` shows every machine's answer and its time.
@@ -1,52 +0,0 @@
---
status: open
opened: 2026-10-05
located-in: []
fixed-by:
amended-design:
---
# 247 — A module cannot put the operator's account in a group
## What was observed
2026-10-05. The module for a peripheral-lighting daemon was assigned to the laptop. Its package
installs the daemon and creates the daemon's group. The daemon then refuses to start: "User is not a
member of the openrazer group". The device files are the group's, so the daemon cannot reach the
devices.
The module's own check names the fix, which is to add the account to the group and log in again. No
module can declare that fix:
- The account is a `user` resource, and the login-shell module already declares it, to set its shell.
- A second module that declares the same account, only to add one group, is refused as a duplicate
name.
- There is no resource for one membership on its own. A whole-account declaration that lists groups
would also take from the account every group it does not list, including the operator's own.
So the step is done by hand, with `sudo`, outside the mesh, and nothing records why the account is in
the group.
## Why it matters beyond this instance
More modules need this than this one: input devices (`input`), serial ports (`uucp`), the container
runtime (`docker`), virtual machines (`libvirt`, `kvm`), and capture or scanner hardware. Each is a
fact a module knows and the operator's account needs. Today every one is a hand step that survives a
reinstall only by memory. A membership added by hand is also never taken away when the module that
needed it is unassigned.
## Open questions
1. Is a membership its own resource (account, group), held by the module that needs it and given
back on undeclare? Or is it a contribution to the account's holder, in the way ADR 0212 lets a
module contribute to a seat?
2. A membership takes effect at the next login. How does the module say so: a finding, or a
moment the power or session seat already knows?
3. What does undeclare do with a membership the account already had before any module declared it?
The host keeps what it found and gives it back, as it does with a whole file it wrote over.
## How it is checked
When fixed, assigning the lighting module to a machine whose account is not in the group puts the
account in the group, says that a new login is needed, and leaves the account's other groups as they
were. Unassigning it removes only a membership the module added.
@@ -1,58 +0,0 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-controller internal/broker]
fixed-by: mesh-controller PR #51
amended-design:
---
# 248 — The controller's event consumer replayed a week, and held every new merge behind it
## What was observed
A merge to the catalogue at 15:17 never reached the controller. No plan was made and no build was asked
for it. Every machine kept running the build from before it. The merge just before, to the record, was
logged twice.
The bus showed why. The controller's durable consumer on the event stream was set to deliver
**everything the stream holds**, not from where it last was:
| | |
|---|---|
| the stream | a week of events, from message 1 514 to message 356 004 |
| the consumer delivered to | message 1 517, later 4 898 |
| acknowledged to | 0, later 1 569 |
| still to deliver | 6 955 merges and build outcomes, a week old |
| allowed outstanding | one at a time ([issue 175](../175-an-announcement-behind-a-long-build-comes-back/00-report.md)) |
So the controller was working through a week of past merges and build outcomes, one at a time, slowly. Every
new one — a merge, a build asked by hand — waited behind them. A build asked by hand finished on its
machine and was never registered. While this went on, the controller's client dropped messages
("slow consumer") several times, because heartbeats and reports share the loop with these events. From 15:17
its event loop did nothing more: no line logged, the one delivered event never acknowledged.
How the consumer came to deliver everything is not certain. The bus and the stores were rebuilt the night
before ([issue 241](../241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md)).
A consumer that is missing is made again by the controller's own assertion, and that made it with the
server's default, which is everything. [Issue 207](../207-a-re-made-worker-replayed-every-ask-the-stream-kept/00-report.md)
closed this for a consumer re-made because its type changed. It did not close it for one that is simply
not there.
## Why it matters
**Delivery stops, and nothing says so.** Status showed every plan done and no machine behind. The merge
that was missed is not "behind", because no plan was ever made for it. A replay of past merges can also
act on them again. Issue 207 records nine modules re-registered from the past the same way.
**The way out was a hand on the bus.** No verb resets a consumer. On the operator's explicit word, the
consumer was re-made from now with a one-off program run as the controller, otherwise unchanged. The
controller's loop still held the old event afterwards, so the new consumer's first delivery went
unacknowledged; the loop needs a restart of the controller to let go of it.
## Noticed alongside, not this issue
- Each merge to the catalogue planned 99 to 100 modules in two tiers and rebuilt modules it did not touch.
The output was byte-identical, so nothing was redeployed, but it costs minutes of the build machine
per merge.
- Three plans for three merges ran over each other. Each sent the build agent to every machine and asked
for the same builds. Nothing supersedes a plan for an older commit.
@@ -1,54 +0,0 @@
# 248 — Diagnosis
## 2026-10-05
1. A merge was not in the controller's log. The forge's own log showed nothing about delivering it, so the
question moved to the bus.
2. The bus's backlog tool named the controller's event consumer: 6 955 pending, redeliveries, one
unacknowledged. Its configuration was read from the server's monitoring endpoint: deliver policy
*all*, one outstanding, 30 seconds to acknowledge, five deliveries.
3. Sampled three times over a minute it did not move, and over the following hour it crawled forward
through week-old events. The controller's log held only its client's warnings: dropped messages,
and one refused reply.
4. The controller's code makes a missing consumer with the configuration it asserts, which sets no
deliver policy, so the server's default applies: everything. Issue 207's fix sets *from now* only on
the path where an existing consumer's type changes.
**Unblocked**, on the operator's explicit word: the consumer re-made from now with its configuration
otherwise unchanged (nothing pending afterwards). The controller's loop still held the old event, so
it takes a restart of the controller to let go of it; that restart waits for the operator's word.
**Fix** (mesh-controller, branch `fix/a-consumer-on-a-history-stream-starts-from-now`):
- a consumer may say it starts **from now** when it is made. The controller's event consumer does. A
consumer that exists keeps where it is. The server would refuse a changed start anyway.
- `broker consumer-reset <stream> <consumer>` re-makes a stuck consumer from now, its configuration
otherwise kept, and refuses a work queue, where what is pending is work. It is the person's act, said by
a command, rather than a one-off program.
Checked by a live test against a throwaway bus:
- made from now, a consumer holds none of the stream's past and does hold the next announcement;
- asserted again, it keeps its place;
- the default replays all of it;
- a reset leaves nothing pending and keeps every other setting;
- a work queue's consumer is refused.
**Not fixed here:** the client dropping messages while the loop acts on a long merge. Reports and
heartbeats are redelivered or replaced, so nothing is lost for good, but the loop holding everything while
it builds is the shape issue 175 already describes.
## 2026-10-05, after the restart
The consumer re-made from now held nothing, and the restarted controller acknowledged what it was handed.
Two merges made right after reached it within two seconds, and the fixed controller was built, delivered and
took over by its own plan within two minutes. The re-asked build of the agent module was registered and
reached all four machines.
**Not explained by this issue:** the record's merge was still logged twice, with the consumer fresh and
nothing replayed. The duplication has a cause of its own, still to be found. It may be that the forge
announces a merge on two paths, or that one event is handled twice.
## Resolved — 2026-10-05
The controller's event consumer is made from now when it is made, and `broker consumer-reset` re-makes a stuck one from now. Live: after the restart, merges reached the controller within seconds. Two tests main then failed, both skipped without a store, were fixed in mesh-controller PR #52. The merge heard twice was a separate cause: [issue 250](../250-a-merge-made-through-the-forges-tool-is-announced-twice/00-report.md).
@@ -1,37 +0,0 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-controller cmd/mesh-controller]
fixed-by: mesh-controller PR #54
amended-design: 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
---
# 249 — A module's new state is refused until a push the merge did not make
## What was observed
A merge gave the agent module a new state, a key-value bucket
([ADR 0201](../../02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)).
The module's new bundle reached all four machines within a minute of the build, at once. The machines'
bus permissions did not include the new state until a push was made by hand afterwards.
On three machines the module's watch of the new state was refused for about two minutes: "claude-code keeps
and reads no state called config". It recovered only because the module asks again with a back-off, and
the hand-made pushes issued the permissions.
## Why it matters
**The code arrives before the right to use it.** A module that does not retry stays broken until someone
pushes. A module whose first act on start is to read its new state fails its start. Nothing in the plan
says the two must travel together.
**No machine went first.** The bundle reached every machine at the same moment. The rollout the operator
was told — one machine first, then the rest — could not be followed, because the merge had already
delivered it everywhere.
## Open questions
- Should a plan send a machine its membership, the grants that come with a module's new
declarations, in the same push as the bundle, and before it?
- Should a merge that changes a module's declarations (state, events, tools) be delivered to one machine
first, and to the rest only once that one reports it healthy?
@@ -1,30 +0,0 @@
# 249 — Diagnosis
## 2026-10-05
**Grants after code.** Both a plan's rollout and a push send every machine its declaration first, and
issue the memberships afterwards. The order is written into the code on purpose, "because the runtime it is
for arrives with it". That reason holds only for a first assignment, and a membership is kept on the bus for
a runtime that connects later anyway. A membership that failed was only printed, and left "until the next
push". The list of bus users travels in the declaration of the machine that holds the bus, which a module's
rollout reaches only if that machine runs the module.
**No machine first.** The module's upgrade policy sends one machine at a time, but a plan's rollout ignored
it and sent every machine at once. One at a time did not wait for the first machine to come up either: it
stopped only if the publish failed.
**Not answered by the open decision on unseen changes** (a removal, a move or a replacement, shown
before it takes effect). That decision leaves an add-only change alone on purpose, and a new state is one.
**Decided** in [ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md):
grants before code, and one machine first unless a module's policy says *together*.
## Resolved — 2026-10-05
[ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md), live the same evening:
- every push now issues memberships before it sends declarations;
- the bus's machine goes first when its list of users moved;
- a plan sent the build agent to one machine first and to the rest once that machine reported.
The first live rollout exposed a fault in the tier gate. The first machine's report, made between the two sends, was read as stale ([issue 256](../256-a-first-machines-report-read-as-stale-between-the-two-sends/00-report.md)).
@@ -1,49 +0,0 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-catalog modules/gitea]
fixed-by: mesh-catalog PR #63, PR #66
amended-design:
---
# 250 — A merge made through the forge's tool is announced twice
## What was observed
The controller logged the record repository's merges twice, seconds apart, with the same commit, even with
its event consumer freshly made ([issue 248](../248-the-controllers-event-consumer-replayed-a-week-and-held-every-merge-behind-it/00-report.md)).
Counted over one day:
- 8 of the record's merges were logged twice, against 4 once;
- 10 of the catalogue's, against 7 once;
- 2 of the controller's.
A code repository's second line reads differently — "it changed nothing any module the mesh holds is built
from" — so it was taken for a different message.
## Diagnosis
The forge's module announces a merge from two places in the same process:
1. its merge tool, the moment it merges;
2. the poll added for [issue 131](../131-nothing-tells-the-mesh-a-source-moved/00-report.md), which
announces every merged pull request it has not recorded as announced.
The tool never records what it announced, so the poll announces it again 0.5 to 16 seconds later.
Merges made in the forge's web interface or by a plain API call are seen by the poll alone, and those are
the ones logged once. The module's own header comments still say merges are announced "from the tools …
one process only".
No harm was done this time, but only by luck. The second event is absorbed because the first one moved
the controller's record of the source. A repository read only by packaging modules has no such record, so
it would get a second plan. The record module synced twice for each merge.
## Fix
The poll is the only emitter: it sees every path and carries the clone address. The tool merges and
answers the merge commit. A merge made through the tool is heard up to thirty seconds later, which the
module already accepts ("an event a minute late is still an event").
## Resolved — 2026-10-05
The forge module's poll is the only announcer of a merge. It asks only the repositories that moved since its last look, and one pass at a time. Live: three merges were each heard once, and a later merge was planned within a minute.
@@ -1,36 +0,0 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-catalog modules/records]
fixed-by: mesh-catalog PR #64
amended-design:
---
# 251 — The record's checkout could not sync after it ran as another account
## What was observed
Every sync of the record module failed, on every merge and every timer: git refused the checkout as
"dubious ownership". The record tools answered all the while, from the checkout as it last stood, and
nothing said it was stale.
## Diagnosis
The module ran in a container, as the superuser, until its code moved into the machine's tool runtime
([ADR 0198](../../02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)).
The runtime launches it as the operator account. The checkout's git directory and 2 598 of its files still
belonged to the superuser. Git refuses a repository owned by another user, and the operator account could
change none of those files.
The module also still declared that it needs a container runtime, a leftover of the same move.
## Fix
The module, ported to Go as part of the fix, clones into a directory of its own inside the one it is given:
a directory it makes, and so owns. What the old layout left behind is removed where it is the module's.
Where it is not, the module names it in its status, with the one command that deletes it. The
container-runtime capability is dropped. A sync that fails is still said in the status, as before.
## Resolved — 2026-10-05
The record module, ported to Go, clones into a directory it makes and owns. Live: it synced to the newest commit, with 692 documents. It names seven leftovers of the old checkout that it cannot remove, with the command that removes them; that is the operator's act.
@@ -1,38 +0,0 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-catalog modules/gitea, mesh-controller cmd/mesh-controller]
fixed-by: mesh-catalog PR #63, mesh-controller PR #54
amended-design: 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
---
# 252 — A merge's changed modules were read wrong, in both directions
## What was observed
Each of three merges to the catalogue within four minutes planned 99 to 100 modules, and rebuilt modules
they did not touch. The outputs were identical, so nothing was redeployed. The rebuilds cost the build
machine minutes for every merge.
## Diagnosis
**Too many.** The controller treats a changed path outside every module directory it knows as shared code,
and rebuilds every module built from the repository. A new module's directory, or one being removed,
counts: the module is registered only after the merge is planned. Each of the three merges added or
removed a module. The build agent, built from the same repository, then joins the set and becomes tier 0.
**Too few.** The forge's module asked for a hundred changed files and was given fifty, the forge's page
size, and reported the list as whole. A merge of 59 files reached the controller with 50. Had the full
rebuild not hidden it, a module whose own files changed would have stayed unbuilt.
## Fix
- The forge's module reads every page of a pull request's files.
- The controller counts a changed path in a sibling of known module directories as that module's own, not
shared, when that module's definition is among the changed paths. A sibling without a definition, such
as a shared library, still means everything, which is the safe direction. Root files still mean
everything.
## Resolved — 2026-10-05
The forge's module reads every page of a pull request's files. The controller counts a new module's own directory as that module's when its definition is among the changed paths. Live: catalogue merges planned one module each.
@@ -1,55 +0,0 @@
---
status: located
opened: 2026-10-05
located-in: [mesh-controller internal/builder, mesh-controller internal/artifacts, mesh-catalog modules/distribution]
fixed-by: mesh-controller PR #53, mesh-host PR #23, mesh-catalog PR #62
amended-design: 03-DESIGN/01-to-be/18-building-a-module.md
---
# 253 — The store's collector would delete every archive the mesh keeps
## What was observed
The store's nightly collector ([ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md))
was installed the same day, its first run due that night. Measured read-only beforehand:
| | |
|---|---|
| blobs in the store | 8 186 |
| blobs a manifest names, kept by the collector | 2 090 |
| blobs it would delete | 6 096 |
| repositories holding only archives, none named by any manifest | 105 |
Among the archives it would delete were the current bundles of the agent module, the machine host, the
controller and the tool runtime. Each was named by no manifest, though the controller's records keep them.
## Why it matters
Machines keep their unpacked copies, so nothing would have stopped at once. But any fresh fetch of an
unchanged module would have failed: a machine joining, a reinstall, an apply that fetches again, the
controller's own next rollout.
## Diagnosis
Images are pushed with manifests. Archives were published as bare blobs, which no manifest names. The
store's stock collector marks only from manifests, so every bare blob is unmarked, kept or not. ADR 0189's
sentence "what the mesh keeps is still a manifest in the store" was true of images only.
Found alongside: the "five most recent builds" reason kept builds of modules the mesh no longer holds,
forever.
## Fix
- **That night, before the first run:** the collector was changed to a dry run, in its module's
definition, and delivered.
- **Then:**
- every archive is published with a manifest that holds it;
- the controller's sweep holds every kept archive before it lets anything go, which backfills those
already published;
- letting an archive go removes its manifest first;
- a forgotten module keeps nothing.
- Real collection returns once the controller reports no kept archive unheld.
## Where it stands — 2026-10-05
Every kept archive is held: the controller's collection command reports 134 of 134 held, none missing. The window the collector needs is no longer reopened by an apply ([issue 224](../224-an-apply-reopens-a-maintenance-window-by-recreating-what-it-held-still/00-report.md)). The collector still runs as a dry run. Turning it to real collection deletes the layers nothing keeps, which is the operator's word to give; this issue resolves when that change lands.
@@ -1,36 +0,0 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-controller cmd/mesh-controller, mesh-controller internal/inventory]
fixed-by: mesh-controller PR #54
amended-design: 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
---
# 254 — Plans for successive merges run over each other, and one was left open
## What was observed
Three merges to the catalogue within four minutes made three plans, and all three ran at once:
- each sent the build agent to every machine;
- each asked for the same tier of builds — one module was asked for 32 seconds apart by two plans;
- the plan for the middle merge still showed "building" hours later, waiting on nothing.
## Diagnosis
A merge's plan is saved without looking at the open plans
([ADR 0162](../../02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md): one plan per merge).
Each open plan advances on its own. Nothing ends a plan whose work a newer merge has taken over, and
nothing lets a person close a plan that waits on nothing.
[Issue 219](../219-an-older-build-that-finishes-later-replaces-a-newer-one/00-report.md) settled only which
build's output wins.
## Fix
[ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md)
§3. A newer merge's plan supersedes the older open plans for the same repository and branch, and takes in
the modules they had not built. A person can close a plan by its id.
## Resolved — 2026-10-05
A newer merge's plan supersedes the older open plans of its repository and branch, and a person can close a plan by its id. Live: the plan left open since the afternoon was closed by hand, and its note said what it had built and never sent. A merge made while the controller's own plan was open took it over: the older plan reads "superseded at tier 0 by" the newer, which planned its three modules again.
@@ -1,36 +0,0 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-catalog modules/systemd]
fixed-by: mesh-catalog PR #65
amended-design:
---
# 255 — The journal verb read nothing for a system service
## What was observed
The service manager seat's `journal` verb answered "-- No entries --" for the controller's service on the
control node, while the service was logging steadily. Twice in one day a session reading the controller's
log fell back to a shell on the machine: the verb that exists for exactly that question answered nothing.
On one machine of four the same verb did answer: the only one whose operator account is in the journal's
group.
## Diagnosis
The module runs as the operator account and escalates the five acts on the system manager with `sudo -n`.
It ran `journalctl` unescalated. journalctl shows an account outside the journal's group only that
account's own entries, and says "-- No entries --" for everything else. That reads as a quiet service, not
as a refusal.
The mesh grants the operator account passwordless escalation on every machine through its own drop-in, as
the sudo module's check confirms.
## Fix
A read of the system journal escalates like an act does. The account's own journal, in the user scope,
does not. The module was ported to Go as part of the fix, its tests with it.
## Resolved — 2026-10-05
The journal verb reads a system unit's journal escalated, in the module ported to Go. Live: the controller's and the host's journals read through the verb on the control node.
@@ -1,31 +0,0 @@
---
status: resolved
opened: 2026-10-05
located-in: [mesh-controller cmd/mesh-controller]
fixed-by: mesh-controller PR #55
amended-design:
---
# 256 — A first machine's report read as stale between the two sends
## What was observed
The first plan to roll out under [ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md)
sent the build agent to one machine first, and to the other two once that machine reported it applied, 27
seconds later. All three applied it within seconds. The plan then said, for seven minutes, that it was
"waiting for build-agent" on the first machine to be applied. It went on only when that machine happened to
report again for another reason.
## Diagnosis
The tier gate asked every machine for a report made after the module's last send. With one machine first,
a module is sent twice, and the second send is the later one. The first machine's report came between the
two sends, so it read as older than the build. The plan waited for that machine's next report, one report
cycle, and never knew why.
## Fix
The gate judges each machine from its own send: the first machine from the first send, the rest from the
second. A plan's wait is also printed to the second rather than the minute, where it had read "0s", and the
first send is printed in the machine's own time rather than in UTC. Checked by the controller's test: a first
machine's report between the two sends opens the gate, and a report from before its send does not.