Compare commits
33
Commits
issues/245-246
...
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
553d0c8893 | ||
|
|
08643a2128 | ||
|
|
4fac135a46 | ||
|
|
f391c5c36c | ||
|
|
e6fbb86373 | ||
|
|
c8af8d5eb1 | ||
|
|
0e193ebc41 | ||
|
|
a0de734ed6 | ||
|
|
55e776a1cd | ||
|
|
c233a1bbac | ||
|
|
035d4a27e3 | ||
|
|
77429488b0 | ||
|
|
3944e4f914 | ||
|
|
ab1fbe3202 | ||
|
|
0fdcb15200 | ||
|
|
91aa481269 | ||
|
|
0627b7f1f3 | ||
|
|
049b50b66c | ||
|
|
44808ac563 | ||
|
|
f000288147 | ||
|
|
80aff0f457 | ||
|
|
400ed8d396 | ||
|
|
df2072aed5 | ||
|
|
ae58570209 | ||
|
|
855c3f9b33 | ||
|
|
f93b02900c | ||
|
|
9306192f93 | ||
|
|
8bd56a2820 | ||
|
|
7497149eed | ||
|
|
1b9196b807 | ||
|
|
b3e7191ddc | ||
|
|
87f80573b8 | ||
|
|
9197b99153 |
@@ -60,6 +60,14 @@ with the server held still for its duration. Plain collection, not `--delete-unt
|
|||||||
mesh keeps is still a manifest in the store, so it is still referenced, so its blobs stay — the
|
mesh keeps is still a manifest in the store, so it is still referenced, so its blobs stay — the
|
||||||
dangerous flag is not needed at all once the mesh is the one deciding.
|
dangerous flag is not needed at all once the mesh is the one deciding.
|
||||||
|
|
||||||
|
> **Progressive insight — 2026-10-05.** "What the mesh keeps is still a manifest in the store" was true
|
||||||
|
> of images and false of archives: the builder published every archive as a bare blob no manifest names,
|
||||||
|
> and the store's collector keeps only what a manifest names. Its first night would have deleted every
|
||||||
|
> archive the mesh keeps ([issue 253](../04-ISSUES/253-the-stores-collector-would-delete-every-archive-the-mesh-keeps/00-report.md)).
|
||||||
|
> The decision stands — the mesh decides, the store reclaims with plain collection. What changes is how an
|
||||||
|
> archive is published: with a manifest that holds it, so the sentence becomes true of archives too. Until
|
||||||
|
> every kept archive is held, the collector runs as a dry run.
|
||||||
|
|
||||||
**3. What the mesh keeps, stated as three reasons rather than a number.** A digest is kept because:
|
**3. What the mesh keeps, stated as three reasons rather than a number.** A digest is kept because:
|
||||||
|
|
||||||
- **a definition names it** — every artifact reference in any module's current recorded manifest,
|
- **a definition names it** — every artifact reference in any module's current recorded manifest,
|
||||||
|
|||||||
+100
@@ -0,0 +1,100 @@
|
|||||||
|
---
|
||||||
|
topic: the mesh
|
||||||
|
status: accepted
|
||||||
|
date: 2026-10-05
|
||||||
|
deciders: jochen
|
||||||
|
reconstructed: false
|
||||||
|
extends: 02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md
|
||||||
|
---
|
||||||
|
|
||||||
|
# 218. A plan sends grants before code, rolls a module out one machine first, and a newer merge takes over an older plan
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
On 2026-10-05 the delivery path was watched through a day of merges, by several sessions at once. Three
|
||||||
|
things went wrong, each recorded as an issue with its evidence.
|
||||||
|
|
||||||
|
- **Code arrived before the right to use it** ([issue 249](../04-ISSUES/249-a-modules-new-state-is-refused-until-a-push-the-merge-did-not-make/00-report.md)).
|
||||||
|
A merge gave a module a new key-value state. The plan sent the new bundle to every machine, and only
|
||||||
|
then issued the memberships that grant the state. On three machines the module's new state was refused
|
||||||
|
for two minutes, until a push made by hand. The order is written into the code on purpose: memberships
|
||||||
|
"after the declaration, because the runtime it is for arrives with it". That reason holds only for a
|
||||||
|
first assignment, and even then a membership is kept on the bus for the runtime that connects later
|
||||||
|
([ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)).
|
||||||
|
- **No machine went first.** The module's upgrade policy sends one machine at a time, but a plan's rollout
|
||||||
|
ignores it and sends every machine running the module at once. One at a time also never waited for the
|
||||||
|
first machine to come up healthy: it stopped only if the publish itself failed. A change was therefore
|
||||||
|
everywhere before anything had seen it run.
|
||||||
|
- **Plans for successive merges ran over each other** ([issue 254](../04-ISSUES/254-plans-for-successive-merges-run-over-each-other-and-one-was-left-open/00-report.md)).
|
||||||
|
Three merges to the catalogue within four minutes made three plans. Each sent the build agent to every
|
||||||
|
machine and asked for the same builds. One was still "building" hours later, with nothing left for it to
|
||||||
|
wait on. [ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md) decides one plan per merge
|
||||||
|
and says nothing about the next merge arriving while one is open. [Issue 219](../04-ISSUES/219-an-older-build-that-finishes-later-replaces-a-newer-one/00-report.md)
|
||||||
|
settled only which build's output wins.
|
||||||
|
|
||||||
|
## Considered Options
|
||||||
|
|
||||||
|
1. **Debounce merges:** wait a window before planning, so close merges make one plan. Rejected: it only
|
||||||
|
delays the overlap, does nothing for merges further apart than the window, and makes every merge slower.
|
||||||
|
2. **Queue plans:** a new plan waits until the older one is done. Rejected: the older plan builds what the
|
||||||
|
newer merge is about to replace, then the newer one builds it again.
|
||||||
|
3. **A newer merge's plan takes over the older plan's unfinished work, a plan rolls a module out one
|
||||||
|
machine first, and grants travel before code.** Chosen.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
**1. Grants before code.** Every send — a plan's rollout and a push alike — issues the memberships for
|
||||||
|
the machines it is about to send to before it sends their declarations, after raising the buckets they
|
||||||
|
name. When the composed list of bus users changes, the machine that holds the bus is sent first, because
|
||||||
|
that list travels in its declaration. A membership that could not be issued fails the send, and the send
|
||||||
|
is tried again. It is never reported as done "until the next push".
|
||||||
|
|
||||||
|
**2. One machine first.** A plan rolls a module out according to the module's upgrade policy
|
||||||
|
([ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md) §3). Unless the policy says
|
||||||
|
*together*:
|
||||||
|
|
||||||
|
- the module is sent to one machine first, the first by name of the machines running it;
|
||||||
|
- the rest are sent only once that machine has reported the new declaration applied and current;
|
||||||
|
- a first machine that reports a failure, or does not report in time, stops the module's rollout there.
|
||||||
|
The plan names the machine and the reason, and the other machines keep what they ran.
|
||||||
|
|
||||||
|
The plan records which machine went first, so a controller replaced mid-rollout resumes from there. A
|
||||||
|
policy of *together* keeps today's behaviour.
|
||||||
|
|
||||||
|
**3. A newer merge takes over an older plan.** When a merge into a repository's branch makes a plan,
|
||||||
|
every open plan for the same repository and branch made before it is superseded, ordered by when each
|
||||||
|
plan was made, never by commit:
|
||||||
|
|
||||||
|
- the modules the older plan had not yet built join the newer plan's set, before its tiers are computed;
|
||||||
|
- the older plan ends in a state of its own, *superseded*, naming the plan that took it over.
|
||||||
|
|
||||||
|
Builds the older plan already asked for still finish and register; issue 219's ordering keeps the newer
|
||||||
|
one current. A person can also close a plan that waits on nothing, by its id. The plan is marked closed
|
||||||
|
by hand and never resumed.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
- A module that gains a state, an event or a tool can use it from its first start on every machine.
|
||||||
|
- A change reaches one machine before the rest. A change that breaks its first machine stops there, with
|
||||||
|
the reason in the plan.
|
||||||
|
- Successive merges build each module once, for the newest commit. The build agent is sent to the
|
||||||
|
machines once per run of merges, not once per merge.
|
||||||
|
- **What got harder:** a rollout takes one machine's report longer than before. A module that must change
|
||||||
|
everywhere at once says *together* in its policy. A plan's record now has a superseded state that
|
||||||
|
readers of the plans must know.
|
||||||
|
|
||||||
|
## How it is checked
|
||||||
|
|
||||||
|
| Rule | Checked by |
|
||||||
|
|---|---|
|
||||||
|
| grants before code | the controller's test: a send records memberships issued before any declaration; the machine holding the bus is sent first when the user list changes; a failed membership fails the send |
|
||||||
|
| one machine first | the controller's test: with a one-at-a-time policy, one machine is sent, the rest only after its applied and current report; a failed first machine stops the module; *together* sends all at once |
|
||||||
|
| a newer merge takes over | the controller's test: an older open plan for the same repository and branch is superseded, its unbuilt modules folded in; a plan for another repository is left alone; a superseded plan is not open |
|
||||||
|
| live | the next merge to the catalogue that gives a module a new state: no refusal of that state on any machine, the first machine named in the plan, one plan open per repository |
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- [Issue 249](../04-ISSUES/249-a-modules-new-state-is-refused-until-a-push-the-merge-did-not-make/00-report.md), [issue 254](../04-ISSUES/254-plans-for-successive-merges-run-over-each-other-and-one-was-left-open/00-report.md)
|
||||||
|
- [ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md) — plans and tiers, extended here
|
||||||
|
- [ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md) — memberships, kept on the bus
|
||||||
|
- [to-be 30](../03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md) — the design this amends
|
||||||
+120
@@ -0,0 +1,120 @@
|
|||||||
|
---
|
||||||
|
topic: the mesh
|
||||||
|
status: accepted
|
||||||
|
date: 2026-10-05
|
||||||
|
deciders: jochen
|
||||||
|
reconstructed: false
|
||||||
|
extends: 02-DECISIONS/0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md
|
||||||
|
---
|
||||||
|
|
||||||
|
# 219. The build queue is controlled through the controller and the build seat
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
The operator asked for tools to control the mesh's builds: see what is queued and running, cancel,
|
||||||
|
clear, stop a build immediately, pause and continue, restart and replay. On the day of the request none
|
||||||
|
existed. The controller could ask for a build and list finished ones. Nothing could see an ask waiting on
|
||||||
|
the build seat's work queue, or one being built. Nothing could take an ask back, and a running build
|
||||||
|
could only be stopped by restarting the machine's build agent, which hands the ask to another holder.
|
||||||
|
|
||||||
|
How builds run today, read from the code and the bus
|
||||||
|
([ADR 0190](0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md),
|
||||||
|
[ADR 0157](0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md)):
|
||||||
|
|
||||||
|
- an ask is a message on the seat's work-queue stream;
|
||||||
|
- every holder pulls one at a time from one shared worker, and acknowledges after it has announced the
|
||||||
|
outcome;
|
||||||
|
- a build says it started, logs its steps, and announces what it built, all on the bus;
|
||||||
|
- an ask delivered five times without an answer stays in the stream for a week, with nothing saying so.
|
||||||
|
|
||||||
|
Three facts constrain any control:
|
||||||
|
|
||||||
|
- **Only the controller may act on the bus's streams and consumers** (design 25 §3, enforced in the bus's
|
||||||
|
user list). Holders may only take from their worker and acknowledge.
|
||||||
|
- **A plan waits for a build's outcome and has no timeout.** An ask that disappears without one leaves
|
||||||
|
its plan waiting, said only as late after half an hour.
|
||||||
|
- **The bus server in use cannot pause a consumer**; that arrived in a later version.
|
||||||
|
|
||||||
|
## Considered Options
|
||||||
|
|
||||||
|
1. **Pause and cancel by remaking the worker consumer.** Rejected: a remade worker delivers from now on,
|
||||||
|
so every queued ask would be skipped silently. Issues 206 and 207 are this mistake both ways round.
|
||||||
|
2. **Upgrade the bus first and use its consumer pause.** Not now: it gives pause and continue, and
|
||||||
|
nothing else on the list, and upgrading the bus is a change of its own.
|
||||||
|
3. **Queue actions are the controller's verbs, process actions are the build seat's verbs, and every
|
||||||
|
action that drops work leaves a failed outcome.** Chosen.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
**1. The queue is the controller's.** Its verbs act on the work-queue stream, the one thing only it may
|
||||||
|
touch:
|
||||||
|
|
||||||
|
| verb | what it does |
|
||||||
|
|---|---|
|
||||||
|
| `queue` | lists every ask: waiting, in flight (with the machine building it, from its started event), and dead (delivered as often as allowed, still in the stream) |
|
||||||
|
| `cancel <id>` | takes back a waiting or dead ask; an ask in flight is refused, and `kill` is named instead |
|
||||||
|
| `clear` | cancels every waiting ask, and with `--dead` every dead one |
|
||||||
|
| `rebuild <module or build>` | asks again for a module's current source, or a past build's source and ref, under a new id |
|
||||||
|
| `replay <build>` | asks again for a past build at its commit, as a **dry run** unless told to register |
|
||||||
|
| `kill <id>`, `pause [node]`, `resume [node]` | pass the request on to the build seat on the right machine, or on every machine |
|
||||||
|
|
||||||
|
**2. The process is the build seat's.** Each holder serves verbs on its own machine:
|
||||||
|
|
||||||
|
- `current`: the build running here, its step and how long, and whether this holder is paused;
|
||||||
|
- `kill`: stops a running build at once. The build's whole process group is ended, along with every
|
||||||
|
container it started. Its outcome is announced as failed, and the ask is acknowledged, so it is not
|
||||||
|
delivered again;
|
||||||
|
- `pause` and `resume`: a paused holder takes nothing new, and a running build finishes. The flag
|
||||||
|
survives the holder's restart.
|
||||||
|
|
||||||
|
A holder restarted mid-build keeps today's behaviour: the ask is not acknowledged, and another holder
|
||||||
|
takes it.
|
||||||
|
|
||||||
|
**3. Nothing dropped is silent.** Cancel, clear and kill each leave a failed outcome for the ask's id,
|
||||||
|
taken in like any other. A plan waiting on that build fails, and says why, instead of waiting. A plan keeps
|
||||||
|
the id it asked for, so it can match its outcome exactly. A holder checks a cancelled ask before it builds
|
||||||
|
it, so an ask taken in the instant it was cancelled is not built.
|
||||||
|
|
||||||
|
**4. A plan says what its builds are waiting on, and a failed plan can go on.**
|
||||||
|
|
||||||
|
- A plan whose build waits on a paused seat says the seat is paused, and on which machines, and is not
|
||||||
|
counted late while it waits.
|
||||||
|
- `plans retry <id>` asks again for the modules a failed plan could not build, under new ids. The plan
|
||||||
|
resumes at that tier and goes on through its later ones. A plan another has superseded, or one
|
||||||
|
already done, is refused.
|
||||||
|
- `rebuild` of a module that an open or failed plan has not yet built joins that plan, so the plan and
|
||||||
|
the build are one thing.
|
||||||
|
|
||||||
|
**5. Replay does not move the mesh backwards unasked.** A replayed build is a dry run: built, its log
|
||||||
|
kept, nothing registered. With `--register` it is registered. If a newer build of the module is already
|
||||||
|
registered, that is refused unless `--older` is said as well: registering an older commit makes it the
|
||||||
|
current one, and the rollout policy sends it to the machines ([issue 207](../04-ISSUES/207-a-re-made-worker-replayed-every-ask-the-stream-kept/00-report.md)).
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
- Every build in the mesh can be seen, taken back, stopped or asked again from the console. None of it
|
||||||
|
needs a shell on a machine.
|
||||||
|
- A cancelled or killed build shows in the build records as failed, with who stopped it. Its plan fails
|
||||||
|
saying the same.
|
||||||
|
- **What got harder:** a holder now serves verbs as well as taking work, and keeps one small flag on
|
||||||
|
disk. Pause is per holder, so "pause the mesh" is the controller asking every holder in turn. A holder
|
||||||
|
away at the time misses it, and the answer names that holder.
|
||||||
|
|
||||||
|
## How it is checked
|
||||||
|
|
||||||
|
| Rule | Checked by |
|
||||||
|
|---|---|
|
||||||
|
| the queue is read and classified right | the controller's test: waiting, in flight and dead asks told apart from the stream and the worker; a live test on a throwaway bus |
|
||||||
|
| nothing dropped is silent | the controller's test: cancel and clear delete the ask and record a failed outcome; a plan asked for that id fails |
|
||||||
|
| a cancelled ask is not built | the holder's test: an ask taken after its cancel is answered failed without building |
|
||||||
|
| kill stops everything it started | the holder's test: the build's process group and its labelled containers are ended; the outcome is failed and the ask acknowledged |
|
||||||
|
| pause survives a restart | the holder's test: the flag is read back at start, and nothing is taken while it is set |
|
||||||
|
| plans follow the queue | the controller's test: a plan waiting on a paused seat says so and is not late; `plans retry` resumes a failed plan and its later tiers are asked; `rebuild` joins the plan that holds the module |
|
||||||
|
| replay is safe | the controller's test: a dry run by default, and `--register` over a newer build refused without `--older` |
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- [ADR 0190](0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md) — holders share the seat's work, one at a time
|
||||||
|
- [ADR 0157](0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md) — a build's own events are its record
|
||||||
|
- [issue 207](../04-ISSUES/207-a-re-made-worker-replayed-every-ask-the-stream-kept/00-report.md) — a remade worker replayed every ask
|
||||||
|
- [to-be 18](../03-DESIGN/01-to-be/18-building-a-module.md) — the design this amends
|
||||||
+157
@@ -0,0 +1,157 @@
|
|||||||
|
---
|
||||||
|
topic: what runs on it
|
||||||
|
status: accepted
|
||||||
|
date: 2026-10-05
|
||||||
|
deciders: jochen
|
||||||
|
reconstructed: false
|
||||||
|
extends: 02-DECISIONS/0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md
|
||||||
|
---
|
||||||
|
|
||||||
|
# 220. What a machine asks needs its uplink held, and the retired resolver pieces go
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
**Three things about a machine's resolver were left half done when the mesh moved to one resolver.**
|
||||||
|
On the production mesh on 2026-10-05, read from the controller's `seats` verb:
|
||||||
|
|
||||||
|
- **`node-dns-resolver` has no holder on any node, and no module in the catalogue claims it.**
|
||||||
|
[ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) retired it
|
||||||
|
and the controller kept its row deliberately, *"deleted once nothing claims it"*, because removing a
|
||||||
|
seat a machine still holds makes that machine unresolvable. That condition now holds. The row still
|
||||||
|
stands in the controller's compiled set and in the store's seat table, and the overview still lists
|
||||||
|
it, unheld, beside the seats a mesh actually has.
|
||||||
|
- **The rule that keeps `/etc/resolv.conf` the mesh's is checked by nothing.**
|
||||||
|
[ADR 0117](0117-a-machines-uplink-is-a-seat.md) found that a network manager rewrites the resolver
|
||||||
|
file on every connectivity change unless it is told not to, and gave that telling to the module
|
||||||
|
holding `node-uplink`. It said the condition *"only if NetworkManager runs"* is expressed by
|
||||||
|
assigning the manager's module. Nothing makes anybody do so: `resolv-conf` can be assigned to a
|
||||||
|
machine with no uplink holder, and the file is then replaced the first time a laptop changes
|
||||||
|
network while every surface of the mesh reads green. Today every node holding
|
||||||
|
`node-resolver-config` also holds `node-uplink` — NetworkManager on the home server, the
|
||||||
|
workstation and the laptop, systemd-networkd on the anchor — by care, not by check.
|
||||||
|
- **The catalogue still carries a systemd-resolved split-DNS module**, `resolved-split-dns`, claiming
|
||||||
|
`node-resolver-config`. [ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md)
|
||||||
|
chose against a stub on every node and says *"There is no `systemd-resolved` module."* It is
|
||||||
|
assigned nowhere. `resolv-conf`'s own resolver file still tells its reader that systemd-resolved or
|
||||||
|
NetworkManager may be assigned *instead* — the opposite of how the roles now divide.
|
||||||
|
|
||||||
|
**A dependency mechanism already exists.** [ADR 0207](0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md)
|
||||||
|
made a module depend on the node seats that apply its resources, derived rather than stated, judged
|
||||||
|
over the node's whole set of assignments, refused at `assign` naming the seat and its possible holders,
|
||||||
|
and refused at composition. [ADR 0210](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)
|
||||||
|
derived a second kind from contributions. What is missing is a dependency that belongs to a *role*
|
||||||
|
rather than to what a module declares.
|
||||||
|
|
||||||
|
## Considered Options
|
||||||
|
|
||||||
|
**For the retired seat:**
|
||||||
|
|
||||||
|
1. **Keep the row until the build seat's retired row goes too, and delete both together.** Rejected:
|
||||||
|
the two have nothing in common but having been retired; one is unclaimed now and the other is not
|
||||||
|
yet known to be.
|
||||||
|
2. **Remove it from the compiled set only.** Rejected: seeding adds a seat a release ships and never
|
||||||
|
removes one ([ADR 0122](0122-a-seat-is-data-a-rename-is-a-database-update.md)), so the store's
|
||||||
|
row, which is the live set, would stay.
|
||||||
|
3. **Remove it from the compiled set and delete the store's row in a numbered migration**, as the
|
||||||
|
artifact store's rename did for its old row. Chosen.
|
||||||
|
|
||||||
|
**For the resolver file and the uplink:**
|
||||||
|
|
||||||
|
1. **Leave it to the operator.** Rejected: it is the failure ADR 0117 describes — the file silently
|
||||||
|
replaced — with the one difference that the operator was told.
|
||||||
|
2. **`resolv-conf` declares the manager's settings itself.** Rejected by ADR 0117 already: which
|
||||||
|
setting depends on which manager runs, and a resolver module that knew about network managers would
|
||||||
|
be the wrong module knowing the wrong thing.
|
||||||
|
3. **A manifest field in `resolv-conf` naming `node-uplink`.** Rejected for the reason ADR 0207
|
||||||
|
rejected its own option 2: a second module claiming the same seat would have to restate it, and
|
||||||
|
one that forgot would pass.
|
||||||
|
4. **The seat carries what its holder needs beside it.** `node-resolver-config` names `node-uplink`;
|
||||||
|
any module claiming the former depends on the latter, derived from the claim and judged exactly as
|
||||||
|
ADR 0207 judges a resource's dependency. Chosen.
|
||||||
|
|
||||||
|
**For the split-DNS module:** keep it for a machine that wants systemd-resolved in charge, or remove it.
|
||||||
|
Kept, it is a second answer to a question ADR 0196 settled, and a claimant the catalogue offers
|
||||||
|
without a record allowing it. Removed.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
**1. `node-dns-resolver` is deleted from the mesh's set.** The controller's compiled set no longer
|
||||||
|
carries it, and a numbered migration of the controller's store deletes its row and any alias naming
|
||||||
|
it. No alias is kept: nothing was renamed, and a manifest still claiming it should be refused at
|
||||||
|
registration, naming the seat. This completes ADR 0194's retirement; nothing it decided changes.
|
||||||
|
|
||||||
|
**2. A seat may name the node seats its holder needs held on the same node.** A module claiming such a
|
||||||
|
seat depends on each of them. The dependency is a third source beside ADR 0207's resources and ADR
|
||||||
|
0210's contributions, and everything ADR 0207 §3 and §4 say of those applies unchanged: met by any
|
||||||
|
module assigned to the node, the claimant included; judged over the node's whole set; refused at
|
||||||
|
`assign` naming the seat and the catalogue's possible holders; refused at composition; and only said,
|
||||||
|
never refused, when no module in the catalogue could hold the needed seat. Unassigning the needed
|
||||||
|
seat's last holder beneath a dependent is refused, naming the dependent. What a seat needs is part of
|
||||||
|
the mesh's definition of the role: compiled with the set, never stored, as ADR 0212 keeps what a seat
|
||||||
|
receives. Adding a need to a seat is a decision, recorded.
|
||||||
|
|
||||||
|
**3. `node-resolver-config` needs `node-uplink`.** The holder that writes the resolver file is right
|
||||||
|
only while the network manager is told to leave it alone, and that telling is the uplink holder's
|
||||||
|
(ADR 0117). Every manager the catalogue knows — NetworkManager, systemd-networkd, dhcpcd — holds
|
||||||
|
`node-uplink`, so the refusal always has a remedy to name.
|
||||||
|
|
||||||
|
**4. `resolved-split-dns` leaves the catalogue.** `resolv-conf` is the only module claiming
|
||||||
|
`node-resolver-config`. Its resolver file's comment says the uplink's holder is required beside it,
|
||||||
|
rather than naming alternatives to assign instead.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
- **The set reads as the mesh is.** Thirty-seven seats in the compiled set; the overview no longer
|
||||||
|
lists a role nothing can fill.
|
||||||
|
- **A machine cannot be given the mesh's resolver file without its network manager being told to keep
|
||||||
|
off it.** A machine with no manager at all — a static configuration — needs the smallest holder,
|
||||||
|
`dhcpcd`, or a new module holding `node-uplink` for its way of configuring the link. That is the
|
||||||
|
point: such a machine has to say what manages its link before the mesh writes a file the manager
|
||||||
|
could overwrite.
|
||||||
|
- **Order of assignment on a new machine**: the uplink holder before or with `resolv-conf`, in one
|
||||||
|
`assign` when together. On the production mesh nothing changes: every node already holds both.
|
||||||
|
- **The uplink becomes harder to take away.** Unassigning a machine's manager module while
|
||||||
|
`resolv-conf` stays is refused; replacing one manager with another is one act assigning the new and
|
||||||
|
unassigning the old, or the dependent goes first.
|
||||||
|
- **A machine wanting systemd-resolved has no module for it.** A future need for one is a new record,
|
||||||
|
not a revival of the removed module.
|
||||||
|
- **Changing `resolv-conf`'s comment rewrites `/etc/resolv.conf` on every node once**, with the same two
|
||||||
|
nameserver lines and options; only the comment differs.
|
||||||
|
- **The merge order matters.** The controller's tests read the catalogue beside them, and the two
|
||||||
|
changes are judged together: the controller's change and the catalogue's removal merge together,
|
||||||
|
the catalogue's first or in the same window, and the controller rolls out only once its test suite
|
||||||
|
passes against the merged catalogue.
|
||||||
|
|
||||||
|
## How it is checked
|
||||||
|
|
||||||
|
| Rule | Checked by |
|
||||||
|
|---|---|
|
||||||
|
| `node-dns-resolver` is not in the set, and the set has thirty-seven seats | mesh-controller's closed-set unit test on the compiled seats |
|
||||||
|
| The store's row goes with it | the migration, and after rollout the controller's `seats` verb listing no `node-dns-resolver` |
|
||||||
|
| `node-resolver-config` needs `node-uplink`, and a module claiming it depends on the uplink with nothing in its manifest | mesh-controller's seat-dependency tests on the seat definition and on a synthetic claimant |
|
||||||
|
| What a seat needs survives loading the set from the store | a unit test loading the store's rows, which carry no such column |
|
||||||
|
| `resolv-conf` without an uplink holder is refused at `assign`, naming `node-uplink` and its possible holders; beside one, or with one in the same act, it passes; a composition without one is refused | the same tests, and `assign` live |
|
||||||
|
| Unassigning the uplink's last holder beneath `resolv-conf` is refused | an unassign test |
|
||||||
|
| In the catalogue, `resolv-conf` depends on the uplink, dhcpcd, NetworkManager and systemd-networkd each hold it, and `resolv-conf` is the only claimant of `node-resolver-config` | a mesh-controller test reading the catalogue beside it |
|
||||||
|
| Two modules deciding what a machine asks are still refused on one node | the resolver test, now with a synthetic second claimant |
|
||||||
|
| Every node of the live mesh holding `node-resolver-config` also holds `node-uplink` | the controller's `seats` verb, read before this was decided and after it rolls out; `status` reports no unheld dependency |
|
||||||
|
|
||||||
|
## References
|
||||||
|
|
||||||
|
- [ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) — retired
|
||||||
|
`node-dns-resolver`; this record deletes it.
|
||||||
|
- [ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md) — no
|
||||||
|
stub, and so no systemd-resolved module.
|
||||||
|
- [ADR 0117](0117-a-machines-uplink-is-a-seat.md) — the uplink's holder keeps the manager off the
|
||||||
|
resolver file; this record makes that a checked dependency.
|
||||||
|
- [ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md) — serving and
|
||||||
|
asking as two seats.
|
||||||
|
- [ADR 0207](0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md),
|
||||||
|
[ADR 0210](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md) — the dependency mechanism
|
||||||
|
this extends.
|
||||||
|
- [ADR 0122](0122-a-seat-is-data-a-rename-is-a-database-update.md) — the set as data, which is why a
|
||||||
|
deletion is a migration.
|
||||||
|
- [The seats](../03-DESIGN/01-to-be/26-the-seats.md) and
|
||||||
|
[connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md), amended alongside.
|
||||||
|
- mesh-controller `internal/catalogue/seats.go`, `internal/catalogue/seat_dependencies.go`, and the
|
||||||
|
store migration deleting the row; mesh-catalog `modules/resolv-conf`.
|
||||||
@@ -194,6 +194,8 @@ python3 00-META/checks/index.py fail if stale
|
|||||||
- **0207** — [A module depends on the node seats that apply its resources](0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md)
|
- **0207** — [A module depends on the node seats that apply its resources](0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md)
|
||||||
- **0210** — [A tool's configuration is its seat holder's, and every other module extends it through the seat](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)
|
- **0210** — [A tool's configuration is its seat holder's, and every other module extends it through the seat](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)
|
||||||
- **0212** — [A seat says what it receives, and the machine's hotkeys are a seat](0212-a-seat-says-what-it-receives-and-the-machines-hotkeys-are-a-seat.md)
|
- **0212** — [A seat says what it receives, and the machine's hotkeys are a seat](0212-a-seat-says-what-it-receives-and-the-machines-hotkeys-are-a-seat.md)
|
||||||
|
- **0218** — [A plan sends grants before code, rolls a module out one machine first, and a newer merge takes over an older plan](0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md)
|
||||||
|
- **0219** — [The build queue is controlled through the controller and the build seat](0219-the-build-queue-is-controlled-through-the-controller-and-the-build-seat.md)
|
||||||
|
|
||||||
### Its tiers, from the bottom up
|
### Its tiers, from the bottom up
|
||||||
|
|
||||||
@@ -316,6 +318,7 @@ python3 00-META/checks/index.py fail if stale
|
|||||||
- **0214** — [Backups guard against mistakes, stay on the machine, and are declared by the module that owns the data](0214-backups-guard-against-mistakes-and-stay-on-the-machine.md)
|
- **0214** — [Backups guard against mistakes, stay on the machine, and are declared by the module that owns the data](0214-backups-guard-against-mistakes-and-stay-on-the-machine.md)
|
||||||
- **0215** — [The machine's message bus is a node seat, and it is never restarted live](0215-the-machines-message-bus-is-a-node-seat-and-is-never-restarted-live.md)
|
- **0215** — [The machine's message bus is a node seat, and it is never restarted live](0215-the-machines-message-bus-is-a-node-seat-and-is-never-restarted-live.md)
|
||||||
- **0216** — [The agent's configuration is registered through its module, at three scopes, and served as one plugin](0216-the-agents-configuration-is-registered-through-its-module-at-three-scopes-and-served-as-one-plugin.md)
|
- **0216** — [The agent's configuration is registered through its module, at three scopes, and served as one plugin](0216-the-agents-configuration-is-registered-through-its-module-at-three-scopes-and-served-as-one-plugin.md)
|
||||||
|
- **0220** — [What a machine asks needs its uplink held, and the retired resolver pieces go](0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md)
|
||||||
|
|
||||||
### How it is built
|
### How it is built
|
||||||
|
|
||||||
|
|||||||
@@ -7,14 +7,16 @@ code:
|
|||||||
- mesh-controller internal/identity/authority.go
|
- mesh-controller internal/identity/authority.go
|
||||||
- mesh-controller cmd/mesh-controller/plan.go (the names the roster publishes)
|
- mesh-controller cmd/mesh-controller/plan.go (the names the roster publishes)
|
||||||
- mesh-controller internal/catalogue/zones.go (the zones a module answers, ADR 0199)
|
- mesh-controller internal/catalogue/zones.go (the zones a module answers, ADR 0199)
|
||||||
- mesh-controller internal/catalogue/seats.go (mesh-dns-resolver, node-hosts-file)
|
- mesh-controller internal/catalogue/seats.go (mesh-dns-resolver, node-hosts-file, node-resolver-config needing node-uplink)
|
||||||
|
- mesh-controller internal/catalogue/seat_dependencies.go (a seat's need checked at assignment, ADR 0220)
|
||||||
- mesh-catalog modules/dnsmasq (the mesh's one resolver)
|
- mesh-catalog modules/dnsmasq (the mesh's one resolver)
|
||||||
- mesh-catalog modules/resolv-conf (what a node asks)
|
- mesh-catalog modules/resolv-conf (what a node asks)
|
||||||
- mesh-catalog modules/hosts (a node's /etc/hosts)
|
- mesh-catalog modules/hosts (a node's /etc/hosts)
|
||||||
- mesh-host internal/identity/serving.go
|
- mesh-host internal/identity/serving.go
|
||||||
- mesh-host internal/apply (the service that reflects a rule set)
|
- mesh-host internal/apply (the service that reflects a rule set)
|
||||||
updated: 2026-10-03
|
updated: 2026-10-05
|
||||||
decisions:
|
decisions:
|
||||||
|
- 02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md
|
||||||
- 02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md
|
- 02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md
|
||||||
- 02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md
|
- 02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md
|
||||||
- 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
|
- 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
|
||||||
@@ -374,10 +376,21 @@ public names keep resolving then, and `.internal` is never asked of a public res
|
|||||||
answers. Containers take the same two from their machine, the runtime copying non-loopback resolvers
|
answers. Containers take the same two from their machine, the runtime copying non-loopback resolvers
|
||||||
into every container, so the runtime is given no `dns` of its own
|
into every container, so the runtime is given no `dns` of its own
|
||||||
([ADR 0196](../../02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md),
|
([ADR 0196](../../02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md),
|
||||||
replacing ADR 0194's per-node `systemd-resolved` stub).
|
replacing ADR 0194's per-node `systemd-resolved` stub). `resolv-conf` is the one module the catalogue
|
||||||
|
offers for it; the systemd-resolved split-DNS module that once claimed the same seat is gone
|
||||||
|
([ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md)).
|
||||||
|
|
||||||
|
**That file stays the mesh's only beside the uplink.** A network manager rewrites `/etc/resolv.conf`
|
||||||
|
on every connectivity change unless it is told not to, and the module holding `node-uplink` is what
|
||||||
|
tells it ([ADR 0117](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md)). So
|
||||||
|
`node-resolver-config` needs `node-uplink` held on the same node: assigning the resolver file to a
|
||||||
|
machine with no uplink holder is refused, naming the managers' modules that could hold it, and taking
|
||||||
|
the last uplink holder from beneath it is refused too
|
||||||
|
([ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md)). *Checked by the controller's seat-dependency tests and by `assign` live.*
|
||||||
|
|
||||||
**No node holds a copy.** The per-node resolver, its zones file and the mesh's region of `/etc/hosts`
|
**No node holds a copy.** The per-node resolver, its zones file and the mesh's region of `/etc/hosts`
|
||||||
go: every resolution fault found on 2026-10-03 was a copy disagreeing with the truth — a hosts file
|
go, and the per-node resolver's seat with them, deleted from the set once nothing claimed it
|
||||||
|
([ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md)): every resolution fault found on 2026-10-03 was a copy disagreeing with the truth — a hosts file
|
||||||
read once at start, an operator's old line beside the mesh's, a node's resolver lent to a LAN. No
|
read once at start, an operator's old line beside the mesh's, a node's resolver lent to a LAN. No
|
||||||
member's resolver answers a LAN; a router pointing at one is moved first. *Checked by each node's
|
member's resolver answers a LAN; a router pointing at one is moved first. *Checked by each node's
|
||||||
`/etc/resolv.conf` naming `mesh-resolver` then a public resolver, by no node but the holder answering
|
`/etc/resolv.conf` naming `mesh-resolver` then a public resolver, by no node but the holder answering
|
||||||
@@ -1035,10 +1048,14 @@ The list is worth having in one place, because it is most of the argument:
|
|||||||
|
|
||||||
## Open
|
## Open
|
||||||
|
|
||||||
- **One resolver ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)).** Not built: every node still runs
|
- **One resolver ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)).** *Correction of fact,
|
||||||
`node-dns-resolver`. The migration's four steps are in the record, in order.
|
2026-10-05:* no node holds `node-dns-resolver` any more, and
|
||||||
Nor are zones or the hosts file's holder ([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)): the
|
[ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md) deletes the seat. What stood here before: *"Not built: every node still runs
|
||||||
workstation moves to the one resolver only once both exist, its lab and operator names depending on them.
|
`node-dns-resolver`. The migration's four steps are in the record, in order."* Nor did it
|
||||||
|
stay true that *"the workstation moves to the one resolver only once"* zones and the hosts file's
|
||||||
|
holder ([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md))
|
||||||
|
exist: every node, the workstation included, asks the one resolver, and every node holds
|
||||||
|
`node-hosts-file`.
|
||||||
|
|
||||||
- ~~**What happens when the hub is down.**~~ **Resolved** by
|
- ~~**What happens when the hub is down.**~~ **Resolved** by
|
||||||
[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md), together with `06`'s
|
[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md), together with `06`'s
|
||||||
|
|||||||
@@ -5,8 +5,9 @@ code:
|
|||||||
- mesh-controller cmd/mesh-builder
|
- mesh-controller cmd/mesh-builder
|
||||||
- mesh-controller internal/builder
|
- mesh-controller internal/builder
|
||||||
- mesh-catalog modules/build-agent
|
- mesh-catalog modules/build-agent
|
||||||
updated: 2026-10-04
|
updated: 2026-10-05
|
||||||
decisions:
|
decisions:
|
||||||
|
- 02-DECISIONS/0219-the-build-queue-is-controlled-through-the-controller-and-the-build-seat.md
|
||||||
- 02-DECISIONS/0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md
|
- 02-DECISIONS/0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md
|
||||||
- 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md
|
- 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md
|
||||||
- 02-DECISIONS/0189-the-store-keeps-what-the-records-name.md
|
- 02-DECISIONS/0189-the-store-keeps-what-the-records-name.md
|
||||||
@@ -346,6 +347,21 @@ scheduled step with the server held still — which is what `while-stopped` exis
|
|||||||
keeps is still a manifest and so still referenced, and the dangerous flag is not needed once the
|
keeps is still a manifest and so still referenced, and the dangerous flag is not needed once the
|
||||||
mesh is the one deciding.
|
mesh is the one deciding.
|
||||||
|
|
||||||
|
**An archive is held by a manifest of its own** (2026-10-05,
|
||||||
|
[issue 253](../../04-ISSUES/253-the-stores-collector-would-delete-every-archive-the-mesh-keeps/00-report.md)).
|
||||||
|
The sentence above held for images and not for archives: an archive was published as a bare blob that no
|
||||||
|
manifest names, and the store's collector keeps only what a manifest names, so a nightly collection
|
||||||
|
would have removed every archive the mesh keeps, the current ones included. So:
|
||||||
|
|
||||||
|
- an archive is published with a manifest that names it and nothing else, built from the archive's digest
|
||||||
|
and size alone so it can be computed again from the record;
|
||||||
|
- the sweep makes sure every archive it keeps is held that way before it lets anything go, and lets go of
|
||||||
|
an archive by removing its manifest first;
|
||||||
|
- the reference a machine fetches is unchanged.
|
||||||
|
|
||||||
|
The collector runs as a dry run until the controller reports no kept archive unheld; only then does it
|
||||||
|
collect for real.
|
||||||
|
|
||||||
A machine behind by more than five builds of a module, recreating a container, cannot pull what it
|
A machine behind by more than five builds of a module, recreating a container, cannot pull what it
|
||||||
was running. It is already a machine the mesh reports as behind, and the answer is the current
|
was running. It is already a machine the mesh reports as behind, and the answer is the current
|
||||||
declaration.
|
declaration.
|
||||||
@@ -379,3 +395,26 @@ not of the recipe: one artifact declared per target, one build each.
|
|||||||
A component's version stops being stamped in at link time. It is unpacked into a directory named for
|
A component's version stops being stamped in at link time. It is unpacked into a directory named for
|
||||||
its version, so it reads its version from its own path, and a build no longer has to know what it will
|
its version, so it reads its version from its own path, and a build no longer has to know what it will
|
||||||
be called.
|
be called.
|
||||||
|
|
||||||
|
## The build queue is controlled
|
||||||
|
|
||||||
|
*2026-10-05 — [ADR 0219](../../02-DECISIONS/0219-the-build-queue-is-controlled-through-the-controller-and-the-build-seat.md).*
|
||||||
|
|
||||||
|
What is asked of the build seat can be seen and controlled from the console. The controller holds the
|
||||||
|
queue, and each machine's build agent holds its own process.
|
||||||
|
|
||||||
|
- **The controller's verbs:**
|
||||||
|
- `queue` lists every ask, waiting, in flight or dead;
|
||||||
|
- `cancel` and `clear` take asks back;
|
||||||
|
- `rebuild` asks again under a new id;
|
||||||
|
- `replay` asks again for a past build at its commit, as a dry run unless told to register, and never
|
||||||
|
over a newer build without saying so;
|
||||||
|
- `kill`, `pause` and `resume` are passed on to the seat on the right machine.
|
||||||
|
- **The build seat's verbs, on each machine:**
|
||||||
|
- `current` says what is building here and whether this holder is paused;
|
||||||
|
- `kill` ends the build's whole process group and every container it started;
|
||||||
|
- `pause` and `resume` set a flag the holder reads before it takes the next ask, which survives its
|
||||||
|
restart.
|
||||||
|
|
||||||
|
Every action that drops work leaves a failed outcome, so a plan waiting on that build fails and says why
|
||||||
|
rather than waiting. A plan keeps the id of every build it asked for. A plan waiting on a paused seat says so and is not counted late; a failed plan can be resumed with `plans retry`, and a `rebuild` of a module a plan holds joins that plan.
|
||||||
|
|||||||
@@ -4,14 +4,16 @@ status: in-progress
|
|||||||
code:
|
code:
|
||||||
- mesh-controller internal/catalogue/seats.go
|
- mesh-controller internal/catalogue/seats.go
|
||||||
- mesh-controller internal/catalogue/resolve.go
|
- mesh-controller internal/catalogue/resolve.go
|
||||||
|
- mesh-controller internal/catalogue/seat_dependencies.go
|
||||||
- mesh-controller internal/inventory/seats.go
|
- mesh-controller internal/inventory/seats.go
|
||||||
- mesh-controller internal/inventory/migrations/0039-a-seat-is-held-by-one-assignment-on-record.sql
|
- mesh-controller internal/inventory/migrations/0039-a-seat-is-held-by-one-assignment-on-record.sql
|
||||||
- mesh-controller cmd/mesh-controller/seats.go
|
- mesh-controller cmd/mesh-controller/seats.go
|
||||||
- mesh-controller cmd/mesh-controller/source.go
|
- mesh-controller cmd/mesh-controller/source.go
|
||||||
- mesh-controller internal/inventory/migrations/0032-a-source-may-live-on-a-seat.sql
|
- mesh-controller internal/inventory/migrations/0032-a-source-may-live-on-a-seat.sql
|
||||||
- mesh-catalog modules/gitea/module.json
|
- mesh-catalog modules/gitea/module.json
|
||||||
updated: 2026-10-03
|
updated: 2026-10-05
|
||||||
decisions:
|
decisions:
|
||||||
|
- 02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md
|
||||||
- 02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md
|
- 02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md
|
||||||
- 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
|
- 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
|
||||||
- 02-DECISIONS/0161-what-deserves-a-seat.md
|
- 02-DECISIONS/0161-what-deserves-a-seat.md
|
||||||
@@ -129,12 +131,12 @@ convention, which later seats departed from.
|
|||||||
| `mesh-git` | `git` | mesh | `git` | the forge |
|
| `mesh-git` | `git` | mesh | `git` | the forge |
|
||||||
| `mesh-build-machine` | `the-build-machine` | node | — | a builder |
|
| `mesh-build-machine` | `the-build-machine` | node | — | a builder |
|
||||||
| `mesh-resolver` | — | mesh | — | the mesh's one resolver, holding every node's internal domain ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)) |
|
| `mesh-resolver` | — | mesh | — | the mesh's one resolver, holding every node's internal domain ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)) |
|
||||||
| ~~`mesh-dns-port`~~ | `the-dns-port` | node | — | retired by [ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md): the local resolver became the mesh's one |
|
| ~~`mesh-dns-port`~~ | `the-dns-port` | node | — | retired by [ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md): the local resolver became the mesh's one; deleted from the set, and from the store's table, once nothing claimed it ([ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md)) |
|
||||||
| `node-hosts-file` | — | node | — | owns `/etc/hosts`: the machine's own lines and the operator's kept region, changed through its verbs `entries`, `add`, `remove` ([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)) |
|
| `node-hosts-file` | — | node | — | owns `/etc/hosts`: the machine's own lines and the operator's kept region, changed through its verbs `entries`, `add`, `remove` ([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)) |
|
||||||
| `mesh-intrusion-prevention` | `the-intrusion-prevention` | node | — | an intrusion-prevention service |
|
| `mesh-intrusion-prevention` | `the-intrusion-prevention` | node | — | an intrusion-prevention service |
|
||||||
| `mesh-packet-filter` | `the-packet-filter` | node | — | the packet filter |
|
| `mesh-packet-filter` | `the-packet-filter` | node | — | the packet filter |
|
||||||
| `mesh-private-network` | `the-private-network` | node | — | the private network the mesh runs over |
|
| `mesh-private-network` | `the-private-network` | node | — | the private network the mesh runs over |
|
||||||
| `mesh-resolver-configuration` | `the-resolver-configuration` | node | — | whichever of the alternative resolver configurations is chosen |
|
| `mesh-resolver-configuration` | `the-resolver-configuration` | node | — | the module writing `/etc/resolv.conf`; its holder needs the uplink's held on the same node ([ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md)) |
|
||||||
| `mesh-showcase` | `the-showcase` | node | — | the showcase module |
|
| `mesh-showcase` | `the-showcase` | node | — | the showcase module |
|
||||||
|
|
||||||
The controller holds **the mesh's own** entries in code, and a test asserts their size and that
|
The controller holds **the mesh's own** entries in code, and a test asserts their size and that
|
||||||
@@ -202,10 +204,28 @@ Nothing reaches that state by accident: an unknown manifest field is refused out
|
|||||||
protocol was written as one. Checked by a registration test accepting a node seat with no protocol
|
protocol was written as one. Checked by a registration test accepting a node seat with no protocol
|
||||||
and by the showcase manifest, which declares one.
|
and by the showcase manifest, which declares one.
|
||||||
|
|
||||||
Most node seats deliver nothing. They say which module is this machine's packet filter, or which of
|
Most node seats deliver nothing. They say which module is this machine's packet filter, or which
|
||||||
two alternative resolver configurations it runs, and a second holder is refused. That is the whole of
|
module writes its resolver file, and a second holder is refused. That is the whole of
|
||||||
their job, and it is a real one: it is the mesh saying what a machine is, in words a person can read.
|
their job, and it is a real one: it is the mesh saying what a machine is, in words a person can read.
|
||||||
|
|
||||||
|
## A seat that needs another beside it
|
||||||
|
|
||||||
|
*2026-10-05* ([ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md)). **A seat may name the node seats its holder needs held on the
|
||||||
|
same node**, because some roles are right only while another is filled beside them. The holder of the
|
||||||
|
resolver configuration writes `/etc/resolv.conf`, and that file stays the mesh's only while the
|
||||||
|
machine's network manager is told to leave it alone — which is what the uplink's holder does
|
||||||
|
([ADR 0117](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md)). So the resolver configuration
|
||||||
|
needs the uplink.
|
||||||
|
|
||||||
|
**A module claiming such a seat depends on each seat it needs**, exactly as a module declaring a
|
||||||
|
service depends on the service manager
|
||||||
|
([ADR 0207](../../02-DECISIONS/0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md)):
|
||||||
|
derived from the claim and never written in a manifest, met by any module assigned to the node, judged
|
||||||
|
over the node's whole set, refused at assignment naming the seat and the modules that could hold it,
|
||||||
|
and refused at composition. Taking the needed seat's last holder from beneath a dependent is refused
|
||||||
|
too. What a seat needs is the mesh's definition of the role, compiled with the set and never stored,
|
||||||
|
and adding a need is a decision.
|
||||||
|
|
||||||
## The overview
|
## The overview
|
||||||
|
|
||||||
The controller lists every seat in the set with its scope, what it delivers, and its holder as a node
|
The controller lists every seat in the set with its scope, what it delivers, and its holder as a node
|
||||||
@@ -259,5 +279,7 @@ checked as their tables say:
|
|||||||
| `secret` has one provider, the holder of `mesh-vault` | 0161: a second claimant of the seat is refused by name (`CanHold`); *correction of fact, 2026-10-01: no parser rule ever reserved the word, the seat does the work*. |
|
| `secret` has one provider, the holder of `mesh-vault` | 0161: a second claimant of the seat is refused by name (`CanHold`); *correction of fact, 2026-10-01: no parser rule ever reserved the word, the seat does the work*. |
|
||||||
| A singular fact about machines is a placement of capacity one, refused by name | 0161: the overlay command's test for a second hub; the store's unique index. |
|
| A singular fact about machines is a placement of capacity one, refused by name | 0161: the overlay command's test for a second hub; the store's unique index. |
|
||||||
| A holder of `node-uplink` is the dialect the machine runs | 0161: the host reports `uplink-<manager>` in its profile with every report; a resolution test refuses the other holder naming the capability. |
|
| A holder of `node-uplink` is the dialect the machine runs | 0161: the host reports `uplink-<manager>` in its profile with every report; a resolution test refuses the other holder naming the capability. |
|
||||||
|
| A seat's holder has the seats it needs beside it: the resolver configuration is refused at `assign` without the uplink held on its node, and the uplink's last holder cannot be taken from beneath it | 0220: seat-dependency tests on the definition, on a synthetic claimant, at assign, at unassign and at composition, and one reading the catalogue for the uplink's possible holders. |
|
||||||
|
| A retired seat leaves the set once nothing claims it, from the compiled set and the store's table both | 0220: the closed-set test's count, and the store migration that deletes the row. |
|
||||||
| Holdings are derived, and the overview lists every seat | 0118: the `seats` command test, including an unheld seat. |
|
| Holdings are derived, and the overview lists every seat | 0118: the `seats` command test, including an unheld seat. |
|
||||||
| A build source on the seat records no address; an unheld seat refuses only self-hosted builds | 0111's tests. |
|
| A build source on the seat records no address; an unheld seat refuses only self-hosted builds | 0111's tests. |
|
||||||
|
|||||||
@@ -2,8 +2,9 @@
|
|||||||
layer: to-be
|
layer: to-be
|
||||||
status: proposed
|
status: proposed
|
||||||
code: []
|
code: []
|
||||||
updated: 2026-10-01
|
updated: 2026-10-05
|
||||||
decisions:
|
decisions:
|
||||||
|
- 02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md
|
||||||
- 02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md
|
- 02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md
|
||||||
- 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md
|
- 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md
|
||||||
- 02-DECISIONS/0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md
|
- 02-DECISIONS/0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md
|
||||||
@@ -129,6 +130,25 @@ controller replaced mid-plan resumes from the store. `status` lists open plans a
|
|||||||
has waited too long. The transition discipline for breaking changes in the list above is still
|
has waited too long. The transition discipline for breaking changes in the list above is still
|
||||||
unwritten, and still the next thing.
|
unwritten, and still the next thing.
|
||||||
|
|
||||||
|
## How a plan sends (2026-10-05)
|
||||||
|
|
||||||
|
Revision, [ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md). Three rules on how a plan delivers what it built.
|
||||||
|
|
||||||
|
- **Grants travel before code.** A send issues the memberships for the machines it is about to send to
|
||||||
|
before their declarations. The machine holding the bus is sent first when the list of bus users changes.
|
||||||
|
A membership that could not be issued fails the send.
|
||||||
|
- **One machine first.** Unless a module's upgrade policy says *together*, a plan sends it to one machine,
|
||||||
|
the first by name, and to the rest only once that machine reports the new declaration applied and
|
||||||
|
current. A first machine that fails stops the module's rollout there, with the reason in the plan.
|
||||||
|
- **A newer merge takes over.** A merge's plan supersedes every older open plan for the same repository
|
||||||
|
and branch, and takes in the modules they had not yet built. A plan that waits on nothing can be closed
|
||||||
|
by hand, by its id.
|
||||||
|
|
||||||
|
What a merge changed is read from the forge whole, page by page
|
||||||
|
([issue 252](../../04-ISSUES/252-a-merges-changed-modules-were-read-wrong/00-report.md)). A changed path in a
|
||||||
|
module directory the mesh does not hold yet is that module's own, not shared code, when its definition is
|
||||||
|
among the changed paths.
|
||||||
|
|
||||||
## Why now, and why not yet
|
## Why now, and why not yet
|
||||||
|
|
||||||
**Why it matters:** self-update is the difference between a mesh a person maintains by typing
|
**Why it matters:** self-update is the difference between a mesh a person maintains by typing
|
||||||
|
|||||||
+8
-2
@@ -1,8 +1,8 @@
|
|||||||
---
|
---
|
||||||
status: open
|
status: resolved
|
||||||
opened: 2026-10-04
|
opened: 2026-10-04
|
||||||
located-in: [mesh-host internal/apply/apply.go, mesh-host internal/apply/schedule.go]
|
located-in: [mesh-host internal/apply/apply.go, mesh-host internal/apply/schedule.go]
|
||||||
fixed-by:
|
fixed-by: mesh-host PR #23
|
||||||
amended-design:
|
amended-design:
|
||||||
---
|
---
|
||||||
|
|
||||||
@@ -62,3 +62,9 @@ is a hole in something new rather than something that broke.
|
|||||||
at the cost of a second source for "is this container meant to be running".
|
at the cost of a second source for "is this container meant to be running".
|
||||||
- Either way: should the *report* say a window is open, so a machine that looks half-stopped at
|
- Either way: should the *report* say a window is open, so a machine that looks half-stopped at
|
||||||
03:31 reads as working rather than broken?
|
03:31 reads as working rather than broken?
|
||||||
|
|
||||||
|
## Resolved — 2026-10-05
|
||||||
|
|
||||||
|
An apply that arrives during a window now leaves the containers the window holds alone, and reports them `held-still`, naming the step. The first apply after the window converges them. The window is recorded under the host's state directory, the one place every applier on a machine shares: the daemon, a hand-run apply and the installer. It is released on every path, and lapses after six hours or when its process is gone. The machine's report lists open windows. Live on all four machines the same evening. Of the issue's two options, the narrower was taken: nothing blocks, so a push is never held for the length of a window.
|
||||||
|
|
||||||
|
Accepted and said in the change: an apply that inspected a container as running in the instant a window opens can still recreate it.
|
||||||
|
|||||||
@@ -1,8 +1,8 @@
|
|||||||
---
|
---
|
||||||
status: located
|
status: resolved
|
||||||
opened: 2026-10-05
|
opened: 2026-10-05
|
||||||
located-in: [mesh-catalog modules/claude-code, mesh-catalog modules/claude-licence-manager]
|
located-in: [mesh-catalog modules/claude-code, mesh-catalog modules/claude-licence-manager]
|
||||||
fixed-by:
|
fixed-by: mesh-catalog PR #45
|
||||||
amended-design:
|
amended-design:
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
@@ -31,3 +31,9 @@
|
|||||||
**Ruled out.** The bus delivered every binding: each machine that took the fourth binding did so in the
|
**Ruled out.** The bus delivered every binding: each machine that took the fourth binding did so in the
|
||||||
same second it was published. The licences themselves were sound. The remaining licence refreshed
|
same second it was published. The licences themselves were sound. The remaining licence refreshed
|
||||||
on every attempt, and the second account's login was adopted from its first report.
|
on every attempt, and the second account's login was adopted from its first report.
|
||||||
|
|
||||||
|
## Resolved — 2026-10-05
|
||||||
|
|
||||||
|
Live on all four machines and the manager the same morning. After the restart every machine reported
|
||||||
|
the generation it held, and the manager's bindings matched them. The manager logged no failure while
|
||||||
|
moving its sequence past them.
|
||||||
|
|||||||
+50
@@ -0,0 +1,50 @@
|
|||||||
|
---
|
||||||
|
status: open
|
||||||
|
opened: 2026-10-05
|
||||||
|
located-in: []
|
||||||
|
fixed-by:
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 245 — `status` calls a module behind when only its repository moved
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
2026-10-05. After three catalogue merges that changed 11 modules, `status` listed **69 modules
|
||||||
|
behind their source**, each as `holds <older commit>, source has <newer commit>`, with the advice
|
||||||
|
"`build --behind` builds them; `push --behind` sends them on".
|
||||||
|
|
||||||
|
Between the two commits, `git diff --name-only` shows changes under 11 module directories only. For
|
||||||
|
58 of the 69 modules listed, for example a Bluetooth module, the container runtime's, the forge's,
|
||||||
|
the package manager's and the bus's, the diff of the module's own path is empty. Their sources did
|
||||||
|
not change; only the repository's commit did.
|
||||||
|
|
||||||
|
The merges' plans were right: they rebuilt the changed modules and those that depend on them, by
|
||||||
|
tier. Only the report was wrong. An agent following the report's own advice ran `build --behind`,
|
||||||
|
which rebuilt all 69. The rebuilt bus module was rolled out, and its container was replaced on the
|
||||||
|
control node. Every node's runtime lost the bus for about a minute.
|
||||||
|
|
||||||
|
## Why it matters beyond this instance
|
||||||
|
|
||||||
|
"Behind" is the word a person and an agent act on, and `status` attaches a command to it. A module
|
||||||
|
is held to a commit of its repository, so after any merge almost every module of that repository
|
||||||
|
reads behind. The list then says nothing about what needs building: it hides the few modules that
|
||||||
|
really are behind among the many that are not, and it invites a rebuild of everything, which is not
|
||||||
|
a harmless act (above).
|
||||||
|
|
||||||
|
## The operator's direction (2026-10-05)
|
||||||
|
|
||||||
|
The only truth is the outcome of the build plan. A plan already decides, from a change, which modules
|
||||||
|
it affects: those whose sources changed and those that depend on them, tier by tier. "Behind" means
|
||||||
|
a module that a plan has decided to rebuild and has not yet rebuilt or rolled out, and nothing else.
|
||||||
|
No second comparison beside the plan is made, whether of commits, of folders or of files, because a
|
||||||
|
second answer to the same question is how the two came to disagree.
|
||||||
|
|
||||||
|
## Open questions
|
||||||
|
|
||||||
|
1. Where does `status` read "behind" from today, and what replaces it: the open plans' remaining
|
||||||
|
tiers?
|
||||||
|
2. What does a module's recorded commit mean once a plan that leaves it untouched has run? Does it
|
||||||
|
move forward, or does the record stop carrying a commit that only says when it was last built?
|
||||||
|
3. Should `build --behind` and `push --behind` take their lists from the same place, so that they can
|
||||||
|
never act on a module no plan named?
|
||||||
+65
@@ -0,0 +1,65 @@
|
|||||||
|
---
|
||||||
|
status: resolved
|
||||||
|
opened: 2026-10-05
|
||||||
|
located-in: [mesh-tools]
|
||||||
|
fixed-by: mesh-tools pull request 13 — discovery waits for every runtime that answered PING, names the ones it missed, and mesh_runtimes says who answered
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 246 — The console says a module runs nowhere when a runtime answers late
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
2026-10-05. On the laptop, the console answered `mesh_machine` for the laptop with no modules and no
|
||||||
|
seats, and a call to one of the laptop's modules with "nothing in the mesh is called slack". The
|
||||||
|
laptop's own runtime said at the same moment that it served 275 tools for 48 modules. A minute
|
||||||
|
earlier, a call to another laptop module had been answered with "it does not run on the laptop; it
|
||||||
|
runs on the workstation", and a retry of the same call worked. Nothing was logged anywhere.
|
||||||
|
|
||||||
|
An agent worked around it by calling the runtime's local MCP port directly, which gives the right
|
||||||
|
answer and bypasses everything the console stands for: one way in, one account, one record of what
|
||||||
|
was called. The operator asked for a tool instead.
|
||||||
|
|
||||||
|
## What was measured
|
||||||
|
|
||||||
|
A read-only probe on the laptop's own runtime credential timed the discovery answers over 25 rounds
|
||||||
|
against the live bus, whose round trip from the laptop was about 40 ms:
|
||||||
|
|
||||||
|
- The laptop's runtime answer was the largest on the mesh, at about 164 kB with 341 endpoints. That
|
||||||
|
is far below the bus's message limit, and it was never shortened.
|
||||||
|
- It arrived last in every round: a median of about 365 ms, and once 813 ms. The other runtimes
|
||||||
|
answered within 180 to 275 ms, and the controller within 50 ms.
|
||||||
|
- The console gathered discovery answers for a fixed 750 ms. Inside the console, the gather runs
|
||||||
|
beside two controller calls and about 570 kB of answers on the same link, and it is slower still
|
||||||
|
while a runtime re-serves after a restart.
|
||||||
|
|
||||||
|
## Root cause
|
||||||
|
|
||||||
|
Discovery decided who was there by who answered in a fixed window. A late answer was not a failure
|
||||||
|
to anyone, so nobody said it. The index simply lacked that runtime, and every answer built on the
|
||||||
|
index then stated as fact that the runtime's modules did not exist, or ran only elsewhere.
|
||||||
|
|
||||||
|
Ruled out by measurement or by reading the code: an answer too large for the bus, subscriptions lost
|
||||||
|
when the bus reconnects, the console not counting its own machine's answer, and the merge of two
|
||||||
|
answers dropping a machine.
|
||||||
|
|
||||||
|
## Resolution
|
||||||
|
|
||||||
|
- Discovery asks who is there (PING, a hundred bytes, answered at once) beside what each serves
|
||||||
|
(INFO). It waits at least the old window, and then up to five seconds for every instance that said
|
||||||
|
it is there, so a large answer is waited for and a quiet mesh costs nothing extra.
|
||||||
|
- A runtime that said it is there and did not say what it serves in time, or that answered recently
|
||||||
|
and not now, is named. While one is unheard, the console never says an address is missing or runs
|
||||||
|
elsewhere: it says which runtime was not heard, and where the controller's records place the module.
|
||||||
|
- A new console tool, `mesh_runtimes`, says for every runtime how long its answer took, how large it
|
||||||
|
was, how many modules and tools it announced, whether it was shortened, and when it was last heard,
|
||||||
|
and which runtimes or machines were not heard.
|
||||||
|
- An announcement still too large after its descriptions are cut to their first line now leaves the
|
||||||
|
descriptions out, and says so.
|
||||||
|
|
||||||
|
## How it is checked
|
||||||
|
|
||||||
|
The fix ships with tests against a real bus: a runtime that answers after the old window is found and
|
||||||
|
called (the same test fails with the fixed window), a runtime that answers PING and never INFO is
|
||||||
|
named, and a restarted runtime, which answers under a new instance, is not reported as missed. Live,
|
||||||
|
`mesh_runtimes` shows every machine's answer and its time.
|
||||||
@@ -0,0 +1,52 @@
|
|||||||
|
---
|
||||||
|
status: open
|
||||||
|
opened: 2026-10-05
|
||||||
|
located-in: []
|
||||||
|
fixed-by:
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 247 — A module cannot put the operator's account in a group
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
2026-10-05. The module for a peripheral-lighting daemon was assigned to the laptop. Its package
|
||||||
|
installs the daemon and creates the daemon's group. The daemon then refuses to start: "User is not a
|
||||||
|
member of the openrazer group". The device files are the group's, so the daemon cannot reach the
|
||||||
|
devices.
|
||||||
|
|
||||||
|
The module's own check names the fix, which is to add the account to the group and log in again. No
|
||||||
|
module can declare that fix:
|
||||||
|
|
||||||
|
- The account is a `user` resource, and the login-shell module already declares it, to set its shell.
|
||||||
|
- A second module that declares the same account, only to add one group, is refused as a duplicate
|
||||||
|
name.
|
||||||
|
- There is no resource for one membership on its own. A whole-account declaration that lists groups
|
||||||
|
would also take from the account every group it does not list, including the operator's own.
|
||||||
|
|
||||||
|
So the step is done by hand, with `sudo`, outside the mesh, and nothing records why the account is in
|
||||||
|
the group.
|
||||||
|
|
||||||
|
## Why it matters beyond this instance
|
||||||
|
|
||||||
|
More modules need this than this one: input devices (`input`), serial ports (`uucp`), the container
|
||||||
|
runtime (`docker`), virtual machines (`libvirt`, `kvm`), and capture or scanner hardware. Each is a
|
||||||
|
fact a module knows and the operator's account needs. Today every one is a hand step that survives a
|
||||||
|
reinstall only by memory. A membership added by hand is also never taken away when the module that
|
||||||
|
needed it is unassigned.
|
||||||
|
|
||||||
|
## Open questions
|
||||||
|
|
||||||
|
1. Is a membership its own resource (account, group), held by the module that needs it and given
|
||||||
|
back on undeclare? Or is it a contribution to the account's holder, in the way ADR 0212 lets a
|
||||||
|
module contribute to a seat?
|
||||||
|
2. A membership takes effect at the next login. How does the module say so: a finding, or a
|
||||||
|
moment the power or session seat already knows?
|
||||||
|
3. What does undeclare do with a membership the account already had before any module declared it?
|
||||||
|
The host keeps what it found and gives it back, as it does with a whole file it wrote over.
|
||||||
|
|
||||||
|
## How it is checked
|
||||||
|
|
||||||
|
When fixed, assigning the lighting module to a machine whose account is not in the group puts the
|
||||||
|
account in the group, says that a new login is needed, and leaves the account's other groups as they
|
||||||
|
were. Unassigning it removes only a membership the module added.
|
||||||
+58
@@ -0,0 +1,58 @@
|
|||||||
|
---
|
||||||
|
status: resolved
|
||||||
|
opened: 2026-10-05
|
||||||
|
located-in: [mesh-controller internal/broker]
|
||||||
|
fixed-by: mesh-controller PR #51
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 248 — The controller's event consumer replayed a week, and held every new merge behind it
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
A merge to the catalogue at 15:17 never reached the controller. No plan was made and no build was asked
|
||||||
|
for it. Every machine kept running the build from before it. The merge just before, to the record, was
|
||||||
|
logged twice.
|
||||||
|
|
||||||
|
The bus showed why. The controller's durable consumer on the event stream was set to deliver
|
||||||
|
**everything the stream holds**, not from where it last was:
|
||||||
|
|
||||||
|
| | |
|
||||||
|
|---|---|
|
||||||
|
| the stream | a week of events, from message 1 514 to message 356 004 |
|
||||||
|
| the consumer delivered to | message 1 517, later 4 898 |
|
||||||
|
| acknowledged to | 0, later 1 569 |
|
||||||
|
| still to deliver | 6 955 merges and build outcomes, a week old |
|
||||||
|
| allowed outstanding | one at a time ([issue 175](../175-an-announcement-behind-a-long-build-comes-back/00-report.md)) |
|
||||||
|
|
||||||
|
So the controller was working through a week of past merges and build outcomes, one at a time, slowly. Every
|
||||||
|
new one — a merge, a build asked by hand — waited behind them. A build asked by hand finished on its
|
||||||
|
machine and was never registered. While this went on, the controller's client dropped messages
|
||||||
|
("slow consumer") several times, because heartbeats and reports share the loop with these events. From 15:17
|
||||||
|
its event loop did nothing more: no line logged, the one delivered event never acknowledged.
|
||||||
|
|
||||||
|
How the consumer came to deliver everything is not certain. The bus and the stores were rebuilt the night
|
||||||
|
before ([issue 241](../241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md)).
|
||||||
|
A consumer that is missing is made again by the controller's own assertion, and that made it with the
|
||||||
|
server's default, which is everything. [Issue 207](../207-a-re-made-worker-replayed-every-ask-the-stream-kept/00-report.md)
|
||||||
|
closed this for a consumer re-made because its type changed. It did not close it for one that is simply
|
||||||
|
not there.
|
||||||
|
|
||||||
|
## Why it matters
|
||||||
|
|
||||||
|
**Delivery stops, and nothing says so.** Status showed every plan done and no machine behind. The merge
|
||||||
|
that was missed is not "behind", because no plan was ever made for it. A replay of past merges can also
|
||||||
|
act on them again. Issue 207 records nine modules re-registered from the past the same way.
|
||||||
|
|
||||||
|
**The way out was a hand on the bus.** No verb resets a consumer. On the operator's explicit word, the
|
||||||
|
consumer was re-made from now with a one-off program run as the controller, otherwise unchanged. The
|
||||||
|
controller's loop still held the old event afterwards, so the new consumer's first delivery went
|
||||||
|
unacknowledged; the loop needs a restart of the controller to let go of it.
|
||||||
|
|
||||||
|
## Noticed alongside, not this issue
|
||||||
|
|
||||||
|
- Each merge to the catalogue planned 99 to 100 modules in two tiers and rebuilt modules it did not touch.
|
||||||
|
The output was byte-identical, so nothing was redeployed, but it costs minutes of the build machine
|
||||||
|
per merge.
|
||||||
|
- Three plans for three merges ran over each other. Each sent the build agent to every machine and asked
|
||||||
|
for the same builds. Nothing supersedes a plan for an older commit.
|
||||||
+54
@@ -0,0 +1,54 @@
|
|||||||
|
# 248 — Diagnosis
|
||||||
|
|
||||||
|
## 2026-10-05
|
||||||
|
|
||||||
|
1. A merge was not in the controller's log. The forge's own log showed nothing about delivering it, so the
|
||||||
|
question moved to the bus.
|
||||||
|
2. The bus's backlog tool named the controller's event consumer: 6 955 pending, redeliveries, one
|
||||||
|
unacknowledged. Its configuration was read from the server's monitoring endpoint: deliver policy
|
||||||
|
*all*, one outstanding, 30 seconds to acknowledge, five deliveries.
|
||||||
|
3. Sampled three times over a minute it did not move, and over the following hour it crawled forward
|
||||||
|
through week-old events. The controller's log held only its client's warnings: dropped messages,
|
||||||
|
and one refused reply.
|
||||||
|
4. The controller's code makes a missing consumer with the configuration it asserts, which sets no
|
||||||
|
deliver policy, so the server's default applies: everything. Issue 207's fix sets *from now* only on
|
||||||
|
the path where an existing consumer's type changes.
|
||||||
|
|
||||||
|
**Unblocked**, on the operator's explicit word: the consumer re-made from now with its configuration
|
||||||
|
otherwise unchanged (nothing pending afterwards). The controller's loop still held the old event, so
|
||||||
|
it takes a restart of the controller to let go of it; that restart waits for the operator's word.
|
||||||
|
|
||||||
|
**Fix** (mesh-controller, branch `fix/a-consumer-on-a-history-stream-starts-from-now`):
|
||||||
|
|
||||||
|
- a consumer may say it starts **from now** when it is made. The controller's event consumer does. A
|
||||||
|
consumer that exists keeps where it is. The server would refuse a changed start anyway.
|
||||||
|
- `broker consumer-reset <stream> <consumer>` re-makes a stuck consumer from now, its configuration
|
||||||
|
otherwise kept, and refuses a work queue, where what is pending is work. It is the person's act, said by
|
||||||
|
a command, rather than a one-off program.
|
||||||
|
|
||||||
|
Checked by a live test against a throwaway bus:
|
||||||
|
|
||||||
|
- made from now, a consumer holds none of the stream's past and does hold the next announcement;
|
||||||
|
- asserted again, it keeps its place;
|
||||||
|
- the default replays all of it;
|
||||||
|
- a reset leaves nothing pending and keeps every other setting;
|
||||||
|
- a work queue's consumer is refused.
|
||||||
|
|
||||||
|
**Not fixed here:** the client dropping messages while the loop acts on a long merge. Reports and
|
||||||
|
heartbeats are redelivered or replaced, so nothing is lost for good, but the loop holding everything while
|
||||||
|
it builds is the shape issue 175 already describes.
|
||||||
|
|
||||||
|
## 2026-10-05, after the restart
|
||||||
|
|
||||||
|
The consumer re-made from now held nothing, and the restarted controller acknowledged what it was handed.
|
||||||
|
Two merges made right after reached it within two seconds, and the fixed controller was built, delivered and
|
||||||
|
took over by its own plan within two minutes. The re-asked build of the agent module was registered and
|
||||||
|
reached all four machines.
|
||||||
|
|
||||||
|
**Not explained by this issue:** the record's merge was still logged twice, with the consumer fresh and
|
||||||
|
nothing replayed. The duplication has a cause of its own, still to be found. It may be that the forge
|
||||||
|
announces a merge on two paths, or that one event is handled twice.
|
||||||
|
|
||||||
|
## Resolved — 2026-10-05
|
||||||
|
|
||||||
|
The controller's event consumer is made from now when it is made, and `broker consumer-reset` re-makes a stuck one from now. Live: after the restart, merges reached the controller within seconds. Two tests main then failed, both skipped without a store, were fixed in mesh-controller PR #52. The merge heard twice was a separate cause: [issue 250](../250-a-merge-made-through-the-forges-tool-is-announced-twice/00-report.md).
|
||||||
+37
@@ -0,0 +1,37 @@
|
|||||||
|
---
|
||||||
|
status: resolved
|
||||||
|
opened: 2026-10-05
|
||||||
|
located-in: [mesh-controller cmd/mesh-controller]
|
||||||
|
fixed-by: mesh-controller PR #54
|
||||||
|
amended-design: 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
|
||||||
|
---
|
||||||
|
|
||||||
|
# 249 — A module's new state is refused until a push the merge did not make
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
A merge gave the agent module a new state, a key-value bucket
|
||||||
|
([ADR 0201](../../02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)).
|
||||||
|
The module's new bundle reached all four machines within a minute of the build, at once. The machines'
|
||||||
|
bus permissions did not include the new state until a push was made by hand afterwards.
|
||||||
|
|
||||||
|
On three machines the module's watch of the new state was refused for about two minutes: "claude-code keeps
|
||||||
|
and reads no state called config". It recovered only because the module asks again with a back-off, and
|
||||||
|
the hand-made pushes issued the permissions.
|
||||||
|
|
||||||
|
## Why it matters
|
||||||
|
|
||||||
|
**The code arrives before the right to use it.** A module that does not retry stays broken until someone
|
||||||
|
pushes. A module whose first act on start is to read its new state fails its start. Nothing in the plan
|
||||||
|
says the two must travel together.
|
||||||
|
|
||||||
|
**No machine went first.** The bundle reached every machine at the same moment. The rollout the operator
|
||||||
|
was told — one machine first, then the rest — could not be followed, because the merge had already
|
||||||
|
delivered it everywhere.
|
||||||
|
|
||||||
|
## Open questions
|
||||||
|
|
||||||
|
- Should a plan send a machine its membership, the grants that come with a module's new
|
||||||
|
declarations, in the same push as the bundle, and before it?
|
||||||
|
- Should a merge that changes a module's declarations (state, events, tools) be delivered to one machine
|
||||||
|
first, and to the rest only once that one reports it healthy?
|
||||||
+30
@@ -0,0 +1,30 @@
|
|||||||
|
# 249 — Diagnosis
|
||||||
|
|
||||||
|
## 2026-10-05
|
||||||
|
|
||||||
|
**Grants after code.** Both a plan's rollout and a push send every machine its declaration first, and
|
||||||
|
issue the memberships afterwards. The order is written into the code on purpose, "because the runtime it is
|
||||||
|
for arrives with it". That reason holds only for a first assignment, and a membership is kept on the bus for
|
||||||
|
a runtime that connects later anyway. A membership that failed was only printed, and left "until the next
|
||||||
|
push". The list of bus users travels in the declaration of the machine that holds the bus, which a module's
|
||||||
|
rollout reaches only if that machine runs the module.
|
||||||
|
|
||||||
|
**No machine first.** The module's upgrade policy sends one machine at a time, but a plan's rollout ignored
|
||||||
|
it and sent every machine at once. One at a time did not wait for the first machine to come up either: it
|
||||||
|
stopped only if the publish failed.
|
||||||
|
|
||||||
|
**Not answered by the open decision on unseen changes** (a removal, a move or a replacement, shown
|
||||||
|
before it takes effect). That decision leaves an add-only change alone on purpose, and a new state is one.
|
||||||
|
|
||||||
|
**Decided** in [ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md):
|
||||||
|
grants before code, and one machine first unless a module's policy says *together*.
|
||||||
|
|
||||||
|
## Resolved — 2026-10-05
|
||||||
|
|
||||||
|
[ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md), live the same evening:
|
||||||
|
|
||||||
|
- every push now issues memberships before it sends declarations;
|
||||||
|
- the bus's machine goes first when its list of users moved;
|
||||||
|
- a plan sent the build agent to one machine first and to the rest once that machine reported.
|
||||||
|
|
||||||
|
The first live rollout exposed a fault in the tier gate. The first machine's report, made between the two sends, was read as stale ([issue 256](../256-a-first-machines-report-read-as-stale-between-the-two-sends/00-report.md)).
|
||||||
@@ -0,0 +1,49 @@
|
|||||||
|
---
|
||||||
|
status: resolved
|
||||||
|
opened: 2026-10-05
|
||||||
|
located-in: [mesh-catalog modules/gitea]
|
||||||
|
fixed-by: mesh-catalog PR #63, PR #66
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 250 — A merge made through the forge's tool is announced twice
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
The controller logged the record repository's merges twice, seconds apart, with the same commit, even with
|
||||||
|
its event consumer freshly made ([issue 248](../248-the-controllers-event-consumer-replayed-a-week-and-held-every-merge-behind-it/00-report.md)).
|
||||||
|
Counted over one day:
|
||||||
|
|
||||||
|
- 8 of the record's merges were logged twice, against 4 once;
|
||||||
|
- 10 of the catalogue's, against 7 once;
|
||||||
|
- 2 of the controller's.
|
||||||
|
|
||||||
|
A code repository's second line reads differently — "it changed nothing any module the mesh holds is built
|
||||||
|
from" — so it was taken for a different message.
|
||||||
|
|
||||||
|
## Diagnosis
|
||||||
|
|
||||||
|
The forge's module announces a merge from two places in the same process:
|
||||||
|
|
||||||
|
1. its merge tool, the moment it merges;
|
||||||
|
2. the poll added for [issue 131](../131-nothing-tells-the-mesh-a-source-moved/00-report.md), which
|
||||||
|
announces every merged pull request it has not recorded as announced.
|
||||||
|
|
||||||
|
The tool never records what it announced, so the poll announces it again 0.5 to 16 seconds later.
|
||||||
|
Merges made in the forge's web interface or by a plain API call are seen by the poll alone, and those are
|
||||||
|
the ones logged once. The module's own header comments still say merges are announced "from the tools …
|
||||||
|
one process only".
|
||||||
|
|
||||||
|
No harm was done this time, but only by luck. The second event is absorbed because the first one moved
|
||||||
|
the controller's record of the source. A repository read only by packaging modules has no such record, so
|
||||||
|
it would get a second plan. The record module synced twice for each merge.
|
||||||
|
|
||||||
|
## Fix
|
||||||
|
|
||||||
|
The poll is the only emitter: it sees every path and carries the clone address. The tool merges and
|
||||||
|
answers the merge commit. A merge made through the tool is heard up to thirty seconds later, which the
|
||||||
|
module already accepts ("an event a minute late is still an event").
|
||||||
|
|
||||||
|
## Resolved — 2026-10-05
|
||||||
|
|
||||||
|
The forge module's poll is the only announcer of a merge. It asks only the repositories that moved since its last look, and one pass at a time. Live: three merges were each heard once, and a later merge was planned within a minute.
|
||||||
+36
@@ -0,0 +1,36 @@
|
|||||||
|
---
|
||||||
|
status: resolved
|
||||||
|
opened: 2026-10-05
|
||||||
|
located-in: [mesh-catalog modules/records]
|
||||||
|
fixed-by: mesh-catalog PR #64
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 251 — The record's checkout could not sync after it ran as another account
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
Every sync of the record module failed, on every merge and every timer: git refused the checkout as
|
||||||
|
"dubious ownership". The record tools answered all the while, from the checkout as it last stood, and
|
||||||
|
nothing said it was stale.
|
||||||
|
|
||||||
|
## Diagnosis
|
||||||
|
|
||||||
|
The module ran in a container, as the superuser, until its code moved into the machine's tool runtime
|
||||||
|
([ADR 0198](../../02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)).
|
||||||
|
The runtime launches it as the operator account. The checkout's git directory and 2 598 of its files still
|
||||||
|
belonged to the superuser. Git refuses a repository owned by another user, and the operator account could
|
||||||
|
change none of those files.
|
||||||
|
|
||||||
|
The module also still declared that it needs a container runtime, a leftover of the same move.
|
||||||
|
|
||||||
|
## Fix
|
||||||
|
|
||||||
|
The module, ported to Go as part of the fix, clones into a directory of its own inside the one it is given:
|
||||||
|
a directory it makes, and so owns. What the old layout left behind is removed where it is the module's.
|
||||||
|
Where it is not, the module names it in its status, with the one command that deletes it. The
|
||||||
|
container-runtime capability is dropped. A sync that fails is still said in the status, as before.
|
||||||
|
|
||||||
|
## Resolved — 2026-10-05
|
||||||
|
|
||||||
|
The record module, ported to Go, clones into a directory it makes and owns. Live: it synced to the newest commit, with 692 documents. It names seven leftovers of the old checkout that it cannot remove, with the command that removes them; that is the operator's act.
|
||||||
@@ -0,0 +1,38 @@
|
|||||||
|
---
|
||||||
|
status: resolved
|
||||||
|
opened: 2026-10-05
|
||||||
|
located-in: [mesh-catalog modules/gitea, mesh-controller cmd/mesh-controller]
|
||||||
|
fixed-by: mesh-catalog PR #63, mesh-controller PR #54
|
||||||
|
amended-design: 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
|
||||||
|
---
|
||||||
|
|
||||||
|
# 252 — A merge's changed modules were read wrong, in both directions
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
Each of three merges to the catalogue within four minutes planned 99 to 100 modules, and rebuilt modules
|
||||||
|
they did not touch. The outputs were identical, so nothing was redeployed. The rebuilds cost the build
|
||||||
|
machine minutes for every merge.
|
||||||
|
|
||||||
|
## Diagnosis
|
||||||
|
|
||||||
|
**Too many.** The controller treats a changed path outside every module directory it knows as shared code,
|
||||||
|
and rebuilds every module built from the repository. A new module's directory, or one being removed,
|
||||||
|
counts: the module is registered only after the merge is planned. Each of the three merges added or
|
||||||
|
removed a module. The build agent, built from the same repository, then joins the set and becomes tier 0.
|
||||||
|
|
||||||
|
**Too few.** The forge's module asked for a hundred changed files and was given fifty, the forge's page
|
||||||
|
size, and reported the list as whole. A merge of 59 files reached the controller with 50. Had the full
|
||||||
|
rebuild not hidden it, a module whose own files changed would have stayed unbuilt.
|
||||||
|
|
||||||
|
## Fix
|
||||||
|
|
||||||
|
- The forge's module reads every page of a pull request's files.
|
||||||
|
- The controller counts a changed path in a sibling of known module directories as that module's own, not
|
||||||
|
shared, when that module's definition is among the changed paths. A sibling without a definition, such
|
||||||
|
as a shared library, still means everything, which is the safe direction. Root files still mean
|
||||||
|
everything.
|
||||||
|
|
||||||
|
## Resolved — 2026-10-05
|
||||||
|
|
||||||
|
The forge's module reads every page of a pull request's files. The controller counts a new module's own directory as that module's when its definition is among the changed paths. Live: catalogue merges planned one module each.
|
||||||
+55
@@ -0,0 +1,55 @@
|
|||||||
|
---
|
||||||
|
status: located
|
||||||
|
opened: 2026-10-05
|
||||||
|
located-in: [mesh-controller internal/builder, mesh-controller internal/artifacts, mesh-catalog modules/distribution]
|
||||||
|
fixed-by: mesh-controller PR #53, mesh-host PR #23, mesh-catalog PR #62
|
||||||
|
amended-design: 03-DESIGN/01-to-be/18-building-a-module.md
|
||||||
|
---
|
||||||
|
|
||||||
|
# 253 — The store's collector would delete every archive the mesh keeps
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
The store's nightly collector ([ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md))
|
||||||
|
was installed the same day, its first run due that night. Measured read-only beforehand:
|
||||||
|
|
||||||
|
| | |
|
||||||
|
|---|---|
|
||||||
|
| blobs in the store | 8 186 |
|
||||||
|
| blobs a manifest names, kept by the collector | 2 090 |
|
||||||
|
| blobs it would delete | 6 096 |
|
||||||
|
| repositories holding only archives, none named by any manifest | 105 |
|
||||||
|
|
||||||
|
Among the archives it would delete were the current bundles of the agent module, the machine host, the
|
||||||
|
controller and the tool runtime. Each was named by no manifest, though the controller's records keep them.
|
||||||
|
|
||||||
|
## Why it matters
|
||||||
|
|
||||||
|
Machines keep their unpacked copies, so nothing would have stopped at once. But any fresh fetch of an
|
||||||
|
unchanged module would have failed: a machine joining, a reinstall, an apply that fetches again, the
|
||||||
|
controller's own next rollout.
|
||||||
|
|
||||||
|
## Diagnosis
|
||||||
|
|
||||||
|
Images are pushed with manifests. Archives were published as bare blobs, which no manifest names. The
|
||||||
|
store's stock collector marks only from manifests, so every bare blob is unmarked, kept or not. ADR 0189's
|
||||||
|
sentence "what the mesh keeps is still a manifest in the store" was true of images only.
|
||||||
|
|
||||||
|
Found alongside: the "five most recent builds" reason kept builds of modules the mesh no longer holds,
|
||||||
|
forever.
|
||||||
|
|
||||||
|
## Fix
|
||||||
|
|
||||||
|
- **That night, before the first run:** the collector was changed to a dry run, in its module's
|
||||||
|
definition, and delivered.
|
||||||
|
- **Then:**
|
||||||
|
- every archive is published with a manifest that holds it;
|
||||||
|
- the controller's sweep holds every kept archive before it lets anything go, which backfills those
|
||||||
|
already published;
|
||||||
|
- letting an archive go removes its manifest first;
|
||||||
|
- a forgotten module keeps nothing.
|
||||||
|
- Real collection returns once the controller reports no kept archive unheld.
|
||||||
|
|
||||||
|
## Where it stands — 2026-10-05
|
||||||
|
|
||||||
|
Every kept archive is held: the controller's collection command reports 134 of 134 held, none missing. The window the collector needs is no longer reopened by an apply ([issue 224](../224-an-apply-reopens-a-maintenance-window-by-recreating-what-it-held-still/00-report.md)). The collector still runs as a dry run. Turning it to real collection deletes the layers nothing keeps, which is the operator's word to give; this issue resolves when that change lands.
|
||||||
+36
@@ -0,0 +1,36 @@
|
|||||||
|
---
|
||||||
|
status: resolved
|
||||||
|
opened: 2026-10-05
|
||||||
|
located-in: [mesh-controller cmd/mesh-controller, mesh-controller internal/inventory]
|
||||||
|
fixed-by: mesh-controller PR #54
|
||||||
|
amended-design: 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
|
||||||
|
---
|
||||||
|
|
||||||
|
# 254 — Plans for successive merges run over each other, and one was left open
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
Three merges to the catalogue within four minutes made three plans, and all three ran at once:
|
||||||
|
|
||||||
|
- each sent the build agent to every machine;
|
||||||
|
- each asked for the same tier of builds — one module was asked for 32 seconds apart by two plans;
|
||||||
|
- the plan for the middle merge still showed "building" hours later, waiting on nothing.
|
||||||
|
|
||||||
|
## Diagnosis
|
||||||
|
|
||||||
|
A merge's plan is saved without looking at the open plans
|
||||||
|
([ADR 0162](../../02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md): one plan per merge).
|
||||||
|
Each open plan advances on its own. Nothing ends a plan whose work a newer merge has taken over, and
|
||||||
|
nothing lets a person close a plan that waits on nothing.
|
||||||
|
[Issue 219](../219-an-older-build-that-finishes-later-replaces-a-newer-one/00-report.md) settled only which
|
||||||
|
build's output wins.
|
||||||
|
|
||||||
|
## Fix
|
||||||
|
|
||||||
|
[ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md)
|
||||||
|
§3. A newer merge's plan supersedes the older open plans for the same repository and branch, and takes in
|
||||||
|
the modules they had not built. A person can close a plan by its id.
|
||||||
|
|
||||||
|
## Resolved — 2026-10-05
|
||||||
|
|
||||||
|
A newer merge's plan supersedes the older open plans of its repository and branch, and a person can close a plan by its id. Live: the plan left open since the afternoon was closed by hand, and its note said what it had built and never sent. A merge made while the controller's own plan was open took it over: the older plan reads "superseded at tier 0 by" the newer, which planned its three modules again.
|
||||||
@@ -0,0 +1,36 @@
|
|||||||
|
---
|
||||||
|
status: resolved
|
||||||
|
opened: 2026-10-05
|
||||||
|
located-in: [mesh-catalog modules/systemd]
|
||||||
|
fixed-by: mesh-catalog PR #65
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 255 — The journal verb read nothing for a system service
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
The service manager seat's `journal` verb answered "-- No entries --" for the controller's service on the
|
||||||
|
control node, while the service was logging steadily. Twice in one day a session reading the controller's
|
||||||
|
log fell back to a shell on the machine: the verb that exists for exactly that question answered nothing.
|
||||||
|
On one machine of four the same verb did answer: the only one whose operator account is in the journal's
|
||||||
|
group.
|
||||||
|
|
||||||
|
## Diagnosis
|
||||||
|
|
||||||
|
The module runs as the operator account and escalates the five acts on the system manager with `sudo -n`.
|
||||||
|
It ran `journalctl` unescalated. journalctl shows an account outside the journal's group only that
|
||||||
|
account's own entries, and says "-- No entries --" for everything else. That reads as a quiet service, not
|
||||||
|
as a refusal.
|
||||||
|
|
||||||
|
The mesh grants the operator account passwordless escalation on every machine through its own drop-in, as
|
||||||
|
the sudo module's check confirms.
|
||||||
|
|
||||||
|
## Fix
|
||||||
|
|
||||||
|
A read of the system journal escalates like an act does. The account's own journal, in the user scope,
|
||||||
|
does not. The module was ported to Go as part of the fix, its tests with it.
|
||||||
|
|
||||||
|
## Resolved — 2026-10-05
|
||||||
|
|
||||||
|
The journal verb reads a system unit's journal escalated, in the module ported to Go. Live: the controller's and the host's journals read through the verb on the control node.
|
||||||
+31
@@ -0,0 +1,31 @@
|
|||||||
|
---
|
||||||
|
status: resolved
|
||||||
|
opened: 2026-10-05
|
||||||
|
located-in: [mesh-controller cmd/mesh-controller]
|
||||||
|
fixed-by: mesh-controller PR #55
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 256 — A first machine's report read as stale between the two sends
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
The first plan to roll out under [ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md)
|
||||||
|
sent the build agent to one machine first, and to the other two once that machine reported it applied, 27
|
||||||
|
seconds later. All three applied it within seconds. The plan then said, for seven minutes, that it was
|
||||||
|
"waiting for build-agent" on the first machine to be applied. It went on only when that machine happened to
|
||||||
|
report again for another reason.
|
||||||
|
|
||||||
|
## Diagnosis
|
||||||
|
|
||||||
|
The tier gate asked every machine for a report made after the module's last send. With one machine first,
|
||||||
|
a module is sent twice, and the second send is the later one. The first machine's report came between the
|
||||||
|
two sends, so it read as older than the build. The plan waited for that machine's next report, one report
|
||||||
|
cycle, and never knew why.
|
||||||
|
|
||||||
|
## Fix
|
||||||
|
|
||||||
|
The gate judges each machine from its own send: the first machine from the first send, the rest from the
|
||||||
|
second. A plan's wait is also printed to the second rather than the minute, where it had read "0s", and the
|
||||||
|
first send is printed in the machine's own time rather than in UTC. Checked by the controller's test: a first
|
||||||
|
machine's report between the two sends opens the gate, and a report from before its send does not.
|
||||||
+49
@@ -0,0 +1,49 @@
|
|||||||
|
---
|
||||||
|
status: located
|
||||||
|
opened: 2026-10-05
|
||||||
|
located-in: [mesh-host]
|
||||||
|
fixed-by: novox/mesh-host#24 (89a7796)
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 257. A plan waited on a declaration its first machine never reported
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
A merge to the controller's repository produced a plan of three tiers, the controller first. The
|
||||||
|
controller was built and sent to its first machine, the anchor, which is also the machine holding the
|
||||||
|
bus. The plan then said, for five minutes and with no end in sight:
|
||||||
|
|
||||||
|
> tier 1 of 3, tier 0 built; waiting for mesh-controller on the anchor, sent first at 20:22 to report
|
||||||
|
> it applied before the rest are sent
|
||||||
|
|
||||||
|
The anchor had applied it. Its host's journal said `applied 470 resource(s)` ten seconds after the
|
||||||
|
send, the new controller was running, and the controller logged the report.
|
||||||
|
|
||||||
|
## Evidence
|
||||||
|
|
||||||
|
The plan waits for a report that names the declaration last sent (ADR 0218, one machine first; the
|
||||||
|
report's `declared` digest against the machine's recorded `sent` digest). Read from the store:
|
||||||
|
|
||||||
|
| | digest | at |
|
||||||
|
|---|---|---|
|
||||||
|
| recorded as sent to the anchor | `09f4d6…` | 20:22:57.699 |
|
||||||
|
| the anchor's report, `declared` | `39b48a…` | 20:23:07.718 |
|
||||||
|
|
||||||
|
The anchor applied and reported a declaration that is not the one the controller recorded sending.
|
||||||
|
So the report never counted, and the plan would have waited until its bound and then failed the
|
||||||
|
rollout at its first machine, for a machine that had applied correctly.
|
||||||
|
|
||||||
|
The other three machines' digests matched their reports at the same moment.
|
||||||
|
|
||||||
|
## What unblocked it
|
||||||
|
|
||||||
|
A push to the anchor, made for another reason (its bus grants), recorded a new send. The anchor's
|
||||||
|
next report named it, the digests matched, and the plan moved on at its next tick and finished.
|
||||||
|
|
||||||
|
## What it costs
|
||||||
|
|
||||||
|
A plan that rolls out the controller itself, to the machine that holds the bus, can stop at its first
|
||||||
|
machine with nothing wrong there. The plan's words are true but useless: they name a machine that
|
||||||
|
already did what was asked. Merges made close together, from several sessions, are the case the
|
||||||
|
plans exist for, and a controller change is among them.
|
||||||
+60
@@ -0,0 +1,60 @@
|
|||||||
|
# Diagnosis
|
||||||
|
|
||||||
|
## 2026-10-05
|
||||||
|
|
||||||
|
**What is known.** The digest the controller recorded for the anchor (`09f4d6…`, at 20:22:57.699) is
|
||||||
|
not the digest of what the anchor applied (`39b48a…`). The anchor's host logged one apply that
|
||||||
|
updated the controller, starting at 20:23:02 and reporting at 20:23:07. The controller that sent it
|
||||||
|
was replaced at 20:23:04 by that same apply. Its successor logged two reports from the anchor, both at
|
||||||
|
20:23:07.
|
||||||
|
|
||||||
|
**Ruled out.**
|
||||||
|
|
||||||
|
- The catalogue's catch-up: that is the catalogue asking the controller, not a machine asking for its
|
||||||
|
declaration.
|
||||||
|
- The host's five-minute check at 20:22:58: it applied nothing new (only "kept" lines). The apply
|
||||||
|
that moved the controller is the one at 20:23:02.
|
||||||
|
- A digest computed differently by the host and the controller: the other three machines matched,
|
||||||
|
and after the push at 20:28:00 the anchor matched too.
|
||||||
|
|
||||||
|
**Not yet confirmed: two sends to one machine in one step.** The anchor is both the plan's first
|
||||||
|
machine and the machine holding the bus. Issue 249 sends the machine holding the bus before the first
|
||||||
|
machine when its user list must change, and this controller change added bus grants. If the anchor
|
||||||
|
was sent twice in that step, each with its own digest, then the host applying the later one while the
|
||||||
|
earlier one's record landed last would give exactly this picture. The next step is to read
|
||||||
|
`sendToEach` for the order of its sends and of their `recordSent` calls, and to check the node's
|
||||||
|
sequence numbers for two sends at 20:22:57.
|
||||||
|
|
||||||
|
**Whatever the cause, the plan's wait has a second fault.** A first machine that reports `applied` for
|
||||||
|
a declaration the controller cannot place says nothing about the build. The plan should say "it
|
||||||
|
applied something other than what was recorded as sent" instead of "waiting", so a reader acts in
|
||||||
|
minutes rather than at the bound.
|
||||||
|
|
||||||
|
## 2026-10-05, later: found
|
||||||
|
|
||||||
|
**The lead above was wrong.** It was not two sends. The machine's own five-minute reconcile did it.
|
||||||
|
|
||||||
|
The host re-applies what it was last told every five minutes. That reconcile read the kept
|
||||||
|
declaration, then waited for any apply in progress. A declaration the link is applying is kept only
|
||||||
|
once its apply ends. On the anchor:
|
||||||
|
|
||||||
|
- the reconcile was due at 20:22:58;
|
||||||
|
- the controller's send arrived at 20:22:57;
|
||||||
|
- the reconcile read the declaration kept before that send, waited, and applied it after the link's
|
||||||
|
apply;
|
||||||
|
- its report named that older declaration.
|
||||||
|
|
||||||
|
That explains every fact above:
|
||||||
|
|
||||||
|
- the report's digest matched nothing recorded as sent;
|
||||||
|
- the other machines, not due, matched;
|
||||||
|
- the next push matched.
|
||||||
|
|
||||||
|
**Seen a second time, with a visible effect** (issue 261): a module assigned on the laptop was
|
||||||
|
applied and then given back 29 seconds later by the reconcile due during that apply.
|
||||||
|
|
||||||
|
**Fixed** in mesh-host: the reconcile reads the kept declaration once it holds the apply lock. A test
|
||||||
|
reproduces the interleaving: it fails with the old order and passes with the fix.
|
||||||
|
|
||||||
|
The plan's "waiting" wording, raised above as a second fault, is unchanged. With the cause gone, a
|
||||||
|
report naming an unrecorded declaration should no longer occur.
|
||||||
@@ -0,0 +1,53 @@
|
|||||||
|
---
|
||||||
|
status: resolved
|
||||||
|
opened: 2026-10-05
|
||||||
|
located-in: [mesh-controller]
|
||||||
|
fixed-by: novox/mesh-controller#59 (4b382ce)
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 258. Every machine bound the mesh's resolver to itself
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
After the mesh moved to one resolver (ADR 0194, 0196), the seat `mesh-dns-resolver` was recorded as
|
||||||
|
held on the anchor, and each other machine was pinned to it for `wildcard-resolution`. Each machine
|
||||||
|
was then pushed. The anchor's resolver configuration named its own private address first. Every
|
||||||
|
other machine's did too: each named **its own** private address, not the anchor's.
|
||||||
|
|
||||||
|
Nothing broke, because each machine still ran a resolver of its own while the move was under way. But
|
||||||
|
the mesh had decided on one resolver, recorded who held it and pinned every machine to it, and no
|
||||||
|
machine used it.
|
||||||
|
|
||||||
|
## Cause
|
||||||
|
|
||||||
|
The controller binds a requirement in one of two branches:
|
||||||
|
|
||||||
|
- **Answered here**, when a module on the same machine provides it.
|
||||||
|
- **Answered elsewhere**, when one on another machine does.
|
||||||
|
|
||||||
|
Only the second branch read the seat's holder (ADR 0110) and a pin. The first took the local provider,
|
||||||
|
and read a pin only to choose between two local ones. Every machine still had its own resolver
|
||||||
|
assigned, so every machine took the first branch.
|
||||||
|
|
||||||
|
The seat's own definition says the requirement "resolves to the holder wherever it is placed". The
|
||||||
|
code did not.
|
||||||
|
|
||||||
|
## Fix
|
||||||
|
|
||||||
|
For a mesh-wide provision, a pin naming another machine, or the seat's holder on another machine,
|
||||||
|
wins over a provider on this one. With neither, the local provider answers as before.
|
||||||
|
|
||||||
|
This is checked by three resolve tests in mesh-controller:
|
||||||
|
|
||||||
|
- the holder elsewhere answers;
|
||||||
|
- a pin elsewhere answers;
|
||||||
|
- a holder here still answers here.
|
||||||
|
|
||||||
|
The tests fail without the fix.
|
||||||
|
|
||||||
|
## Verified
|
||||||
|
|
||||||
|
2026-10-05, after the fix rolled out: all four machines' resolver configuration names the anchor's
|
||||||
|
resolver first and the public one second. Each resolves the mesh's machine names, a wildcard name
|
||||||
|
under a machine, and a public name.
|
||||||
@@ -0,0 +1,41 @@
|
|||||||
|
---
|
||||||
|
status: open
|
||||||
|
opened: 2026-10-05
|
||||||
|
located-in: [mesh-controller]
|
||||||
|
fixed-by:
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 259. A push to one machine sent every machine
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
A change to the resolver modules was merged with the upgrade policy `record`, so that each machine
|
||||||
|
would get it only when pushed. The plan was to go one machine at a time: the anchor first, then each
|
||||||
|
of the others, each checked before the next.
|
||||||
|
|
||||||
|
`push <anchor>` sent all four machines, each a new declaration carrying the change. A later
|
||||||
|
`push <laptop>` did the same. A further fault in the change (issue 260) was therefore met on every
|
||||||
|
machine at once, not on one.
|
||||||
|
|
||||||
|
## Cause
|
||||||
|
|
||||||
|
A named push ends by flushing every other machine whose declaration differs from what it was last
|
||||||
|
sent (issue 057, ADR 0083). This was meant for consequences of the push, such as a provider's grant
|
||||||
|
list after a consumer was assigned. It cannot tell a consequence from a change the upgrade policy is
|
||||||
|
holding back. Under `record`, every machine running the module differs, so every machine is flushed.
|
||||||
|
|
||||||
|
Issue 249 met the same confusion for the machine holding the bus, and narrowed that check to the user
|
||||||
|
list alone. The flush has no such narrowing.
|
||||||
|
|
||||||
|
## What it costs
|
||||||
|
|
||||||
|
`record` is the policy for a change that must be walked through the mesh by hand. A named push cannot
|
||||||
|
do that, so the policy does not hold the change back from anything but the merge. The command's name
|
||||||
|
says one machine, and the command's output is the only place that says otherwise.
|
||||||
|
|
||||||
|
## Open
|
||||||
|
|
||||||
|
- Should the flush send a machine only what changed as a consequence: grants, user lists and bound
|
||||||
|
facts, not module versions held by a policy?
|
||||||
|
- Or should it list such machines as behind and leave them, as `push --behind` would find them?
|
||||||
@@ -0,0 +1,40 @@
|
|||||||
|
---
|
||||||
|
status: open
|
||||||
|
opened: 2026-10-05
|
||||||
|
located-in: [mesh-host, mesh-catalog]
|
||||||
|
fixed-by:
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 260. The resolver was restarted before the file it reads existed
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
The first apply of the resolver's new configuration failed on every machine. Each host's journal:
|
||||||
|
|
||||||
|
> failed dnsmasq.service: restarting dnsmasq.service: starting it again: systemctl exited 1
|
||||||
|
|
||||||
|
The resolver's own journal:
|
||||||
|
|
||||||
|
> dnsmasq: cannot read /etc/mesh-resolver/zones.conf: No such file or directory
|
||||||
|
|
||||||
|
The service manager's automatic restart started it again in the same second, and it answered. The
|
||||||
|
host still reported the apply as failed, on every machine.
|
||||||
|
|
||||||
|
## Cause
|
||||||
|
|
||||||
|
The new configuration names a second file, the mesh's zones. The host wrote the configuration and
|
||||||
|
restarted the service on it **before** it created the zones file. The zones file was created next in
|
||||||
|
the same apply. The service's `restart-on` names both files, but nothing orders the restart after
|
||||||
|
every file it reads has been written.
|
||||||
|
|
||||||
|
## What it costs
|
||||||
|
|
||||||
|
A configuration that adds a file it reads fails its first start everywhere. Here the service manager's
|
||||||
|
restart policy hid it within a second. A service without one would have stayed down, holding the
|
||||||
|
machine's name resolution with it.
|
||||||
|
|
||||||
|
## Open
|
||||||
|
|
||||||
|
- Should the host run every restart and reload after all files of the apply are written?
|
||||||
|
- Or should it order each restart after every resource the service names in `restart-on`?
|
||||||
@@ -0,0 +1,40 @@
|
|||||||
|
---
|
||||||
|
status: located
|
||||||
|
opened: 2026-10-05
|
||||||
|
located-in: [mesh-host]
|
||||||
|
fixed-by: novox/mesh-host#24 (89a7796)
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 261. A module was applied and given back half a minute later
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
The `hosts` module was assigned to the laptop and pushed. The laptop's host applied it: the module's
|
||||||
|
block was written at the start of `/etc/hosts`, its tools were installed, and the host reported the
|
||||||
|
declaration it had been sent. Twenty-nine seconds later the same host gave the block and the tools
|
||||||
|
back. Nothing on the controller had sent a declaration without the module. Five minutes later the
|
||||||
|
module was back.
|
||||||
|
|
||||||
|
The host's journal on the laptop, in order:
|
||||||
|
|
||||||
|
- `updated hosts.own (/etc/hosts): the mesh's region added at the start`
|
||||||
|
- `applied 342 resource(s)`
|
||||||
|
- `restored hosts.own (/etc/hosts)`
|
||||||
|
- `removed hosts.bundle-tools`
|
||||||
|
|
||||||
|
## Cause
|
||||||
|
|
||||||
|
The host's five-minute reconcile was due during that apply. It read the declaration kept before the
|
||||||
|
push, waited for the push's apply to finish, and then applied the older declaration over it. That
|
||||||
|
older declaration did not have the module, so the module was given back. The next reconcile read the
|
||||||
|
newer declaration, which had been kept by then, and applied the module again.
|
||||||
|
|
||||||
|
This is the same fault as [issue 257](../257-a-plan-waited-on-a-declaration-its-first-machine-never-reported/00-report.md),
|
||||||
|
where it showed only as a report naming a declaration nobody had recorded sending.
|
||||||
|
|
||||||
|
## Fix
|
||||||
|
|
||||||
|
The reconcile reads the kept declaration once it holds the apply lock, so it always applies the
|
||||||
|
latest declaration the mesh sent. This is checked by mesh-host's test
|
||||||
|
`TestAReconcileAppliesWhatWasKeptWhenItsTurnComes`, which fails with the old order.
|
||||||
Reference in New Issue
Block a user