Compare commits
69
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
213dcbafe9 | ||
|
|
0f163a3e7c | ||
|
|
09fa192cac | ||
|
|
3e55277aa9 | ||
|
|
8fade781a7 | ||
|
|
f6f5563de2 | ||
|
|
97ff655e34 | ||
|
|
202f2aa144 | ||
|
|
d77399055c | ||
|
|
3ba75a312b | ||
|
|
64c2d3a0a1 | ||
|
|
14dabf60ee | ||
|
|
1d8745b54d | ||
|
|
f9e0517e0f | ||
|
|
88f7f79fbb | ||
|
|
3ee27b870d | ||
|
|
9bcd714ee6 | ||
|
|
3a8c23ef1a | ||
|
|
fd4df96c21 | ||
|
|
1e02593adf | ||
|
|
aa5a735994 | ||
|
|
6db919f789 | ||
|
|
9b7fac3084 | ||
|
|
947965e750 | ||
|
|
00a0d75ced | ||
|
|
178994ee0d | ||
|
|
4f1300fd48 | ||
|
|
6e59281806 | ||
|
|
dae33af8b0 | ||
|
|
48e4ac9aa7 | ||
|
|
eb7356360f | ||
|
|
6719ab029c | ||
|
|
93bed2147f | ||
|
|
95ce92c62f | ||
|
|
072489e653 | ||
|
|
eb4b4256a1 | ||
|
|
553d0c8893 | ||
|
|
08643a2128 | ||
|
|
4fac135a46 | ||
|
|
f391c5c36c | ||
|
|
e6fbb86373 | ||
|
|
c8af8d5eb1 | ||
|
|
0e193ebc41 | ||
|
|
a0de734ed6 | ||
|
|
55e776a1cd | ||
|
|
c233a1bbac | ||
|
|
035d4a27e3 | ||
|
|
77429488b0 | ||
|
|
3944e4f914 | ||
|
|
ab1fbe3202 | ||
|
|
0fdcb15200 | ||
|
|
91aa481269 | ||
|
|
0627b7f1f3 | ||
|
|
049b50b66c | ||
|
|
44808ac563 | ||
|
|
f000288147 | ||
|
|
80aff0f457 | ||
|
|
400ed8d396 | ||
|
|
df2072aed5 | ||
|
|
ae58570209 | ||
|
|
855c3f9b33 | ||
|
|
f93b02900c | ||
|
|
9306192f93 | ||
|
|
8bd56a2820 | ||
|
|
7497149eed | ||
|
|
1b9196b807 | ||
|
|
b3e7191ddc | ||
|
|
87f80573b8 | ||
|
|
9197b99153 |
+10
-2
@@ -7,7 +7,7 @@ another — and a mesh you cannot name precisely is a mesh two people describe d
|
||||
|
||||
## The mesh and its machines
|
||||
|
||||
- **node** — a machine in the mesh. There are 0..n of them, and each runs the host agent. A node is
|
||||
- **node** — a machine in the mesh. There are 0..n of them, and each runs the node-engine. A node is
|
||||
just a machine that has joined; being one implies nothing about what it runs.
|
||||
- **operator account** — the login name of the person who works on a node, stated on the node
|
||||
record; empty for a machine nobody logs into. Everything the mesh places under a person's home is
|
||||
@@ -28,6 +28,13 @@ another — and a mesh you cannot name precisely is a mesh two people describe d
|
||||
control-plane/data-plane, and opaque here).
|
||||
- **mesh-controller** — the module that runs the controller. It **claims** the `mesh-controller`
|
||||
seat at mesh scope, which is what makes it singular. Replaces the module name **`mesh-control`**.
|
||||
- **node-engine** — the program on every node that applies what the controller declares: it receives the
|
||||
node's declaration, writes the files, runs the services and containers, and reports what it did. It
|
||||
is the engine, not a module: it owns no file's content, and every file it writes belongs to the module
|
||||
that declared it. Replaces **"host agent"**, **"the host"** and **`mesh-host`**. *Agent* is avoided
|
||||
because the word already means two other things here, the build agent and the operator's coding
|
||||
agent. The code still carries the old name (the `mesh-host` repository, its binary and service
|
||||
unit) until the rename is made there; a record written before the rename keeps the old name.
|
||||
(The git repository has been renamed `mesh-control` -> `mesh-controller` on the forge; the module,
|
||||
container and image it produces are `mesh-controller`.)
|
||||
- **foundation** — the store and the broker, raised at genesis before any module system exists.
|
||||
@@ -73,7 +80,8 @@ another — and a mesh you cannot name precisely is a mesh two people describe d
|
||||
The set, with who holds each seat, is the overview of what a mesh has
|
||||
([26 — The seats](../03-DESIGN/01-to-be/26-the-seats.md)). A seat has a **capacity**: a
|
||||
capacity-1 seat is exclusive (one holder); a higher-capacity seat is a **bench** (several holders
|
||||
coexist).
|
||||
coexist). The first bench is `mesh-dns-resolver`, a *replicated* mesh seat: one holder per machine,
|
||||
each on record and each answering the same names ([ADR 0223](../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)).
|
||||
- **claim** — a module taking a spot on a seat. `claims: [{name, scope}]` in a manifest. A
|
||||
mesh-scoped exclusive claim is how the mesh says "there is one of me". A foundation seat is
|
||||
named after the server it guards: the `mesh-controller`, `postgres` and `lavinmq` modules claim
|
||||
|
||||
@@ -0,0 +1,66 @@
|
||||
---
|
||||
status: graduated
|
||||
initiated: 2026-10-06
|
||||
touches:
|
||||
- 00-META/how-we-build.md
|
||||
- 01-RESEARCH/017-a-mesh-that-heals-itself/00-overview.md
|
||||
- 01-RESEARCH/028-the-meshs-output-channel/00-overview.md
|
||||
- 01-RESEARCH/019-a-warm-twin-of-the-running-mesh/00-overview.md
|
||||
- 02-DECISIONS/0141-the-host-delivers-its-own-successor.md
|
||||
- 02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md
|
||||
- 02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md
|
||||
- 02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md
|
||||
- 03-DESIGN/01-to-be/06-the-controller.md
|
||||
- 03-DESIGN/01-to-be/09-the-node-lifecycle.md
|
||||
- 03-DESIGN/01-to-be/25-the-bus-on-nats.md
|
||||
- 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
|
||||
- 04-ISSUES/187-the-mesh-tells-nobody-when-it-stops-working/00-report.md
|
||||
became:
|
||||
- 02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md
|
||||
- 03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md
|
||||
---
|
||||
|
||||
# 031 — A core that cannot fail silently
|
||||
|
||||
**What.** The principles the mesh's core must hold — the controller, the machine host, the bus, the
|
||||
console and the path a change takes through them — so that it is fully diagnosable, monitors itself,
|
||||
heals what it knows how to heal, and upgrades itself without a person standing by. And the mechanisms
|
||||
and the order in which to build them.
|
||||
|
||||
**Why.** The operator, 2026-10-06: *"I still notice a lot of race issues, and commands being ignored, or
|
||||
no feedback, no logs, no monitoring. Our mesh core setup must be fully diagnosable, with active
|
||||
monitoring, self-healing, self-upgradeable, self-monitoring. The core principles must be very sturdy, no
|
||||
ambiguities, clear plan of execution, fail-proof setup."*
|
||||
|
||||
The record bears it out. In the six days to 2026-10-06, 92 issue reports were opened. Of the 48 read here
|
||||
as core failures, **every one was noticed because a person or an agent looked**, and **none was raised by
|
||||
the mesh unasked**. Four faults came back through a different door after their first fix, because each
|
||||
fix closed an instance and left its class open. One merge was skipped by the bus, and twenty-three over three
|
||||
days have no matching action; a provider failed for twenty-three hours with only its own
|
||||
journal saying so; seven databases were dropped on one unreadable file.
|
||||
|
||||
**What it touches.** The controller (its verbs, `status`, plans, a lease), the host (its apply and
|
||||
report), the bus (its advisories and its upgrade), the console, the build path, the output channel of
|
||||
research 028, and the self-healing intent of research 017, which this effort extends from the loops that
|
||||
converge modules to the core that runs those loops. 017 deferred heartbeats, conditions and advisories
|
||||
until the bus was NATS; it is now.
|
||||
|
||||
**Documents.**
|
||||
|
||||
- [01 — The evidence](01-evidence.md): 48 issues classified by class of failure (races, dropped
|
||||
commands, no feedback, logs only, two writers, manual repair, self-upgrade, CI-versus-live, third
|
||||
party), with time to detect and how each was noticed.
|
||||
- [02 — Principles](02-principles.md): nine, each with what exists, what is missing and **how it is
|
||||
checked**; and the candidates weighed and not kept.
|
||||
- [03 — Mechanisms and roadmap](03-mechanisms-and-roadmap.md): conditions, watchdogs from a signals table,
|
||||
a self-check (`doctor`), the output channel, healers with a hand-act log, a controller lease and report
|
||||
sequences, staged core upgrades with rollback, a facts snapshot for merge checks, a lab replay of every
|
||||
incident; five phases, ordered by risk removed; the five largest risks today.
|
||||
|
||||
**Graduated 2026-10-06.** The operator approved the conclusion the same day (*"do the research and
|
||||
implement it"*). The nine principles and the plan became
|
||||
[ADR 0227](../../02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md);
|
||||
the mechanisms, the tables and the six phases became
|
||||
[to-be 45](../../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md). The two measurements owed
|
||||
— the signals table's bounds and a week of the hand-act log — are Phase 0's work there, not
|
||||
preconditions of the decision.
|
||||
@@ -0,0 +1,202 @@
|
||||
# 01 — The evidence, classified by class of failure
|
||||
|
||||
Every issue report opened between 2026-09-30 and 2026-10-06 that bears on the mesh's core — the
|
||||
controller, the machine host, the bus, the console, the build path — read in full and classified by
|
||||
**the class of failure**, not by the component it was found in. A component view says "fix the host";
|
||||
a class view says "the same thing is wrong in four places", which is what a principle is for.
|
||||
|
||||
## The count
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| issue reports opened 2026-10-01 to 2026-10-06 | **92** (about fifteen a day) |
|
||||
| of those (and a few from the days before), classified below as core failures | **48** distinct issues |
|
||||
| classified in more than one class | 17 of 48 |
|
||||
| a fault that **came back** after a fix of the same symptom | 4 chains: 200 → 265, 230 → 264, 257 → 261 → 267, 175 → 184 → 248 |
|
||||
| noticed because a person or an agent looked — at a stalled plan, a wrong outcome, a journal, a test run by hand, a review | **48 of 48** |
|
||||
| of those, the mesh's own answer carried the fact for whoever asked (a refusal, a `status` line, a push's output) | 4 (233, 244, 259, 263) |
|
||||
| raised by the mesh to anyone, unasked | **0 of 48** |
|
||||
|
||||
The recurrences matter most. Each fix was correct for its instance and left the class standing, so the
|
||||
same symptom came back through a different door days later. That is the measurement behind the
|
||||
operator's mandate: point fixes are converging on the instances, not on the class.
|
||||
|
||||
## The classes
|
||||
|
||||
Nine classes, as the mandate frames them. The table under each is the evidence; *detected* is the time
|
||||
from the fault's start to the moment anyone knew; *noticed by* is how.
|
||||
|
||||
### (a) Races: concurrent actors without an ordering
|
||||
|
||||
Two actors act on the same thing, and the order of arrival — not an explicit order — decides the outcome.
|
||||
|
||||
| Issue | The two actors | Detected | Noticed by |
|
||||
|---|---|---|---|
|
||||
| 204 | an outgoing and an incoming controller both sent declarations | 2 min | a person saw a module undone |
|
||||
| 201 | a plan's push carried a controller digest older than its successor had written | 10 min crash loop | a person, nothing answered |
|
||||
| 214 | the controller rebuilding itself; the outcome reached the old one or neither | 27 min | a person asked the plan twice |
|
||||
| 219 | an older build finishing later replaced a newer one | hours | a person reading builds |
|
||||
| 234 | seven declarations arrived during an eleven-minute apply; the newest was composed from a stale view and undeclared four modules | 8 min of removals | a person, the modules were gone |
|
||||
| 254 | three plans for three merges ran at once, each asking the same builds | hours | a person, one plan stuck "building" |
|
||||
| 256 | a first machine's report landed between a module's two sends and read as stale | 7 min | a person |
|
||||
| 257, 261 | the machine's five-minute reconcile and a delivery took the apply lock in the wrong order | 5 min (257), 29 s visible undo (261) | a person |
|
||||
| 267 | a reconcile's report queued behind a delivery's apply overtook it at the controller | until a hand push | a person |
|
||||
| 265 | a push reloaded the bus's permissions while its own answer was still owed | 54 of 103 pushes over two days | a person, "did not answer in time" |
|
||||
|
||||
**What they share.** Every one is a receiver that kept *the last thing written* rather than *the
|
||||
newest thing by an explicit order*. Declarations carry a sequence (issue 107); **reports do not**
|
||||
(267 says so: "the report does not carry the declaration's sequence number, so the digest decides").
|
||||
Plans are ordered by when they were made only since ADR 0218. Builds are ordered since issue 219.
|
||||
Controllers have no epoch, so two instances can both act (204). Ordering was added one message kind at
|
||||
a time, each after a race in it was seen.
|
||||
|
||||
### (b) Commands silently ignored, arguments dropped
|
||||
|
||||
| Issue | What was dropped | Effect |
|
||||
|---|---|---|
|
||||
| 244 | the console removed `node` from every mesh-seat verb's schema and call | `plan` could only refuse; **`push <one machine>` arrived empty and pushed every machine** |
|
||||
| 259 | a named push's flush sent every machine a build a policy held back | a fault met on every machine at once, not one |
|
||||
| 202 | a module whose setting was unset was *left out* of the machine | the resolver vanished from a declaration, no error |
|
||||
| 188 | a refusal inside "who is on the network" dropped a machine | 40 min, every symptom pointed elsewhere |
|
||||
| 231 | a misspelled placeholder written to a file as literal text | passes every check |
|
||||
| 241 | an unreadable contributions file read as "nobody asks" | **seven databases dropped and recreated empty** |
|
||||
| 255 | the journal verb read nothing and said "-- No entries --" | a refusal that reads as a quiet service |
|
||||
| 246 | a runtime that answered late was treated as absent | the console said modules "run nowhere" |
|
||||
|
||||
**What they share.** A receiver that could not tell *nothing was asked* from *something was lost on the
|
||||
way*, and chose a default. In 241 and 244 the default was the most destructive reading available.
|
||||
|
||||
### (c) Outcomes not fed back to the caller
|
||||
|
||||
| Issue | What the caller was told | What happened |
|
||||
|---|---|---|
|
||||
| 200, 265 | "did not answer in time" | the push ran; the answer was refused by the bus |
|
||||
| 176 | the console's build tool neither waits nor registers | — |
|
||||
| 229 | `plans` answers once in prose; nothing waits for a plan | an agent went round the mesh with `curl` |
|
||||
| 230, 264 | a host stood aside for its successor and its report was cancelled | the plan waited for ever, reading `late: false` |
|
||||
| 186 | the build machine dropped 26 of 43 asks; nothing counts asks against outcomes | inferred two hours later |
|
||||
| 237 | `assign` answered "held" and "does not resolve for lack of it" in one breath | a person or agent would loop |
|
||||
|
||||
Since 2026-10-06 the controller answers within ten seconds and keeps every call's outcome under an id
|
||||
(`calls`, issue 265). Read live the same night: **that log holds the last hundred calls in the
|
||||
controller's memory**, so a controller restart — which every merge to the controller's own repository
|
||||
causes — forgets every outcome it held. And `status`, a read-only verb, took **18 seconds** to answer,
|
||||
twice in a row, so even the health question is answered only through the "still running, ask `calls`"
|
||||
path.
|
||||
|
||||
### (d) Failures visible only as log lines
|
||||
|
||||
| Issue | Where it was said | For how long |
|
||||
|---|---|---|
|
||||
| 179 (recurred) | the identity provider's journal, every five seconds | **23 hours**, about 31 000 refused logins |
|
||||
| 184 | the controller's log: slow consumer, heartbeats dropped | 24 min deaf |
|
||||
| 187 | five faults in one day, each found by reading a container's log hours later | hours each |
|
||||
| 183, 217, 265 | a `Permissions Violation` line from the bus client library | days |
|
||||
| 233 | `status` said `refused`, correctly; nothing said it had lasted | 1.5 h |
|
||||
| 243 | nothing: machines silently ignored lower licence generations | until a login waited three minutes |
|
||||
| 248 | the controller's event loop stopped logging at 15:17 | hours; "status showed every plan done" |
|
||||
| 266 | nothing: a merge was skipped by the bus | **23 unmatched merges over three days** |
|
||||
| 238 | a ban list of 400 entries | the operator's own address banned for four weeks |
|
||||
|
||||
`status` printed its all-well sentence through 179, 248 and 266. ADR 0224 made the first of those break
|
||||
it. The other two have no signal that `status` reads.
|
||||
|
||||
### (e) State that two writers own
|
||||
|
||||
| Issue | The two writers |
|
||||
|---|---|
|
||||
| 190, 222 | the runtime's configuration written by modules that are not the runtime, and by the controller |
|
||||
| 201 | the controller seat's row written by a successor, read by a predecessor pushed back in |
|
||||
| 239 | two definitions, in two repositories, held one module name |
|
||||
| 245 | "behind" answered by a commit comparison beside the plan that already knows |
|
||||
| 257, 261, 267 | the machine's state written by both the delivery and the five-minute reconcile |
|
||||
| 250 | a merge announced by the forge's tool and by its poll |
|
||||
| 179 | the identity provider's admin password: the mesh minted one, the database kept another |
|
||||
|
||||
The operator's direction on 245 is the principle in their own words: *"a second answer to the same
|
||||
question is how the two came to disagree."*
|
||||
|
||||
### (f) Manual repair needed
|
||||
|
||||
Counted from the reports' own "what unblocked it" sections:
|
||||
|
||||
| Repair by hand | Issues | Times |
|
||||
|---|---|---|
|
||||
| a push by hand to unstick a plan waiting on a report | 230, 257, 264, 267 | at least 4 |
|
||||
| a controller restart to recreate a missing object or let go of a held message | 208, 248 | 2 |
|
||||
| a one-off program run as the controller, outside the service | 201, 248 | 2 |
|
||||
| the identity provider's admin reset through its bootstrap command | 179 | 2 |
|
||||
| a kept file restored on a machine by hand | 233 | 1 |
|
||||
| a stuck plan closed by hand | 214, 254 | 2 |
|
||||
| a ban lifted by hand | 238 | 1 |
|
||||
| a consumer remade from now | 248 | 1 |
|
||||
|
||||
ADR 0224 (*detected automatically, repaired where safe, loud where not*) is the first rule that turns
|
||||
one of these into a mechanism. 248's `broker consumer-reset` and 254's `plans close` turned two into
|
||||
verbs a person runs. Every other row is still a hand on a machine.
|
||||
|
||||
### (g) Self-upgrade fragility
|
||||
|
||||
The core updates itself: the controller rebuilds and replaces itself, the host delivers its own
|
||||
successor (ADR 0141), the runtime and the console are modules, and the bus is a module on the control
|
||||
node.
|
||||
|
||||
| Issue | What the self-upgrade broke |
|
||||
|---|---|
|
||||
| 201 | the controller pushed back to an older build than its own successor's row |
|
||||
| 204 | two controllers both sending during a handover |
|
||||
| 213, 223 | the controller ran as a container, and a new mesh installed it so |
|
||||
| 214 | the plan that rebuilds the controller lost track of it |
|
||||
| 230, 264 | the host that stands aside loses the report of the apply that delivered it (fixed twice) |
|
||||
| 245 | a rebuild of everything replaced the bus's container: **every runtime lost the bus for a minute** |
|
||||
| 248, 266 | a controller restart is where merges go missing: most of 266's 23 lie in such windows |
|
||||
| 217 | a refused announcement crash-looped nine runtimes, closing the path that would merge the fix |
|
||||
|
||||
A machine's first-in-line rollout (ADR 0218) protects modules. It does not protect the core from itself:
|
||||
the health a plan waits for is "reported applied", which a controller that cannot plan, a host that
|
||||
cannot report or a bus that drops messages can each satisfy. Nothing rolls back. The recovery in 201
|
||||
was the mesh's own binary run by hand from the newer image.
|
||||
|
||||
### (h) Checks that pass in CI and fail live
|
||||
|
||||
| Issue | The environmental fact the check did not have |
|
||||
|---|---|
|
||||
| 177 | the store-backed tests are skipped by the quick check, and the mesh runs none of a module's tests |
|
||||
| 236 | the host's declaration validation is not run by the catalogue check |
|
||||
| 262 | musl takes an NXDOMAIN for IPv6 as final; glibc does not |
|
||||
| 263 | the real machine names make a consumer's identity 23–26 characters against a bound of 20 |
|
||||
| 202 | the controller's test against the real catalogue, run by nobody until that day |
|
||||
| 228 | the host's removal has no case for a `user` — found by a review, not a test |
|
||||
|
||||
Each check was right about the world it was given. None of them was given the mesh's world: its machine
|
||||
names, its catalogue, its host's validation, its C libraries.
|
||||
|
||||
### (i) Third-party bugs
|
||||
|
||||
| Issue | |
|
||||
|---|---|
|
||||
| 266 | the bus server's 2.10 release skips messages on a consumer with several filter subjects |
|
||||
| 265 | a reload of the bus's authorization forgets every reply permission already granted (documented server behaviour, read from its source) |
|
||||
| 262 | a resolver answering NXDOMAIN where NODATA is correct, met by musl's stricter reading |
|
||||
|
||||
The lesson of 266 is not "upgrade the bus" — it is that nothing compared *what was announced* with
|
||||
*what was acted on*, so a dependency's bug was silent for three days. A defence in depth (watch the
|
||||
outcome, not the transport) would have caught it whichever layer was wrong.
|
||||
|
||||
## What would have prevented or caught each class
|
||||
|
||||
| Class | Would have been prevented by | Would have been caught by |
|
||||
|---|---|---|
|
||||
| (a) races | every message ordered by its writer, stale refused by every receiver | a lab replay of the interleaving |
|
||||
| (b) dropped | refusing an unknown or unreadable input by name | a schema walk over every verb |
|
||||
| (c) no feedback | answer at once with an id; outcome kept durably | a watchdog on "asked and never finished" |
|
||||
| (d) logs only | — | a condition in `status` and a notification |
|
||||
| (e) two writers | one writer per piece of state | a registry of writers checked at composition |
|
||||
| (f) manual repair | a healer for every repair done twice | a counter of hand acts |
|
||||
| (g) self-upgrade | one machine first, health-gated, rolled back | a lab upgrade with a deliberately broken build |
|
||||
| (h) CI vs live | checks fed the real mesh's facts | the same, before merge |
|
||||
| (i) third party | pinning and testing the version that runs | an end-to-end count of announced vs acted |
|
||||
|
||||
The two columns are the principles of [02](02-principles.md). Read by count, **the "caught by" column
|
||||
is the cheapest and widest**: a watchdog and a condition would have shortened most of the 48 from "a
|
||||
person noticed" to minutes, whatever the cause.
|
||||
@@ -0,0 +1,247 @@
|
||||
# 02 — Principles for the core, each with how it is checked
|
||||
|
||||
Nine principles. Each is stated as a rule a reviewer can refuse a change against, carries the classes of
|
||||
[01](01-evidence.md) it answers, says what exists already, and says **how it is checked** — the
|
||||
repository's own rule ([how-we-build §5](../../00-META/how-we-build.md)): a rule that states no check is
|
||||
indistinguishable from a wrong one.
|
||||
|
||||
They extend, not replace, the six of [research 017](../017-a-mesh-that-heals-itself/01-the-intended-behaviour.md)
|
||||
(a loop compares with what is; healing is the ordinary path again; a repair never destroys; nothing fails
|
||||
silently; what cannot be fixed goes to an agent; correctness, not only liveness). Those are about the
|
||||
loops that converge modules. These are about **the core that runs those loops**: the controller, the
|
||||
host, the bus, the console and the path a change takes through them. 017 deferred heartbeats, conditions
|
||||
and advisories until the bus was NATS. It is now, so that deferral has expired.
|
||||
|
||||
**The core**, for this document: the controller, the machine host and its launcher, the bus server, the
|
||||
tool runtime and the console, the build seat, and the forge's announcer of merges. Everything a change
|
||||
passes through before a module's own code runs.
|
||||
|
||||
---
|
||||
|
||||
## P1 — One writer per piece of state
|
||||
|
||||
Every piece of state the mesh keeps has exactly one writer, named. Anyone else who wants it changed asks
|
||||
that writer; nobody writes beside it, and nobody computes a second answer to a question it already
|
||||
answers.
|
||||
|
||||
- **Answers:** (e), most of (a). Issues 190, 201, 204, 222, 239, 245, 250, 257/261/267.
|
||||
- **Exists:** the controller is the only writer of stream definitions (to-be 25); ADR 0222 §3 (the
|
||||
controller writes no file a seat's holder owns); the collision check at composition (no two modules
|
||||
declare one path, unit, name or package); the operator's direction on 245.
|
||||
- **Missing:** a single writer for a *machine's applied state* (the delivery and the five-minute
|
||||
reconcile both apply and both report — three issues in two days); a single *controller* (two instances
|
||||
can both act during a handover, 204: nothing holds a lease); a single announcer per event kind (250
|
||||
was found by counting duplicates by hand).
|
||||
- **How it is checked:**
|
||||
1. A **writers table** — state kind, its writer, where it is kept — is part of the to-be design, and a
|
||||
test in each core repository asserts that the code paths that write each kind are the one named
|
||||
(by a lint over the store's write calls and the bus subjects each component publishes on, the
|
||||
latter read from the grants the controller composes: a subject two components may publish on is
|
||||
refused at composition unless the table says it is shared).
|
||||
2. **Live:** the controller holds a **lease** (a key-value entry with a revision) and every message it
|
||||
sends carries the lease's epoch; a host refuses a declaration from an older epoch and says so. A
|
||||
probe ([03](03-mechanisms-and-roadmap.md), the self-check) asserts one lease holder and no message
|
||||
from a stale epoch in the last interval.
|
||||
|
||||
## P2 — Everything that changes state carries its writer's order, and every receiver refuses what is older
|
||||
|
||||
Declarations, reports, plans, builds, calls and announcements each carry `(writer, epoch, sequence)`.
|
||||
Every receiver keeps the highest it has accepted per writer, and **refuses** — with a line in the mesh's
|
||||
own words and a counter — anything older. Arrival order never decides.
|
||||
|
||||
- **Answers:** (a). Issues 201, 204, 214, 219, 234, 256, 257, 261, 264, 267.
|
||||
- **Exists:** a declaration carries a sequence and a host keeps the newest (issue 107, to-be 25 §3); a
|
||||
newer build wins over an older one finishing later (219); a newer merge's plan supersedes an older
|
||||
one (ADR 0218 §3); a report about a declaration the mesh has moved past no longer replaces the stored
|
||||
account (267).
|
||||
- **Missing:** a **report carries no sequence** — the digest last recorded as sent decides, which is why
|
||||
each of 256, 257 and 267 needed its own rule. No epoch on the controller. A plan's state carries no
|
||||
revision, so two instances can both advance it (214).
|
||||
- **How it is checked:**
|
||||
1. A contract test per message kind, in the receiver's repository: deliver `n`, then `n−1`; the state
|
||||
names `n` and a refusal is counted. Deliver from epoch `e−1` after `e`: refused. A new message kind
|
||||
without such a test fails a check that lists every subject the component consumes against the
|
||||
tests that name it.
|
||||
2. **Live:** the refusals counter is a signal (P5): zero is normal, a burst is a condition naming the
|
||||
writer that sent stale.
|
||||
|
||||
## P3 — Every command is acknowledged at once, and its outcome is kept where it can be read later
|
||||
|
||||
A call is answered within a bound the caller can rely on — with its result, or with an id. Its outcome is
|
||||
kept **durably**, outlives the process that ran it, and can be read by id or waited on. Nothing is fired
|
||||
and forgotten, and no answer depends on what the command does to the transport carrying it.
|
||||
|
||||
- **Answers:** (c). Issues 176, 186, 200, 229, 230, 237, 264, 265.
|
||||
- **Exists:** since 265, a seat's call answers within ten seconds or says "running" with an id, `push`
|
||||
answers before it acts, and `calls` keeps the last hundred calls and their answers; the build path
|
||||
says "asked, not waited for" and ADR 0219 §3 makes every cancel/kill leave an outcome.
|
||||
- **Missing:** `calls` lives in the controller's memory, so **a controller restart forgets every
|
||||
outcome** — and the controller restarts on every merge to its own repository. Nothing waits on a plan
|
||||
(229). Observed the night this effort began: the read-only `status` took 18 s, so the health question
|
||||
itself is answered through the "still running" path.
|
||||
- **How it is checked:**
|
||||
1. A test that walks every verb the controller announces: each answers within `AnswerWithin`, with a
|
||||
result or an id (already partly built for 265).
|
||||
2. A test that restarts the controller between a call and the read of its outcome: the outcome is
|
||||
still there.
|
||||
3. **Live:** a probe calls `status` and asserts it answers *in full* within the bound; a call
|
||||
`running` for longer than its verb's declared bound is a condition (P5).
|
||||
|
||||
## P4 — Nothing is dropped silently: an input that is unknown, unreadable or unmet is refused by name
|
||||
|
||||
A receiver that cannot read, place or honour an input refuses it and says which, where and why. It never
|
||||
substitutes a default — above all never "empty" — for "I could not tell". An unknown argument, key,
|
||||
placeholder, seat or subject is refused, naming it.
|
||||
|
||||
- **Answers:** (b). Issues 188, 202, 231, 241, 244, 246, 255, 259.
|
||||
- **Exists:** the host's parser refuses unknown keys (ADR 0007, 0045); an undeclared setting or endpoint
|
||||
is refused (ADR 0164, 0138); a verb takes only its declared arguments (ADR 0154) and, since 244, the
|
||||
controller and the console refuse an undeclared one by name; a seat the mesh does not answer is
|
||||
refused by name (ADR 0222 §1); ADR 0219 §3, "nothing dropped is silent".
|
||||
- **Missing:** the rule is in a dozen records and in no principle, so each new reader re-meets it. The
|
||||
destructive form — **an unreadable input read as "nothing asked", then acted on** (241) — has no
|
||||
general guard.
|
||||
- **How it is checked:**
|
||||
1. Per component, a test that feeds each input reader an unreadable, malformed and foreign input and
|
||||
asserts a refusal, never an empty result. A reader whose error path returns an empty value fails a
|
||||
lint that the core repositories run (the shape is mechanical: an error branch that returns the
|
||||
zero value of a collection).
|
||||
2. Schema walks for every verb (244's tests) and every placeholder namespace (231).
|
||||
3. **Destructive deltas are braked**: a reconcile that would withdraw more than a bound of what it
|
||||
holds (one consumer, one module, a fraction set per provider) stops and raises a condition instead
|
||||
(P7). Checked by a test that empties the input and asserts nothing is withdrawn.
|
||||
|
||||
## P5 — Every expected signal has a watchdog: absence is itself a condition
|
||||
|
||||
Every signal the core expects on a cadence or after an act — a heartbeat, a report after a send, a
|
||||
plan's progress, a build's outcome after its ask, the controller's event loop taking something in, a
|
||||
provider's standing, an announcement turning into an action — has a declared bound. Silence past the
|
||||
bound is raised as a condition, naming what was expected, from whom, since when.
|
||||
|
||||
- **Answers:** (c), (d), (i). Issues 179, 184, 186, 187, 230, 243, 248, 257, 264, 266, 267.
|
||||
- **Exists:** a machine's "last heard from — out of touch N m" in `node show`; ADR 0090's stuck machine
|
||||
(three identical reports); ADR 0224's provider standing, shown with its silence after thirty minutes;
|
||||
266's catch-up of merges not acted on after ten minutes (the first true *announced-versus-acted*
|
||||
watchdog); ADR 0162's bound on a plan, which 230 found never fires (`late: false` for ever).
|
||||
- **Missing:** a list of the signals, their bounds and their owners; the plan bound working; the event
|
||||
loop's last-taken age (187); asks counted against outcomes (186); the bus's own advisories (slow
|
||||
consumer, maximum deliveries, permission violations) read as observations instead of log lines.
|
||||
- **How it is checked:**
|
||||
1. A **signals table** — signal, emitter, cadence or trigger, bound, condition raised — kept in the
|
||||
to-be design, and compiled into the controller. A test generated from it suppresses each signal in
|
||||
turn and asserts the named condition is raised within its bound and cleared when the signal
|
||||
returns.
|
||||
2. **Live:** the self-check (P6) reports, for every row, the age of the newest signal, so a row that
|
||||
never fires is itself visible.
|
||||
|
||||
## P6 — The mesh checks itself continuously, against live facts, and says what it found outward
|
||||
|
||||
The invariants the design states are probed **against the running mesh** on a schedule, not only in unit
|
||||
tests. A violation is a **condition** — durable, with since-when, evidence and who can resolve it —
|
||||
shown in `status` and **sent to the operator** through the output channel. The checker's own heartbeat is
|
||||
watched from somewhere it does not run.
|
||||
|
||||
- **Answers:** (d), (h). Issues 177, 187, 238, 245, 253, 262.
|
||||
- **Exists:** `status` itself (assembled from reports, as how-we-build §5 requires); 017's *condition*
|
||||
shape; 028's output seat, researched only; individual live checks done by hand (262: "the live check
|
||||
stays by hand, on each resolver"); 253's measurement before the collector's first run — the one time
|
||||
in the window a check ran before the damage.
|
||||
- **Missing:** a scheduled runner, a condition store, a channel out, and a watcher's watcher.
|
||||
- **How it is checked:**
|
||||
1. Every invariant in the to-be design that names a live check is a **probe** in the self-check's
|
||||
registry; a check over the design documents counts invariants with a stated live probe against
|
||||
those without, and the number may only go down.
|
||||
2. The self-check publishes a heartbeat; a second machine's watcher raises "the self-check is silent"
|
||||
through a channel that does not depend on the control node.
|
||||
3. **Live, once:** a deliberately broken invariant on a lab mesh appears in `status` and as a
|
||||
notification within one probe interval.
|
||||
|
||||
## P7 — A known failure heals itself, under a brake, and every repair is said
|
||||
|
||||
A failure that has been repaired by hand twice is a failure the mesh must repair itself: by running the
|
||||
ordinary path again (017 P2), never by destroying (017 P3), with a budget and a back-off, and with one
|
||||
line and one event saying what it did and why. When the budget is spent, or the only repair destroys,
|
||||
it is a condition and a notification, not a retry.
|
||||
|
||||
- **Answers:** (f). Issues 179, 208, 214, 230, 233, 248, 254, 257, 264, 267.
|
||||
- **Exists:** ADR 0224 §5, *detected automatically, repaired where safe, loud where not*, applied to the
|
||||
identity provider's admin; the host's launcher rolls back once to known-good and then halts (ADR
|
||||
0141); the provisioner's `holds` re-provisions what a backend lost (017/02); `broker consumer-reset`
|
||||
and `plans close` as verbs.
|
||||
- **Missing:** the general mechanism. The most frequent hand act in the window — **a push by hand to
|
||||
unstick a plan waiting on a report** — has no healer: the controller could ask the machine to report
|
||||
again (the machine knows what it applied) before waiting longer.
|
||||
- **How it is checked:**
|
||||
1. Every hand act on the core is done through a verb that records it (who, what, why) — the **hand-act
|
||||
log**. Its count per week is a reported number; an act recorded twice for the same cause is a
|
||||
condition asking for a healer.
|
||||
2. Each healer ships with a test that induces its failure, asserts the repair and the event, and
|
||||
asserts the brake after the budget.
|
||||
|
||||
## P8 — The core upgrades itself one machine at a time, health-gated, and rolls back on its own
|
||||
|
||||
A new controller, host, runtime or bus reaches one machine first; it is **healthy** only when the
|
||||
self-check's probes for that component pass there (not merely when it "reported applied"); the rest
|
||||
follow only then. A component that does not become healthy within its bound is rolled back to the last
|
||||
known good **by something other than itself**, and the rollback is said. The component being replaced is
|
||||
never the only witness of its successor's success.
|
||||
|
||||
- **Answers:** (g). Issues 201, 204, 213, 214, 217, 230, 245, 248, 264, 266.
|
||||
- **Exists:** ADR 0218's one machine first, for modules, with "applied and current" as the gate; ADR
|
||||
0141/0142's side-by-side host versions, known-good and the launcher's single rollback; ADR 0185's
|
||||
controller serving what it can when it is behind its seat's row; 264's report kept on disk across the
|
||||
hand-over.
|
||||
- **Missing:** a health definition per core component; a gate stronger than "reported"; rollback for
|
||||
the controller, the runtime and the bus; a lease hand-over between controllers (P1); a planned,
|
||||
rehearsed path for the bus, which is still one process on one machine and whose next upgrade (266:
|
||||
2.10 → 2.11) is one-way.
|
||||
- **How it is checked:**
|
||||
1. **Lab:** a deliberately broken build of each core component (one that starts and does nothing; one
|
||||
that crashes; one that cannot reach the bus) is merged on a lab mesh. Each is rolled back without a
|
||||
hand, the mesh ends on the previous build, and a condition and a notification say so.
|
||||
2. **Live:** every core rollout leaves a record — first machine, health verdict, time to verdict,
|
||||
rolled back or not — readable through `plans`.
|
||||
|
||||
## P9 — A check is fed the real mesh's facts before a change is merged
|
||||
|
||||
A check whose verdict depends on the environment — names and their lengths, the machines that exist,
|
||||
the catalogue as it is, the host's validation, the C library, the server versions — runs against **the
|
||||
mesh's real facts**, exported and anonymised, before merge. A dependency's version that the mesh runs is
|
||||
the version its tests run.
|
||||
|
||||
- **Answers:** (h), (i). Issues 177, 202, 228, 236, 262, 263, 266.
|
||||
- **Exists:** 266's `TestTheImageIsTheServerTestedHere` (the bus image's release equals the tested
|
||||
server's); ADR 0223's composition test that renders the resolver's machine list; 202's test against the real
|
||||
catalogue (run by hand).
|
||||
- **Missing:** the export of facts; a merge gate that composes every real machine's declaration with the
|
||||
change and runs the host's validation over it (which would have refused 236, 263 and 202 in their own
|
||||
pull requests); a resolver test under musl as well as glibc.
|
||||
- **How it is checked:**
|
||||
1. The controller exports a **facts snapshot** (machines, names, assignments, seats, catalogue
|
||||
commit — no secrets, no addresses) daily; the core repositories' merge check composes every machine
|
||||
from it with the change applied and runs the host's validation; a pull request that makes any
|
||||
machine fail to compose or validate fails its check, naming the machine's role and the module.
|
||||
2. The snapshot's age is a signal (P5).
|
||||
|
||||
---
|
||||
|
||||
## Candidates weighed and not kept as principles
|
||||
|
||||
- **"Every invariant has a live probe, not only a unit test"** — merged into P6; it is how P6 is built.
|
||||
- **"Environment-dependent checks run against the real mesh's facts"** — kept as P9; the third-party
|
||||
case (i) folded into it, because pinning and testing the version that runs is the same act.
|
||||
- **"Self-healing is the default"** — kept as P7 but narrowed to *known* failures, those repaired by
|
||||
hand twice. A default of healing everything heals what is not understood, which is how a repair
|
||||
destroys (241's reconcile was, in its own terms, healing).
|
||||
- **"No loop blocks on long work"** (184, 248, 175) — not a separate principle: a blocked loop is a
|
||||
signal gone silent (P5, the loop's last-taken age) and a design defect each owner fixes; stating it
|
||||
as a principle adds a rule with no mesh-wide check.
|
||||
- **"Clear plan of execution"** from the mandate — not a principle about the mesh; it is the roadmap in
|
||||
[03](03-mechanisms-and-roadmap.md), and each phase there states its own verification.
|
||||
|
||||
## How the principles relate
|
||||
|
||||
P2 and P1 **prevent** the races. P4 **prevents** the silent drops. P3, P5 and P6 **catch** whatever the
|
||||
first three miss, at the cost of minutes, not hours. P7 and P8 **repair**. P9 **moves** the catching
|
||||
before merge. The order of the roadmap follows from that: catching first, because it is cheapest and
|
||||
covers every class, including the ones nobody has met yet.
|
||||
@@ -0,0 +1,221 @@
|
||||
# 03 — Mechanisms and a phased roadmap
|
||||
|
||||
The principles of [02](02-principles.md) need few new things. Most of the parts exist in some form;
|
||||
what is missing is the connective tissue that makes a fact the mesh already has reach someone without
|
||||
being asked. This document names the mechanisms, then orders them into phases **by risk removed per unit
|
||||
of effort**, each phase with deliverables and a verification that says it is done.
|
||||
|
||||
Effort is given in **focused working days** of one agent-and-operator pair, at the pace the record shows
|
||||
(a located issue to a merged fix in under a day is common). The figures are for ordering, not promises.
|
||||
|
||||
## The mechanisms
|
||||
|
||||
### M1 — Conditions (P5, P6)
|
||||
|
||||
017's *condition*, built: a durable fact about something the mesh owns — what is wrong, since when, the
|
||||
evidence, what was tried, who can resolve it, and whether it may clear itself. Kept in a key-value
|
||||
bucket the controller writes (one writer, P1), keyed by subject (`plan/<id>`, `machine/<role>`,
|
||||
`provider/<module>/<consumer>`, `core/<component>`). Raised and cleared by observation only; a person
|
||||
can **silence** one for a stated time, never resolve it. `status` becomes, first, the list of open
|
||||
conditions; the all-well sentence is "no open conditions". ADR 0224's provider standing is the first
|
||||
condition kind and moves into it unchanged.
|
||||
|
||||
### M2 — Watchdogs from a signals table (P5)
|
||||
|
||||
One table, compiled into the controller, of every signal the core expects. Its first rows, each from an
|
||||
issue in [01](01-evidence.md):
|
||||
|
||||
| Signal | Bound (to be measured, then set) | Condition raised | Issue |
|
||||
|---|---|---|---|
|
||||
| machine heartbeat | 3 × interval | machine silent (asleep is a declared state, ADR 0211) | 187 |
|
||||
| report after a send | the machine's last apply duration × 3, at least 2 min | sent, not reported — then **ask the machine to report again** (M5) | 230, 257, 264, 267 |
|
||||
| plan tier progress | per tier, from build and apply durations | plan stalled at tier N, waiting on X | 214, 230, 254 |
|
||||
| controller event loop took something | 2 min while the stream has pending | controller deaf | 184, 248 |
|
||||
| merge announced → plan made or "nothing reads it" | 10 min (exists, 266) | merge never acted on | 248, 266 |
|
||||
| build asked → outcome | build's own declared timeout | ask lost | 186 |
|
||||
| call `running` → finished | the verb's declared bound | call hung | 265 |
|
||||
| provider standing repeated | 30 min (exists, ADR 0224) | provider silent | 179 |
|
||||
| bus advisories: slow consumer, maximum deliveries, permission violation | any | bus refused or dropped X for Y | 183, 187, 217, 265 |
|
||||
| self-check heartbeat | 2 × its interval, watched from a second machine | the watcher is silent | — |
|
||||
| facts snapshot age | 2 days | merge checks run on stale facts | 263 |
|
||||
|
||||
The bus advisories are the cheapest row: the server already publishes them on its system subjects, and
|
||||
the controller only has to subscribe (read-only) and translate each into the mesh's words, naming the
|
||||
call or the consumer, as 265 now does for a refused reply.
|
||||
|
||||
### M3 — The self-check: `doctor` (P6)
|
||||
|
||||
A controller verb, `doctor`, and the same code run every few minutes by the controller itself. Each run
|
||||
executes the **probe registry** — the live form of the design's invariants — and raises or clears
|
||||
conditions. First probes, each an invariant that a person checked by hand in the window:
|
||||
|
||||
- every machine's declaration composes, and every host would accept it (236, 263);
|
||||
- every holder of the mesh's resolver answers a machine name for IPv4 and NODATA for IPv6 (262);
|
||||
- every seat on record has a live holder that answers (208, 218);
|
||||
- every kept archive is held by a manifest (253 — the controller's collection command already reports it);
|
||||
- exactly one controller holds the lease (P1);
|
||||
- every durable consumer's position is near its stream's head (248's replay, 266's skip);
|
||||
- no address the mesh owns is in a ban list (238);
|
||||
- `status` answers in full within its bound (P3).
|
||||
|
||||
`doctor` with no argument answers the last run's verdict at once (P3) and, with `run`, runs now under an
|
||||
id. Its own heartbeat is a signal (M2), watched from a second machine.
|
||||
|
||||
### M4 — The output channel (P6)
|
||||
|
||||
[Research 028](../028-the-meshs-output-channel/00-overview.md)'s seat, built minimally first: **one**
|
||||
channel the operator chose (028 records it), plus the desktop notifier where the operator is. A condition
|
||||
is sent when raised, once more if it lasts past a bound, and when it clears. Deduplicated by the
|
||||
condition's key. The watcher's watcher (028's open question) is the second-machine watchdog of M3,
|
||||
sending through a channel that does not pass through the control node.
|
||||
|
||||
### M5 — Healers (P7)
|
||||
|
||||
A healer is a registered response to one condition kind: its repair (the ordinary path again), its
|
||||
budget, its brake, and the event it emits. First healers, all from hand acts in [01](01-evidence.md) §(f):
|
||||
|
||||
| Condition | Repair | Brake |
|
||||
|---|---|---|
|
||||
| sent, not reported | ask the machine to report what it last applied (it keeps it since 264); if that names another declaration, send again | twice, then condition |
|
||||
| plan stalled on a superseded or finished wait | close the plan with its note (`plans close`, done by the mesh) | once |
|
||||
| seat holder without its worker | raise the seat's objects again (208) | once per holder |
|
||||
| consumer far behind on a history stream | `broker consumer-reset` (248) — **only** for the consumers the table marks resettable | once, then condition |
|
||||
| provider admin refuses the mesh's secret | ADR 0224 §5 (exists) | exists |
|
||||
|
||||
And the **hand-act log**: every repair a person makes on the core goes through a verb that records who,
|
||||
what and why. Its weekly count is the measure of P7.
|
||||
|
||||
### M6 — Order and epochs (P1, P2)
|
||||
|
||||
- The controller takes a **lease** in a key-value bucket before it acts and renews it; its revision is
|
||||
the **epoch** every declaration and plan write carries. A starting controller waits for the lease;
|
||||
the outgoing one stops sending when it loses it. That closes 204 and makes 201/214 detectable.
|
||||
- A **report carries the sequence** of the declaration it is about; the controller keeps the highest
|
||||
per machine and refuses older accounts by sequence, not by a digest lookup (256, 257, 267 become one
|
||||
rule).
|
||||
- **The host has one apply queue.** A delivery and the reconcile are two reasons to enqueue the same
|
||||
act; the queue applies the newest declaration once, and makes one report (257/261/267 become
|
||||
impossible rather than handled).
|
||||
- Every receiver's refusal of something stale is a counted line (P2's live check).
|
||||
|
||||
### M7 — Staged, reversible core upgrades (P8)
|
||||
|
||||
- **A health definition per core component**, written as probes in M3's registry: the controller
|
||||
answers `status` in bound and holds the lease; the host has reported its current declaration; the
|
||||
runtime has announced and answers a PING; the bus has every stream and every durable consumer and
|
||||
passes a request/reply round trip.
|
||||
- **The gate:** ADR 0218's first machine is judged by those probes, not by "applied".
|
||||
- **Rollback by a witness that is not the new build:** the host's launcher for the host (exists); for
|
||||
the controller, the previous controller's process kept installed beside it, re-started by the host
|
||||
when the new one does not take the lease in bound; for the runtime, the host's known-good the same
|
||||
way. Each rollback is a condition, so it is said.
|
||||
- **The bus is planned, not rolled.** A bus upgrade is a declared maintenance step: streams snapshotted,
|
||||
the step announced as a condition while it runs, every consumer's position checked after (M3). The
|
||||
one-way 2.10 → 2.11 upgrade of 266 is the first. Whether the bus should become a cluster of three so
|
||||
that it can be upgraded live is a question for its own effort.
|
||||
|
||||
### M8 — The facts snapshot and the merge gate (P9)
|
||||
|
||||
The controller exports a facts snapshot (machines by role and name length, assignments, seats, catalogue
|
||||
commit) to a place the build seat reads. The core repositories' and the catalogue's merge checks compose
|
||||
every machine with the change and run the host's validation. The resolver module's tests run under musl
|
||||
and glibc. A bus, store or library version the mesh runs is the one its tests run (266's pattern,
|
||||
generalised).
|
||||
|
||||
### M9 — The lab replay (all)
|
||||
|
||||
[Research 019](../019-a-warm-twin-of-the-running-mesh/00-overview.md)'s warm twin, or a throwaway lab of
|
||||
three containers, running **scripted replays of each incident** in the window: a reconcile due during a
|
||||
push; a host self-update during its report; a bus reload during a call; a consumer with several filters
|
||||
under mixed traffic; a missing consumer made with the server's default; an unreadable contributions
|
||||
file; a controller rebuilding itself mid-plan; two controllers at once. Each replay asserts the
|
||||
principle's outcome (refused stale, condition raised, healed, rolled back). They run on every merge to a
|
||||
core repository.
|
||||
|
||||
---
|
||||
|
||||
## The roadmap
|
||||
|
||||
Ordered by **risk removed per day**. Detection comes first because it covers every class at once,
|
||||
including the ones not met yet; prevention second; repair and staging third; the pre-merge and lab work
|
||||
last because they are larger and pay off over months.
|
||||
|
||||
### Phase 0 — Finish what is in flight (2–3 days)
|
||||
|
||||
- Land and roll out the located fixes: 244, 264, 265, 266 (including the bus's planned 2.11 upgrade, done
|
||||
as M7's first planned bus step), 267, 257/261.
|
||||
- Make `calls` durable (a key-value bucket, bounded by count and age), and bring `status` inside its own
|
||||
answer bound.
|
||||
- Start the **hand-act log** now, before anything else, so every later phase is measured against a
|
||||
baseline.
|
||||
- **Done when:** the four located core issues resolve with their live checks; a controller restart
|
||||
keeps `calls`; `status` answers in full within ten seconds.
|
||||
|
||||
### Phase 1 — The mesh says when it is wrong (5–8 days)
|
||||
|
||||
- M1 conditions (ADR 0224's standing moved into them); M2 watchdogs for the first eight rows; the bus
|
||||
advisories subscribed and translated; M3 `doctor` with the first probes; M4 with one channel and the
|
||||
second-machine watcher.
|
||||
- **Done when:** on a lab mesh, suppressing each signal in the table raises its condition within its
|
||||
bound and sends a notification; clearing it clears both. Live: a week of conditions read back, every
|
||||
one either real or a bound corrected.
|
||||
- **Risk removed:** every class in [01](01-evidence.md) moves from "noticed by a person, hours later" to
|
||||
"said by the mesh, minutes later".
|
||||
|
||||
### Phase 2 — Order and one writer (5–7 days)
|
||||
|
||||
- M6: the controller's lease and epoch; a report's sequence; the host's single apply queue; stale
|
||||
refusals counted.
|
||||
- The writers table and the signals table written into a to-be design, with their compile-time checks.
|
||||
- P4's lint for empty-on-error readers in the core repositories, and the destructive-delta brake in the
|
||||
provisioner harness.
|
||||
- **Done when:** the lab replays of 204, 257/261/267 and 241 end with "refused stale", "one report" and
|
||||
"withdrawal braked" respectively; the contract test for every consumed subject exists.
|
||||
- **Risk removed:** class (a), the largest by count, and the destructive half of (b).
|
||||
|
||||
### Phase 3 — Healers (3–5 days)
|
||||
|
||||
- M5's first healers and the rule that a repair done by hand twice asks for one (from the hand-act log).
|
||||
- **Done when:** a lab mesh recovers from each induced failure in the healers table with no hand, says so,
|
||||
and brakes after its budget. Live: a week in which the hand-act log has no repeat.
|
||||
|
||||
### Phase 4 — Core upgrades that roll back (8–12 days)
|
||||
|
||||
- M7: health definitions; the gate; rollback for the controller and the runtime; the bus as a planned
|
||||
step.
|
||||
- **Done when:** on a lab mesh, a deliberately broken build of the controller, the host and the runtime
|
||||
is each rolled back without a hand, the mesh ends on the previous build, and the rollback is a
|
||||
condition and a notification. Live: the next three core rollouts each record a health verdict.
|
||||
- **Risk removed:** class (g) — the failures that take the control path itself down.
|
||||
|
||||
### Phase 5 — Checks before merge, and the replay suite (8–12 days, then ongoing)
|
||||
|
||||
- M8: the facts snapshot and the compose-and-validate merge gate; the libc matrix; versions tested as
|
||||
run.
|
||||
- M9: every incident in the window as a scripted replay, run on every core merge; each new core issue
|
||||
adds its replay as its "how it is checked".
|
||||
- **Done when:** the replays of 236, 262, 263 and 266 fail on the commit before their fix and pass after;
|
||||
a new core issue cannot resolve without a replay or a stated reason why none is possible.
|
||||
|
||||
**Total, roughly six to eight weeks of focused work**, with the first visible change — the mesh saying
|
||||
when it is wrong — inside the first two.
|
||||
|
||||
## The five largest risks today
|
||||
|
||||
Ranked by likelihood × damage, from the window's evidence and the mesh as read the night this effort
|
||||
began:
|
||||
|
||||
1. **A stall nobody is told about.** A plan, a merge or a report that stops is noticed only by someone
|
||||
looking (230, 248, 257, 264, 266, 267; issue 187 still open). Every release goes through a plan.
|
||||
2. **A destructive act on a misread input.** 241 dropped seven databases on one unreadable file; 234
|
||||
undeclared four modules from a stale composition; 245's advice rebuilt everything and cut the bus;
|
||||
253's collector would have deleted every archive. There is no general brake on a large withdrawal.
|
||||
3. **The core replacing itself with nothing to roll it back.** A controller build that starts but cannot
|
||||
plan (201, 214) halts every later change, because the controller is what plans the fix. Only the
|
||||
host has a launcher rollback.
|
||||
4. **The bus as a single, un-upgradeable process.** It runs a release that skips messages (266), its
|
||||
reload drops owed replies (265), replacing it cuts every machine off (245), and its next upgrade is
|
||||
one-way and needs a restart.
|
||||
5. **Concurrent actors with no lease or report order.** Several agent sessions, plans and reconciles act
|
||||
at once; on the night this began, four named pushes, one per machine, started within two seconds. The digests
|
||||
catch most of it now, one rule per message kind; the next message kind will not have one.
|
||||
@@ -108,6 +108,14 @@ declared slug is a strictly better escape hatch than an opaque hash. **D stays r
|
||||
|
||||
## Consequences (of E)
|
||||
|
||||
> **The mechanism changed — 2026-10-06, by [ADR 0225](0225-a-consumers-identity-is-bounded-by-the-provision-it-requires.md).**
|
||||
> Option C, named above as the later refinement, is taken: each offer states the longest identity its
|
||||
> backend keeps, and a consumer is bounded by the provision it requires rather than by 20 everywhere.
|
||||
> 20 stays the bound of the object store and of a provider that is told its consumers and does not
|
||||
> say. The overflow is refused before merge by the catalogue check, and at composition the consumer
|
||||
> is left out of its provider's grants and reported — never the provider's machine refused. What
|
||||
> stands: the identity is said once, the slug is the remedy, nothing is hashed or truncated.
|
||||
|
||||
- A module manifest gains an optional `slug`; a node may carry one too. `ConsumerIdentity` prefers
|
||||
the slug over the cleaned name for each half. `identityLimit` becomes 20 (the true minimum), and
|
||||
`CheckIdentity` refuses at `module add` / assignment — now with a message naming the slug to set.
|
||||
|
||||
@@ -51,6 +51,14 @@ the builder) — cost with no new property.
|
||||
restarts the runtime when that file changes — the `/etc/hosts` pattern for the content, the
|
||||
nftables pattern for the reload. No module author is involved; being on the network is what
|
||||
grants the trust, because being on the network is what the trust *is*.
|
||||
|
||||
> **The mechanism changed — 2026-10-05, by [ADR 0222](0222-a-module-is-told-where-a-mesh-seats-holder-is-reached-and-the-controller-writes-no-file-a-seats-holder-owns.md).**
|
||||
> What stands: being on the network grants the trust, the overlay is the transport security, no
|
||||
> module author chooses it, and the trust is written into the runtime's file rather than over it
|
||||
> and reloaded rather than restarted (ADR 0102). What moved: the controller no longer injects the
|
||||
> file or the service. The container runtime's own module writes `insecure-registries`, told where
|
||||
> this machine reaches the store by `${seat:mesh-artifact-store:reach}`, because the controller
|
||||
> writes no file a seat's holder owns ([issue 190](../04-ISSUES/190-the-runtimes-configuration-is-written-by-modules-that-are-not-the-runtime/00-report.md)).
|
||||
3. **No accounts (issue 042), recorded as the position it always was.** Reading and pushing
|
||||
require presence on the overlay and nothing else. The boundary is enforced, not assumed: the
|
||||
registry's `listens` is `from: mesh`, the firewall derives from it, and the overlay admits only
|
||||
|
||||
@@ -45,6 +45,13 @@ an operator obligation, and an obligation enforced by nothing is issue 057 resta
|
||||
declaration is computed from the whole mesh; delivering a mesh that is knowingly inconsistent and
|
||||
merely saying so would make "push succeeded" mean less than it says.
|
||||
|
||||
> **The mechanism changed — 2026-10-05, by [ADR 0221](0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md).**
|
||||
> A named push still flushes every other machine that is behind, compared against what each was last
|
||||
> sent. A machine whose modules would move to a build their upgrade policy records, or that an open plan
|
||||
> has not sent it yet, is no longer flushed: the push names it and leaves it for `push <node>`. "Behind
|
||||
> for an unrelated reason" no longer covers a held upgrade
|
||||
> ([issue 259](../04-ISSUES/259-a-named-push-sent-every-machine/00-report.md)).
|
||||
|
||||
## Consequences
|
||||
|
||||
- One push is sufficient for a cross-node consumer: the provider's grants arrive from the same
|
||||
|
||||
@@ -64,6 +64,15 @@ node therefore trusts the mesh's registry as soon as it is on the private networ
|
||||
networking no longer touches the runtime. The hosts file networking writes is still written whole
|
||||
and stays held until networking is taken; a converge preview names it among the files it replaces.
|
||||
|
||||
> **The mechanism changed — 2026-10-05, by [ADR 0222](0222-a-module-is-told-where-a-mesh-seats-holder-is-reached-and-the-controller-writes-no-file-a-seats-holder-owns.md).**
|
||||
> Writing into a shared file, adding to a list and reloading rather than restarting all stand. What
|
||||
> moved is the writer: the networking module no longer writes the runtime's trust or declares its
|
||||
> service. The container runtime's own module writes it into its own file and reloads its own
|
||||
> service ([issue 190](../04-ISSUES/190-the-runtimes-configuration-is-written-by-modules-that-are-not-the-runtime/00-report.md)).
|
||||
> The controller test named under *How it is checked* below, which held the networking module to
|
||||
> declaring the runtime's file, is replaced by one holding the networking module to declaring neither,
|
||||
> and by one holding the runtime's module to the trust.
|
||||
|
||||
## Consequences
|
||||
|
||||
- The runtime's file on a machine in use keeps its data directory, its logging settings and
|
||||
|
||||
@@ -60,6 +60,14 @@ with the server held still for its duration. Plain collection, not `--delete-unt
|
||||
mesh keeps is still a manifest in the store, so it is still referenced, so its blobs stay — the
|
||||
dangerous flag is not needed at all once the mesh is the one deciding.
|
||||
|
||||
> **Progressive insight — 2026-10-05.** "What the mesh keeps is still a manifest in the store" was true
|
||||
> of images and false of archives: the builder published every archive as a bare blob no manifest names,
|
||||
> and the store's collector keeps only what a manifest names. Its first night would have deleted every
|
||||
> archive the mesh keeps ([issue 253](../04-ISSUES/253-the-stores-collector-would-delete-every-archive-the-mesh-keeps/00-report.md)).
|
||||
> The decision stands — the mesh decides, the store reclaims with plain collection. What changes is how an
|
||||
> archive is published: with a manifest that holds it, so the sentence becomes true of archives too. Until
|
||||
> every kept archive is held, the collector runs as a dry run.
|
||||
|
||||
**3. What the mesh keeps, stated as three reasons rather than a number.** A digest is kept because:
|
||||
|
||||
- **a definition names it** — every artifact reference in any module's current recorded manifest,
|
||||
|
||||
+5
@@ -11,6 +11,11 @@ extends: 0191-the-meshs-resolver-holds-only-the-meshs-own-names.md
|
||||
|
||||
# 194. The mesh has one resolver, and every node asks it for the mesh's names
|
||||
|
||||
> **The mechanism changed — 2026-10-05, by [ADR 0223](0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md).** `mesh-resolver` (named `mesh-dns-resolver` in the
|
||||
> set) is no longer of capacity one: it is a replicated seat, held on the anchor and on the home
|
||||
> server, each holding every node's internal domain from the same roster. One resolving module, the
|
||||
> mesh's names held only by its holders, and the retirement of every per-node copy stand.
|
||||
|
||||
> **Narrowed, not replaced — 2026-10-03.** How a node asks is decided again by
|
||||
> [ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md): every
|
||||
> node and container asks `mesh-resolver` first and a public resolver only when it is silent. There is
|
||||
|
||||
+7
@@ -10,6 +10,13 @@ supersedes-in-part:
|
||||
|
||||
# 196. A node asks the mesh's resolver first, and a public one only when it is silent
|
||||
|
||||
> **Narrowed, not replaced — 2026-10-05, by [ADR 0223](0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md).** The public resolver listed second goes: musl asks
|
||||
> every listed server at once and takes the first reply, so a public "no such name" for a mesh name
|
||||
> won it, and every Alpine build on the home server failed. Every machine now lists the mesh's
|
||||
> resolvers — two holders of `mesh-dns-resolver`, its own first on a holder — and nothing else. Every
|
||||
> node and container asking the mesh's resolver for every name, with no stub and no runtime `dns`,
|
||||
> stands. The "fallback" consequence and check below describe what this record decided, not what runs.
|
||||
|
||||
## Context
|
||||
|
||||
**[ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) gave the
|
||||
|
||||
+4
@@ -9,6 +9,10 @@ extends: 0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-nam
|
||||
|
||||
# 199. A module that answers names declares its zone, and a node's hosts file is one module's
|
||||
|
||||
> **Decided to change — 2026-10-05, by [ADR 0223](0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md), not yet built.** The `hosts` module is renamed `hostname`
|
||||
> and its seat `node-hostname`, and it owns `/etc/hostname` as well as `/etc/hosts`; the operator's
|
||||
> lines are kept as this record decides. Until that is built, everything below stands as decided.
|
||||
|
||||
## Context
|
||||
|
||||
**[ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) and
|
||||
|
||||
@@ -64,6 +64,13 @@ instead of `_`, which is the whole of the difference between the mesh's identifi
|
||||
the one buckets, vhosts and hostnames use. A provider that needs a prefix or a suffix writes it
|
||||
around the placeholder, because a served value is a string.
|
||||
|
||||
> **The mechanism changed — 2026-10-06, by [ADR 0225](0225-a-consumers-identity-is-bounded-by-the-provision-it-requires.md).**
|
||||
> The identity is no longer capped at twenty characters for every provision: each offer states the
|
||||
> bound its backend keeps, and twenty is the bound of the object store and of a provider that does not
|
||||
> say. What stands: the identity is still the mesh's, and `dns` is still that name with its separator
|
||||
> written `-`. An offer serving `${consumer:as:dns}` now bounds its consumers at 63 or less, so the
|
||||
> label still fits.
|
||||
|
||||
The rejected alternative is **the provider returning values from provisioning** — the natural
|
||||
channel, since the provider is what derived them. It is rejected for three reasons, in order of
|
||||
weight. It inverts the delivery the mesh is built on: a grant would carry data the provider wrote
|
||||
|
||||
+100
@@ -0,0 +1,100 @@
|
||||
---
|
||||
topic: the mesh
|
||||
status: accepted
|
||||
date: 2026-10-05
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md
|
||||
---
|
||||
|
||||
# 218. A plan sends grants before code, rolls a module out one machine first, and a newer merge takes over an older plan
|
||||
|
||||
## Context
|
||||
|
||||
On 2026-10-05 the delivery path was watched through a day of merges, by several sessions at once. Three
|
||||
things went wrong, each recorded as an issue with its evidence.
|
||||
|
||||
- **Code arrived before the right to use it** ([issue 249](../04-ISSUES/249-a-modules-new-state-is-refused-until-a-push-the-merge-did-not-make/00-report.md)).
|
||||
A merge gave a module a new key-value state. The plan sent the new bundle to every machine, and only
|
||||
then issued the memberships that grant the state. On three machines the module's new state was refused
|
||||
for two minutes, until a push made by hand. The order is written into the code on purpose: memberships
|
||||
"after the declaration, because the runtime it is for arrives with it". That reason holds only for a
|
||||
first assignment, and even then a membership is kept on the bus for the runtime that connects later
|
||||
([ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md)).
|
||||
- **No machine went first.** The module's upgrade policy sends one machine at a time, but a plan's rollout
|
||||
ignores it and sends every machine running the module at once. One at a time also never waited for the
|
||||
first machine to come up healthy: it stopped only if the publish itself failed. A change was therefore
|
||||
everywhere before anything had seen it run.
|
||||
- **Plans for successive merges ran over each other** ([issue 254](../04-ISSUES/254-plans-for-successive-merges-run-over-each-other-and-one-was-left-open/00-report.md)).
|
||||
Three merges to the catalogue within four minutes made three plans. Each sent the build agent to every
|
||||
machine and asked for the same builds. One was still "building" hours later, with nothing left for it to
|
||||
wait on. [ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md) decides one plan per merge
|
||||
and says nothing about the next merge arriving while one is open. [Issue 219](../04-ISSUES/219-an-older-build-that-finishes-later-replaces-a-newer-one/00-report.md)
|
||||
settled only which build's output wins.
|
||||
|
||||
## Considered Options
|
||||
|
||||
1. **Debounce merges:** wait a window before planning, so close merges make one plan. Rejected: it only
|
||||
delays the overlap, does nothing for merges further apart than the window, and makes every merge slower.
|
||||
2. **Queue plans:** a new plan waits until the older one is done. Rejected: the older plan builds what the
|
||||
newer merge is about to replace, then the newer one builds it again.
|
||||
3. **A newer merge's plan takes over the older plan's unfinished work, a plan rolls a module out one
|
||||
machine first, and grants travel before code.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**1. Grants before code.** Every send — a plan's rollout and a push alike — issues the memberships for
|
||||
the machines it is about to send to before it sends their declarations, after raising the buckets they
|
||||
name. When the composed list of bus users changes, the machine that holds the bus is sent first, because
|
||||
that list travels in its declaration. A membership that could not be issued fails the send, and the send
|
||||
is tried again. It is never reported as done "until the next push".
|
||||
|
||||
**2. One machine first.** A plan rolls a module out according to the module's upgrade policy
|
||||
([ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md) §3). Unless the policy says
|
||||
*together*:
|
||||
|
||||
- the module is sent to one machine first, the first by name of the machines running it;
|
||||
- the rest are sent only once that machine has reported the new declaration applied and current;
|
||||
- a first machine that reports a failure, or does not report in time, stops the module's rollout there.
|
||||
The plan names the machine and the reason, and the other machines keep what they ran.
|
||||
|
||||
The plan records which machine went first, so a controller replaced mid-rollout resumes from there. A
|
||||
policy of *together* keeps today's behaviour.
|
||||
|
||||
**3. A newer merge takes over an older plan.** When a merge into a repository's branch makes a plan,
|
||||
every open plan for the same repository and branch made before it is superseded, ordered by when each
|
||||
plan was made, never by commit:
|
||||
|
||||
- the modules the older plan had not yet built join the newer plan's set, before its tiers are computed;
|
||||
- the older plan ends in a state of its own, *superseded*, naming the plan that took it over.
|
||||
|
||||
Builds the older plan already asked for still finish and register; issue 219's ordering keeps the newer
|
||||
one current. A person can also close a plan that waits on nothing, by its id. The plan is marked closed
|
||||
by hand and never resumed.
|
||||
|
||||
## Consequences
|
||||
|
||||
- A module that gains a state, an event or a tool can use it from its first start on every machine.
|
||||
- A change reaches one machine before the rest. A change that breaks its first machine stops there, with
|
||||
the reason in the plan.
|
||||
- Successive merges build each module once, for the newest commit. The build agent is sent to the
|
||||
machines once per run of merges, not once per merge.
|
||||
- **What got harder:** a rollout takes one machine's report longer than before. A module that must change
|
||||
everywhere at once says *together* in its policy. A plan's record now has a superseded state that
|
||||
readers of the plans must know.
|
||||
|
||||
## How it is checked
|
||||
|
||||
| Rule | Checked by |
|
||||
|---|---|
|
||||
| grants before code | the controller's test: a send records memberships issued before any declaration; the machine holding the bus is sent first when the user list changes; a failed membership fails the send |
|
||||
| one machine first | the controller's test: with a one-at-a-time policy, one machine is sent, the rest only after its applied and current report; a failed first machine stops the module; *together* sends all at once |
|
||||
| a newer merge takes over | the controller's test: an older open plan for the same repository and branch is superseded, its unbuilt modules folded in; a plan for another repository is left alone; a superseded plan is not open |
|
||||
| live | the next merge to the catalogue that gives a module a new state: no refusal of that state on any machine, the first machine named in the plan, one plan open per repository |
|
||||
|
||||
## References
|
||||
|
||||
- [Issue 249](../04-ISSUES/249-a-modules-new-state-is-refused-until-a-push-the-merge-did-not-make/00-report.md), [issue 254](../04-ISSUES/254-plans-for-successive-merges-run-over-each-other-and-one-was-left-open/00-report.md)
|
||||
- [ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md) — plans and tiers, extended here
|
||||
- [ADR 0160](0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md) — memberships, kept on the bus
|
||||
- [to-be 30](../03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md) — the design this amends
|
||||
+120
@@ -0,0 +1,120 @@
|
||||
---
|
||||
topic: the mesh
|
||||
status: accepted
|
||||
date: 2026-10-05
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md
|
||||
---
|
||||
|
||||
# 219. The build queue is controlled through the controller and the build seat
|
||||
|
||||
## Context
|
||||
|
||||
The operator asked for tools to control the mesh's builds: see what is queued and running, cancel,
|
||||
clear, stop a build immediately, pause and continue, restart and replay. On the day of the request none
|
||||
existed. The controller could ask for a build and list finished ones. Nothing could see an ask waiting on
|
||||
the build seat's work queue, or one being built. Nothing could take an ask back, and a running build
|
||||
could only be stopped by restarting the machine's build agent, which hands the ask to another holder.
|
||||
|
||||
How builds run today, read from the code and the bus
|
||||
([ADR 0190](0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md),
|
||||
[ADR 0157](0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md)):
|
||||
|
||||
- an ask is a message on the seat's work-queue stream;
|
||||
- every holder pulls one at a time from one shared worker, and acknowledges after it has announced the
|
||||
outcome;
|
||||
- a build says it started, logs its steps, and announces what it built, all on the bus;
|
||||
- an ask delivered five times without an answer stays in the stream for a week, with nothing saying so.
|
||||
|
||||
Three facts constrain any control:
|
||||
|
||||
- **Only the controller may act on the bus's streams and consumers** (design 25 §3, enforced in the bus's
|
||||
user list). Holders may only take from their worker and acknowledge.
|
||||
- **A plan waits for a build's outcome and has no timeout.** An ask that disappears without one leaves
|
||||
its plan waiting, said only as late after half an hour.
|
||||
- **The bus server in use cannot pause a consumer**; that arrived in a later version.
|
||||
|
||||
## Considered Options
|
||||
|
||||
1. **Pause and cancel by remaking the worker consumer.** Rejected: a remade worker delivers from now on,
|
||||
so every queued ask would be skipped silently. Issues 206 and 207 are this mistake both ways round.
|
||||
2. **Upgrade the bus first and use its consumer pause.** Not now: it gives pause and continue, and
|
||||
nothing else on the list, and upgrading the bus is a change of its own.
|
||||
3. **Queue actions are the controller's verbs, process actions are the build seat's verbs, and every
|
||||
action that drops work leaves a failed outcome.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**1. The queue is the controller's.** Its verbs act on the work-queue stream, the one thing only it may
|
||||
touch:
|
||||
|
||||
| verb | what it does |
|
||||
|---|---|
|
||||
| `queue` | lists every ask: waiting, in flight (with the machine building it, from its started event), and dead (delivered as often as allowed, still in the stream) |
|
||||
| `cancel <id>` | takes back a waiting or dead ask; an ask in flight is refused, and `kill` is named instead |
|
||||
| `clear` | cancels every waiting ask, and with `--dead` every dead one |
|
||||
| `rebuild <module or build>` | asks again for a module's current source, or a past build's source and ref, under a new id |
|
||||
| `replay <build>` | asks again for a past build at its commit, as a **dry run** unless told to register |
|
||||
| `kill <id>`, `pause [node]`, `resume [node]` | pass the request on to the build seat on the right machine, or on every machine |
|
||||
|
||||
**2. The process is the build seat's.** Each holder serves verbs on its own machine:
|
||||
|
||||
- `current`: the build running here, its step and how long, and whether this holder is paused;
|
||||
- `kill`: stops a running build at once. The build's whole process group is ended, along with every
|
||||
container it started. Its outcome is announced as failed, and the ask is acknowledged, so it is not
|
||||
delivered again;
|
||||
- `pause` and `resume`: a paused holder takes nothing new, and a running build finishes. The flag
|
||||
survives the holder's restart.
|
||||
|
||||
A holder restarted mid-build keeps today's behaviour: the ask is not acknowledged, and another holder
|
||||
takes it.
|
||||
|
||||
**3. Nothing dropped is silent.** Cancel, clear and kill each leave a failed outcome for the ask's id,
|
||||
taken in like any other. A plan waiting on that build fails, and says why, instead of waiting. A plan keeps
|
||||
the id it asked for, so it can match its outcome exactly. A holder checks a cancelled ask before it builds
|
||||
it, so an ask taken in the instant it was cancelled is not built.
|
||||
|
||||
**4. A plan says what its builds are waiting on, and a failed plan can go on.**
|
||||
|
||||
- A plan whose build waits on a paused seat says the seat is paused, and on which machines, and is not
|
||||
counted late while it waits.
|
||||
- `plans retry <id>` asks again for the modules a failed plan could not build, under new ids. The plan
|
||||
resumes at that tier and goes on through its later ones. A plan another has superseded, or one
|
||||
already done, is refused.
|
||||
- `rebuild` of a module that an open or failed plan has not yet built joins that plan, so the plan and
|
||||
the build are one thing.
|
||||
|
||||
**5. Replay does not move the mesh backwards unasked.** A replayed build is a dry run: built, its log
|
||||
kept, nothing registered. With `--register` it is registered. If a newer build of the module is already
|
||||
registered, that is refused unless `--older` is said as well: registering an older commit makes it the
|
||||
current one, and the rollout policy sends it to the machines ([issue 207](../04-ISSUES/207-a-re-made-worker-replayed-every-ask-the-stream-kept/00-report.md)).
|
||||
|
||||
## Consequences
|
||||
|
||||
- Every build in the mesh can be seen, taken back, stopped or asked again from the console. None of it
|
||||
needs a shell on a machine.
|
||||
- A cancelled or killed build shows in the build records as failed, with who stopped it. Its plan fails
|
||||
saying the same.
|
||||
- **What got harder:** a holder now serves verbs as well as taking work, and keeps one small flag on
|
||||
disk. Pause is per holder, so "pause the mesh" is the controller asking every holder in turn. A holder
|
||||
away at the time misses it, and the answer names that holder.
|
||||
|
||||
## How it is checked
|
||||
|
||||
| Rule | Checked by |
|
||||
|---|---|
|
||||
| the queue is read and classified right | the controller's test: waiting, in flight and dead asks told apart from the stream and the worker; a live test on a throwaway bus |
|
||||
| nothing dropped is silent | the controller's test: cancel and clear delete the ask and record a failed outcome; a plan asked for that id fails |
|
||||
| a cancelled ask is not built | the holder's test: an ask taken after its cancel is answered failed without building |
|
||||
| kill stops everything it started | the holder's test: the build's process group and its labelled containers are ended; the outcome is failed and the ask acknowledged |
|
||||
| pause survives a restart | the holder's test: the flag is read back at start, and nothing is taken while it is set |
|
||||
| plans follow the queue | the controller's test: a plan waiting on a paused seat says so and is not late; `plans retry` resumes a failed plan and its later tiers are asked; `rebuild` joins the plan that holds the module |
|
||||
| replay is safe | the controller's test: a dry run by default, and `--register` over a newer build refused without `--older` |
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0190](0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md) — holders share the seat's work, one at a time
|
||||
- [ADR 0157](0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md) — a build's own events are its record
|
||||
- [issue 207](../04-ISSUES/207-a-re-made-worker-replayed-every-ask-the-stream-kept/00-report.md) — a remade worker replayed every ask
|
||||
- [to-be 18](../03-DESIGN/01-to-be/18-building-a-module.md) — the design this amends
|
||||
+161
@@ -0,0 +1,161 @@
|
||||
---
|
||||
topic: what runs on it
|
||||
status: accepted
|
||||
date: 2026-10-05
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md
|
||||
---
|
||||
|
||||
# 220. What a machine asks needs its uplink held, and the retired resolver pieces go
|
||||
|
||||
> **Decided to go — 2026-10-05, by [ADR 0223](0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md), not yet built.** `/etc/resolv.conf` becomes the `node-uplink`
|
||||
> holder's file, and `resolv-conf`, `node-resolver-config` and this record's dependency of it on
|
||||
> `node-uplink` retire with it. Until that is built, everything below stands as decided.
|
||||
|
||||
## Context
|
||||
|
||||
**Three things about a machine's resolver were left half done when the mesh moved to one resolver.**
|
||||
On the production mesh on 2026-10-05, read from the controller's `seats` verb:
|
||||
|
||||
- **`node-dns-resolver` has no holder on any node, and no module in the catalogue claims it.**
|
||||
[ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) retired it
|
||||
and the controller kept its row deliberately, *"deleted once nothing claims it"*, because removing a
|
||||
seat a machine still holds makes that machine unresolvable. That condition now holds. The row still
|
||||
stands in the controller's compiled set and in the store's seat table, and the overview still lists
|
||||
it, unheld, beside the seats a mesh actually has.
|
||||
- **The rule that keeps `/etc/resolv.conf` the mesh's is checked by nothing.**
|
||||
[ADR 0117](0117-a-machines-uplink-is-a-seat.md) found that a network manager rewrites the resolver
|
||||
file on every connectivity change unless it is told not to, and gave that telling to the module
|
||||
holding `node-uplink`. It said the condition *"only if NetworkManager runs"* is expressed by
|
||||
assigning the manager's module. Nothing makes anybody do so: `resolv-conf` can be assigned to a
|
||||
machine with no uplink holder, and the file is then replaced the first time a laptop changes
|
||||
network while every surface of the mesh reads green. Today every node holding
|
||||
`node-resolver-config` also holds `node-uplink` — NetworkManager on the home server, the
|
||||
workstation and the laptop, systemd-networkd on the anchor — by care, not by check.
|
||||
- **The catalogue still carries a systemd-resolved split-DNS module**, `resolved-split-dns`, claiming
|
||||
`node-resolver-config`. [ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md)
|
||||
chose against a stub on every node and says *"There is no `systemd-resolved` module."* It is
|
||||
assigned nowhere. `resolv-conf`'s own resolver file still tells its reader that systemd-resolved or
|
||||
NetworkManager may be assigned *instead* — the opposite of how the roles now divide.
|
||||
|
||||
**A dependency mechanism already exists.** [ADR 0207](0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md)
|
||||
made a module depend on the node seats that apply its resources, derived rather than stated, judged
|
||||
over the node's whole set of assignments, refused at `assign` naming the seat and its possible holders,
|
||||
and refused at composition. [ADR 0210](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)
|
||||
derived a second kind from contributions. What is missing is a dependency that belongs to a *role*
|
||||
rather than to what a module declares.
|
||||
|
||||
## Considered Options
|
||||
|
||||
**For the retired seat:**
|
||||
|
||||
1. **Keep the row until the build seat's retired row goes too, and delete both together.** Rejected:
|
||||
the two have nothing in common but having been retired; one is unclaimed now and the other is not
|
||||
yet known to be.
|
||||
2. **Remove it from the compiled set only.** Rejected: seeding adds a seat a release ships and never
|
||||
removes one ([ADR 0122](0122-a-seat-is-data-a-rename-is-a-database-update.md)), so the store's
|
||||
row, which is the live set, would stay.
|
||||
3. **Remove it from the compiled set and delete the store's row in a numbered migration**, as the
|
||||
artifact store's rename did for its old row. Chosen.
|
||||
|
||||
**For the resolver file and the uplink:**
|
||||
|
||||
1. **Leave it to the operator.** Rejected: it is the failure ADR 0117 describes — the file silently
|
||||
replaced — with the one difference that the operator was told.
|
||||
2. **`resolv-conf` declares the manager's settings itself.** Rejected by ADR 0117 already: which
|
||||
setting depends on which manager runs, and a resolver module that knew about network managers would
|
||||
be the wrong module knowing the wrong thing.
|
||||
3. **A manifest field in `resolv-conf` naming `node-uplink`.** Rejected for the reason ADR 0207
|
||||
rejected its own option 2: a second module claiming the same seat would have to restate it, and
|
||||
one that forgot would pass.
|
||||
4. **The seat carries what its holder needs beside it.** `node-resolver-config` names `node-uplink`;
|
||||
any module claiming the former depends on the latter, derived from the claim and judged exactly as
|
||||
ADR 0207 judges a resource's dependency. Chosen.
|
||||
|
||||
**For the split-DNS module:** keep it for a machine that wants systemd-resolved in charge, or remove it.
|
||||
Kept, it is a second answer to a question ADR 0196 settled, and a claimant the catalogue offers
|
||||
without a record allowing it. Removed.
|
||||
|
||||
## Decision
|
||||
|
||||
**1. `node-dns-resolver` is deleted from the mesh's set.** The controller's compiled set no longer
|
||||
carries it, and a numbered migration of the controller's store deletes its row and any alias naming
|
||||
it. No alias is kept: nothing was renamed, and a manifest still claiming it should be refused at
|
||||
registration, naming the seat. This completes ADR 0194's retirement; nothing it decided changes.
|
||||
|
||||
**2. A seat may name the node seats its holder needs held on the same node.** A module claiming such a
|
||||
seat depends on each of them. The dependency is a third source beside ADR 0207's resources and ADR
|
||||
0210's contributions, and everything ADR 0207 §3 and §4 say of those applies unchanged: met by any
|
||||
module assigned to the node, the claimant included; judged over the node's whole set; refused at
|
||||
`assign` naming the seat and the catalogue's possible holders; refused at composition; and only said,
|
||||
never refused, when no module in the catalogue could hold the needed seat. Unassigning the needed
|
||||
seat's last holder beneath a dependent is refused, naming the dependent. What a seat needs is part of
|
||||
the mesh's definition of the role: compiled with the set, never stored, as ADR 0212 keeps what a seat
|
||||
receives. Adding a need to a seat is a decision, recorded.
|
||||
|
||||
**3. `node-resolver-config` needs `node-uplink`.** The holder that writes the resolver file is right
|
||||
only while the network manager is told to leave it alone, and that telling is the uplink holder's
|
||||
(ADR 0117). Every manager the catalogue knows — NetworkManager, systemd-networkd, dhcpcd — holds
|
||||
`node-uplink`, so the refusal always has a remedy to name.
|
||||
|
||||
**4. `resolved-split-dns` leaves the catalogue.** `resolv-conf` is the only module claiming
|
||||
`node-resolver-config`. Its resolver file's comment says the uplink's holder is required beside it,
|
||||
rather than naming alternatives to assign instead.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **The set reads as the mesh is.** Thirty-seven seats in the compiled set; the overview no longer
|
||||
lists a role nothing can fill.
|
||||
- **A machine cannot be given the mesh's resolver file without its network manager being told to keep
|
||||
off it.** A machine with no manager at all — a static configuration — needs the smallest holder,
|
||||
`dhcpcd`, or a new module holding `node-uplink` for its way of configuring the link. That is the
|
||||
point: such a machine has to say what manages its link before the mesh writes a file the manager
|
||||
could overwrite.
|
||||
- **Order of assignment on a new machine**: the uplink holder before or with `resolv-conf`, in one
|
||||
`assign` when together. On the production mesh nothing changes: every node already holds both.
|
||||
- **The uplink becomes harder to take away.** Unassigning a machine's manager module while
|
||||
`resolv-conf` stays is refused; replacing one manager with another is one act assigning the new and
|
||||
unassigning the old, or the dependent goes first.
|
||||
- **A machine wanting systemd-resolved has no module for it.** A future need for one is a new record,
|
||||
not a revival of the removed module.
|
||||
- **Changing `resolv-conf`'s comment rewrites `/etc/resolv.conf` on every node once**, with the same two
|
||||
nameserver lines and options; only the comment differs.
|
||||
- **The merge order matters.** The controller's tests read the catalogue beside them, and the two
|
||||
changes are judged together: the controller's change and the catalogue's removal merge together,
|
||||
the catalogue's first or in the same window, and the controller rolls out only once its test suite
|
||||
passes against the merged catalogue.
|
||||
|
||||
## How it is checked
|
||||
|
||||
| Rule | Checked by |
|
||||
|---|---|
|
||||
| `node-dns-resolver` is not in the set, and the set has thirty-seven seats | mesh-controller's closed-set unit test on the compiled seats |
|
||||
| The store's row goes with it | the migration, and after rollout the controller's `seats` verb listing no `node-dns-resolver` |
|
||||
| `node-resolver-config` needs `node-uplink`, and a module claiming it depends on the uplink with nothing in its manifest | mesh-controller's seat-dependency tests on the seat definition and on a synthetic claimant |
|
||||
| What a seat needs survives loading the set from the store | a unit test loading the store's rows, which carry no such column |
|
||||
| `resolv-conf` without an uplink holder is refused at `assign`, naming `node-uplink` and its possible holders; beside one, or with one in the same act, it passes; a composition without one is refused | the same tests, and `assign` live |
|
||||
| Unassigning the uplink's last holder beneath `resolv-conf` is refused | an unassign test |
|
||||
| In the catalogue, `resolv-conf` depends on the uplink, dhcpcd, NetworkManager and systemd-networkd each hold it, and `resolv-conf` is the only claimant of `node-resolver-config` | a mesh-controller test reading the catalogue beside it |
|
||||
| Two modules deciding what a machine asks are still refused on one node | the resolver test, now with a synthetic second claimant |
|
||||
| Every node of the live mesh holding `node-resolver-config` also holds `node-uplink` | the controller's `seats` verb, read before this was decided and after it rolls out; `status` reports no unheld dependency |
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) — retired
|
||||
`node-dns-resolver`; this record deletes it.
|
||||
- [ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md) — no
|
||||
stub, and so no systemd-resolved module.
|
||||
- [ADR 0117](0117-a-machines-uplink-is-a-seat.md) — the uplink's holder keeps the manager off the
|
||||
resolver file; this record makes that a checked dependency.
|
||||
- [ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md) — serving and
|
||||
asking as two seats.
|
||||
- [ADR 0207](0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md),
|
||||
[ADR 0210](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md) — the dependency mechanism
|
||||
this extends.
|
||||
- [ADR 0122](0122-a-seat-is-data-a-rename-is-a-database-update.md) — the set as data, which is why a
|
||||
deletion is a migration.
|
||||
- [The seats](../03-DESIGN/01-to-be/26-the-seats.md) and
|
||||
[connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md), amended alongside.
|
||||
- mesh-controller `internal/catalogue/seats.go`, `internal/catalogue/seat_dependencies.go`, and the
|
||||
store migration deleting the row; mesh-catalog `modules/resolv-conf`.
|
||||
+122
@@ -0,0 +1,122 @@
|
||||
---
|
||||
topic: the mesh
|
||||
status: accepted
|
||||
date: 2026-10-05
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
|
||||
---
|
||||
|
||||
# 221. A push sends no build a policy or a plan holds back, except to the machine it names
|
||||
|
||||
## Context
|
||||
|
||||
[ADR 0083](0083-one-push-leaves-the-mesh-consistent.md) makes a named push finish what it starts. After
|
||||
the named machine is sent, every other machine whose declaration differs from what it was last sent is
|
||||
sent too. The case it was written for is a grant: assigning a consumer changes the provider's
|
||||
declaration on another machine. ADR 0083 accepted that a machine behind *for an unrelated reason* is
|
||||
flushed as well, and called that correct rather than a cost.
|
||||
|
||||
On 2026-10-05 that reasoning met an upgrade policy
|
||||
([issue 259](../04-ISSUES/259-a-named-push-sent-every-machine/00-report.md)). A change to the resolver
|
||||
modules was merged with the policy `record`, so that each machine would take it only when pushed, one
|
||||
at a time: the anchor first, each checked before the next. `push <anchor>` sent all four machines the
|
||||
new build, and so did a later `push <laptop>`. A fault in the change
|
||||
([issue 260](../04-ISSUES/260-the-resolver-started-before-its-zones-file-existed/00-report.md)) was met
|
||||
on every machine at once.
|
||||
|
||||
The cascade compares one digest per machine, and a digest cannot say why a machine differs. Under
|
||||
`record`, every machine running the module differs from the merge on, so every machine is flushed. The
|
||||
same holds for a plan's rollout under [ADR 0218](0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md):
|
||||
while the plan waits on its first machine, the rest differ, and any named push elsewhere sends them the
|
||||
build the plan is holding. [Issue 249](../04-ISSUES/249-a-modules-new-state-is-refused-until-a-push-the-merge-did-not-make/00-report.md)
|
||||
met the same confusion for the machine holding the bus, which a send adds when its user list changed.
|
||||
It narrowed that check to the user list, but the machine, once added, is still sent its whole
|
||||
declaration.
|
||||
|
||||
So `record` held a change back from nothing but the merge, and "one machine first" held it back only
|
||||
from the plan's own sends.
|
||||
|
||||
## Considered Options
|
||||
|
||||
1. **Report instead of cascade:** list every machine that is behind and send none. Rejected: it is the
|
||||
option ADR 0083 rejected, and for the same reason. [Issue 057](../04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md)'s provider would wait for a second push
|
||||
nothing tells anyone to make.
|
||||
2. **Send a held machine its consequences and keep its held modules at their last build.** The cascade
|
||||
would compose the machine with each held module pinned at the build it was last sent, so a grant
|
||||
still reaches it and the upgrade does not. Rejected: the catalogue resolves every machine against
|
||||
one manifest per module, the current one. Pinning means resolving a machine's set against a mix of
|
||||
current and older manifests, read back from the build records, beside the current settings, seats
|
||||
and grants. The result matches neither what the machine runs nor what the mesh would send it, so
|
||||
neither `plan` nor `status` could show it. A grant composed for the new build may not fit the old
|
||||
one. It is a second composition path to get one edge case right, and the edge case has a one-word
|
||||
remedy: name the machine.
|
||||
3. **Keep, with every send, which build of each module the machine was sent, and leave a machine any
|
||||
of whose modules a policy or a plan holds back.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**1. A send records the builds it carried.** With the digest of every declaration it sends, the mesh
|
||||
keeps which build of each module the declaration carried: the commit the module's current build was
|
||||
made from. A module left out of the declaration keeps the build it was last sent. A declaration sent by
|
||||
hand records that what it carried is not known.
|
||||
|
||||
**2. A push does not send a machine it did not name a held build.** A machine reached by a named push's
|
||||
cascade, or added because it holds the bus, is not sent when any module it runs would move to a build
|
||||
that:
|
||||
|
||||
- its upgrade policy records rather than rolls out, or
|
||||
- an open plan has not yet sent it: the plan is still building the module, or has sent it to its first
|
||||
machine and this is not that machine.
|
||||
|
||||
A module the machine was never sent counts as a move. A machine whose last send's builds are not known
|
||||
counts as held. It was sent before this was kept, or by hand, so a held upgrade cannot be told apart
|
||||
from anything else.
|
||||
|
||||
**3. It is named, not hidden.** The push says which machine it left, which module and which builds, why,
|
||||
and that `push <node>` sends it. For the machine holding the bus it also says that the bus may refuse
|
||||
what this push's machines were newly granted until that machine is sent.
|
||||
|
||||
**4. Everything else stays as it is.** A machine with nothing held is flushed exactly as ADR 0083
|
||||
decides: a grant, a peer, a setting. The machine a push names is sent everything, held builds included.
|
||||
A push that names no machine sends every machine. `push --behind` still sends every machine that is
|
||||
behind, held or not. It is the remedy `record` names when an upgrade is announced ("`push --behind`
|
||||
when you want them"), and the one command for taking a recorded upgrade everywhere. Narrowing it would
|
||||
leave `record` with no way to say "now".
|
||||
|
||||
## Consequences
|
||||
|
||||
- `record` and "one machine first" hold a change back from every push that does not name the machine.
|
||||
Walking a change through the mesh is `push <anchor>`, check, `push <next>`.
|
||||
- **The cost:** a held machine that is also owed a grant from this push waits for its own push. The
|
||||
provider in issue 057's case, if a held upgrade is pending on it, is not sent its new grant, and its
|
||||
consumer is refused until the provider is pushed. The push names that machine and the remedy, so the
|
||||
wait is announced, not silent. When the holder of the bus is held, a new grant may be refused by the
|
||||
bus until it is sent, and the push says so.
|
||||
- On the first push after this ships, no machine's last send has its builds recorded yet. Each is held
|
||||
from cascades until it is pushed once: by name, in a whole-mesh push, or by `push --behind`.
|
||||
- A manifest handed over by hand does not change the commit a module records, so a cascade still sends
|
||||
its change. Only builds have a commit to compare.
|
||||
- ADR 0083's consequence that a machine behind for an unrelated reason is flushed is narrowed. It still
|
||||
holds for every reason except a build held back.
|
||||
|
||||
## How it is checked
|
||||
|
||||
| Rule | Checked by |
|
||||
|---|---|
|
||||
| a send records the builds it carried, and "not known" | the controller's test: the builds kept with a send round-trip, an empty send is known and empty, a send by hand reads as not known |
|
||||
| a left-out module keeps its last build | the controller's test: a module left out of a declaration records the build it was last sent |
|
||||
| a recorded upgrade holds a machine a push did not name | the controller's test: with a module under `record` moved, the cascade of a push naming the anchor leaves the laptop's last send unchanged and names it, the module, both builds and `push laptop` |
|
||||
| held and owed something else: not sent, both said | the same test: a newly placed machine changes the laptop's peers while the upgrade is held; the laptop is not sent and the output says why and that what else it is owed waits |
|
||||
| a roll-out policy is not held | the same test: with the policy set to roll out, the laptop is sent and its new build recorded |
|
||||
| a consequence nothing holds is still sent (issue 057) | the controller's test: a placed machine's peers reach the others by cascade; a machine whose last send is not known is held |
|
||||
| a plan waiting on its first machine holds the rest | the controller's test: a machine the plan has not reached is held, the first machine is not; a module sent everywhere, failed, or in a closed plan holds nothing |
|
||||
| live | the next change merged under `record`: `push <anchor>` sends the anchor alone and names every other machine running the module |
|
||||
|
||||
## References
|
||||
|
||||
- [Issue 259](../04-ISSUES/259-a-named-push-sent-every-machine/00-report.md), [issue 260](../04-ISSUES/260-the-resolver-started-before-its-zones-file-existed/00-report.md), [issue 249](../04-ISSUES/249-a-modules-new-state-is-refused-until-a-push-the-merge-did-not-make/00-report.md)
|
||||
- [ADR 0083](0083-one-push-leaves-the-mesh-consistent.md): the cascade, narrowed here
|
||||
- [ADR 0218](0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md): one machine first, now held from a push as well
|
||||
- [ADR 0010](0010-delivery.md): delivery, and what `push --behind` answers
|
||||
- [to-be 30](../03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md): the design this amends
|
||||
+150
@@ -0,0 +1,150 @@
|
||||
---
|
||||
topic: the mesh
|
||||
status: accepted
|
||||
date: 2026-10-05
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md
|
||||
---
|
||||
|
||||
# 222. A module is told where a mesh seat's holder is reached, and the controller writes no file a seat's holder owns
|
||||
|
||||
## Context
|
||||
|
||||
[Issue 190](../04-ISSUES/190-the-runtimes-configuration-is-written-by-modules-that-are-not-the-runtime/00-report.md)
|
||||
found the container runtime's configuration file written by modules that are not the runtime's. The
|
||||
first half is closed: the resolver module no longer writes into that file, and the catalogue's runtime
|
||||
module (the holder of `node-container-runtime`) writes `live-restore` and reloads its own service. The
|
||||
second half stands. The controller's private network still generates two resources on every machine
|
||||
on the network: one writes `insecure-registries`, naming the mesh's artifact store, into the runtime's
|
||||
file ([ADR 0082](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md) §2,
|
||||
[ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md)), and the other declares the
|
||||
runtime's service, reloaded on that file. So two parties declare one path and one unit on every
|
||||
machine, and nothing refuses it, because the collision check runs over catalogue manifests and the
|
||||
private network's resources only exist once its generator has answered for a machine.
|
||||
|
||||
On 2026-10-05 the operator made the rule general: **the controller never writes a file a seat's holder
|
||||
owns; it tells the owner.** The hosts file had already moved the same way
|
||||
([ADR 0199](0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)).
|
||||
|
||||
What the runtime's module needs is one fact: the address this machine reaches the mesh's artifact
|
||||
store at. The controller already composes that address for every machine at every push. It puts it
|
||||
into every image and archive reference the mesh built, and never stores it. Today a module has no way
|
||||
to ask for it. A binding gives a consumer the provider's address and a credential. The `${seat:…:<port>}`
|
||||
placeholder gives only a port, and only for the store and the broker the controller itself dials.
|
||||
|
||||
Two facts about the move itself, read from the host's code:
|
||||
|
||||
- The private network writes one member into a list. The host adds a list member and records it per
|
||||
resource, so between the two writers declaring it and one of them going, the member is owned by
|
||||
whichever record added it first.
|
||||
- The host removes every resource no longer declared before it applies anything, so in the apply
|
||||
where the private network's record goes and the runtime module's member arrives, the member leaves
|
||||
and comes back with no reload in between.
|
||||
|
||||
## Considered Options
|
||||
|
||||
1. **Keep the private network writing the trust.** Rejected: it is the defect issue 190 names. Three
|
||||
writers became two, and the second is computed code no check can see.
|
||||
2. **An operator setting on the runtime module naming the registry**, as issue 190 first proposed
|
||||
under [ADR 0164](0164-a-setting-is-declared-with-its-default-its-meaning-and-what-changing-it-costs.md).
|
||||
Rejected: where the store is, is a fact the mesh holds, not a choice an operator makes. A setting
|
||||
would be a copy that goes wrong the day the store moves.
|
||||
3. **The runtime module requires `artifact-store` and reads the address from its binding
|
||||
(`${bound:…}`).** Rejected. A binding makes the module the store's consumer, and that mints a
|
||||
credential for every machine's runtime, which needs none: reading from the store needs presence on
|
||||
the private network and nothing else (ADR 0082 §3). It also makes a cycle: the store runs as a
|
||||
container in the runtime, and the runtime would require the store before it could be assigned.
|
||||
4. **A placeholder that answers where a mesh seat's holder is reached, with no binding.** Chosen. It
|
||||
is the reasoning of `${seat:…:<port>}` taken one step further: nothing is required, nothing is
|
||||
granted, and the answer is an address the mesh already holds in the clear.
|
||||
|
||||
## Decision
|
||||
|
||||
**1. `${seat:<mesh seat>:reach}` is where this machine reaches the holder of a seat the mesh holds once,
|
||||
as host:port.** It is filled where every placeholder is, in a file's content and in an environment
|
||||
value, from the address the controller composes for that machine at that push. It requires nothing
|
||||
and grants nothing, and no credential comes with it.
|
||||
|
||||
- **Only `mesh-artifact-store` is answered.** Another seat is refused by name, never answered with
|
||||
nothing, so a module asking a question the mesh does not answer learns that at composition. A seat
|
||||
is added to what is answered by a record saying why its address is needed.
|
||||
- **The answer may be empty**: no machine on the private network holds the seat yet, which is how
|
||||
genesis begins. In a file written into as JSON, an empty member is dropped from its list, and a list
|
||||
left with no members is dropped, so the software is never told an empty name.
|
||||
|
||||
**2. The runtime's module states the registry's trust.** Its `daemon` resource writes
|
||||
`insecure-registries`, naming `${seat:mesh-artifact-store:reach}`, beside `live-restore`, and it
|
||||
reloads its own service. ADR 0082's decision stands: being on the private network is what grants the
|
||||
trust, the private network is the transport security, and no module author chooses it. Only who writes
|
||||
it moves, from the private network to the runtime's module.
|
||||
|
||||
**3. The controller writes no file a seat's holder owns.** The private network generates neither the
|
||||
runtime's file nor its service.
|
||||
|
||||
**4. What the mesh computes is held to the collision check.** At composition, each computed module's
|
||||
resources, as its generator answers for that machine, are checked beside the other modules' resources.
|
||||
The check is the same one: no two modules on a machine declare one path, unit, name or package. A
|
||||
collision is refused by name. The generator is asked once, and what is checked is exactly what is
|
||||
declared.
|
||||
|
||||
**5. A unit still held is not given back when another record of it goes.** When the host removes a
|
||||
service resource whose unit another declared resource still gives a state, it forgets that record and
|
||||
leaves the unit as it is. What the record going had found is not the machine's to restore while
|
||||
another resource holds the unit, and restoring it could stop the runtime only for the remaining
|
||||
resource to start it again in the same apply.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **Rollout has a fixed order, in three steps.**
|
||||
1. The controller learns the placeholder. Nothing uses it yet.
|
||||
2. The catalogue's runtime module uses it. A controller that does not know the placeholder would
|
||||
send it through unfilled, so this waits until step 1 is deployed.
|
||||
3. The controller stops generating the private network's two resources and checks generated
|
||||
resources for collisions. With step 3 before step 2, machines would lose the trust. The host
|
||||
change in Decision 5 is deployed before step 3.
|
||||
|
||||
This replaces issue 190's single push. That was needed when two writers of one scalar key would
|
||||
have been refused. A list member written by two records is tolerated by the host.
|
||||
- **In the apply of step 3**, the private network's records go first: the member leaves the list and
|
||||
the service record is forgotten. Then the runtime module's file is applied, and the member is added
|
||||
back, recorded as the runtime module's. The runtime is reloaded once. The address is unchanged, so
|
||||
the trust never lapses for a pull.
|
||||
- **A machine holding the store, before any network exists**, is answered with the loopback address
|
||||
it reaches the store at. A runtime already trusts loopback, so this adds a redundant member, not a
|
||||
wrong one.
|
||||
- **A second writer cannot return through generated code.** It is refused at composition, naming
|
||||
both modules and what they share.
|
||||
- **A module can now learn where the artifact store is without being its consumer.** That is a
|
||||
capability, and it is bounded: one seat today, and widened only by a record.
|
||||
|
||||
## How it is checked
|
||||
|
||||
| Rule | Checked by |
|
||||
|---|---|
|
||||
| `${seat:mesh-artifact-store:reach}` is filled with this machine's address for the store, in a file and in an environment value | mesh-controller `TestASeatsReachIsWhereThisMachineReachesItsHolder`, and through the whole composition `TestTheRuntimesTrustIsComposedFromTheSeatsReach` |
|
||||
| An empty answer adds nothing to a list, and other members stay | mesh-controller `TestAnUnansweredReachAddsNothingToAList` |
|
||||
| Any other seat is refused by name | mesh-controller `TestAReachTheMeshDoesNotAnswerIsRefused` |
|
||||
| The catalogue's runtime module writes `live-restore` and the trust, and writes no trust when no store is reachable | mesh-controller `TestTheRuntimesModuleTrustsTheMeshsRegistry`, which composes the catalogue's manifest beside it |
|
||||
| The private network declares neither the runtime's file nor its service | mesh-controller `TestTheNetworkWritesNothingOfTheRuntimes` |
|
||||
| A generated resource colliding with a module's is refused by name; disjoint ones compose; the generator is asked once | mesh-controller `TestAGeneratedResourceCollidingWithAModulesIsRefused`, `TestAGeneratedResourceBesideAModulesOwnIsComposed`, `TestAComputedModuleIsAskedAboutTheNodeItIsFor` |
|
||||
| The member moves between records in one apply without leaving the list; one reload; a unit still held is forgotten, never stopped or disabled; the plan says so | mesh-host `TestTheRegistryMovesToTheRuntimesModuleWithoutLeavingTheList`, `TestThePlanForgetsARecordOfAUnitStillHeld` |
|
||||
| After rollout, every machine's runtime trusts the store's address | the runtime module's `docker_daemon_config` tool on each machine, read before step 3 and after it |
|
||||
|
||||
## References
|
||||
|
||||
- [Issue 190](../04-ISSUES/190-the-runtimes-configuration-is-written-by-modules-that-are-not-the-runtime/00-report.md),
|
||||
steps 2 and 5.
|
||||
- [ADR 0082](0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md): the trust and why it
|
||||
is plain HTTP; its mechanism moves here.
|
||||
- [ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md): writing into a shared file;
|
||||
its writer of the runtime's trust moves here.
|
||||
- [ADR 0079](0079-the-foundation-seats-are-named-after-their-servers.md): a mesh seat names the
|
||||
server, which is what lets a placeholder ask about it.
|
||||
- [ADR 0166](0166-the-container-runtime-is-a-node-seat-and-the-host-creates-containers-through-its-holder.md):
|
||||
the runtime module given the registry as a value.
|
||||
- [ADR 0199](0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md):
|
||||
the hosts file left the controller the same way.
|
||||
- mesh-controller `internal/catalogue/seat_into.go`, `internal/catalogue/declaration.go`
|
||||
(`generatedHere`), `internal/overlay/generator.go`; mesh-catalog `modules/docker`; mesh-host
|
||||
`internal/apply/apply.go` (orphan removal).
|
||||
@@ -0,0 +1,204 @@
|
||||
---
|
||||
topic: the tiers
|
||||
status: accepted
|
||||
date: 2026-10-05
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
supersedes-in-part:
|
||||
- 0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md
|
||||
extends: 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
|
||||
---
|
||||
|
||||
# 223. The mesh has two resolvers, and a machine lists only them
|
||||
|
||||
## Context
|
||||
|
||||
**[ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md) wrote
|
||||
every machine's `/etc/resolv.conf` as the mesh's resolver first and a public resolver second.** Its
|
||||
reasoning is the C library's: servers are asked in order, the next only when one does not answer, and
|
||||
an answer from the first — "no such name" included — is final. That holds for glibc. It does not hold
|
||||
for musl, the C library of every Alpine image: musl sends the query to every listed server at once
|
||||
and takes the first reply.
|
||||
|
||||
**On the production mesh on 2026-10-05 that made the anchor's own name unresolvable from the home
|
||||
server's builds.** The build agent on the home server runs its steps in Alpine containers on the host
|
||||
network, so they read the machine's `resolv.conf` as it is. Asked for `<anchor>.internal`, the public
|
||||
resolver — which has no such name and is the nearer of the two — answered NXDOMAIN first, and musl
|
||||
took it. Reproduced six times out of six inside the build agent's own container; every build on that
|
||||
machine failed fetching from the anchor. The same lookup from glibc on the same machine answered every
|
||||
time. [Issue 262](../04-ISSUES/262-an-alpine-container-could-not-find-a-machine-by-its-mesh-name/00-report.md)
|
||||
was the same library failing on a different answer (NXDOMAIN for a missing IPv6 record); fixing that
|
||||
did not touch this one, because here the wrong answer comes from a server that should never have been
|
||||
asked.
|
||||
|
||||
**The public line was there for one case: the mesh's resolver unreachable.** ADR 0196 kept public
|
||||
names resolving with the anchor down, the tunnel down, or a laptop behind a captive portal. That case
|
||||
is real and rare; the musl case is every lookup of a mesh name from every Alpine container on any
|
||||
machine that is not the anchor.
|
||||
|
||||
**A seat's work can already be shared by several holders.** [ADR 0190](0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md)
|
||||
made the build role node-scoped with every holder pulling from one queue. A mesh-scoped seat has
|
||||
always had exactly one holder: by derivation, or on record since
|
||||
[ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md), where a handover replaces the
|
||||
holder in one write and every other eligible assignment stands beside it, silent.
|
||||
|
||||
## Considered Options
|
||||
|
||||
1. **The mesh's resolver alone.** Every mesh name and every public name answered consistently, by one
|
||||
server. Rejected alone: with the anchor down, no machine resolves anything — the reason ADR 0194
|
||||
gave against it, and still true.
|
||||
2. **A local forwarder on every machine** — a small resolver on loopback that asks the mesh's resolver
|
||||
for the mesh's suffix and public resolvers for the rest, with `resolv.conf` naming only it. Rejected:
|
||||
it is ADR 0194's per-node stub again, which ADR 0196 removed — a daemon and a module on every node,
|
||||
and a container cannot use a loopback resolver, so containers would need a second configuration.
|
||||
Every resolution fault found on 2026-10-03 was a per-node copy disagreeing with the truth.
|
||||
3. **Split by kind of machine** — servers list the mesh's resolver alone, laptops keep a public
|
||||
fallback. Rejected: a laptop runs Alpine containers too, and a rule that differs by machine is a
|
||||
rule nobody can state about the mesh.
|
||||
4. **Two mesh resolvers, and nothing else listed.** The same module, the same roster and the same zones
|
||||
on two machines, and every machine lists both. Whichever answers first gives the same answer, so
|
||||
musl's race is harmless; a public name still resolves through either. Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**1. Now: the mesh has two resolvers.** The seat `mesh-dns-resolver` may be held on more than one
|
||||
machine — the anchor and the home server — each holder answering the same mesh names: the same
|
||||
module, the same machine list, the same zones, rendered by the controller into each.
|
||||
|
||||
- **A seat may be replicated.** It is an attribute of the seat in the mesh's definition, compiled with
|
||||
the set and never stored, as what a seat needs ([ADR 0220](0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md))
|
||||
and what it receives are. `mesh-dns-resolver` is the only replicated seat. Making another one is a
|
||||
decision, recorded. In the glossary's words it is the mesh's first **bench** — a seat whose
|
||||
holders coexist — and *replicated* says what kind: every holder answers the same thing.
|
||||
- **Every holder is on record, and each is added by an act**: `seat mesh-dns-resolver --add
|
||||
<node>/<module>` records one more holder beside those on record. Assigning the module is not enough:
|
||||
an assignment not on record stands beside the holders, eligible and silent, exactly as ADR 0131 says,
|
||||
and two claimants with nothing on record are refused as for any mesh seat. `--to` still hands the
|
||||
seat over, leaving exactly one holder. `--add` on a seat held once is refused, naming `--to`.
|
||||
- **One per machine still**: two modules on one machine claiming it are refused. And a seat held once
|
||||
stays held once: a second claimant on another machine is refused while nothing is on record, and a
|
||||
store recording two holders of such a seat is refused, naming the seat.
|
||||
- **Every machine's `/etc/resolv.conf` lists every holder's private address, the holder on the machine
|
||||
itself first if it is one, then the rest in name order, and no public resolver.** The controller
|
||||
gives a module's roster template the holders of each replicated seat, ordered so; the module holding
|
||||
`node-resolver-config` writes the file from it. On a holder, its own private address is first: the
|
||||
resolver listens there and on loopback, and the private address is the one a container on that
|
||||
machine can reach.
|
||||
- **The requirement stays.** What writes the file still requires `wildcard-resolution`, so a machine is
|
||||
refused when nothing in the mesh resolves, rather than given a file listing nothing. A holder answers
|
||||
its own requirement; any other machine is bound to the first holder by name. Nothing reads that
|
||||
binding's address any more — the file lists every holder — and the binding is kept for the refusal
|
||||
and the order of delivery.
|
||||
|
||||
**2. Next, decided and not yet built: `/etc/resolv.conf` belongs to the uplink's holder.** The program
|
||||
that manages the machine's network already has to be told to keep off the file
|
||||
([ADR 0117](0117-a-machines-uplink-is-a-seat.md)); instead, the `node-uplink` holder writes it, given
|
||||
the resolvers by the mesh: NetworkManager through its global DNS configuration, dhcpcd through static
|
||||
nameservers, and systemd-networkd's module declaring the file itself. `resolv-conf` and the seat
|
||||
`node-resolver-config` then retire, and ADR 0220's dependency of `node-resolver-config` on `node-uplink`
|
||||
goes with them. One owner for the file, and it is the program that would otherwise rewrite it.
|
||||
|
||||
> **Progressive insight — 2026-10-05.** Built, the managers' own mechanisms turned out not to write
|
||||
> the mesh's file: NetworkManager's global DNS configuration and dhcpcd's resolv.conf hook each write
|
||||
> `/etc/resolv.conf` in their own form — their own header, their own options line — and dhcpcd reads
|
||||
> its configuration only at its next start, so a change of resolvers would wait for one. The record
|
||||
> said NetworkManager and dhcpcd would be given the resolvers "through its global DNS configuration"
|
||||
> and "through static nameservers"; instead every one of the three modules declares the file itself,
|
||||
> from one template, and keeps its manager off it as before (`dns=none`, `nohook resolv.conf`, and
|
||||
> nothing for systemd-networkd). The decision — the uplink's holder owns the file, and `resolv-conf`,
|
||||
> `node-resolver-config` and ADR 0220's dependency retire — stands. How parts 2 and 3 are checked is
|
||||
> in [connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md).
|
||||
|
||||
**3. Next, decided and not yet built: a machine's names are one seat's.** `/etc/hosts` and
|
||||
`/etc/hostname` belong to one seat for the machine's identity. The `hosts` module, holding
|
||||
`node-hosts-file` ([ADR 0199](0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)),
|
||||
is renamed **`hostname`** and holds the seat renamed **`node-hostname`**; it writes `/etc/hostname`
|
||||
and the machine's own `127.0.1.1` line, and keeps every operator's line as ADR 0199 does today.
|
||||
|
||||
- **Why one seat for both files**: the mesh's only content in `/etc/hosts` is the machine's own name,
|
||||
and nothing owns `/etc/hostname` today. Two files saying one fact belong to one owner.
|
||||
- **Why `node-hostname`**: a system seat is named `node-` for its scope and then for its role
|
||||
([ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md)), and the role
|
||||
is the machine's name. `node-identity` was considered and rejected: a machine's identity in the mesh
|
||||
already means its key and certificate. `node-host` was rejected for the reason the module name
|
||||
`host` was: "the host" is what the [node-engine](../00-META/glossary.md) was called until now, and
|
||||
its repository and binary still carry `mesh-host`.
|
||||
- **Why a module and not part of the node-engine**: the node-engine applies every module's resources
|
||||
and owns no file's content; every file it writes belongs to the module that declared it. A file's
|
||||
content belongs to a seat's holder, and a seat's holder stays replaceable — another module can hold
|
||||
`node-hostname` on a machine that names itself another way, and the node-engine does not change.
|
||||
- The rename goes through the seat set's rename (an alias keeps `node-hosts-file` resolving,
|
||||
[ADR 0122](0122-a-seat-is-data-a-rename-is-a-database-update.md)), so nothing claiming the old name
|
||||
breaks in between.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **musl and glibc agree.** Every listed server answers a mesh name the same way, so the race musl
|
||||
runs has one possible outcome; a public name resolves through either holder's upstreams.
|
||||
- **Accepted cost: with no mesh resolver reachable, a machine has no DNS at all** until it can reach
|
||||
one. The anchor is the hub of the private network, so with it down the private network is down too,
|
||||
and the second resolver is reachable only on its own machine and from machines on its own LAN. The
|
||||
mesh assumes the anchor is up about 99.9% of the time. The same holds for a laptop behind a captive
|
||||
portal that keeps the tunnel down: it resolves nothing, the portal's own name included, until the
|
||||
tunnel is up.
|
||||
- **The resolver's options change from one attempt to two**, one second each. With no public resolver
|
||||
to fall back to, a single dropped datagram would otherwise fail a lookup on a machine that reaches
|
||||
only one resolver.
|
||||
- **Holdings are keyed by seat and assignment.** The store's holding table takes a numbered migration;
|
||||
every existing row is one per seat and satisfies the new key. A handover replaces every holder in one
|
||||
transaction.
|
||||
- **Unassigning a holder takes its own row only**: the other resolver keeps holding. Removing the
|
||||
second resolver is unassigning it, then `push --behind`.
|
||||
- **The rollout order matters.** The controller that knows replicated seats and renders the holders
|
||||
rolls out first; the catalogue's `resolv-conf`, which reads the holders, second — a controller without
|
||||
them cannot render it; and only then is the second holder added, because an older `resolv-conf`
|
||||
reading its one binding would name the first holder by name order, which may be the new one, and
|
||||
still list the public resolver.
|
||||
- **A container keeps the resolvers it started with.** A container on the default bridge copies its
|
||||
machine's `resolv.conf` when it starts; one on a user-defined network is answered by the runtime's
|
||||
embedded resolver, which forwards to the servers it read at start. Either keeps the public resolver
|
||||
until restarted.
|
||||
- **Parts 2 and 3 change who writes two files on every machine**, each a handover between modules on
|
||||
the same path; they wait for their own build handoff.
|
||||
|
||||
## How it is checked
|
||||
|
||||
| Rule | Checked by |
|
||||
|---|---|
|
||||
| `mesh-dns-resolver` is replicated, and no other seat is; it survives loading the set from the store | mesh-controller unit test over the compiled set and over a store row without the attribute |
|
||||
| Two holders on record compose on both, with no refusal, and each lists itself first | mesh-controller resolution and composition test with the catalogue's `dnsmasq` and `resolv-conf` |
|
||||
| A machine holding no resolver lists both holders, by name, and no public resolver | the same test on a third machine, asserting no public address appears |
|
||||
| A holder answers its own requirement although another holder sorts first | a resolution test on the second holder (issue 258's case kept) |
|
||||
| A seat held once still refuses a second claimant, and refuses two holders on record | resolution tests on `mesh-store` |
|
||||
| Two claimants of the replicated seat with nothing on record are refused, naming the handover | a resolution test |
|
||||
| Every consumer is bound to the same holder whatever order the mesh was resolved in | a unit test on the holder among providers |
|
||||
| The store keeps several holders of one seat, once each; a handover leaves one; unassigning takes only its own row | a store test against a live database, through the migration |
|
||||
| `seat --add` is refused for a seat held once | the controller's `seat` command |
|
||||
| The resolver's machine list has one host record per machine (issue 262's missing check) | a mesh-controller composition test counting host records |
|
||||
| `resolv-conf` lists only the seat's holders, with two short attempts | a mesh-controller test reading the catalogue's `resolv-conf` |
|
||||
| Live: every machine's `/etc/resolv.conf` lists both holders' private addresses, its own first on a holder, and nothing else; an Alpine container on the host network of the home server resolves `<anchor>.internal` every time | after rollout, read through each machine's tools, and `getent hosts` in an Alpine container repeated ten times |
|
||||
| Parts 2 and 3 | not yet built; their checks are written with their handoff |
|
||||
|
||||
## References
|
||||
|
||||
- [ADR 0194](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md) — the
|
||||
mesh's resolver; now held on two machines, each holding every node's internal domain.
|
||||
- [ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md) —
|
||||
replaced in part: its public resolver second goes. Every node and container asking the mesh's
|
||||
resolvers for every name, with no stub and no runtime `dns`, stands.
|
||||
- [ADR 0190](0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md) —
|
||||
several holders sharing a role, for a node seat.
|
||||
- [ADR 0131](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md) — holders on record, handover,
|
||||
eligible and silent.
|
||||
- [ADR 0220](0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md) — the
|
||||
uplink's dependency, which part 2 retires.
|
||||
- [ADR 0117](0117-a-machines-uplink-is-a-seat.md), [ADR 0199](0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md),
|
||||
[ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md),
|
||||
[ADR 0122](0122-a-seat-is-data-a-rename-is-a-database-update.md).
|
||||
- [Issue 258](../04-ISSUES/258-every-machine-bound-the-resolver-to-itself/00-report.md),
|
||||
[issue 262](../04-ISSUES/262-an-alpine-container-could-not-find-a-machine-by-its-mesh-name/00-report.md).
|
||||
- [Connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md) and [the seats](../03-DESIGN/01-to-be/26-the-seats.md),
|
||||
amended alongside.
|
||||
- mesh-controller `internal/catalogue/seats.go`, `resolve.go`, `roster.go`,
|
||||
`cmd/mesh-controller/seats.go`, `holdings.go`, and the store migration keying a holding by seat and
|
||||
assignment; mesh-catalog `modules/resolv-conf`, `modules/dnsmasq`.
|
||||
+134
@@ -0,0 +1,134 @@
|
||||
---
|
||||
topic: the mesh
|
||||
status: accepted
|
||||
date: 2026-10-06
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0040-what-a-module-is.md
|
||||
---
|
||||
|
||||
# 224. A provider that keeps failing a consumer is a problem the controller reports
|
||||
|
||||
## Context
|
||||
|
||||
**A provider's provisioner is the only thing in the mesh that knows whether a provision was made.**
|
||||
The controller composes a grant, delivers the contributions file and the minted secret, and the node
|
||||
reports that it applied every resource. Whether the provider then made the consumer's database, client
|
||||
or bucket is known to the provisioner loop alone ([ADR 0040](0040-what-a-module-is.md),
|
||||
[ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)), and until now it
|
||||
said so only in its own journal.
|
||||
|
||||
**On 2026-10-05 that cost a day.** The identity provider's database was moved that morning. Its
|
||||
realm's administrator kept the password it had before, the module's minted one was inert, and its
|
||||
provisioner failed every consumer every five seconds — about 31,000 refused logins from shortly after
|
||||
midnight until it was repaired by hand that night. Every surface the mesh has said the mesh was well:
|
||||
the machines applied what they were sent, the modules were current, `status` printed its all-well
|
||||
sentence. The same fault had been found and repaired by hand four days earlier
|
||||
([issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)).
|
||||
|
||||
[Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
|
||||
already taught that "the mesh and the machines agree" is not "it works", and `status` says so under its
|
||||
all-well line. This is the narrower, checkable half of that gap: not whether a consumer can reach its
|
||||
provider, which only dialling answers, but whether the provider has told the mesh it cannot do its job.
|
||||
|
||||
## Considered Options
|
||||
|
||||
1. **Leave it to the journal, and to a person reading it.** Rejected: that is what happened, twice.
|
||||
2. **The controller dials every provision.** Rejected for now: it is issue 145's open question, it
|
||||
needs the controller to hold or borrow every consumer's credential, and it would still not say
|
||||
*why* a provider fails.
|
||||
3. **The provider reports its standing through the node's report.** Rejected: the node-engine applies
|
||||
resources and knows nothing of what a module's code does after it starts; the provisioner runs in the
|
||||
node's tool runtime, which speaks to the bus, not to the host.
|
||||
4. **The provider announces, as an event, a consumer it keeps failing; the controller keeps the newest
|
||||
word and `status` names it.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**1. A provider announces a consumer it keeps failing.** When a provisioner has failed one consumer —
|
||||
its create, its periodic check, or reading the consumer's minted secret — for five minutes without a
|
||||
single success in between, it emits `provisioner.failing`, naming the consumer, the consumer's
|
||||
machine, the provision, the class of error and the error's first words, since when, and how many
|
||||
attempts. It says it again every fifteen minutes while it lasts. The first success after that is
|
||||
`provisioner.recovered`; so is a consumer the mesh stopped asking for, and so is the first success
|
||||
for each consumer after the provider starts, because a provider restarted after announcing a failure
|
||||
has forgotten it.
|
||||
|
||||
- **The classes** are `credentials-rejected`, `unreachable`, `secret-unreadable` and `refused` (any
|
||||
other refusal), read from the error's words unless the provider's own code classes it better. They are
|
||||
what a person reading `status` needs before opening the journal.
|
||||
- **No secret travels.** The error is the provider's own text with the consumer's password removed, as
|
||||
its log line already is.
|
||||
|
||||
**2. Every provider may say it, whatever its manifest lists.** The permission to publish the two events
|
||||
is derived for every module that receives contributions; no manifest declares them. A provider whose
|
||||
manifest forgot them would otherwise have its announcement refused by the bus and fail as silently as
|
||||
before.
|
||||
|
||||
**3. The controller follows both events from every module, keeps the newest failing word per provider
|
||||
module, its machine and the consumer, and removes it on recovery.** Who said it is read from the subject
|
||||
the bus let the provider publish on, never from the body. A recovery that arrives while the store is
|
||||
away is held and delivered again, because it is said once. The controller's subscription names the two
|
||||
events with a wildcard for the emitter — the one pattern on its list — rather than a list of providers
|
||||
somebody would have to extend.
|
||||
|
||||
**4. `status` names every consumer a provider still assigned where it ran says it keeps failing**, and
|
||||
such a consumer breaks the all-well sentence. Its JSON carries them as `failing`; `node show` lists those
|
||||
whose provider or consumer is on that machine. A provider no longer assigned there has nothing running
|
||||
to fail anybody, and is not asked about. A standing not said again for thirty minutes is shown with how
|
||||
long it has been silent: the provider stopped saying anything, and its last word is all the mesh has.
|
||||
|
||||
**5. Where a provider's failure has one known cause it can repair safely, it repairs it.** The identity
|
||||
provider's administrator refusing the mesh's secret is the first: the module now checks the
|
||||
administrator's login on start, every five minutes and whenever its provisioner is refused, and repairs
|
||||
a refusal through the server's own bootstrap command — a temporary administrator, the real one's
|
||||
password set to the mesh's, the temporary one removed, nothing printed — then checks again and says
|
||||
what it did. A repair that fails is braked, from ten minutes doubling to six hours, and announced; while
|
||||
the administrator is refused the provisioner stops asking the server, so a lockout policy is not
|
||||
provoked, and its consumers are still announced failing. The mechanism is the module's
|
||||
([issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)),
|
||||
the rule is this record's: **detected automatically, repaired where safe, loud where not.**
|
||||
|
||||
## Consequences
|
||||
|
||||
- **The loop is carried by every Go provider until the Go SDK has one.** The provisioner loop is the
|
||||
SDK's ([ADR 0039](0039-what-the-sdk-holds-and-refuses.md)); the Go SDK has none yet, so the two Go
|
||||
providers carry identical copies and a test in each fails when they differ. A TypeScript provider
|
||||
announces nothing until the TypeScript SDK's loop does the same; until then its consumers fail as
|
||||
silently as before, and the next provider to be ported to Go closes that gap for itself.
|
||||
- **The controller's event consumer takes one message at a time**, so a standing can wait behind a
|
||||
build being acted on for some minutes. Fifteen-minute repetition makes that harmless for a failure;
|
||||
a recovery is held, never dropped.
|
||||
- **The store keeps one row per provider module, its machine and consumer**, through a numbered
|
||||
migration.
|
||||
- **The rollout order matters.** The controller that derives the permission and follows the events
|
||||
first; the catalogue's providers second — an older controller refuses nothing that matters, but
|
||||
every announcement is then refused by the bus and logged by the provider.
|
||||
|
||||
## How it is checked
|
||||
|
||||
| Rule | Checked by |
|
||||
|---|---|
|
||||
| A consumer failed for five minutes without a success is announced, again every fifteen, and recovered on the first success | provider loop tests in mesh-catalog, postgres and keycloak (`standing_test.go`) |
|
||||
| A failing periodic check and an unreadable secret count, not only a failing create | the same tests |
|
||||
| A withdrawn consumer and the first success after a start are announced recovered | the same tests |
|
||||
| The two Go providers carry the same loop | `harness_same_test.go` in each, comparing the files |
|
||||
| Every module that receives contributions is granted the two events, and no other module is | mesh-controller inventory test over the derived declaration and the bus permissions |
|
||||
| The controller may hear the two events from any module, and not every event | the same test over the controller's permissions |
|
||||
| The emitter is read from the subject; a failing word is kept, a recovery cleared, and a recovery held while the store is away | mesh-controller link tests with a fake store |
|
||||
| One row per provider, machine and consumer; recovered removes only its own | an inventory test against a live database, through the migration |
|
||||
| `status` names it and is not all-well; the JSON carries it; `node show` names it on both machines; an unassigned provider's word is not a problem | mesh-controller command tests against a live database |
|
||||
| The identity provider repairs a refused administrator, verifies, brakes a failed repair, never puts a secret in a command line, and stops asking the server while refused | keycloak module tests with a fake server and a fake container runtime; a live test against a throwaway server when asked for |
|
||||
| Live: after rollout, `status` shows nothing failing on a healthy mesh; a provider made to fail a consumer for five minutes appears, and disappears on its next success | read through the console after rollout |
|
||||
|
||||
## References
|
||||
|
||||
- [Issue 179](../04-ISSUES/179-an-adopted-identity-providers-admin-never-took-the-minted-secret/00-report.md)
|
||||
— the failure, twice, and the repair this record makes automatic.
|
||||
- [Issue 145](../04-ISSUES/145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md)
|
||||
— the scope of the all-well sentence, which this narrows and does not close.
|
||||
- [ADR 0040](0040-what-a-module-is.md), [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md),
|
||||
[ADR 0039](0039-what-the-sdk-holds-and-refuses.md), [ADR 0042](0042-the-shape-of-an-event-on-the-wire.md).
|
||||
- [The module protocol](../03-DESIGN/01-to-be/19-the-module-protocol.md), provisioning, amended alongside.
|
||||
- mesh-controller `internal/link/standing.go`, `internal/inventory/standing.go` and its migration,
|
||||
`cmd/mesh-controller/standing.go`; mesh-catalog `modules/postgres` and `modules/keycloak`.
|
||||
@@ -0,0 +1,148 @@
|
||||
---
|
||||
topic: what runs on it
|
||||
status: accepted
|
||||
date: 2026-10-06
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md
|
||||
---
|
||||
|
||||
# 225. A consumer's identity is bounded by the provision it requires, judged before merge, and never refuses its provider
|
||||
|
||||
## Context
|
||||
|
||||
[ADR 0049](0049-a-consumers-identity-fits-the-tightest-backend.md) bounds every consumer's identity,
|
||||
`mesh_<machine>_<slug-or-module>`, by the tightest backend anywhere in the mesh: an S3 access key's 20
|
||||
characters. It chose that over per-interface bounds (its option C) because one constant unblocked the
|
||||
object store, and named C as the refinement "if non-S3 consumers are paying for S3's limit often
|
||||
enough to mind". [Issue 263](../04-ISSUES/263-every-consumer-pays-for-the-tightest-backends-name-limit/00-report.md)
|
||||
is that point. The operator: "the 20 character limit has bitten us multiple times".
|
||||
|
||||
What happened on 2026-10-06, and what it shows:
|
||||
|
||||
- **The bound applied where it meant nothing.** A change made the network-manager modules require the
|
||||
mesh's resolver provision. The resolver mints no credential and keeps no name: its provider is not
|
||||
even told who its consumers are (it receives nothing). `networkmanager`'s identity, 23 to 26
|
||||
characters on real machine names, was refused all the same.
|
||||
- **It was found late and far from its cause.** The catalogue's module check and the controller's
|
||||
tests passed. ADR 0049 says the refusal comes at assignment; this was a new requirement on modules
|
||||
already assigned, so no assignment saw it. It surfaced when the provider composed its grants.
|
||||
- **It refused the wrong machine, wholly.** The controller judged the identity while composing the
|
||||
*provider's* declaration, and one refusal there fails the whole composition. The anchor holds the
|
||||
resolver, so the anchor — every module on it — could not be pushed until a slug was changed
|
||||
elsewhere.
|
||||
|
||||
The catalogue's providers keep very different names. Read from each provider's code: the object store
|
||||
keeps the identity as an access key (20) and a bucket name; PostgreSQL as a role and a database (63);
|
||||
MongoDB as a database (63); SQL Server as a login and a database (128); the identity provider as a
|
||||
client id (255); the forge as a user name (40); the mail server as a mailbox's local part (64); the
|
||||
public DNS provider as a label (63); the cache, the vault, the message broker and the time-series
|
||||
store as names with no limit worth stating. The resolver and both route providers keep no name of
|
||||
their consumers at all.
|
||||
|
||||
## Considered Options
|
||||
|
||||
1. **Keep one bound, raise or lower it.** Any single number is wrong for most provisions: 20 refuses
|
||||
a database consumer for an object store's key, 63 lets the object store fail at provision time
|
||||
again ([issue 034](../04-ISSUES/034-mesh-login-exceeds-s3-access-key-limit/00-report.md)). Rejected —
|
||||
it is the cause.
|
||||
2. **Per-interface bounds in a mesh-wide table of provisions.** Puts the fact in the controller, which
|
||||
would then know what an S3 key is. The mesh is name-agnostic about what a provider does with an
|
||||
identity (`ConsumerIdentity`); a table would make it learn every backend. Rejected.
|
||||
3. **Each offer states its own bound; the bound a consumer meets is that of the provision it
|
||||
requires** (ADR 0049's option C, placed where [ADR 0202](0202-a-provider-declares-what-it-derives-for-each-consumer.md)
|
||||
already places what a provider derives: in its own definition). Chosen.
|
||||
4. **Derive a different identity per provision** (C's own "against": one consumer, several names).
|
||||
Not needed: the identity stays one name, said once, and must fit every provision the module
|
||||
requires. Only the *bound* is per provision. Rejected as unnecessary.
|
||||
|
||||
And for where the refusal lands:
|
||||
|
||||
5. **Keep refusing the provider's composition.** Rejected — it makes one consumer's name a reason
|
||||
no push reaches a machine that did nothing wrong.
|
||||
6. **Refuse the consumer's whole machine.** Right machine, still too wide, and still late. Not chosen;
|
||||
the consumer is named instead, and refused in its own pull request (below).
|
||||
|
||||
## Decision
|
||||
|
||||
**1. An offer states the longest consumer identity its backend keeps.** In the provider's
|
||||
definition, beside the provision: `identity` with a `max` and the words for what keeps it (`in`), so a
|
||||
refusal can say "an S3 access key keeps 20"; `in` alone for a backend with no limit worth stating; or
|
||||
`false` for a provision that keeps no name derived from its consumer. The bound applied to a consumer
|
||||
is that of the provision it requires, from the offer of the module answering it — not the tightest
|
||||
backend in the mesh.
|
||||
|
||||
**Unsaid, the bound follows from what the provider is told.** A provider that receives the provision,
|
||||
or serves its consumers a value built from their identity, is told who each consumer is and may make
|
||||
a name of it in a backend nobody measured: it keeps ADR 0049's 20. A provider told neither keeps
|
||||
nothing of its consumers: no bound. A provider whose definition is not at hand is held to 20.
|
||||
|
||||
**2. An overflow is refused before merge.** The catalogue check (`module check`, and the controller's
|
||||
test over the real catalogue) judges every module's identity, built on the longest machine name of
|
||||
the mesh, against the bound of every provision it wants that a module in the catalogue offers — the
|
||||
tightest where several offer it. The pull request that introduces an overflow — a new requirement, a
|
||||
lowered bound, a longer module name — is the one that fails, naming the module and the longest slug
|
||||
that would fit. The longest machine name is a parameter of the check; its default is the mesh's own
|
||||
longest, raised in the same change that names a longer machine.
|
||||
|
||||
**3. One consumer's identity never refuses its provider's machine.** When the provider's declaration
|
||||
is composed, a consumer whose identity overflows the bound is left out of the grants and returned
|
||||
beside them; every other consumer is granted and the declaration composes. The consumer is named
|
||||
where an operator looks: on the push and plan of the provider, on the plan of the consumer's own
|
||||
machine, and in `status` (and its document), which does not call the mesh well while one stands.
|
||||
|
||||
**4. An identity served as a DNS label stays inside one.** The mesh writes an identity into a label
|
||||
without truncating it ([ADR 0202](0202-a-provider-declares-what-it-derives-for-each-consumer.md)),
|
||||
so an offer that serves `${consumer:as:dns}` must bound its consumers at 63 or less.
|
||||
|
||||
What does not change: the identity is still derived once and said to both ends (ADR 0049, issue 023);
|
||||
the slug is still the remedy; truncation and hashing are still refused.
|
||||
|
||||
## Consequences
|
||||
|
||||
- A consumer of a database or the identity provider keeps a legible name on a long machine name; only
|
||||
a consumer of the object store still needs a slug of a few characters. Requiring a keyless
|
||||
provision costs nothing in name length.
|
||||
- The catalogue's providers each state a bound in their offer, and the check reads it there. A new
|
||||
provider that is told its consumers and says nothing keeps the old 20: safe, and visible in review.
|
||||
- A consumer left out of the grants holds a binding and credential its provider never created. It
|
||||
fails to authenticate; that is reported by name in `status` until a slug fixes it, rather than
|
||||
hidden behind a machine nobody could push.
|
||||
- The check's default machine-name length is a fact about one mesh carried in code. A mesh that
|
||||
names a longer machine must raise it, or pass its own to `module check`; nothing reminds it to.
|
||||
- Old controllers refuse the new field (manifests are parsed strictly), so the controller ships before
|
||||
the catalogue that states bounds.
|
||||
- ADR 0049's consequence "`identityLimit` becomes 20 … `CheckIdentity` refuses at `module add` /
|
||||
assignment" now holds only for a provision that keeps 20, and is checked before merge and reported
|
||||
at composition rather than refused at assignment. ADR 0049 carries a note saying so.
|
||||
|
||||
**How each rule is checked.**
|
||||
|
||||
- *Rule 1:* catalogue unit tests — an offer's stated bound is the one applied; an unstated one is 20
|
||||
for a provider told its consumers and none for one told nothing; `false` is none; the field parses
|
||||
strictly and a bound without `in`, or shorter than any identity, is refused when the definition is
|
||||
parsed. Over the real catalogue: the resolver provision bounds nothing and the object store 20.
|
||||
- *Rule 2:* `module check` refuses a module whose identity overflows what it requires, naming the
|
||||
slug length, and passes a long name requiring a keyless provision; a test runs the same judgement
|
||||
over every manifest in the real catalogue on the default machine-name length, so the catalogue's
|
||||
own pull request fails on an overflow.
|
||||
- *Rule 3:* a controller test against a real store reproduces the night it was found — the network
|
||||
manager, no slug, on a six-character machine, requiring the resolver provision, beside a consumer
|
||||
that overflows an object store: the provider's declaration composes, the network manager is granted,
|
||||
the overflowing consumer is left out, named by the push, carried by `status` and its document, and
|
||||
the mesh is not called well.
|
||||
- *Rule 4:* a test that a DNS-label identity under a bound over 63 is refused when the definition is
|
||||
parsed.
|
||||
|
||||
## References
|
||||
|
||||
- [Issue 263](../04-ISSUES/263-every-consumer-pays-for-the-tightest-backends-name-limit/00-report.md) —
|
||||
the observation.
|
||||
- [ADR 0049](0049-a-consumers-identity-fits-the-tightest-backend.md) — the bound this refines; its
|
||||
option C.
|
||||
- [ADR 0202](0202-a-provider-declares-what-it-derives-for-each-consumer.md) —
|
||||
a provider declares what it derives; the bound is declared beside it.
|
||||
- `mesh-controller` `internal/catalogue/identity.go` (`IdentityBound`, `CheckIdentityWithin`,
|
||||
`IdentityProblems`, `Resolution.Overflowing`), `manifest.go` (`Offer.Identity`,
|
||||
`IdentityBoundOf`), `cmd/mesh-controller/plan.go` (`grantsFor`), `check.go`, `status.go`.
|
||||
- `mesh-catalog` — each provider's offer states its bound.
|
||||
+217
@@ -0,0 +1,217 @@
|
||||
---
|
||||
topic: the mesh
|
||||
status: accepted
|
||||
date: 2026-10-06
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md
|
||||
---
|
||||
|
||||
# 227. The core holds nine rules, each checked, and is built to them in six phases
|
||||
|
||||
## Context
|
||||
|
||||
**The core is the controller, the node-engine and its launcher, the bus server, the node tools and the
|
||||
console, the build seat, and the forge's announcer of merges** — everything a change passes through
|
||||
before a module's own code runs. On 2026-10-06 the operator asked for it to be *"fully diagnosable, with
|
||||
active monitoring, self-healing, self-upgradeable, self-monitoring … very sturdy, no ambiguities, clear
|
||||
plan of execution, fail-proof setup"*, and approved [research 031](../01-RESEARCH/031-a-core-that-cannot-fail-silently/00-overview.md)'s
|
||||
conclusion the same day: *"do the research and implement it"*.
|
||||
|
||||
The evidence is [031/01](../01-RESEARCH/031-a-core-that-cannot-fail-silently/01-evidence.md):
|
||||
|
||||
- **92 issue reports in six days; 48 of them core failures.** Every one of the 48 was noticed because a
|
||||
person or an agent looked. **None was raised by the mesh unasked.** In four the mesh's own answer
|
||||
carried the fact for whoever asked; in none did it tell anybody.
|
||||
- **Four faults came back through a different door after their first fix** (200 → 265, 230 → 264,
|
||||
257 → 261 → 267, 175 → 184 → 248): each fix closed an instance and left the class open.
|
||||
- **The classes, by count:** races between actors with no explicit order (10 issues); commands or
|
||||
arguments dropped with a default chosen in their place (8, two of them destructive —
|
||||
[241](../04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md)
|
||||
dropped seven databases on one unreadable file,
|
||||
[244](../04-ISSUES/244-a-verb-whose-schema-is-empty-cannot-be-called-through-the-console/00-report.md)
|
||||
pushed every machine when one was named); outcomes never fed back (8); failures visible only as log
|
||||
lines (9, the longest twenty-three hours); state with two writers (7); repairs by hand (at least
|
||||
fifteen acts, the commonest a push by hand to unstick a plan waiting on a report); self-upgrade
|
||||
breaking the core (10); checks that pass in CI and fail on the mesh's real facts (6); third-party
|
||||
faults nobody compared against outcomes (3, one of them
|
||||
[266](../04-ISSUES/266-a-merge-on-the-bus-was-never-handed-to-the-controller/00-report.md):
|
||||
23 merges unacted on over three days).
|
||||
- **The parts mostly exist, one rule per message kind.** Declarations carry a sequence; reports do not.
|
||||
Builds and plans are ordered; controllers have no epoch. `calls` keeps outcomes, in memory, and a
|
||||
controller restart — every merge to its own repository — forgets them. ADR 0224 made one failure
|
||||
kind a problem `status` reports and repairs where safe; nothing generalises it.
|
||||
|
||||
**Checked against GENESIS.** The mission's core value *failure must be loud — prefer failing to lying*
|
||||
is the rule this record enforces on the core. The context says *human agents are few, often one, and
|
||||
usually asleep; anything requiring a human to notice it will be noticed late* — the 48-of-48 count is
|
||||
that sentence measured. The effect promises that *the mesh notices when something is wrong before you
|
||||
do* and asks *with the context, not a log line*. Nothing in 031 conflicts with GENESIS; the effort is
|
||||
the gap between the effect and the as-is, counted.
|
||||
|
||||
## Considered Options
|
||||
|
||||
1. **Keep fixing issues one at a time.** Rejected: that is what the window did, and four classes
|
||||
recurred through a different door. Point fixes converge on instances, not classes, and the next
|
||||
message kind or the next reader of an unreadable file starts from nothing.
|
||||
2. **Monitoring from outside: a metrics stack with alert rules** (an exporter per component, a time
|
||||
series store, an alert manager). Rejected: it observes symptoms the core would still keep to itself
|
||||
(a report never sent is not a metric), it is a second answer to questions the controller already
|
||||
answers (how-we-build §5: a report is read from the system), it adds three services to a mesh with one
|
||||
operator, and it repairs nothing. The mesh's facts are in the controller; the watcher belongs there,
|
||||
with one watcher of the watcher outside it.
|
||||
3. **Prevention first: order and one writer before anything else.** Rejected as the *first* step, kept
|
||||
as the second. Prevention covers the classes already met; detection covers every class including
|
||||
those not met yet, and is cheaper per day. Phase 1 makes the mesh say when it is wrong; Phase 2
|
||||
removes the largest class.
|
||||
4. **Heal everything by default.** Rejected: a default of healing heals what is not understood, which is
|
||||
how a repair destroys — 241's reconcile was, in its own terms, healing. Healing is narrowed to
|
||||
*known* failures, ones repaired by hand twice, under a brake.
|
||||
5. **A three-server bus cluster, so the bus can be upgraded live.** Not decided here: it is a change of
|
||||
the foundation's shape and needs its own effort. Until then a bus upgrade is a planned, announced
|
||||
step (rule 8).
|
||||
6. **The lab as the place every core change is proven.** Rejected by [ADR 0149](0149-the-live-mesh-is-the-test-bed.md),
|
||||
which stands: the live mesh is the test bed. Lab replays are kept for what must not be done to the
|
||||
live mesh on purpose — two controllers at once, a deliberately broken controller build, a suppressed
|
||||
signal — and each rule still has a live check.
|
||||
7. **Nine principles as stated, the mechanisms M1–M9 and the phased roadmap of 031/03.** Chosen.
|
||||
|
||||
## Decision
|
||||
|
||||
**The nine rules below hold for the core. A change to a core repository is refused in review if it
|
||||
breaks one, and each rule is checked as its row in *How it is checked* says.** They extend, and do not
|
||||
replace, [research 017](../01-RESEARCH/017-a-mesh-that-heals-itself/01-the-intended-behaviour.md)'s
|
||||
six for the loops that converge modules.
|
||||
|
||||
1. **One writer per piece of state.** Every kind of state the core keeps has exactly one writer, named
|
||||
in the writers table of [to-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md).
|
||||
Anyone else asks that writer. Nobody computes a second answer to a question it already answers. The
|
||||
controller holds a **lease**; only the holder acts.
|
||||
2. **Everything that changes state carries its writer's order, and every receiver refuses what is
|
||||
older.** Declarations, reports, plans, builds, calls and announcements carry writer, epoch and
|
||||
sequence. A receiver keeps the highest it accepted per writer and refuses anything older, with a line
|
||||
in the mesh's words and a counter. Arrival order never decides.
|
||||
3. **Every command is answered at once, and its outcome is kept where it can be read later.** Within
|
||||
the verb's declared bound, with a result or a call id. The outcome is durable: it outlives the process
|
||||
that ran it, and is read or waited on by id. No answer depends on what the command does to the
|
||||
transport carrying it.
|
||||
4. **Nothing is dropped silently.** An input that is unknown, unreadable or unmet is refused by name —
|
||||
which input, where, why. A default is never substituted for *I could not tell*, above all never
|
||||
"empty". A reconcile that would withdraw more than a bound of what it holds stops and raises a
|
||||
condition instead.
|
||||
5. **Every expected signal has a watchdog; absence is a condition.** Every signal the core expects on
|
||||
a cadence or after an act has a row in the signals table: emitter, trigger, bound, condition raised.
|
||||
Silence past the bound is a condition naming what was expected, from whom, since when.
|
||||
6. **The mesh checks itself continuously against live facts and says what it found outward.** The
|
||||
design's invariants are probes run on a schedule against the running mesh (`doctor`). A violation is
|
||||
a **condition** — durable, with since-when, evidence and who can resolve it — shown in `status` and
|
||||
sent to the operator through the output channel. The checker's heartbeat is watched from a machine
|
||||
that is not the control node, through a channel that does not pass through it.
|
||||
7. **A known failure heals itself, under a brake, and every repair is said.** A failure repaired by
|
||||
hand twice gets a healer: the ordinary path again, never a destructive act, with a budget and a
|
||||
back-off, and one event saying what it did. A spent budget, or a repair that could only destroy, is
|
||||
a condition, not a retry. Every repair a person makes on the core goes through a verb that records
|
||||
who, what and why (the **hand-act log**).
|
||||
8. **The core upgrades itself one machine at a time, health-gated, and rolls back on its own.** A new
|
||||
controller, node-engine or node tools build is judged on its first machine by that component's
|
||||
health probes, not by "reported applied". One not healthy within its bound is rolled back to the
|
||||
last known good **by something other than itself**, and the rollback is a condition. The component
|
||||
being replaced is never the only witness of its successor. A bus upgrade is a planned, announced
|
||||
maintenance step, never a plain rollout.
|
||||
9. **A check is fed the real mesh's facts before a change is merged.** A check whose verdict depends
|
||||
on the environment runs against a facts snapshot exported by the controller — every machine
|
||||
composed with the change and validated by the node-engine's validator — and a dependency's version
|
||||
the mesh runs is the version its tests run.
|
||||
|
||||
**The plan of execution is the six phases of to-be 45**, in that order, each ending at its own *done
|
||||
when*:
|
||||
|
||||
| Phase | What it delivers | Rules |
|
||||
|---|---|---|
|
||||
| 0 — finish what is in flight | the located core fixes rolled out; `calls` durable; `status` inside its bound; the hand-act log; the bus's planned upgrade as the first maintenance step; the durations the bounds are measured from | 3, 7, 8 |
|
||||
| 1 — the mesh says when it is wrong | conditions; watchdogs from the signals table; the bus's advisories; `doctor`; the output channel in its minimal form and the second-machine watcher | 5, 6 |
|
||||
| 2 — order and one writer | the controller's lease and epoch; a report's sequence; one apply queue on every machine; stale refusals counted; the writers table enforced; the empty-on-error lint; the withdrawal brake | 1, 2, 4 |
|
||||
| 3 — healers | the first healers, each braked; a repeated hand act asks for one | 7 |
|
||||
| 4 — core upgrades that roll back | a health definition per core component; the gate; rollback by a witness; the bus as a planned step | 8 |
|
||||
| 5 — checks before merge, and replays | the facts snapshot and the compose-and-validate merge gate; versions tested as run; every core incident a replay | 9, and all |
|
||||
|
||||
**The output channel is built in the smallest form research 028 allows.** One mesh seat,
|
||||
`operator-channel`, accepting `notify` as a work queue and holding its open messages in its own state
|
||||
([028 Q1](../01-RESEARCH/028-the-meshs-output-channel/03-open-questions.md), option a); the controller
|
||||
emits condition events and the holder decides what is sent (028 Q5, option a); the Telegram channel the
|
||||
operator required and the desktop notifier as the two channels; deduplicated by the condition's key; no
|
||||
answering back except through the mesh's own verbs; a message carries roles and words, never an
|
||||
address, a path or a secret, refused by the holder otherwise (028 Q8). Routing by presence, quiet hours,
|
||||
answering back and the external dead-man service stay open in 028, whose graduation amends to-be 45.
|
||||
|
||||
**ADR 0224's provider standing becomes the first condition kind**, unchanged in what it says and
|
||||
when; its storage moves into the condition store.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **`status` changes meaning.** It becomes, first, the list of open conditions; the all-well sentence
|
||||
is "no open conditions", silenced ones included. A silenced condition is still open; silence stops
|
||||
only its messages, for a stated time, with a reason, recorded as a hand act. Nobody resolves a
|
||||
condition by hand: it clears when observation says so.
|
||||
- **The bounds are measured, not guessed.** The signals table's first bounds are provisional; Phase 0
|
||||
records the durations they are set from, and Phase 1's first live week corrects every bound that
|
||||
raised a condition that was not real. A corrected bound is a change to the table, reviewed like code.
|
||||
- **The controller's restart stops being a forgetting.** Calls, conditions, the hand-act log and the
|
||||
lease live in key-value buckets ([ADR 0201](0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md));
|
||||
a new holder of the lease marks a call the old one left running as abandoned, which is said.
|
||||
- **The node-engine gains an order it did not have.** One apply queue replaces the delivery path and
|
||||
the five-minute reconcile as two appliers; a report carries the declaration's sequence; a declaration
|
||||
from an older controller epoch is refused. Issues 257, 261 and 267 become impossible rather than
|
||||
handled. The controller–node-engine wire changes, so the rollout order is the node-engine first
|
||||
(it accepts both shapes), the controller second.
|
||||
- **More is checked before merge, and merges get slower.** The compose-and-validate gate runs every
|
||||
machine through the validator on every core and catalogue merge. That is minutes, against the hours
|
||||
each of 202, 236 and 263 cost.
|
||||
- **A third party now carries operational words.** Telegram is not end-to-end encrypted for bots, so the
|
||||
holder's content rule is the only thing between a condition's evidence and someone else's server. It
|
||||
is enforced by the holder and tested there, not trusted to each source.
|
||||
- **The watcher outside the control node holds the channel's secret.** That is one more machine with a
|
||||
bot token, accepted for the one case nothing else covers: the control node or the bus is what failed.
|
||||
- **The roadmap is six to eight weeks of focused work.** The first visible change — the mesh saying
|
||||
when it is wrong — is inside the first two. Until Phase 2 lands, races are still caught by watchdogs,
|
||||
not prevented.
|
||||
- **The rules are not yet in how-we-build.** They are this record's until they are carried into
|
||||
how-we-build §2 through playbook 05 (constitution sync), which is a separate change.
|
||||
|
||||
## How it is checked
|
||||
|
||||
| Rule | Checked by | From |
|
||||
|---|---|---|
|
||||
| 1. one writer | the writers table in to-be 45; a test per core repository that the code paths writing each kind are the named writer's; the controller refusing, at composition, a bus subject two components may publish on unless the table says it is shared; live, the `doctor` probe *exactly one lease holder, no stale-epoch message in the last interval* | Phase 2 |
|
||||
| 2. order | a contract test per consumed message kind in the receiver's repository (deliver *n*, then *n−1*: refused and counted; epoch *e−1* after *e*: refused); a check listing every consumed subject against the tests that name it; live, the stale-refusals signal | Phase 2 |
|
||||
| 3. answered, kept | a test walking every verb the controller announces (answers within its bound, with a result or an id); a test restarting the controller between a call and the read of its outcome; live, the `doctor` probe *status answers in full within ten seconds* and the *call hung* signal | Phase 0 |
|
||||
| 4. nothing dropped | per component, a test feeding each input reader an unreadable, malformed and foreign input (a refusal, never an empty result); a lint in each core repository refusing an error branch that returns an empty collection; schema walks over every verb and placeholder namespace; a test emptying a provider's input and asserting nothing is withdrawn | Phase 2 |
|
||||
| 5. watchdogs | a test generated from the signals table that suppresses each signal in turn and asserts its condition is raised within its bound and cleared when it returns; live, `doctor` reporting the age of the newest signal of every row | Phase 1 |
|
||||
| 6. self-check, outward | a check over the to-be designs counting invariants with a live probe against those without, which may only go down; the second-machine watcher raising *self-check silent* when the controller is stopped; once, on a lab mesh, a broken invariant appearing in `status` and as a message within one probe interval | Phase 1 |
|
||||
| 7. healers | a test per healer inducing its failure, asserting the repair, the event and the brake after the budget; the hand-act log's weekly count in `status`; a cause recorded twice raising *healer wanted* | Phase 3 |
|
||||
| 8. staged upgrades | on a lab mesh, a broken build of the controller, the node-engine and the node tools (one that starts and does nothing, one that crashes, one that cannot reach the bus) each rolled back with no hand, ending on the previous build, said as a condition and a message; live, every core rollout's record (first machine, verdict, time to verdict, rolled back or not) readable through `plans` | Phase 4 |
|
||||
| 9. real facts | the merge gate in mesh-controller, mesh-host and mesh-catalog failing a change that makes any machine of the snapshot fail to compose or validate, naming the machine's role and the module; a test that a pinned dependency's version equals the version in the snapshot; the replays of 236, 262, 263 and 266 failing on the commit before their fix | Phase 5 |
|
||||
| the plan | each phase's *done when* in to-be 45, recorded in that design when met, with the design's status following | each phase |
|
||||
|
||||
## References
|
||||
|
||||
- [Research 031](../01-RESEARCH/031-a-core-that-cannot-fail-silently/00-overview.md) — the evidence,
|
||||
the principles with the candidates not kept, the mechanisms and the roadmap.
|
||||
- [To-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) — the design this record
|
||||
authorises.
|
||||
- [Research 017](../01-RESEARCH/017-a-mesh-that-heals-itself/00-overview.md) — the condition, and
|
||||
the six principles for the loops; [research 028](../01-RESEARCH/028-the-meshs-output-channel/00-overview.md)
|
||||
— the output channel, of which the minimal form is taken here.
|
||||
- [ADR 0224](0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md) —
|
||||
*detected automatically, repaired where safe, loud where not*, generalised here.
|
||||
- [ADR 0141](0141-the-host-delivers-its-own-successor.md), [ADR 0162](0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md),
|
||||
[ADR 0218](0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md),
|
||||
[ADR 0149](0149-the-live-mesh-is-the-test-bed.md), [ADR 0201](0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md).
|
||||
- Issues [187](../04-ISSUES/187-the-mesh-tells-nobody-when-it-stops-working/00-report.md),
|
||||
[204](../04-ISSUES/204-a-controller-handover-re-sent-every-node-a-stale-declaration/00-report.md),
|
||||
[241](../04-ISSUES/241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md),
|
||||
[248](../04-ISSUES/248-the-controllers-event-consumer-replayed-a-week-and-held-every-merge-behind-it/00-report.md),
|
||||
[264](../04-ISSUES/264-a-self-updating-engine-lost-the-report-of-the-apply-that-delivered-it/00-report.md),
|
||||
[265](../04-ISSUES/265-a-push-outlived-its-caller-and-its-answer-was-refused/00-report.md),
|
||||
[266](../04-ISSUES/266-a-merge-on-the-bus-was-never-handed-to-the-controller/00-report.md),
|
||||
[267](../04-ISSUES/267-a-reconciles-report-overtook-the-apply-that-followed-it/00-report.md).
|
||||
@@ -194,6 +194,12 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0207** — [A module depends on the node seats that apply its resources](0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md)
|
||||
- **0210** — [A tool's configuration is its seat holder's, and every other module extends it through the seat](0210-a-tools-configuration-is-its-seat-holders-and-every-other-module-extends-it-through-the-seat.md)
|
||||
- **0212** — [A seat says what it receives, and the machine's hotkeys are a seat](0212-a-seat-says-what-it-receives-and-the-machines-hotkeys-are-a-seat.md)
|
||||
- **0218** — [A plan sends grants before code, rolls a module out one machine first, and a newer merge takes over an older plan](0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md)
|
||||
- **0219** — [The build queue is controlled through the controller and the build seat](0219-the-build-queue-is-controlled-through-the-controller-and-the-build-seat.md)
|
||||
- **0221** — [A push sends no build a policy or a plan holds back, except to the machine it names](0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md)
|
||||
- **0222** — [A module is told where a mesh seat's holder is reached, and the controller writes no file a seat's holder owns](0222-a-module-is-told-where-a-mesh-seats-holder-is-reached-and-the-controller-writes-no-file-a-seats-holder-owns.md)
|
||||
- **0224** — [A provider that keeps failing a consumer is a problem the controller reports](0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md)
|
||||
- **0227** — [The core holds nine rules, each checked, and is built to them in six phases](0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md)
|
||||
|
||||
### Its tiers, from the bottom up
|
||||
|
||||
@@ -231,6 +237,7 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0194** — [The mesh has one resolver, and every node asks it for the mesh's names](0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)
|
||||
- **0196** — [A node asks the mesh's resolver first, and a public one only when it is silent](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md)
|
||||
- **0199** — [A module that answers names declares its zone, and a node's hosts file is one module's](0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)
|
||||
- **0223** — [The mesh has two resolvers, and a machine lists only them](0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)
|
||||
|
||||
### What runs on them, and how it gets there
|
||||
|
||||
@@ -316,6 +323,8 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0214** — [Backups guard against mistakes, stay on the machine, and are declared by the module that owns the data](0214-backups-guard-against-mistakes-and-stay-on-the-machine.md)
|
||||
- **0215** — [The machine's message bus is a node seat, and it is never restarted live](0215-the-machines-message-bus-is-a-node-seat-and-is-never-restarted-live.md)
|
||||
- **0216** — [The agent's configuration is registered through its module, at three scopes, and served as one plugin](0216-the-agents-configuration-is-registered-through-its-module-at-three-scopes-and-served-as-one-plugin.md)
|
||||
- **0220** — [What a machine asks needs its uplink held, and the retired resolver pieces go](0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md)
|
||||
- **0225** — [A consumer's identity is bounded by the provision it requires, judged before merge, and never refuses its provider](0225-a-consumers-identity-is-bounded-by-the-provision-it-requires.md)
|
||||
|
||||
### How it is built
|
||||
|
||||
|
||||
@@ -7,14 +7,19 @@ code:
|
||||
- mesh-controller internal/identity/authority.go
|
||||
- mesh-controller cmd/mesh-controller/plan.go (the names the roster publishes)
|
||||
- mesh-controller internal/catalogue/zones.go (the zones a module answers, ADR 0199)
|
||||
- mesh-controller internal/catalogue/seats.go (mesh-dns-resolver, node-hosts-file)
|
||||
- mesh-catalog modules/dnsmasq (the mesh's one resolver)
|
||||
- mesh-catalog modules/resolv-conf (what a node asks)
|
||||
- mesh-catalog modules/hosts (a node's /etc/hosts)
|
||||
- mesh-controller internal/catalogue/seats.go (mesh-dns-resolver, node-hostname, node-uplink)
|
||||
- mesh-controller internal/catalogue/resolve.go (one owner per path, a rendered fact's included, ADR 0223)
|
||||
- mesh-controller cmd/mesh-controller/holdings.go (the holders of a replicated seat, ADR 0223)
|
||||
- mesh-controller internal/catalogue/roster.go (each replicated seat's holders, for a template)
|
||||
- mesh-catalog modules/dnsmasq (the mesh's resolvers)
|
||||
- mesh-catalog modules/networkmanager, modules/systemd-networkd, modules/dhcpcd (what a node asks, written by its uplink's holder)
|
||||
- mesh-catalog modules/hostname (a node's /etc/hostname and /etc/hosts)
|
||||
- mesh-host internal/identity/serving.go
|
||||
- mesh-host internal/apply (the service that reflects a rule set)
|
||||
updated: 2026-10-03
|
||||
- mesh-host internal/apply (the service that reflects a rule set; a whole file handed to its new owner)
|
||||
updated: 2026-10-05
|
||||
decisions:
|
||||
- 02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md
|
||||
- 02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md
|
||||
- 02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md
|
||||
- 02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md
|
||||
- 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
|
||||
@@ -298,10 +303,10 @@ expensively enough to be worth restating:
|
||||
- **A node must not pin its own public name locally.** The duplicate record breaks resolution of
|
||||
that name for everything else that needs it.
|
||||
|
||||
**What the host receives:** what to ask, not what to answer. The mesh has **one resolver**, holding
|
||||
every node's internal domain; a node asks it first and a public resolver only when it is silent
|
||||
**What the host receives:** what to ask, not what to answer. The mesh has **two resolvers**, each
|
||||
holding every node's internal domain; a node lists both and nothing else
|
||||
([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md),
|
||||
[ADR 0196](../../02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md)).
|
||||
[ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)).
|
||||
|
||||
**What goes away:** the `/etc/hosts` floor. It exists because a node had to reach the mesh
|
||||
database before its own DNS existed; with [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)
|
||||
@@ -359,36 +364,83 @@ not a list of containers.
|
||||
somebody starts by hand is not the mesh's to configure, and reaching into every container on a
|
||||
machine — declared or not — is what a nameserver in `resolv.conf` would be for.
|
||||
|
||||
### One resolver for the mesh
|
||||
### The mesh's resolvers
|
||||
|
||||
*2026-10-03.* **The mesh's names live in one place: the module holding `mesh-resolver`**, a mesh-scoped
|
||||
seat of capacity one, placed on the node every tunnel converges on. It holds one wildcard per node —
|
||||
`<node>.internal` and everything under it — and listens on the private network only. It answers the
|
||||
mesh's names from what it holds and forwards every other name, giving the public answer.
|
||||
*2026-10-03, revised 2026-10-05.* **The mesh's names live in one module, held on two machines: the
|
||||
holders of `mesh-dns-resolver`**, a mesh-scoped seat that is *replicated* — held on the anchor and on
|
||||
the home server, each running the same module with the same machine list and the same zones, rendered
|
||||
by the controller into each. Each holds one wildcard per node — `<node>.internal` and everything
|
||||
under it — and one host record per node, listens on its private address and loopback only, answers
|
||||
the mesh's names from what it holds and forwards every other name, giving the public answer
|
||||
([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md),
|
||||
[ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)).
|
||||
Each holder is on record, added by `seat mesh-dns-resolver --add <node>/<module>`; an assignment of
|
||||
the module not on record stands beside them, eligible and silent, and a seat held once stays held
|
||||
once. *Checked by the controller's resolution tests with two holders on record, with two claimants
|
||||
and nothing on record, and on a seat held once.*
|
||||
|
||||
**Every node asks it for everything, and a public resolver only when it is silent.** The module
|
||||
holding `node-resolver-config` writes `/etc/resolv.conf` naming `mesh-resolver` first and a public
|
||||
resolver second, with a short timeout and one attempt: the C library moves to the second only when the
|
||||
first does not answer — the anchor or the tunnel down, a captive portal holding the tunnel back — so
|
||||
public names keep resolving then, and `.internal` is never asked of a public resolver while the mesh's
|
||||
answers. Containers take the same two from their machine, the runtime copying non-loopback resolvers
|
||||
into every container, so the runtime is given no `dns` of its own
|
||||
**Every node asks them for everything, and nothing else.** `/etc/resolv.conf` lists every holder's
|
||||
private address — the holder on the machine itself first if it is one, then the rest by name — with a
|
||||
short timeout and two attempts, and **no public resolver**. ADR 0196 listed a public resolver second,
|
||||
for the anchor being unreachable; a C library that asks every listed server at once and takes the
|
||||
first reply — musl, so every Alpine container — took the public resolver's "no such name" for a mesh
|
||||
name, and every build on the home server failed. With only the mesh's resolvers listed, whichever
|
||||
answers first gives the one answer. The cost is stated: a machine that reaches no mesh resolver has no
|
||||
names until it does, and with the anchor down the second resolver is reachable only on its own machine
|
||||
and its own LAN. Containers take the same lines from their machine, the runtime copying non-loopback
|
||||
resolvers into every container, so the runtime is given no `dns` of its own
|
||||
([ADR 0196](../../02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md),
|
||||
replacing ADR 0194's per-node `systemd-resolved` stub).
|
||||
replacing ADR 0194's per-node `systemd-resolved` stub). *Checked by the controller's composition tests
|
||||
on both holders and on a third machine, for every uplink module, and live by each machine's
|
||||
`/etc/resolv.conf` and an Alpine container on the home server resolving the anchor's name every
|
||||
time.*
|
||||
|
||||
**No node holds a copy.** The per-node resolver, its zones file and the mesh's region of `/etc/hosts`
|
||||
go: every resolution fault found on 2026-10-03 was a copy disagreeing with the truth — a hosts file
|
||||
**The file is the uplink's holder's** ([ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)).
|
||||
A network manager rewrites `/etc/resolv.conf` on every connectivity change unless it is told not to
|
||||
([ADR 0117](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md)), so the module holding
|
||||
`node-uplink` — the one telling it — writes the file itself, and there is one owner for it, the one
|
||||
whose program would otherwise overwrite it. Each of the catalogue's three managers' modules
|
||||
(NetworkManager, systemd-networkd, dhcpcd) renders the same template from the resolver's holders and
|
||||
requires the mesh's resolver, so a machine is refused when nothing in the mesh resolves rather than
|
||||
given a file listing nothing. The managers' own mechanisms were weighed and not used: NetworkManager's
|
||||
global DNS and dhcpcd's static nameservers each write the file in their own form — their own header,
|
||||
their own options line — so neither can write the mesh's file byte for byte, and dhcpcd reads its
|
||||
configuration only at its next start; each manager is told to keep off the file and the module
|
||||
declares it. No other module may write that path, as a file or as a rendered fact: two modules on one
|
||||
node declaring one path are refused. The module that wrote the file before, its seat
|
||||
`node-resolver-config` and that seat's need of the uplink beside it
|
||||
([ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md))
|
||||
retire. The file changes owner in one apply on each machine: the node-engine hands a whole file to the
|
||||
resource declaring its path now rather than removing it first, so a machine is never without it.
|
||||
*Checked by the controller's tests that the three modules carry one identical template and that
|
||||
nothing else in the catalogue writes the path, a resolution test refusing a second writer, and the
|
||||
node-engine's handover test, in which the file is present at every step of the apply.*
|
||||
|
||||
**A machine's names are one seat's** ([ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)).
|
||||
The `hosts` module is renamed `hostname` and holds `node-hostname`, the seat once named
|
||||
`node-hosts-file`, whose former name resolves to it. It writes `/etc/hostname` from its `hostname`
|
||||
setting — with no default: what a machine calls itself is the operator's, and the mesh's name for a
|
||||
machine and its own need not agree — and the machine's `127.0.1.1` line, keeping the operator's lines.
|
||||
A new name takes effect at the next boot; nothing sets it live, because a graphical session's X
|
||||
authority is keyed by the name the session started under. *Checked by a resolution test that a
|
||||
module claiming the old name and one claiming the new are one seat on one machine, and a composition
|
||||
test that `/etc/hostname` is the setting, left out naming the key when nothing sets it.*
|
||||
|
||||
**No node holds a copy of its own.** The two holders hold the same rendering of one roster, never a
|
||||
list anyone edits. The per-node resolver, its zones file and the mesh's region of `/etc/hosts`
|
||||
go, and the per-node resolver's seat with them, deleted from the set once nothing claimed it
|
||||
([ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md)): every resolution fault found on 2026-10-03 was a copy disagreeing with the truth — a hosts file
|
||||
read once at start, an operator's old line beside the mesh's, a node's resolver lent to a LAN. No
|
||||
member's resolver answers a LAN; a router pointing at one is moved first. *Checked by each node's
|
||||
`/etc/resolv.conf` naming `mesh-resolver` then a public resolver, by no node but the holder answering
|
||||
`/etc/resolv.conf` listing the resolver's holders and nothing else, by no node but a holder answering
|
||||
DNS on any address, and by the router's DHCP DNS option naming the router.*
|
||||
|
||||
**Names that are neither a node nor a route.** A module that answers names declares a zone (a
|
||||
setting) and the listen that answers it; the controller hands the `mesh-dns-resolver` holder every
|
||||
zone with its module's node address and published port, and the holder forwards that zone there and
|
||||
zone with its module's node address and published port, and each holder forwards that zone there and
|
||||
answers nothing in it itself — the lab answers `<machine>.incus` for its running scenarios this way.
|
||||
An operator's own names, unrelated to the mesh, live in `/etc/hosts`'s kept region, held per node by
|
||||
the `node-hosts-file` seat's holder and changed through its tools; the controller holds none of them
|
||||
the `node-hostname` seat's holder and changed through its tools; the controller holds none of them
|
||||
([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)).
|
||||
*Checked by the holder's configuration carrying one forwarding rule per declared zone, and by a push
|
||||
leaving the hosts file's operator region byte for byte.*
|
||||
@@ -1035,10 +1087,14 @@ The list is worth having in one place, because it is most of the argument:
|
||||
|
||||
## Open
|
||||
|
||||
- **One resolver ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)).** Not built: every node still runs
|
||||
`node-dns-resolver`. The migration's four steps are in the record, in order.
|
||||
Nor are zones or the hosts file's holder ([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)): the
|
||||
workstation moves to the one resolver only once both exist, its lab and operator names depending on them.
|
||||
- **One resolver ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)).** *Correction of fact,
|
||||
2026-10-05:* no node holds `node-dns-resolver` any more, and
|
||||
[ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md) deletes the seat. What stood here before: *"Not built: every node still runs
|
||||
`node-dns-resolver`. The migration's four steps are in the record, in order."* Nor did it
|
||||
stay true that *"the workstation moves to the one resolver only once"* zones and the hosts file's
|
||||
holder ([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md))
|
||||
exist: every node, the workstation included, asks the one resolver, and every node holds
|
||||
`node-hosts-file`.
|
||||
|
||||
- ~~**What happens when the hub is down.**~~ **Resolved** by
|
||||
[ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md), together with `06`'s
|
||||
|
||||
@@ -5,8 +5,9 @@ code:
|
||||
- mesh-controller cmd/mesh-builder
|
||||
- mesh-controller internal/builder
|
||||
- mesh-catalog modules/build-agent
|
||||
updated: 2026-10-04
|
||||
updated: 2026-10-05
|
||||
decisions:
|
||||
- 02-DECISIONS/0219-the-build-queue-is-controlled-through-the-controller-and-the-build-seat.md
|
||||
- 02-DECISIONS/0190-a-seats-work-is-shared-by-its-holders-and-building-is-the-first-such-role.md
|
||||
- 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md
|
||||
- 02-DECISIONS/0189-the-store-keeps-what-the-records-name.md
|
||||
@@ -346,6 +347,21 @@ scheduled step with the server held still — which is what `while-stopped` exis
|
||||
keeps is still a manifest and so still referenced, and the dangerous flag is not needed once the
|
||||
mesh is the one deciding.
|
||||
|
||||
**An archive is held by a manifest of its own** (2026-10-05,
|
||||
[issue 253](../../04-ISSUES/253-the-stores-collector-would-delete-every-archive-the-mesh-keeps/00-report.md)).
|
||||
The sentence above held for images and not for archives: an archive was published as a bare blob that no
|
||||
manifest names, and the store's collector keeps only what a manifest names, so a nightly collection
|
||||
would have removed every archive the mesh keeps, the current ones included. So:
|
||||
|
||||
- an archive is published with a manifest that names it and nothing else, built from the archive's digest
|
||||
and size alone so it can be computed again from the record;
|
||||
- the sweep makes sure every archive it keeps is held that way before it lets anything go, and lets go of
|
||||
an archive by removing its manifest first;
|
||||
- the reference a machine fetches is unchanged.
|
||||
|
||||
The collector runs as a dry run until the controller reports no kept archive unheld; only then does it
|
||||
collect for real.
|
||||
|
||||
A machine behind by more than five builds of a module, recreating a container, cannot pull what it
|
||||
was running. It is already a machine the mesh reports as behind, and the answer is the current
|
||||
declaration.
|
||||
@@ -379,3 +395,26 @@ not of the recipe: one artifact declared per target, one build each.
|
||||
A component's version stops being stamped in at link time. It is unpacked into a directory named for
|
||||
its version, so it reads its version from its own path, and a build no longer has to know what it will
|
||||
be called.
|
||||
|
||||
## The build queue is controlled
|
||||
|
||||
*2026-10-05 — [ADR 0219](../../02-DECISIONS/0219-the-build-queue-is-controlled-through-the-controller-and-the-build-seat.md).*
|
||||
|
||||
What is asked of the build seat can be seen and controlled from the console. The controller holds the
|
||||
queue, and each machine's build agent holds its own process.
|
||||
|
||||
- **The controller's verbs:**
|
||||
- `queue` lists every ask, waiting, in flight or dead;
|
||||
- `cancel` and `clear` take asks back;
|
||||
- `rebuild` asks again under a new id;
|
||||
- `replay` asks again for a past build at its commit, as a dry run unless told to register, and never
|
||||
over a newer build without saying so;
|
||||
- `kill`, `pause` and `resume` are passed on to the seat on the right machine.
|
||||
- **The build seat's verbs, on each machine:**
|
||||
- `current` says what is building here and whether this holder is paused;
|
||||
- `kill` ends the build's whole process group and every container it started;
|
||||
- `pause` and `resume` set a flag the holder reads before it takes the next ask, which survives its
|
||||
restart.
|
||||
|
||||
Every action that drops work leaves a failed outcome, so a plan waiting on that build fails and says why
|
||||
rather than waiting. A plan keeps the id of every build it asked for. A plan waiting on a paused seat says so and is not counted late; a failed plan can be resumed with `plans retry`, and a `rebuild` of a module a plan holds joins that plan.
|
||||
|
||||
@@ -5,8 +5,11 @@ code:
|
||||
- mesh-sdk src
|
||||
- mesh-tools src/broker-nats.ts (and broker-amqp.ts until the rollout)
|
||||
- mesh-controller internal/link
|
||||
updated: 2026-09-26
|
||||
- mesh-controller cmd/mesh-controller (status)
|
||||
- mesh-catalog modules/postgres, modules/keycloak (the Go provisioner loop)
|
||||
updated: 2026-10-06
|
||||
decisions:
|
||||
- 02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md
|
||||
- 02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md
|
||||
- 02-DECISIONS/0106-the-bus-is-nats.md
|
||||
- 02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md
|
||||
@@ -232,6 +235,17 @@ A provider ships the provisioner that creates instances of what it offers
|
||||
same provision, so a grant addressed to a node alone does not name a consumer, and withdrawing one
|
||||
would take another's away.
|
||||
|
||||
### A provider says which consumer it keeps failing
|
||||
|
||||
Whether a provision was made is known to the provider's loop alone. A consumer the loop has failed for
|
||||
five minutes without one success — its create, its periodic check, or reading the secret the mesh
|
||||
minted for it — is announced as `provisioner.failing`, naming the consumer, its machine, the class of
|
||||
error and since when, and again every fifteen minutes while it lasts; the first success, a withdrawal,
|
||||
and the first success after the provider restarts are `provisioner.recovered`. Every module that
|
||||
receives contributions may publish both, derived and never declared. The controller keeps the newest
|
||||
failing word per provider, machine and consumer, and `status` names each one until it recovers
|
||||
([ADR 0224](../../02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md)).
|
||||
|
||||
### Checked, and it agrees (2026-09-16)
|
||||
|
||||
This looked like the sharpest disagreement and was not one. The live wire is the contributions file
|
||||
@@ -255,4 +269,5 @@ two implementations drift from while both believe they conform.
|
||||
| The envelope is the envelope | An emitted event is compared header by header against a fixture; a missing required header fails, an unknown `x-` header is accepted. |
|
||||
| Delivery is at-least-once | A fixture delivered twice is handled once. |
|
||||
| A grant names a consumer | A grant fixture is read by every implementation and yields the same module and the same node. |
|
||||
| A provider that keeps failing a consumer says so | The loop's tests drive five minutes of failure to one `provisioner.failing` and a success to `provisioner.recovered`; the controller's tests carry it from the bus to `status`. |
|
||||
| A partial SDK is legitimate | An implementation claiming the floor and events passes those suites and is listed for them; a module using tools in that language is refused at build time with the reason. |
|
||||
|
||||
@@ -4,14 +4,21 @@ status: in-progress
|
||||
code:
|
||||
- mesh-controller internal/catalogue/seats.go
|
||||
- mesh-controller internal/catalogue/resolve.go
|
||||
- mesh-controller internal/catalogue/seat_dependencies.go
|
||||
- mesh-controller internal/inventory/seats.go
|
||||
- mesh-controller internal/inventory/migrations/0039-a-seat-is-held-by-one-assignment-on-record.sql
|
||||
- mesh-controller internal/inventory/migrations/0062-a-replicated-seat-has-several-holders-on-record.sql
|
||||
- mesh-controller cmd/mesh-controller/holdings.go
|
||||
- mesh-controller cmd/mesh-controller/seats.go
|
||||
- mesh-controller cmd/mesh-controller/source.go
|
||||
- mesh-controller internal/inventory/migrations/0032-a-source-may-live-on-a-seat.sql
|
||||
- mesh-catalog modules/gitea/module.json
|
||||
updated: 2026-10-03
|
||||
- mesh-controller internal/inventory/migrations/0063-a-machines-names-are-one-seats.sql
|
||||
- mesh-controller internal/inventory/migrations/0064-the-resolver-file-is-the-uplinks.sql
|
||||
updated: 2026-10-05
|
||||
decisions:
|
||||
- 02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md
|
||||
- 02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md
|
||||
- 02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md
|
||||
- 02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
|
||||
- 02-DECISIONS/0161-what-deserves-a-seat.md
|
||||
@@ -36,7 +43,7 @@ A seat has four properties, fixed by the mesh rather than by any module:
|
||||
| property | is |
|
||||
|---|---|
|
||||
| name | what a definition names and an assignment holds, and what a person reads in the list |
|
||||
| scope | node, site or mesh: where its capacity applies. Every seat in the set has a capacity of one, so one holder per scope. A bench, a seat with several holders, is a word the glossary keeps and no seat uses yet |
|
||||
| scope | node, site or mesh: where its capacity applies. Every seat in the set has a capacity of one, so one holder per scope — except a *replicated* mesh seat, which may be held on several machines at once, one holder per machine, each on record ([ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)). `mesh-dns-resolver` is the only one. Replicated is part of the seat's definition, compiled with the set and never stored |
|
||||
| delivers | the provision its holder answers for, or nothing |
|
||||
| decision | the record that made it a seat |
|
||||
|
||||
@@ -67,6 +74,17 @@ needs ([28](28-building-the-bus.md), task 5.3). **And a holding is the assignmen
|
||||
holder takes the row with it, so a seat never points at something that is not running anywhere, and
|
||||
the seat falls back to derivation rather than to nothing.
|
||||
|
||||
**A replicated seat has several holders on record** (revision, 2026-10-05,
|
||||
[ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)).
|
||||
`seat <name> --add <node>/<module>` records one more holder beside those on record, judged exactly
|
||||
as a handover is; `--to` still leaves exactly one. Each holder is added by that act and never by being
|
||||
assigned: an assignment not on record is eligible and silent, and two claimants with nothing on record
|
||||
are refused, as for any mesh seat. `--add` on a seat held once is refused, naming `--to`, and a store
|
||||
recording two holders of a seat held once is refused at resolution, naming the seat. A requirement the
|
||||
seat delivers is answered on a holder by itself, and elsewhere by the first holder in name order. What
|
||||
the holders are is given to a module's roster template, its own machine first — how every machine's
|
||||
resolver file lists both of the mesh's resolvers.
|
||||
|
||||
The handover refuses what would make the new holder wrong before anything is written: the seat must
|
||||
exist, the assignment must exist, and the module must be able to hold the seat — claim it at its scope
|
||||
and provide what it delivers, judged against the store's row and not against anything compiled into a
|
||||
@@ -128,13 +146,13 @@ convention, which later seats departed from.
|
||||
| `mesh-npm-package-registry` | `npm-package-registry` | mesh | `npm-package-registry` | the forge |
|
||||
| `mesh-git` | `git` | mesh | `git` | the forge |
|
||||
| `mesh-build-machine` | `the-build-machine` | node | — | a builder |
|
||||
| `mesh-resolver` | — | mesh | — | the mesh's one resolver, holding every node's internal domain ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)) |
|
||||
| ~~`mesh-dns-port`~~ | `the-dns-port` | node | — | retired by [ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md): the local resolver became the mesh's one |
|
||||
| `node-hosts-file` | — | node | — | owns `/etc/hosts`: the machine's own lines and the operator's kept region, changed through its verbs `entries`, `add`, `remove` ([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)) |
|
||||
| `mesh-resolver` | — | mesh | — | the mesh's resolvers, each holding every node's internal domain ([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)); replicated, held on the anchor and the home server ([ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)) |
|
||||
| ~~`mesh-dns-port`~~ | `the-dns-port` | node | — | retired by [ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md): the local resolver became the mesh's one; deleted from the set, and from the store's table, once nothing claimed it ([ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md)) |
|
||||
| `node-hostname` | `node-hosts-file` (renamed by [ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md); the former name an alias) | node | — | owns a machine's names: `/etc/hostname`, written from its holder's `hostname` setting with no default and taking effect at the next boot, and in `/etc/hosts` the machine's own lines and the operator's kept region, changed through its verbs `entries`, `add`, `remove` ([ADR 0199](../../02-DECISIONS/0199-a-module-that-answers-names-declares-its-zone-and-a-nodes-hosts-file-is-one-modules.md)). Held by the module `hostname`, formerly `hosts` |
|
||||
| `mesh-intrusion-prevention` | `the-intrusion-prevention` | node | — | an intrusion-prevention service |
|
||||
| `mesh-packet-filter` | `the-packet-filter` | node | — | the packet filter |
|
||||
| `mesh-private-network` | `the-private-network` | node | — | the private network the mesh runs over |
|
||||
| `mesh-resolver-configuration` | `the-resolver-configuration` | node | — | whichever of the alternative resolver configurations is chosen |
|
||||
| ~~`mesh-resolver-configuration`~~ | `the-resolver-configuration` | node | — | retired by [ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md): `/etc/resolv.conf` is written by the holder of `node-uplink`, the program that would otherwise rewrite it, and the need of the uplink beside it ([ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md)) goes with it; deleted from the set, and from the store's table, once nothing claims it |
|
||||
| `mesh-showcase` | `the-showcase` | node | — | the showcase module |
|
||||
|
||||
The controller holds **the mesh's own** entries in code, and a test asserts their size and that
|
||||
@@ -202,10 +220,28 @@ Nothing reaches that state by accident: an unknown manifest field is refused out
|
||||
protocol was written as one. Checked by a registration test accepting a node seat with no protocol
|
||||
and by the showcase manifest, which declares one.
|
||||
|
||||
Most node seats deliver nothing. They say which module is this machine's packet filter, or which of
|
||||
two alternative resolver configurations it runs, and a second holder is refused. That is the whole of
|
||||
Most node seats deliver nothing. They say which module is this machine's packet filter, or which
|
||||
module writes its resolver file, and a second holder is refused. That is the whole of
|
||||
their job, and it is a real one: it is the mesh saying what a machine is, in words a person can read.
|
||||
|
||||
## A seat that needs another beside it
|
||||
|
||||
*2026-10-05* ([ADR 0220](../../02-DECISIONS/0220-what-a-machine-asks-needs-its-uplink-held-and-the-retired-resolver-pieces-go.md)). **A seat may name the node seats its holder needs held on the
|
||||
same node**, because some roles are right only while another is filled beside them. The holder of the
|
||||
resolver configuration writes `/etc/resolv.conf`, and that file stays the mesh's only while the
|
||||
machine's network manager is told to leave it alone — which is what the uplink's holder does
|
||||
([ADR 0117](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md)). So the resolver configuration
|
||||
needs the uplink.
|
||||
|
||||
**A module claiming such a seat depends on each seat it needs**, exactly as a module declaring a
|
||||
service depends on the service manager
|
||||
([ADR 0207](../../02-DECISIONS/0207-a-module-depends-on-the-node-seats-that-apply-its-resources.md)):
|
||||
derived from the claim and never written in a manifest, met by any module assigned to the node, judged
|
||||
over the node's whole set, refused at assignment naming the seat and the modules that could hold it,
|
||||
and refused at composition. Taking the needed seat's last holder from beneath a dependent is refused
|
||||
too. What a seat needs is the mesh's definition of the role, compiled with the set and never stored,
|
||||
and adding a need is a decision.
|
||||
|
||||
## The overview
|
||||
|
||||
The controller lists every seat in the set with its scope, what it delivers, and its holder as a node
|
||||
@@ -259,5 +295,8 @@ checked as their tables say:
|
||||
| `secret` has one provider, the holder of `mesh-vault` | 0161: a second claimant of the seat is refused by name (`CanHold`); *correction of fact, 2026-10-01: no parser rule ever reserved the word, the seat does the work*. |
|
||||
| A singular fact about machines is a placement of capacity one, refused by name | 0161: the overlay command's test for a second hub; the store's unique index. |
|
||||
| A holder of `node-uplink` is the dialect the machine runs | 0161: the host reports `uplink-<manager>` in its profile with every report; a resolution test refuses the other holder naming the capability. |
|
||||
| A seat's holder has the seats it needs beside it: the resolver configuration is refused at `assign` without the uplink held on its node, and the uplink's last holder cannot be taken from beneath it | 0220: seat-dependency tests on the definition, on a synthetic claimant, at assign, at unassign and at composition, and one reading the catalogue for the uplink's possible holders. |
|
||||
| A replicated seat has every holder on record and composes on each; a seat held once still refuses a second holder, on record or not; `--add` is refused for it | 0223: resolution and composition tests with two holders on record and a third machine, with two claimants and nothing on record, and on `mesh-store`; a unit test that only `mesh-dns-resolver` is replicated, surviving the store's rows; store tests keeping several holders once each, a handover leaving one, and unassigning one taking only its row. |
|
||||
| A retired seat leaves the set once nothing claims it, from the compiled set and the store's table both | 0220: the closed-set test's count, and the store migration that deletes the row. |
|
||||
| Holdings are derived, and the overview lists every seat | 0118: the `seats` command test, including an unheld seat. |
|
||||
| A build source on the seat records no address; an unheld seat refuses only self-hosted builds | 0111's tests. |
|
||||
|
||||
@@ -2,8 +2,9 @@
|
||||
layer: to-be
|
||||
status: in-progress
|
||||
code: [mesh-controller internal/catalogue]
|
||||
updated: 2026-10-02
|
||||
updated: 2026-10-06
|
||||
decisions:
|
||||
- 02-DECISIONS/0225-a-consumers-identity-is-bounded-by-the-provision-it-requires.md
|
||||
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
|
||||
- 02-DECISIONS/0155-a-definition-names-no-installation-and-how-that-is-checked.md
|
||||
- 02-DECISIONS/0115-one-assignment-of-a-module-per-node.md
|
||||
@@ -250,7 +251,11 @@ value in a container's environment is refused when the definition is parsed, wit
|
||||
**A module is assigned at most once to a node**, and that pair is the assignment's identity
|
||||
([ADR 0112](../../02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md)). Its directories,
|
||||
containers, login, broker account and settings are keyed by it, as today, and a login still fits the
|
||||
tightest backend ([ADR 0049](../../02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md)).
|
||||
backend that keeps it ([ADR 0049](../../02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md)):
|
||||
the bound is the one the offer of each provision it requires states, a keyless provision states none,
|
||||
and an overflow is refused by the catalogue check before merge and left out of the provider's grants,
|
||||
reported, at composition — never a refusal of the provider's machine
|
||||
([ADR 0225](../../02-DECISIONS/0225-a-consumers-identity-is-bounded-by-the-provision-it-requires.md)).
|
||||
|
||||
**A module may run on many nodes, and one assignment may hold a seat**
|
||||
([ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) (superseding [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md))). The definition
|
||||
|
||||
@@ -2,8 +2,10 @@
|
||||
layer: to-be
|
||||
status: proposed
|
||||
code: []
|
||||
updated: 2026-10-01
|
||||
updated: 2026-10-05
|
||||
decisions:
|
||||
- 02-DECISIONS/0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md
|
||||
- 02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md
|
||||
- 02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md
|
||||
- 02-DECISIONS/0157-a-build-says-what-it-does-on-the-bus-as-it-happens.md
|
||||
- 02-DECISIONS/0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md
|
||||
@@ -129,6 +131,40 @@ controller replaced mid-plan resumes from the store. `status` lists open plans a
|
||||
has waited too long. The transition discipline for breaking changes in the list above is still
|
||||
unwritten, and still the next thing.
|
||||
|
||||
## How a plan sends (2026-10-05)
|
||||
|
||||
Revision, [ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md). Three rules on how a plan delivers what it built.
|
||||
|
||||
- **Grants travel before code.** A send issues the memberships for the machines it is about to send to
|
||||
before their declarations. The machine holding the bus is sent first when the list of bus users changes.
|
||||
A membership that could not be issued fails the send.
|
||||
- **One machine first.** Unless a module's upgrade policy says *together*, a plan sends it to one machine,
|
||||
the first by name, and to the rest only once that machine reports the new declaration applied and
|
||||
current. A first machine that fails stops the module's rollout there, with the reason in the plan.
|
||||
- **A newer merge takes over.** A merge's plan supersedes every older open plan for the same repository
|
||||
and branch, and takes in the modules they had not yet built. A plan that waits on nothing can be closed
|
||||
by hand, by its id.
|
||||
|
||||
What a merge changed is read from the forge whole, page by page
|
||||
([issue 252](../../04-ISSUES/252-a-merges-changed-modules-were-read-wrong/00-report.md)). A changed path in a
|
||||
module directory the mesh does not hold yet is that module's own, not shared code, when its definition is
|
||||
among the changed paths.
|
||||
|
||||
## What a push does not send (2026-10-05)
|
||||
|
||||
Revision, [ADR 0221](../../02-DECISIONS/0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md).
|
||||
A named push still sends every other machine its work left behind, but no longer a build that an upgrade
|
||||
policy or a plan is holding back
|
||||
([issue 259](../../04-ISSUES/259-a-named-push-sent-every-machine/00-report.md)).
|
||||
|
||||
- **A send records the builds it carried**: which build of each module the machine was sent.
|
||||
- **A machine a push did not name is left** when a module it runs would move to a build its policy
|
||||
records rather than rolls out, or that an open plan has not sent it yet. A machine whose last send's
|
||||
builds are not known is left too. The push names the machine, the module, the builds and the reason,
|
||||
and says `push <node>` sends it.
|
||||
- The machine a push names, a push of every machine, and `push --behind` send held builds as before.
|
||||
Walking a change through the mesh under `record` is one `push <node>` per machine.
|
||||
|
||||
## Why now, and why not yet
|
||||
|
||||
**Why it matters:** self-update is the difference between a mesh a person maintains by typing
|
||||
|
||||
@@ -0,0 +1,436 @@
|
||||
---
|
||||
layer: to-be
|
||||
status: designed
|
||||
code: []
|
||||
updated: 2026-10-06
|
||||
decisions:
|
||||
- 02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md
|
||||
- 02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md
|
||||
- 02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md
|
||||
- 02-DECISIONS/0141-the-host-delivers-its-own-successor.md
|
||||
- 02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md
|
||||
---
|
||||
|
||||
# 45 — A core that cannot fail silently
|
||||
|
||||
**The core says when it is wrong, refuses what is stale or unreadable, repairs what it knows how to
|
||||
repair, replaces itself one machine at a time with something other than itself watching, and is checked
|
||||
against the mesh's real facts before a change merges** ([ADR 0227](../../02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md),
|
||||
from [research 031](../../01-RESEARCH/031-a-core-that-cannot-fail-silently/00-overview.md)).
|
||||
|
||||
**The core**, here: the controller, the node-engine and its launcher, the bus server, the node tools and
|
||||
the console, the build seat, and the forge's announcer of merges.
|
||||
|
||||
This document is the build's specification. §1–§9 are the parts; §10 is the order they are built in,
|
||||
each phase with what it delivers, in which repository, and when it is done. Every bound marked
|
||||
*provisional* is set from the durations Phase 0 records and corrected in Phase 1's first live week.
|
||||
|
||||
## The parts, and how a fact reaches the operator
|
||||
|
||||
```
|
||||
signals ──────────► watchdogs ─┐ ┌─► status (open conditions first)
|
||||
(heartbeats, reports, │ │
|
||||
plan progress, advisories) ├─► CONDITION STORE ───┼─► condition events ─► operator-channel ─► Telegram
|
||||
doctor probes ─────────────────┤ (controller is │ holder desktop notifier
|
||||
(live invariants, every 5 min) │ its only writer) └─► healers ─► repair, braked ─► event
|
||||
events (provider failing, …) ──┘ │
|
||||
└─► hand-act log (people)
|
||||
|
||||
second machine: watcher ── hears doctor's heartbeat? ── no, past bound ──► Telegram directly (not the bus)
|
||||
```
|
||||
|
||||
Owning repositories: **mesh-controller** (the condition store, watchdogs, `doctor`, healers, the lease,
|
||||
calls, the hand-act log, the facts snapshot), **mesh-host** (the node-engine: the apply queue, report
|
||||
order, epoch refusal, the `report` verb, rollback witnessing), **mesh-tools** (the node tools and the
|
||||
console: their heartbeat, passing every argument, health answers), **mesh-catalog** (the
|
||||
`operator-channel` seat's holder and channels, the watcher, the provider brake, the catalogue's merge
|
||||
gate), **mesh-lab** (the replays and the induced-failure scenarios).
|
||||
|
||||
---
|
||||
|
||||
## 1. The writers table (rule 1)
|
||||
|
||||
Every kind of state the core keeps, and the one component that writes it. Anyone else asks.
|
||||
|
||||
| State | Writer | Kept in | Others |
|
||||
|---|---|---|---|
|
||||
| a machine's declaration | controller (lease holder) | the bus, last per subject | read |
|
||||
| a machine's applied state and its report | the node-engine's apply queue (§6) | the machine; the report on the bus | the reconcile and a delivery *enqueue*, never apply |
|
||||
| the controller lease | the controller instance holding it | key-value `mesh-controller_lease` | a candidate waits |
|
||||
| plans and their tiers | controller (lease holder), compare-and-set on the plan's revision | the controller's store | read through `plans` |
|
||||
| conditions | controller | key-value `mesh-controller_conditions` | raise or clear only through observations the controller reads |
|
||||
| calls and their outcomes | controller | key-value `mesh-controller_calls` | read by id |
|
||||
| the hand-act log | controller, through the verbs that act | key-value `mesh-controller_hand-acts` | — |
|
||||
| stream definitions and bus permissions | controller | the bus | — |
|
||||
| builds and their outcomes | the build seat's holder | its own state | the controller asks |
|
||||
| a merge announced | **one** announcer per forge (the hook, or the poll when the hook is absent — never both) | the bus | — |
|
||||
| a provider's standing | the provider | the provider's events | the controller keeps the newest word as a condition |
|
||||
| the operator-channel's open messages | the seat's holder | its own key-value state | — |
|
||||
| the facts snapshot | controller | the artifact store, `facts/latest` | the build seat reads |
|
||||
|
||||
A bus subject two components may publish on is refused when the controller composes grants, unless
|
||||
this table marks it shared. Changing a writer is a change to this table, through a decision.
|
||||
|
||||
## 2. The condition store (rules 5, 6)
|
||||
|
||||
**A condition is a durable fact about something the mesh owns that is wrong.** It is raised and cleared
|
||||
by observation only.
|
||||
|
||||
**Where.** A key-value bucket, `mesh-controller_conditions`, written by the controller alone, one entry
|
||||
per open condition. Every transition — raised, changed, silenced, cleared — is also appended to a
|
||||
history kept ninety days, read through `conditions history`.
|
||||
|
||||
**The key** names the thing and the kind, so the same fault said again is the same condition:
|
||||
`<scope>.<id>.<kind>`, where scope is one of `machine`, `plan`, `call`, `build`, `merge`, `provider`,
|
||||
`seat`, `bus`, `core`, `probe`, `mesh`. Examples of the shape: `machine.<node>.silent`,
|
||||
`plan.<id>.stalled`, `provider.<module>.<node>.<consumer>.failing`, `core.controller.<node>.rolled-back`,
|
||||
`bus.<consumer>.slow-consumer`.
|
||||
|
||||
**The fields:**
|
||||
|
||||
| Field | Holds |
|
||||
|---|---|
|
||||
| kind | the condition kind, from the signals table, the probe registry or an event kind |
|
||||
| subject | scope, id, and the machine it concerns when there is one |
|
||||
| severity | `urgent` (needs the operator now) or `warning` (when they can) — two levels, no more |
|
||||
| summary | one line in the mesh's words |
|
||||
| evidence | the newest observations, at most ten, each with its time |
|
||||
| source | the signals-table row, probe or event that raised it |
|
||||
| raised, last observed | times; and how many observations since raised |
|
||||
| tried | each healer attempt: when, what, outcome |
|
||||
| resolver | `self` (clears on observation), `healer:<name>`, `operator`, or `agent` |
|
||||
| silenced | until when, by whom, why — empty when not silenced |
|
||||
| epoch | the controller lease epoch that last wrote it |
|
||||
|
||||
**The life of one:**
|
||||
|
||||
```
|
||||
(absent) ──observation past bound──► OPEN ──healer tries──► OPEN (tried += …)
|
||||
│ ▲ │ budget spent
|
||||
│ └─ seen again ◄───────┘──► OPEN, resolver: operator, severity: urgent
|
||||
│
|
||||
silence (verb, ≤ 7 days, reason, hand act) ──► OPEN, silenced (no messages)
|
||||
│
|
||||
observation says resolved ──► CLEARED: entry removed, transition kept in history
|
||||
```
|
||||
|
||||
- **Nobody resolves a condition by hand.** It clears when the signal returns or the probe passes.
|
||||
- **A condition cleared and raised again within ten minutes** reopens with its count increased; it is
|
||||
not a new message.
|
||||
- **The verbs:** `conditions` (open ones, filtered by scope, severity or machine), `conditions show
|
||||
<key>`, `conditions silence <key> --for <duration> --why <text>`, `conditions history`.
|
||||
- **The events**, emitted by the controller on its own subjects: `condition-raised`,
|
||||
`condition-changed` (severity or resolver changed; not every observation), `condition-cleared`. The
|
||||
operator-channel's holder and any other surface consume them; the controller learns nothing about
|
||||
telling.
|
||||
- **`status`** lists open conditions first, urgent before warning, oldest first, silenced ones marked
|
||||
with their expiry. The all-well sentence requires none open, silenced included.
|
||||
- **ADR 0224's provider standing** is the first kind: `provider-failing`, raised by the provider's
|
||||
event, cleared by its recovery, its silence after thirty minutes the row S8 below.
|
||||
|
||||
## 3. The signals table (rule 5)
|
||||
|
||||
Compiled into the controller. A watchdog per row; the test generated from the table suppresses each
|
||||
signal and asserts its condition. `doctor signals` shows, for every row, the age of its newest signal.
|
||||
|
||||
| Row | Signal | Emitter | Trigger | Bound | Condition kind | Severity | Healer |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| S1 | machine heartbeat | node-engine | its interval | 3 × interval, *provisional* | `silent` — not raised while the machine has declared itself asleep or shut down (ADR 0211) | warning; urgent after 30 min for the control node | — |
|
||||
| S2 | report after a send | node-engine | each declaration sent | max(2 min, 3 × that machine's last apply duration), *provisional* | `sent-not-reported` | warning | H1 |
|
||||
| S3 | plan tier progress | controller's plan | each tier entered | from the tier's build and apply durations, *provisional* | `stalled` (names the tier and what it waits on) | warning | H2 |
|
||||
| S4 | the controller's event loop takes a message | controller | while its consumer has pending messages | 2 min | `controller-deaf` | urgent | — |
|
||||
| S5 | a merge announced becomes a plan or *nothing reads it* | announcer → controller | each merge | 10 min (exists, issue 266) | `merge-not-acted` | urgent | — |
|
||||
| S6 | a build asked → its outcome | build seat | each ask | the build's declared timeout | `ask-lost` | warning | — |
|
||||
| S7 | a call running → finished | controller | each call | the verb's declared bound | `call-hung` | warning | — |
|
||||
| S8 | a provider's failing word repeated | provider | every 15 min while failing (ADR 0224) | 30 min | `provider-silent` | warning | — |
|
||||
| S9 | bus advisories: slow consumer, maximum deliveries, permission violation, consumer deleted | bus server's system subjects | any | any occurrence | `slow-consumer`, `max-deliveries`, `refused`, `consumer-lost` — each naming the call, consumer or module in the mesh's words | warning | H4 for `slow-consumer` on a resettable consumer |
|
||||
| S10 | the self-check's heartbeat | controller's `doctor` | every run | 2 × its interval, watched **from the second machine** (§5) | `self-check-silent` | urgent | — |
|
||||
| S11 | node tools heartbeat | node tools | its interval | 3 × interval, *provisional* | `tools-silent` | warning | — |
|
||||
| S12 | the controller lease renewed | controller | every 5 s | 15 s | `lease-lost` | urgent | — |
|
||||
| S13 | stale refusals | every receiver (rule 2) | each refusal | more than 5 from one writer in 5 min | `stale-writer` (names the writer) | warning | — |
|
||||
| S14 | facts snapshot exported | controller | daily | 2 days | `facts-stale` | warning | — |
|
||||
| S15 | a hand act with a cause already recorded | hand-act log | each act | the second within 14 days | `healer-wanted` | warning | — |
|
||||
|
||||
The bus advisories (S9) cost one read-only subscription: the server already publishes them. The
|
||||
controller translates each into a condition naming the thing in the mesh's words, as the refused reply
|
||||
of issue 265 is translated today.
|
||||
|
||||
## 4. The self-check: `doctor` (rule 6)
|
||||
|
||||
A **probe registry** in the controller: each probe is a live invariant of a design, with an id, an
|
||||
interval, a timeout and the condition kind it raises. The controller runs the registry every five
|
||||
minutes; each probe has thirty seconds. A probe that errors or times out raises `probe-failed` for
|
||||
itself — an unanswered probe is never a pass.
|
||||
|
||||
| Probe | Asserts | From |
|
||||
|---|---|---|
|
||||
| D1 | every machine's declaration composes, and passes the node-engine's validation (the validator is a package of mesh-host the controller and the merge gate import — one validator) | 236, 263 |
|
||||
| D2 | every holder of the mesh's resolver answers a machine name for IPv4, and NODATA for IPv6 | 262 |
|
||||
| D3 | every seat on record has a live holder that answers | 208, 218 |
|
||||
| D4 | every kept archive is held by a manifest | 253 |
|
||||
| D5 | exactly one lease holder; no message from a stale epoch in the last interval | 204 |
|
||||
| D6 | every durable consumer exists with its definition and is near its stream's head | 248, 266 |
|
||||
| D7 | every stream the controller defines exists with its definition | 208 |
|
||||
| D8 | no address the mesh owns is in a ban list | 238 |
|
||||
| D9 | `status` answers in full within ten seconds | 265 |
|
||||
| D10 | every machine runs the node-engine and node tools builds its plan says, or is inside a plan's window | version split |
|
||||
| H-* | the health probes of §8, run for every core component on every machine | rule 8 |
|
||||
|
||||
- **`doctor`** answers the last run's verdict at once: per probe, pass, fail or failed-to-run, and age.
|
||||
**`doctor run`** runs now under a call id. **`doctor probes`** lists the registry; **`doctor
|
||||
signals`** the table's ages.
|
||||
- **Every run ends with a heartbeat** event carrying the run's id and counts. That is S10.
|
||||
- **The registry is the design's live form.** A check over the to-be designs counts invariants that
|
||||
name a probe against those that do not; the number without may only go down.
|
||||
|
||||
## 5. The output channel, minimal form (rule 6)
|
||||
|
||||
From [research 028](../../01-RESEARCH/028-the-meshs-output-channel/00-overview.md), the smallest form
|
||||
that works; its open questions stay open and its graduation amends this section.
|
||||
|
||||
- **One mesh seat, `operator-channel`**, held once. It **accepts** `notify` as a work queue, so a
|
||||
message waits for a holder; it keeps its **open messages** in its own key-value state, so a restart
|
||||
forgets nothing; it **serves** `open` and `history`.
|
||||
- **The holder consumes the controller's condition events** and decides what is sent. The controller
|
||||
calls nobody.
|
||||
- **Two channels**, each a module contributing itself to the seat: **Telegram** (a bot to the
|
||||
operator's chat; its token and chat id are the channel module's secrets) and the **desktop
|
||||
notifier** of the machine the operator is at.
|
||||
- **A message** is: the condition's key, its subject (a machine's role, a module, a plan), kind,
|
||||
severity, the one-line summary, since when, and the verb that shows more. **Deduplicated by the key.**
|
||||
- **When:** on `condition-raised`; once more if still open after 1 hour (urgent) or 12 hours
|
||||
(warning); on `condition-cleared`, by editing the first message where the channel can. A silenced
|
||||
condition sends nothing. Urgent goes to both channels; warning to the desktop notifier when the
|
||||
operator's session is there, otherwise to Telegram.
|
||||
- **Rate:** at most twenty messages an hour; the excess is folded into one message naming them all.
|
||||
- **What may leave the mesh:** roles and words. A message carrying an address, a path or anything
|
||||
shaped like a secret is refused by the holder and raises `channel-refused` instead.
|
||||
- **No answering back** in this form; acknowledging is `conditions silence`, through the mesh.
|
||||
- **The watcher's watcher.** A module, `mesh-watcher`, assigned to one machine that is not the control
|
||||
node, holds the Telegram channel's secret too. It listens for the self-check heartbeat (S10) and for
|
||||
the bus itself; when either has been silent past its bound it sends to Telegram **directly over
|
||||
HTTPS, not through the bus**, and says so again when they return. It is the only sender that does not
|
||||
pass through the control node.
|
||||
|
||||
## 6. Order: lease, epoch, report sequence, one apply queue (rules 1, 2)
|
||||
|
||||
**The lease.** A controller instance acts — sends a declaration, writes a plan, a condition or a call
|
||||
— only while it holds the key `holder` in `mesh-controller_lease`: written with compare-and-set, a
|
||||
fifteen-second time to live, renewed every five seconds. The **epoch** is the bucket revision at which
|
||||
it was taken. A starting controller waits for the key to be absent or expired. One that fails a renewal
|
||||
stops acting at once and exits, so its service manager restarts it as a waiting candidate.
|
||||
|
||||
```
|
||||
controller A (epoch 41) ──renew──renew──╳ (renewal refused)──► stops sending, exits
|
||||
controller B ──wait──────────────take (epoch 57)──► acts; marks A's running calls abandoned
|
||||
node-engine accepts 41 … then 57; refuses anything from 41 after 57, counted, reported
|
||||
```
|
||||
|
||||
**What carries the order:**
|
||||
|
||||
| Message | Carries | The receiver keeps, per writer | Refuses |
|
||||
|---|---|---|---|
|
||||
| declaration | epoch, sequence | highest epoch, then highest sequence | older epoch; same epoch, lower sequence |
|
||||
| report | machine, the declaration's epoch and sequence it is about, the node-engine's own report sequence (kept on disk, increasing across restarts and self-updates) | highest declaration sequence, then report sequence | an account of an older declaration; an older report |
|
||||
| plan write | the plan's revision, the epoch | — (compare-and-set) | a write against a revision already moved |
|
||||
| build outcome | the build's ask id and order | newest per module | an older build finishing later (exists, 219) |
|
||||
| call | call id, epoch | — | a finish from an epoch that is not the holder's, recorded as abandoned |
|
||||
| merge announcement | forge, repository, commit, the announcer's sequence | newest per repository | a duplicate (one announcer) |
|
||||
|
||||
**Every refusal** is one line in the receiver's log in the mesh's words, a counter, and — from the
|
||||
node-engine — a report naming the refused declaration, so the controller sees it. The counter feeds
|
||||
S13.
|
||||
|
||||
**One apply queue on every machine.** The node-engine has one worker that applies; a delivery, the
|
||||
five-minute reconcile, a self-update hand-over and a manual `apply` are four reasons to **enqueue**.
|
||||
The queue holds at most one pending request, coalesced. When the worker starts it takes the newest
|
||||
declaration held at that moment, applies it once, and sends one report naming the sequence it applied.
|
||||
Nothing else in the node-engine applies.
|
||||
|
||||
**The `report` verb.** The node-engine answers, on request, the report of the last declaration it
|
||||
applied, from what it keeps on disk. It is what healer H1 asks.
|
||||
|
||||
**Durable calls.** Every call's record — verb, arguments with secrets removed, caller, started, epoch,
|
||||
state (`running`, `finished`, `failed`, `abandoned`), the answer bounded in size, finished — lives in
|
||||
`mesh-controller_calls`, the last thousand or fourteen days. A new lease holder marks the previous
|
||||
holder's running calls `abandoned`, which is said.
|
||||
|
||||
## 7. Healers and the hand-act log (rule 7)
|
||||
|
||||
**A healer** is a registered response to one condition kind: its repair (the ordinary path again), its
|
||||
budget, its back-off, its brake, and the `healer-acted` event naming the condition, the act and the
|
||||
outcome. A healer may not withdraw, delete or recreate data; such a repair is a condition for the
|
||||
operator.
|
||||
|
||||
| Healer | Condition | Repair | Budget, then |
|
||||
|---|---|---|---|
|
||||
| H1 | `sent-not-reported` | ask the machine's node-engine for `report`; if it names an older declaration, send the current one again | twice per condition, then resolver `operator`, urgent |
|
||||
| H2 | `stalled` on a wait that is superseded or already finished | close the plan with its note, as `plans close` does | once |
|
||||
| H3 | a seat holder without its worker (D3) | raise the seat's objects again (208) | once per holder per hour |
|
||||
| H4 | `slow-consumer` far behind on a stream marked *resettable* in the stream table | `broker consumer-reset` (248) | once a day per consumer, then condition |
|
||||
| H5 | `provider-failing` with the administrator refusing the mesh's secret | ADR 0224 §5 (exists, in the module) | as ADR 0224 |
|
||||
|
||||
**The hand-act log.** Every verb that repairs by hand — a named `push` outside a plan, `plans close`,
|
||||
`broker consumer-reset`, `conditions silence`, and `hand-act record` for an act done outside the mesh —
|
||||
takes a required `--why` and writes an entry: who, which verb and arguments, why, when, the condition
|
||||
key it addresses if any, and a **cause** (the condition kind, or a word the person gives). `status`
|
||||
shows the week's count. A cause recorded twice within fourteen days raises `healer-wanted` (S15).
|
||||
|
||||
## 8. Staged core upgrades and rollback (rule 8)
|
||||
|
||||
**Health, per core component**, as probes in the registry (§4):
|
||||
|
||||
| Component | Healthy when | Witness that rolls it back |
|
||||
|---|---|---|
|
||||
| controller | holds the lease within 60 s of starting; `status` answers in full within 10 s; `doctor` ran once | the node-engine on the control node, which keeps the previous controller build installed beside the new one and reads the lease bucket |
|
||||
| node-engine | has reported its current declaration under its own build | its launcher (ADR 0141), which keeps the known-good |
|
||||
| node tools | announced, and answer a ping within 5 s | the node-engine, which keeps the known-good |
|
||||
| bus | every stream and durable consumer present (D6, D7); a request/reply round trip from every machine | none — a planned step, below |
|
||||
|
||||
**The gate.** ADR 0218's first machine for a core component is judged by that component's health
|
||||
probes passing three consecutive times over two minutes, within ten minutes of the apply. Only then do
|
||||
the other machines follow. "Reported applied" is not enough.
|
||||
|
||||
```
|
||||
merge ─► plan ─► first machine applies new build ─► health probes ×3 within 10 min?
|
||||
├─ yes ─► the rest follow, one tier at a time
|
||||
└─ no ──► witness restores previous build
|
||||
─► condition core.<component>.<node>.rolled-back (urgent)
|
||||
─► plan halts that component, records why
|
||||
```
|
||||
|
||||
- **Every core rollout leaves a record** in its plan: component, first machine, from and to build,
|
||||
verdict, time to verdict, rolled back or not. Read through `plans`.
|
||||
- **A rolled-back build is not retried** by the same plan. A newer merge makes a new plan.
|
||||
|
||||
**The bus is a planned step.** A bus upgrade is a maintenance step a person starts through the
|
||||
controller: the streams are snapshotted; a `bus-maintenance` condition is open for the step's duration;
|
||||
the bus is replaced; afterwards D6, D7 and the round trip must pass, or the step is reported failed and
|
||||
the snapshot is the way back. A step whose new version cannot be reverted (the bus's 2.10 → 2.11 is one)
|
||||
says so before it starts and runs only on a person's explicit word, recorded as a hand act. Whether the
|
||||
bus becomes a cluster that can be upgraded live is left to its own effort.
|
||||
|
||||
## 9. Before merge: facts and replays (rule 9)
|
||||
|
||||
**The facts snapshot.** The controller exports daily, and after any change of machines, assignments or
|
||||
seats: every machine (a stable pseudonym of the same length as its name, its role, operating system,
|
||||
C library, architecture, node-engine and node tools builds), assignments, settings keys and their
|
||||
non-secret values, seats and their holders, the catalogue commit, and the versions the mesh runs of
|
||||
the bus server, the store and the node-engine. No secret and no address: an address is replaced by one
|
||||
from a documentation range. It is kept in the artifact store as `facts/latest`, where the build seat
|
||||
reads it.
|
||||
|
||||
**The merge gate.** In mesh-controller, mesh-host and mesh-catalog, a check composes every machine of
|
||||
the snapshot with the change applied and runs the node-engine's validator over each. A change that
|
||||
makes any machine fail to compose or validate fails its check, naming the machine's role and the
|
||||
module. The resolver module's tests run under both C libraries the snapshot lists. A test in each core
|
||||
repository asserts that a pinned dependency's version equals the version the snapshot says runs.
|
||||
|
||||
**The replays.** mesh-lab carries a scripted scenario for each core incident, asserting the rule's
|
||||
outcome, run on every merge to mesh-controller, mesh-host and mesh-tools:
|
||||
|
||||
| Replay | Incident | Asserts |
|
||||
|---|---|---|
|
||||
| R1 | two controllers at once (204) | the second waits; nothing from the stale epoch is applied |
|
||||
| R2 | a reconcile due during a push (257, 261, 267) | one apply, one report, the newest sequence |
|
||||
| R3 | the node-engine self-updates during its report (230, 264) | the report arrives under the new build |
|
||||
| R4 | the bus's authorization reloads during a call (265) | the call's outcome is readable by id |
|
||||
| R5 | a consumer with several filter subjects under mixed traffic (266) | no announcement is skipped; S5 fires if one is |
|
||||
| R6 | an unreadable contributions file (241) | refused by name; nothing withdrawn; a condition |
|
||||
| R7 | the controller rebuilds itself mid-plan (214) | the plan continues under the new epoch |
|
||||
| R8 | a broken controller, node-engine and node tools build | each rolled back with no hand; condition and message |
|
||||
| R9 | each signal of §3 suppressed | its condition within its bound, cleared on return |
|
||||
|
||||
A new core issue resolves with its replay added, or with a stated reason none is possible. The live
|
||||
mesh stays the test bed ([ADR 0149](../../02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md)): a
|
||||
replay covers what must not be done to it on purpose, and every rule keeps a live check.
|
||||
|
||||
---
|
||||
|
||||
## 10. The phases
|
||||
|
||||
Ordered by risk removed per day: detection first, because it covers every class including those not
|
||||
met yet. A phase is done when its *done when* holds; this section records the date when it does.
|
||||
|
||||
### Phase 0 — Finish what is in flight
|
||||
|
||||
| Repository | Delivers |
|
||||
|---|---|
|
||||
| mesh-controller | the located fixes of 244, 265, 266, 267 rolled out; `calls` moved into `mesh-controller_calls`; `status` answered from a summary the event loop keeps current, inside ten seconds; the hand-act log with `--why` on the repairing verbs and `hand-act record`; recording the durations the bounds come from (apply duration per machine, heartbeat gaps, plan tier durations, build durations) |
|
||||
| mesh-host | the fixes of 264 and 257/261 rolled out to every machine |
|
||||
| mesh-tools | the console passing a mesh seat's `node` (244) rolled out |
|
||||
| mesh-catalog | the bus's 2.10 → 2.11 upgrade, done as the first planned bus step by hand (snapshot, announced, checked after), recorded as a hand act |
|
||||
|
||||
**Done when:** the four located core issues resolve with their live checks; a controller restart keeps
|
||||
every call's outcome; `status` answers in full within ten seconds five times in a row; the hand-act log
|
||||
has a week of entries; the durations are recorded for every machine.
|
||||
|
||||
### Phase 1 — The mesh says when it is wrong
|
||||
|
||||
| Repository | Delivers |
|
||||
|---|---|
|
||||
| mesh-controller | the condition store, its verbs, history and events (§2); ADR 0224's standing moved into it; watchdogs for S1–S10 and S12–S14 with bounds set from Phase 0's durations; the bus advisories subscribed and translated (S9); `doctor` with D1–D10 and its heartbeat (§4); `status` led by open conditions; the test generated from the signals table |
|
||||
| mesh-host | the heartbeat carries its interval; D1's validator published as a package the controller imports |
|
||||
| mesh-tools | the node tools' heartbeat (S11) |
|
||||
| mesh-catalog | the `operator-channel` seat and its holder; the Telegram channel and the desktop notifier contributing to it; `mesh-watcher` on a machine other than the control node (§5) |
|
||||
| mesh-lab | R9: each signal suppressed in turn |
|
||||
|
||||
**Done when:** on a lab mesh, suppressing each signal raises its condition within its bound and sends a
|
||||
message; restoring it clears both. Stopping the controller makes the watcher send *self-check silent*
|
||||
within twice the self-check's interval. Live: a week of conditions read back, every one real or its
|
||||
bound corrected in the table.
|
||||
|
||||
### Phase 2 — Order and one writer
|
||||
|
||||
| Repository | Delivers |
|
||||
|---|---|
|
||||
| mesh-controller | the lease and epoch (§6); plans written by compare-and-set; a report kept by sequence; abandoned calls marked; S12, S13, D5; the writers table enforced at grant composition; contract tests for every consumed subject and the check listing them; the empty-on-error lint |
|
||||
| mesh-host | one apply queue; the report sequence kept on disk; epoch refusal reported; the `report` verb; the contract tests and lint |
|
||||
| mesh-tools | refusing an unreadable or unknown input by name; the lint |
|
||||
| mesh-catalog | the withdrawal brake in the providers' loop: a reconcile that would withdraw more than one consumer, or a set fraction, stops and raises a condition |
|
||||
| mesh-lab | R1, R2, R6 |
|
||||
|
||||
**Done when:** R1 ends with nothing from the stale epoch applied, R2 with one report of the newest
|
||||
sequence, R6 with the withdrawal braked; every consumed subject has its contract test.
|
||||
|
||||
### Phase 3 — Healers
|
||||
|
||||
| Repository | Delivers |
|
||||
|---|---|
|
||||
| mesh-controller | the healer registry and H1–H4, each braked and said; S15 |
|
||||
| mesh-lab | an induced failure per healer |
|
||||
|
||||
**Done when:** a lab mesh recovers from each induced failure with no hand, says so, and brakes after
|
||||
its budget. Live: a week with no cause repeated in the hand-act log.
|
||||
|
||||
### Phase 4 — Core upgrades that roll back
|
||||
|
||||
| Repository | Delivers |
|
||||
|---|---|
|
||||
| mesh-controller | the health probes of §8; the gate on a core component's first machine; the rollout record in the plan; the bus maintenance step as a verb |
|
||||
| mesh-host | keeping the previous controller and node tools builds; restoring one when its health is not met in bound, watching the lease bucket for the controller |
|
||||
| mesh-tools | answering the health ping |
|
||||
| mesh-lab | R3, R7, R8 |
|
||||
|
||||
**Done when:** on a lab mesh, a broken build of the controller, the node-engine and the node tools —
|
||||
one that starts and does nothing, one that crashes, one that cannot reach the bus — is each rolled
|
||||
back with no hand, the mesh ends on the previous build, and a condition and a message say so. Live: the
|
||||
next three core rollouts each record a health verdict.
|
||||
|
||||
### Phase 5 — Checks before merge, and the replays
|
||||
|
||||
| Repository | Delivers |
|
||||
|---|---|
|
||||
| mesh-controller | the facts snapshot and S14 |
|
||||
| mesh-controller, mesh-host, mesh-catalog | the compose-and-validate merge gate; versions tested as run; the resolver's tests under both C libraries |
|
||||
| mesh-lab | R4, R5 and the rest of the window's incidents; running the replays on every core merge |
|
||||
|
||||
**Done when:** the replays of 236, 262, 263 and 266 fail on the commit before their fix and pass after;
|
||||
a new core issue cannot resolve without a replay or a stated reason.
|
||||
|
||||
## What is not decided here
|
||||
|
||||
- The bus as a cluster of three, to upgrade it live.
|
||||
- Routing by presence, quiet hours, answering back through a channel, and an external dead-man
|
||||
service — research 028.
|
||||
- A condition that needs judgement handed to an agent as work — research 017.
|
||||
@@ -45,6 +45,7 @@ document is written and this one's status becomes `implemented`.
|
||||
| [`38-building-the-operators-machine.md`](38-building-the-operators-machine.md) | **In progress.** The work of design 37 as packages: the runtime serves many modules, the controller composes one per node, the console becomes its serving mode, the packet filter moves first, then the shell and the service manager — tested on the live mesh by the operator's decision | [ADR 0175](../../02-DECISIONS/0175-one-tool-runtime-per-node-serves-every-modules-tools-on-the-host-side.md), [0160](../../02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md), [0149](../../02-DECISIONS/0149-the-live-mesh-is-the-test-bed.md) |
|
||||
| [`41-the-shell-and-the-accounts-environment.md`](41-the-shell-and-the-accounts-environment.md) | **In progress.** The shell and the account's environment as modules: an environment module every module contributes variables and `PATH` entries to, shell code contributed to the login shell in named slots, the prompt and plugins as modules, the host giving a login back, and the service manager's module finished | [ADR 0203](../../02-DECISIONS/0203-the-accounts-environment-is-one-modules-and-every-module-contributes-to-it.md), [ADR 0204](../../02-DECISIONS/0204-a-module-contributes-shell-code-to-the-login-shell-in-named-slots.md), [ADR 0205](../../02-DECISIONS/0205-software-the-distribution-does-not-package-ships-as-a-pinned-archive-of-the-module.md) |
|
||||
| [`42-the-machines-modules-in-order.md`](42-the-machines-modules-in-order.md) | **In progress.** The order the machines' modules of research 026 and 027 are built and rolled out: every machine's first (sudo, localization, time sync, pacman, logrotate, avahi, systemd, docker, `~/.ssh`, scripts, kernel), then both workstations', then one machine model's; each proven on one workstation before the rest | [ADR 0173](../../02-DECISIONS/0173-the-operators-machine-is-the-meshs-and-a-module-is-what-it-declares.md), [ADR 0182](../../02-DECISIONS/0182-inside-a-home-the-mesh-owns-what-it-places-and-holds-the-rest-as-found.md), [ADR 0205](../../02-DECISIONS/0205-software-the-distribution-does-not-package-ships-as-a-pinned-archive-of-the-module.md) |
|
||||
| [`45-a-core-that-cannot-fail-silently.md`](45-a-core-that-cannot-fail-silently.md) | **Designed.** The core says when it is wrong, refuses what is stale or unreadable, heals what it knows, upgrades one machine at a time with a witness that rolls it back, and is checked against the real mesh before merge: the writers and signals tables, the condition store, `doctor`, the minimal output channel, healers, the lease and report order, staged upgrades, the facts snapshot and replays, in six phases | [ADR 0227](../../02-DECISIONS/0227-the-core-holds-nine-rules-each-checked-and-is-built-to-them-in-six-phases.md), [ADR 0224](../../02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md), [ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md) |
|
||||
|
||||
## Not yet written
|
||||
|
||||
|
||||
+58
-4
@@ -1,9 +1,9 @@
|
||||
---
|
||||
status: resolved
|
||||
status: located
|
||||
opened: 2026-10-01
|
||||
located-in: [the identity provider's assignment on the control node (an adopted database whose admin predates the mesh), mesh-catalog modules/keycloak (the minted `admin` own-secret, applied by the server only when it creates its master realm)]
|
||||
fixed-by: done by hand on 2026-10-01 through the server's own bootstrap command — the admin's password set to the value the mesh minted; no code changed
|
||||
amended-design:
|
||||
located-in: [the identity provider's assignment on the control node (an adopted database whose admin predates the mesh, and the same database moved on 2026-10-05), mesh-catalog modules/keycloak (the minted `admin` own-secret, applied by the server only when it creates its master realm), every provider's provisioner loop (a consumer failed for a day said so only in a journal), mesh-controller status (nothing read what a provider could not do)]
|
||||
fixed-by: twice by hand through the server's own bootstrap command (2026-10-01, 2026-10-05); the safety nets in mesh-controller PR #70, mesh-host PR #28 and mesh-catalog PR #80 (ADR 0224), resolved when they are merged and rolled out
|
||||
amended-design: 03-DESIGN/01-to-be/19-the-module-protocol.md
|
||||
---
|
||||
|
||||
# 179 — An adopted identity provider's admin never took the secret the mesh minted
|
||||
@@ -44,6 +44,60 @@ only the download client's: a credential the mesh cannot make is the operator's
|
||||
*How it was checked:* `keycloak_list_realms` through the console answers; `keycloak_list_clients`
|
||||
on the realm lists both mesh-named clients; the server's log stops the five-second login error.
|
||||
|
||||
## Recurred, 2026-10-05
|
||||
|
||||
**The identity provider's database was moved that morning, and the admin's old password came back
|
||||
with it.** The record the 2026-10-01 repair wrote lived in the database; the database that came up
|
||||
on the new store was taken from before that repair, so the realm's `admin` again held a password that
|
||||
predates the mesh, and the server — which applies the minted variable only when it creates its master
|
||||
realm — did not touch it. From shortly after midnight the provisioner failed every consumer every five
|
||||
seconds with *401 invalid_grant, Invalid user credentials*: about 31,000 refused logins until it was
|
||||
repaired by hand, the same way as the first time, at about 23:55. In those twenty-three hours nothing
|
||||
anywhere said so except the provider's own journal. The machines applied what they were sent, the
|
||||
module was current, and `status` printed its all-well sentence.
|
||||
|
||||
**The repair, as done both times**, inside the server's container and with nothing printed: the
|
||||
server's `bootstrap-admin user` command creates a temporary administrator, from a password generated
|
||||
in the container and handed over in an environment variable, on a management port other than the
|
||||
default (the running server holds that one); the admin client logs in as it and sets `admin`'s password
|
||||
to the value the mesh minted; the temporary administrator and every temporary file are removed.
|
||||
|
||||
## What makes it not happen silently again (ADR 0224)
|
||||
|
||||
Twice by hand is a pattern, and the operator's direction was that it never happen again: detected
|
||||
automatically, repaired automatically where safe, loud where not, and tested.
|
||||
[ADR 0224](../../02-DECISIONS/0224-a-provider-that-keeps-failing-a-consumer-is-a-problem-the-controller-reports.md)
|
||||
records the rule; three safety nets implement it.
|
||||
|
||||
1. **The identity provider repairs its own admin.** Its code — ported from TypeScript to Go with this
|
||||
change — checks that `admin` logs in with the mesh's secret once the server answers, every five
|
||||
minutes after, and at once whenever the provisioner is refused. A refusal is repaired exactly as
|
||||
above, by the module, inside the server's container through the container runtime it already
|
||||
declares, with the secrets passed on standard input and never on a command line; then the login is
|
||||
checked again and one line and an `admin.repaired` event say what was done and why. A repair that
|
||||
fails is announced as `admin.unrepaired`, said loudly in the journal with a pointer here, and braked —
|
||||
ten minutes, doubling to six hours — and while the admin is refused the provisioner stops asking the
|
||||
server, so a lockout policy is never provoked. `keycloak_admin_check` reports the admin's state and,
|
||||
with `repair`, repairs it now.
|
||||
2. **A provider that keeps failing a consumer says so.** The provisioner loop announces a consumer it
|
||||
has failed for five minutes without a success — create, the periodic check or an unreadable secret —
|
||||
as `provisioner.failing`, with the class of error, and again every fifteen minutes; its recovery as
|
||||
`provisioner.recovered`.
|
||||
3. **`status` names it.** The controller keeps each provider's newest failing word per consumer, and
|
||||
`status`, its JSON and `node show` list it; it breaks the all-well sentence. On 2026-10-05 the first
|
||||
line of `status` would have been the identity provider failing both consumers with
|
||||
`credentials-rejected`, five minutes after midnight.
|
||||
|
||||
The first is specific to this module and is the repair; the second and third are for every provider,
|
||||
and are what would have said it on the day even if the repair had not existed.
|
||||
|
||||
*How it is checked:* the module's tests drive the repair against a fake server and a fake container
|
||||
runtime — repaired and verified on a refusal, never while the server is unreachable, braked after a
|
||||
failed repair, no secret on any command line, the provisioner short-circuited while refused; the loop's
|
||||
tests drive the announcements; the controller's tests take a failing word from the bus to `status` and
|
||||
`node show` and back out on recovery. Live, after rollout: a deliberately wrong admin password on a
|
||||
throwaway server is repaired within five minutes, and `status` stays all-well while it is.
|
||||
|
||||
## Open
|
||||
|
||||
The manifest still says the admin password is the mesh's to mint, which is true of a fresh install
|
||||
|
||||
+36
-3
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: located
|
||||
status: resolved
|
||||
opened: 2026-10-01
|
||||
located-in: [mesh-catalog modules/dnsmasq, mesh-controller internal/overlay/generator.go, mesh-controller internal/catalogue/resolve.go (checkResources)]
|
||||
fixed-by:
|
||||
located-in: [mesh-catalog modules/dnsmasq, mesh-catalog modules/docker, mesh-controller internal/overlay/generator.go, mesh-controller internal/catalogue/resolve.go (checkResources), mesh-controller internal/catalogue/seat_into.go, mesh-host internal/apply/apply.go (orphan removal)]
|
||||
fixed-by: novox/mesh-controller#62 (d84c969), novox/mesh-catalog#72 (5c89eaf), novox/mesh-host#26 (a192bcc), novox/mesh-controller#63 (c34b937)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -71,6 +71,32 @@ gives that module declared settings with defaults. The fix, once both are accept
|
||||
5. The collision check sees a computed module's resources as well, so a second writer cannot come
|
||||
back through generated code.
|
||||
|
||||
> **Where it stands, 2026-10-05.** Step 1 is done: the resolver module writes nothing into the
|
||||
> runtime's file, and the runtime's module (the holder of `node-container-runtime`) writes
|
||||
> `live-restore` and reloads its own service. No module writes `dns`
|
||||
> ([ADR 0196](../../02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md)).
|
||||
> That answers the first two open questions below.
|
||||
>
|
||||
> Steps 2, 3 and 5 are decided in
|
||||
> [ADR 0222](../../02-DECISIONS/0222-a-module-is-told-where-a-mesh-seats-holder-is-reached-and-the-controller-writes-no-file-a-seats-holder-owns.md),
|
||||
> and the operator's rule is general: the controller never writes a file a seat's holder owns; it
|
||||
> tells the owner. The registry reaches the runtime's module as a value the mesh holds, not as a
|
||||
> setting: `${seat:mesh-artifact-store:reach}`, with no binding. The ADR 0082 and ADR 0102 notes
|
||||
> step 2 asks for are written.
|
||||
>
|
||||
> Step 4's single push is replaced by an order. A list member written by two records is tolerated
|
||||
> by the host, which a scalar key in two modules was not. Four pull requests, merged in this order:
|
||||
>
|
||||
> 1. The controller learns the placeholder (mesh-controller #62).
|
||||
> 2. The runtime's module states `insecure-registries` with it (mesh-catalog #72).
|
||||
> 3. The host leaves a unit alone when another declared service still holds it. It is deployed
|
||||
> before 4. Without it, removing the private network's record of the runtime's service gives back
|
||||
> what that record found (mesh-host #26).
|
||||
> 4. The controller stops generating the private network's two resources, and checks generated
|
||||
> resources for collisions (mesh-controller #63).
|
||||
>
|
||||
> `fixed-by:` is filled when they merge.
|
||||
|
||||
## Open questions
|
||||
|
||||
- **How the resolver's address reaches the runtime.** Either the resolver seat (`node-dns-resolver`)
|
||||
@@ -83,3 +109,10 @@ gives that module declared settings with defaults. The fix, once both are accept
|
||||
- **The adopted machine's predecessor values.** The runtime module adopting a file with a
|
||||
hand-written `dns` and `live-restore: false` replaces both. That is intended, and is the one
|
||||
restart the operator must make on that machine.
|
||||
|
||||
## Verified
|
||||
|
||||
2026-10-05: on every machine `insecure-registries` is now the runtime's module's own, filled from the
|
||||
store's seat; the private network's two generated resources are gone. The handover kept the list
|
||||
member throughout, and the old reload record was forgotten, not given back ("docker.runtime still
|
||||
holds the unit"), so the runtime was never stopped.
|
||||
|
||||
+8
-2
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: open
|
||||
status: resolved
|
||||
opened: 2026-10-04
|
||||
located-in: [mesh-host internal/apply/apply.go, mesh-host internal/apply/schedule.go]
|
||||
fixed-by:
|
||||
fixed-by: mesh-host PR #23
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -62,3 +62,9 @@ is a hole in something new rather than something that broke.
|
||||
at the cost of a second source for "is this container meant to be running".
|
||||
- Either way: should the *report* say a window is open, so a machine that looks half-stopped at
|
||||
03:31 reads as working rather than broken?
|
||||
|
||||
## Resolved — 2026-10-05
|
||||
|
||||
An apply that arrives during a window now leaves the containers the window holds alone, and reports them `held-still`, naming the step. The first apply after the window converges them. The window is recorded under the host's state directory, the one place every applier on a machine shares: the daemon, a hand-run apply and the installer. It is released on every path, and lapses after six hours or when its process is gone. The machine's report lists open windows. Live on all four machines the same evening. Of the issue's two options, the narrower was taken: nothing blocks, so a push is never held for the length of a window.
|
||||
|
||||
Accepted and said in the change: an apply that inspected a container as running in the instant a window opens can still recreate it.
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: located
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-catalog modules/claude-code, mesh-catalog modules/claude-licence-manager]
|
||||
fixed-by:
|
||||
fixed-by: mesh-catalog PR #45
|
||||
amended-design:
|
||||
---
|
||||
|
||||
|
||||
@@ -31,3 +31,9 @@
|
||||
**Ruled out.** The bus delivered every binding: each machine that took the fourth binding did so in the
|
||||
same second it was published. The licences themselves were sound. The remaining licence refreshed
|
||||
on every attempt, and the second account's login was adopted from its first report.
|
||||
|
||||
## Resolved — 2026-10-05
|
||||
|
||||
Live on all four machines and the manager the same morning. After the restart every machine reported
|
||||
the generation it held, and the manager's bindings matched them. The manager logged no failure while
|
||||
moving its sequence past them.
|
||||
|
||||
+21
-1
@@ -1,5 +1,5 @@
|
||||
---
|
||||
status: open
|
||||
status: located
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-tools, mesh-controller]
|
||||
fixed-by:
|
||||
@@ -60,6 +60,26 @@ and it is wrong.
|
||||
schema describes — a test that calls each with its schema's required set and asserts the answer
|
||||
is not "needs" an argument the schema does not have.
|
||||
|
||||
## Seen again, 2026-10-05 and 2026-10-06 — and once it did harm
|
||||
|
||||
The same picture for more of the mesh's own verbs, read through the console:
|
||||
|
||||
```
|
||||
mesh_describe mesh-controller.push → {"properties": {}}
|
||||
mesh_describe mesh-controller.plan → {"properties": {}}
|
||||
mesh_describe mesh-controller.assign → module only, no node
|
||||
mesh_describe mesh-controller.pin → provision, from, module — no node
|
||||
mesh_describe mesh-controller.settings → module, values, clear — no node
|
||||
```
|
||||
|
||||
`build` and `queue`, which take no machine, described in full. For `plan` this is the refusal above.
|
||||
For **`push` it was worse than a refusal**: a push takes no required argument, so a call naming one
|
||||
machine arrived empty, ran as a push of every machine behind, and answered "N node(s) told" — an
|
||||
operator who believed they had pushed one machine had pushed the mesh. Nothing in the answer said the
|
||||
machine had been lost.
|
||||
|
||||
The diagnosis is in [`01-diagnosis.md`](01-diagnosis.md).
|
||||
|
||||
## Worked around, for now
|
||||
|
||||
Through `<node>/node-login-shell.execute`, running the controller's own command line on the
|
||||
|
||||
+63
@@ -0,0 +1,63 @@
|
||||
# Diagnosis
|
||||
|
||||
## 2026-10-06
|
||||
|
||||
**The schema is not empty where it is kept.** The controller's verb table declares `node` for every
|
||||
verb that acts on a machine, and so does the seat's record: asked through the console,
|
||||
`mesh-controller.tools` answers `push`, `plan`, `assign`, `pin` and `settings` each with `node` among
|
||||
their properties. The controller's announcement on the bus carries the same schemas. So the
|
||||
controller publishes them whole, and the loss is after it.
|
||||
|
||||
**The console takes `node` out of every schema and every call.** The console's address grammar
|
||||
(ADR 0195) puts the machine in the address for a seat held on every machine (`<node>/<seat>.<verb>`)
|
||||
and for a module on one machine (`<node>/<module>.<tool>`), and so describes those tools without a
|
||||
`node` argument and deletes one from their calls. It did both for **every** address — including a
|
||||
seat held once for the mesh, where there is no machine in the address and `node` is the verb's own
|
||||
argument: the machine the verb acts on. That is exactly the set observed: each verb lost `node` and
|
||||
nothing else, `push` and `plan` had nothing else and described as empty, and the verbs that take no
|
||||
machine (`build`, `queue`) were untouched. The older flat listing of the same console had this right
|
||||
(it kept a mesh seat's schema as declared); the address form did not carry the distinction over.
|
||||
|
||||
**Why the verb then refused, or did the wrong thing.** The controller's verbs read the arguments they
|
||||
know and ignore the rest. Given nothing, `plan` and `node` said they needed `node`; `push`, whose
|
||||
`node` is optional, read its absence as "every machine behind" and did that. A verb that ignores what
|
||||
it does not read cannot tell "nothing was asked" from "something was lost on the way".
|
||||
|
||||
**Ruled out.**
|
||||
|
||||
- A stale seat record. The record is older than the controller's current source (some verbs lack
|
||||
arguments added since), but it carries `node` wherever the source does.
|
||||
- The controller's announcement. It sends the record's schema per verb, unchanged.
|
||||
|
||||
**Located in** `mesh-tools` (the console: the schema it describes and the arguments it sends) and
|
||||
`mesh-controller` (verbs that pass over what they are given).
|
||||
|
||||
## What the fix does, and how it is checked
|
||||
|
||||
Pull requests: novox/mesh-tools#14 and novox/mesh-controller#71 — open, not merged.
|
||||
|
||||
- **The console** takes `node` out only where the address names the machine. There, a `node` naming
|
||||
the same machine is redundant and dropped; one naming another machine is refused. A call to a seat's
|
||||
verb carrying an argument the verb does not declare is refused, naming it — the seat's record is the
|
||||
mesh's own description of the verb, so the console can judge it.
|
||||
- **The controller** refuses, naming the argument: one the verb does not declare; one it declares but
|
||||
did not use for the command it composed (two arguments where one wins, or half of a two-argument
|
||||
shape); a switch that is neither "true" nor "false". A push that names no machine says, as the
|
||||
first line of its answer, that it is a push of the whole mesh and which machines it is sending, and
|
||||
its last line names the machines told. `plan` gains `files`, `push` gains `behind` (the whole mesh
|
||||
said outright, refused beside `node`), and the listing verbs a `limit`.
|
||||
- **Checked by tests that walk every verb the controller announces:** for every combination of a
|
||||
verb's declared arguments, the verb either refuses or every argument given changes the command it
|
||||
runs; every verb refuses an argument it does not declare, and `node` where it takes none; and every
|
||||
flag of the command a verb runs — read from the command's own source — is an argument of the verb's
|
||||
schema or is listed, with the reason, as set by the verb or withheld from it. A verb added later,
|
||||
or a flag added to a command, fails these until its schema says how a caller reaches it.
|
||||
|
||||
What the tests do not check is the console's own describe-and-call path against a live bus; its unit
|
||||
test covers the choice of when `node` is the address's and when it is the verb's.
|
||||
|
||||
## Related
|
||||
|
||||
Issue 259 is another way a named push reached every machine: there the machine arrived and the
|
||||
controller's flush after it sent the others. Different cause, same reader's harm — the answer of a
|
||||
push is the only place that says how far it went, which is why the whole-mesh line is in the answer.
|
||||
+50
@@ -0,0 +1,50 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-05
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 245 — `status` calls a module behind when only its repository moved
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-05. After three catalogue merges that changed 11 modules, `status` listed **69 modules
|
||||
behind their source**, each as `holds <older commit>, source has <newer commit>`, with the advice
|
||||
"`build --behind` builds them; `push --behind` sends them on".
|
||||
|
||||
Between the two commits, `git diff --name-only` shows changes under 11 module directories only. For
|
||||
58 of the 69 modules listed, for example a Bluetooth module, the container runtime's, the forge's,
|
||||
the package manager's and the bus's, the diff of the module's own path is empty. Their sources did
|
||||
not change; only the repository's commit did.
|
||||
|
||||
The merges' plans were right: they rebuilt the changed modules and those that depend on them, by
|
||||
tier. Only the report was wrong. An agent following the report's own advice ran `build --behind`,
|
||||
which rebuilt all 69. The rebuilt bus module was rolled out, and its container was replaced on the
|
||||
control node. Every node's runtime lost the bus for about a minute.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
"Behind" is the word a person and an agent act on, and `status` attaches a command to it. A module
|
||||
is held to a commit of its repository, so after any merge almost every module of that repository
|
||||
reads behind. The list then says nothing about what needs building: it hides the few modules that
|
||||
really are behind among the many that are not, and it invites a rebuild of everything, which is not
|
||||
a harmless act (above).
|
||||
|
||||
## The operator's direction (2026-10-05)
|
||||
|
||||
The only truth is the outcome of the build plan. A plan already decides, from a change, which modules
|
||||
it affects: those whose sources changed and those that depend on them, tier by tier. "Behind" means
|
||||
a module that a plan has decided to rebuild and has not yet rebuilt or rolled out, and nothing else.
|
||||
No second comparison beside the plan is made, whether of commits, of folders or of files, because a
|
||||
second answer to the same question is how the two came to disagree.
|
||||
|
||||
## Open questions
|
||||
|
||||
1. Where does `status` read "behind" from today, and what replaces it: the open plans' remaining
|
||||
tiers?
|
||||
2. What does a module's recorded commit mean once a plan that leaves it untouched has run? Does it
|
||||
move forward, or does the record stop carrying a commit that only says when it was last built?
|
||||
3. Should `build --behind` and `push --behind` take their lists from the same place, so that they can
|
||||
never act on a module no plan named?
|
||||
+65
@@ -0,0 +1,65 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-tools]
|
||||
fixed-by: mesh-tools pull request 13 — discovery waits for every runtime that answered PING, names the ones it missed, and mesh_runtimes says who answered
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 246 — The console says a module runs nowhere when a runtime answers late
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-05. On the laptop, the console answered `mesh_machine` for the laptop with no modules and no
|
||||
seats, and a call to one of the laptop's modules with "nothing in the mesh is called slack". The
|
||||
laptop's own runtime said at the same moment that it served 275 tools for 48 modules. A minute
|
||||
earlier, a call to another laptop module had been answered with "it does not run on the laptop; it
|
||||
runs on the workstation", and a retry of the same call worked. Nothing was logged anywhere.
|
||||
|
||||
An agent worked around it by calling the runtime's local MCP port directly, which gives the right
|
||||
answer and bypasses everything the console stands for: one way in, one account, one record of what
|
||||
was called. The operator asked for a tool instead.
|
||||
|
||||
## What was measured
|
||||
|
||||
A read-only probe on the laptop's own runtime credential timed the discovery answers over 25 rounds
|
||||
against the live bus, whose round trip from the laptop was about 40 ms:
|
||||
|
||||
- The laptop's runtime answer was the largest on the mesh, at about 164 kB with 341 endpoints. That
|
||||
is far below the bus's message limit, and it was never shortened.
|
||||
- It arrived last in every round: a median of about 365 ms, and once 813 ms. The other runtimes
|
||||
answered within 180 to 275 ms, and the controller within 50 ms.
|
||||
- The console gathered discovery answers for a fixed 750 ms. Inside the console, the gather runs
|
||||
beside two controller calls and about 570 kB of answers on the same link, and it is slower still
|
||||
while a runtime re-serves after a restart.
|
||||
|
||||
## Root cause
|
||||
|
||||
Discovery decided who was there by who answered in a fixed window. A late answer was not a failure
|
||||
to anyone, so nobody said it. The index simply lacked that runtime, and every answer built on the
|
||||
index then stated as fact that the runtime's modules did not exist, or ran only elsewhere.
|
||||
|
||||
Ruled out by measurement or by reading the code: an answer too large for the bus, subscriptions lost
|
||||
when the bus reconnects, the console not counting its own machine's answer, and the merge of two
|
||||
answers dropping a machine.
|
||||
|
||||
## Resolution
|
||||
|
||||
- Discovery asks who is there (PING, a hundred bytes, answered at once) beside what each serves
|
||||
(INFO). It waits at least the old window, and then up to five seconds for every instance that said
|
||||
it is there, so a large answer is waited for and a quiet mesh costs nothing extra.
|
||||
- A runtime that said it is there and did not say what it serves in time, or that answered recently
|
||||
and not now, is named. While one is unheard, the console never says an address is missing or runs
|
||||
elsewhere: it says which runtime was not heard, and where the controller's records place the module.
|
||||
- A new console tool, `mesh_runtimes`, says for every runtime how long its answer took, how large it
|
||||
was, how many modules and tools it announced, whether it was shortened, and when it was last heard,
|
||||
and which runtimes or machines were not heard.
|
||||
- An announcement still too large after its descriptions are cut to their first line now leaves the
|
||||
descriptions out, and says so.
|
||||
|
||||
## How it is checked
|
||||
|
||||
The fix ships with tests against a real bus: a runtime that answers after the old window is found and
|
||||
called (the same test fails with the fixed window), a runtime that answers PING and never INFO is
|
||||
named, and a restarted runtime, which answers under a new instance, is not reported as missed. Live,
|
||||
`mesh_runtimes` shows every machine's answer and its time.
|
||||
@@ -0,0 +1,52 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-05
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 247 — A module cannot put the operator's account in a group
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-05. The module for a peripheral-lighting daemon was assigned to the laptop. Its package
|
||||
installs the daemon and creates the daemon's group. The daemon then refuses to start: "User is not a
|
||||
member of the openrazer group". The device files are the group's, so the daemon cannot reach the
|
||||
devices.
|
||||
|
||||
The module's own check names the fix, which is to add the account to the group and log in again. No
|
||||
module can declare that fix:
|
||||
|
||||
- The account is a `user` resource, and the login-shell module already declares it, to set its shell.
|
||||
- A second module that declares the same account, only to add one group, is refused as a duplicate
|
||||
name.
|
||||
- There is no resource for one membership on its own. A whole-account declaration that lists groups
|
||||
would also take from the account every group it does not list, including the operator's own.
|
||||
|
||||
So the step is done by hand, with `sudo`, outside the mesh, and nothing records why the account is in
|
||||
the group.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
More modules need this than this one: input devices (`input`), serial ports (`uucp`), the container
|
||||
runtime (`docker`), virtual machines (`libvirt`, `kvm`), and capture or scanner hardware. Each is a
|
||||
fact a module knows and the operator's account needs. Today every one is a hand step that survives a
|
||||
reinstall only by memory. A membership added by hand is also never taken away when the module that
|
||||
needed it is unassigned.
|
||||
|
||||
## Open questions
|
||||
|
||||
1. Is a membership its own resource (account, group), held by the module that needs it and given
|
||||
back on undeclare? Or is it a contribution to the account's holder, in the way ADR 0212 lets a
|
||||
module contribute to a seat?
|
||||
2. A membership takes effect at the next login. How does the module say so: a finding, or a
|
||||
moment the power or session seat already knows?
|
||||
3. What does undeclare do with a membership the account already had before any module declared it?
|
||||
The host keeps what it found and gives it back, as it does with a whole file it wrote over.
|
||||
|
||||
## How it is checked
|
||||
|
||||
When fixed, assigning the lighting module to a machine whose account is not in the group puts the
|
||||
account in the group, says that a new login is needed, and leaves the account's other groups as they
|
||||
were. Unassigning it removes only a membership the module added.
|
||||
+58
@@ -0,0 +1,58 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-controller internal/broker]
|
||||
fixed-by: mesh-controller PR #51
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 248 — The controller's event consumer replayed a week, and held every new merge behind it
|
||||
|
||||
## What was observed
|
||||
|
||||
A merge to the catalogue at 15:17 never reached the controller. No plan was made and no build was asked
|
||||
for it. Every machine kept running the build from before it. The merge just before, to the record, was
|
||||
logged twice.
|
||||
|
||||
The bus showed why. The controller's durable consumer on the event stream was set to deliver
|
||||
**everything the stream holds**, not from where it last was:
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| the stream | a week of events, from message 1 514 to message 356 004 |
|
||||
| the consumer delivered to | message 1 517, later 4 898 |
|
||||
| acknowledged to | 0, later 1 569 |
|
||||
| still to deliver | 6 955 merges and build outcomes, a week old |
|
||||
| allowed outstanding | one at a time ([issue 175](../175-an-announcement-behind-a-long-build-comes-back/00-report.md)) |
|
||||
|
||||
So the controller was working through a week of past merges and build outcomes, one at a time, slowly. Every
|
||||
new one — a merge, a build asked by hand — waited behind them. A build asked by hand finished on its
|
||||
machine and was never registered. While this went on, the controller's client dropped messages
|
||||
("slow consumer") several times, because heartbeats and reports share the loop with these events. From 15:17
|
||||
its event loop did nothing more: no line logged, the one delivered event never acknowledged.
|
||||
|
||||
How the consumer came to deliver everything is not certain. The bus and the stores were rebuilt the night
|
||||
before ([issue 241](../241-one-unreadable-grants-file-dropped-every-database-on-the-control-node/00-report.md)).
|
||||
A consumer that is missing is made again by the controller's own assertion, and that made it with the
|
||||
server's default, which is everything. [Issue 207](../207-a-re-made-worker-replayed-every-ask-the-stream-kept/00-report.md)
|
||||
closed this for a consumer re-made because its type changed. It did not close it for one that is simply
|
||||
not there.
|
||||
|
||||
## Why it matters
|
||||
|
||||
**Delivery stops, and nothing says so.** Status showed every plan done and no machine behind. The merge
|
||||
that was missed is not "behind", because no plan was ever made for it. A replay of past merges can also
|
||||
act on them again. Issue 207 records nine modules re-registered from the past the same way.
|
||||
|
||||
**The way out was a hand on the bus.** No verb resets a consumer. On the operator's explicit word, the
|
||||
consumer was re-made from now with a one-off program run as the controller, otherwise unchanged. The
|
||||
controller's loop still held the old event afterwards, so the new consumer's first delivery went
|
||||
unacknowledged; the loop needs a restart of the controller to let go of it.
|
||||
|
||||
## Noticed alongside, not this issue
|
||||
|
||||
- Each merge to the catalogue planned 99 to 100 modules in two tiers and rebuilt modules it did not touch.
|
||||
The output was byte-identical, so nothing was redeployed, but it costs minutes of the build machine
|
||||
per merge.
|
||||
- Three plans for three merges ran over each other. Each sent the build agent to every machine and asked
|
||||
for the same builds. Nothing supersedes a plan for an older commit.
|
||||
+54
@@ -0,0 +1,54 @@
|
||||
# 248 — Diagnosis
|
||||
|
||||
## 2026-10-05
|
||||
|
||||
1. A merge was not in the controller's log. The forge's own log showed nothing about delivering it, so the
|
||||
question moved to the bus.
|
||||
2. The bus's backlog tool named the controller's event consumer: 6 955 pending, redeliveries, one
|
||||
unacknowledged. Its configuration was read from the server's monitoring endpoint: deliver policy
|
||||
*all*, one outstanding, 30 seconds to acknowledge, five deliveries.
|
||||
3. Sampled three times over a minute it did not move, and over the following hour it crawled forward
|
||||
through week-old events. The controller's log held only its client's warnings: dropped messages,
|
||||
and one refused reply.
|
||||
4. The controller's code makes a missing consumer with the configuration it asserts, which sets no
|
||||
deliver policy, so the server's default applies: everything. Issue 207's fix sets *from now* only on
|
||||
the path where an existing consumer's type changes.
|
||||
|
||||
**Unblocked**, on the operator's explicit word: the consumer re-made from now with its configuration
|
||||
otherwise unchanged (nothing pending afterwards). The controller's loop still held the old event, so
|
||||
it takes a restart of the controller to let go of it; that restart waits for the operator's word.
|
||||
|
||||
**Fix** (mesh-controller, branch `fix/a-consumer-on-a-history-stream-starts-from-now`):
|
||||
|
||||
- a consumer may say it starts **from now** when it is made. The controller's event consumer does. A
|
||||
consumer that exists keeps where it is. The server would refuse a changed start anyway.
|
||||
- `broker consumer-reset <stream> <consumer>` re-makes a stuck consumer from now, its configuration
|
||||
otherwise kept, and refuses a work queue, where what is pending is work. It is the person's act, said by
|
||||
a command, rather than a one-off program.
|
||||
|
||||
Checked by a live test against a throwaway bus:
|
||||
|
||||
- made from now, a consumer holds none of the stream's past and does hold the next announcement;
|
||||
- asserted again, it keeps its place;
|
||||
- the default replays all of it;
|
||||
- a reset leaves nothing pending and keeps every other setting;
|
||||
- a work queue's consumer is refused.
|
||||
|
||||
**Not fixed here:** the client dropping messages while the loop acts on a long merge. Reports and
|
||||
heartbeats are redelivered or replaced, so nothing is lost for good, but the loop holding everything while
|
||||
it builds is the shape issue 175 already describes.
|
||||
|
||||
## 2026-10-05, after the restart
|
||||
|
||||
The consumer re-made from now held nothing, and the restarted controller acknowledged what it was handed.
|
||||
Two merges made right after reached it within two seconds, and the fixed controller was built, delivered and
|
||||
took over by its own plan within two minutes. The re-asked build of the agent module was registered and
|
||||
reached all four machines.
|
||||
|
||||
**Not explained by this issue:** the record's merge was still logged twice, with the consumer fresh and
|
||||
nothing replayed. The duplication has a cause of its own, still to be found. It may be that the forge
|
||||
announces a merge on two paths, or that one event is handled twice.
|
||||
|
||||
## Resolved — 2026-10-05
|
||||
|
||||
The controller's event consumer is made from now when it is made, and `broker consumer-reset` re-makes a stuck one from now. Live: after the restart, merges reached the controller within seconds. Two tests main then failed, both skipped without a store, were fixed in mesh-controller PR #52. The merge heard twice was a separate cause: [issue 250](../250-a-merge-made-through-the-forges-tool-is-announced-twice/00-report.md).
|
||||
+37
@@ -0,0 +1,37 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-controller cmd/mesh-controller]
|
||||
fixed-by: mesh-controller PR #54
|
||||
amended-design: 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
|
||||
---
|
||||
|
||||
# 249 — A module's new state is refused until a push the merge did not make
|
||||
|
||||
## What was observed
|
||||
|
||||
A merge gave the agent module a new state, a key-value bucket
|
||||
([ADR 0201](../../02-DECISIONS/0201-a-module-keeps-its-current-state-in-key-value-buckets-it-declares-and-reaches-through-the-runtime.md)).
|
||||
The module's new bundle reached all four machines within a minute of the build, at once. The machines'
|
||||
bus permissions did not include the new state until a push was made by hand afterwards.
|
||||
|
||||
On three machines the module's watch of the new state was refused for about two minutes: "claude-code keeps
|
||||
and reads no state called config". It recovered only because the module asks again with a back-off, and
|
||||
the hand-made pushes issued the permissions.
|
||||
|
||||
## Why it matters
|
||||
|
||||
**The code arrives before the right to use it.** A module that does not retry stays broken until someone
|
||||
pushes. A module whose first act on start is to read its new state fails its start. Nothing in the plan
|
||||
says the two must travel together.
|
||||
|
||||
**No machine went first.** The bundle reached every machine at the same moment. The rollout the operator
|
||||
was told — one machine first, then the rest — could not be followed, because the merge had already
|
||||
delivered it everywhere.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should a plan send a machine its membership, the grants that come with a module's new
|
||||
declarations, in the same push as the bundle, and before it?
|
||||
- Should a merge that changes a module's declarations (state, events, tools) be delivered to one machine
|
||||
first, and to the rest only once that one reports it healthy?
|
||||
+30
@@ -0,0 +1,30 @@
|
||||
# 249 — Diagnosis
|
||||
|
||||
## 2026-10-05
|
||||
|
||||
**Grants after code.** Both a plan's rollout and a push send every machine its declaration first, and
|
||||
issue the memberships afterwards. The order is written into the code on purpose, "because the runtime it is
|
||||
for arrives with it". That reason holds only for a first assignment, and a membership is kept on the bus for
|
||||
a runtime that connects later anyway. A membership that failed was only printed, and left "until the next
|
||||
push". The list of bus users travels in the declaration of the machine that holds the bus, which a module's
|
||||
rollout reaches only if that machine runs the module.
|
||||
|
||||
**No machine first.** The module's upgrade policy sends one machine at a time, but a plan's rollout ignored
|
||||
it and sent every machine at once. One at a time did not wait for the first machine to come up either: it
|
||||
stopped only if the publish failed.
|
||||
|
||||
**Not answered by the open decision on unseen changes** (a removal, a move or a replacement, shown
|
||||
before it takes effect). That decision leaves an add-only change alone on purpose, and a new state is one.
|
||||
|
||||
**Decided** in [ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md):
|
||||
grants before code, and one machine first unless a module's policy says *together*.
|
||||
|
||||
## Resolved — 2026-10-05
|
||||
|
||||
[ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md), live the same evening:
|
||||
|
||||
- every push now issues memberships before it sends declarations;
|
||||
- the bus's machine goes first when its list of users moved;
|
||||
- a plan sent the build agent to one machine first and to the rest once that machine reported.
|
||||
|
||||
The first live rollout exposed a fault in the tier gate. The first machine's report, made between the two sends, was read as stale ([issue 256](../256-a-first-machines-report-read-as-stale-between-the-two-sends/00-report.md)).
|
||||
@@ -0,0 +1,49 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-catalog modules/gitea]
|
||||
fixed-by: mesh-catalog PR #63, PR #66
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 250 — A merge made through the forge's tool is announced twice
|
||||
|
||||
## What was observed
|
||||
|
||||
The controller logged the record repository's merges twice, seconds apart, with the same commit, even with
|
||||
its event consumer freshly made ([issue 248](../248-the-controllers-event-consumer-replayed-a-week-and-held-every-merge-behind-it/00-report.md)).
|
||||
Counted over one day:
|
||||
|
||||
- 8 of the record's merges were logged twice, against 4 once;
|
||||
- 10 of the catalogue's, against 7 once;
|
||||
- 2 of the controller's.
|
||||
|
||||
A code repository's second line reads differently — "it changed nothing any module the mesh holds is built
|
||||
from" — so it was taken for a different message.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
The forge's module announces a merge from two places in the same process:
|
||||
|
||||
1. its merge tool, the moment it merges;
|
||||
2. the poll added for [issue 131](../131-nothing-tells-the-mesh-a-source-moved/00-report.md), which
|
||||
announces every merged pull request it has not recorded as announced.
|
||||
|
||||
The tool never records what it announced, so the poll announces it again 0.5 to 16 seconds later.
|
||||
Merges made in the forge's web interface or by a plain API call are seen by the poll alone, and those are
|
||||
the ones logged once. The module's own header comments still say merges are announced "from the tools …
|
||||
one process only".
|
||||
|
||||
No harm was done this time, but only by luck. The second event is absorbed because the first one moved
|
||||
the controller's record of the source. A repository read only by packaging modules has no such record, so
|
||||
it would get a second plan. The record module synced twice for each merge.
|
||||
|
||||
## Fix
|
||||
|
||||
The poll is the only emitter: it sees every path and carries the clone address. The tool merges and
|
||||
answers the merge commit. A merge made through the tool is heard up to thirty seconds later, which the
|
||||
module already accepts ("an event a minute late is still an event").
|
||||
|
||||
## Resolved — 2026-10-05
|
||||
|
||||
The forge module's poll is the only announcer of a merge. It asks only the repositories that moved since its last look, and one pass at a time. Live: three merges were each heard once, and a later merge was planned within a minute.
|
||||
+36
@@ -0,0 +1,36 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-catalog modules/records]
|
||||
fixed-by: mesh-catalog PR #64
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 251 — The record's checkout could not sync after it ran as another account
|
||||
|
||||
## What was observed
|
||||
|
||||
Every sync of the record module failed, on every merge and every timer: git refused the checkout as
|
||||
"dubious ownership". The record tools answered all the while, from the checkout as it last stood, and
|
||||
nothing said it was stale.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
The module ran in a container, as the superuser, until its code moved into the machine's tool runtime
|
||||
([ADR 0198](../../02-DECISIONS/0198-a-modules-long-running-code-is-launched-by-the-node-runtime-and-reaches-the-bus-through-it.md)).
|
||||
The runtime launches it as the operator account. The checkout's git directory and 2 598 of its files still
|
||||
belonged to the superuser. Git refuses a repository owned by another user, and the operator account could
|
||||
change none of those files.
|
||||
|
||||
The module also still declared that it needs a container runtime, a leftover of the same move.
|
||||
|
||||
## Fix
|
||||
|
||||
The module, ported to Go as part of the fix, clones into a directory of its own inside the one it is given:
|
||||
a directory it makes, and so owns. What the old layout left behind is removed where it is the module's.
|
||||
Where it is not, the module names it in its status, with the one command that deletes it. The
|
||||
container-runtime capability is dropped. A sync that fails is still said in the status, as before.
|
||||
|
||||
## Resolved — 2026-10-05
|
||||
|
||||
The record module, ported to Go, clones into a directory it makes and owns. Live: it synced to the newest commit, with 692 documents. It names seven leftovers of the old checkout that it cannot remove, with the command that removes them; that is the operator's act.
|
||||
@@ -0,0 +1,38 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-catalog modules/gitea, mesh-controller cmd/mesh-controller]
|
||||
fixed-by: mesh-catalog PR #63, mesh-controller PR #54
|
||||
amended-design: 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
|
||||
---
|
||||
|
||||
# 252 — A merge's changed modules were read wrong, in both directions
|
||||
|
||||
## What was observed
|
||||
|
||||
Each of three merges to the catalogue within four minutes planned 99 to 100 modules, and rebuilt modules
|
||||
they did not touch. The outputs were identical, so nothing was redeployed. The rebuilds cost the build
|
||||
machine minutes for every merge.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
**Too many.** The controller treats a changed path outside every module directory it knows as shared code,
|
||||
and rebuilds every module built from the repository. A new module's directory, or one being removed,
|
||||
counts: the module is registered only after the merge is planned. Each of the three merges added or
|
||||
removed a module. The build agent, built from the same repository, then joins the set and becomes tier 0.
|
||||
|
||||
**Too few.** The forge's module asked for a hundred changed files and was given fifty, the forge's page
|
||||
size, and reported the list as whole. A merge of 59 files reached the controller with 50. Had the full
|
||||
rebuild not hidden it, a module whose own files changed would have stayed unbuilt.
|
||||
|
||||
## Fix
|
||||
|
||||
- The forge's module reads every page of a pull request's files.
|
||||
- The controller counts a changed path in a sibling of known module directories as that module's own, not
|
||||
shared, when that module's definition is among the changed paths. A sibling without a definition, such
|
||||
as a shared library, still means everything, which is the safe direction. Root files still mean
|
||||
everything.
|
||||
|
||||
## Resolved — 2026-10-05
|
||||
|
||||
The forge's module reads every page of a pull request's files. The controller counts a new module's own directory as that module's when its definition is among the changed paths. Live: catalogue merges planned one module each.
|
||||
+55
@@ -0,0 +1,55 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-controller internal/builder, mesh-controller internal/artifacts, mesh-catalog modules/distribution]
|
||||
fixed-by: mesh-controller PR #53, mesh-host PR #23, mesh-catalog PR #62
|
||||
amended-design: 03-DESIGN/01-to-be/18-building-a-module.md
|
||||
---
|
||||
|
||||
# 253 — The store's collector would delete every archive the mesh keeps
|
||||
|
||||
## What was observed
|
||||
|
||||
The store's nightly collector ([ADR 0189](../../02-DECISIONS/0189-the-store-keeps-what-the-records-name.md))
|
||||
was installed the same day, its first run due that night. Measured read-only beforehand:
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| blobs in the store | 8 186 |
|
||||
| blobs a manifest names, kept by the collector | 2 090 |
|
||||
| blobs it would delete | 6 096 |
|
||||
| repositories holding only archives, none named by any manifest | 105 |
|
||||
|
||||
Among the archives it would delete were the current bundles of the agent module, the machine host, the
|
||||
controller and the tool runtime. Each was named by no manifest, though the controller's records keep them.
|
||||
|
||||
## Why it matters
|
||||
|
||||
Machines keep their unpacked copies, so nothing would have stopped at once. But any fresh fetch of an
|
||||
unchanged module would have failed: a machine joining, a reinstall, an apply that fetches again, the
|
||||
controller's own next rollout.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
Images are pushed with manifests. Archives were published as bare blobs, which no manifest names. The
|
||||
store's stock collector marks only from manifests, so every bare blob is unmarked, kept or not. ADR 0189's
|
||||
sentence "what the mesh keeps is still a manifest in the store" was true of images only.
|
||||
|
||||
Found alongside: the "five most recent builds" reason kept builds of modules the mesh no longer holds,
|
||||
forever.
|
||||
|
||||
## Fix
|
||||
|
||||
- **That night, before the first run:** the collector was changed to a dry run, in its module's
|
||||
definition, and delivered.
|
||||
- **Then:**
|
||||
- every archive is published with a manifest that holds it;
|
||||
- the controller's sweep holds every kept archive before it lets anything go, which backfills those
|
||||
already published;
|
||||
- letting an archive go removes its manifest first;
|
||||
- a forgotten module keeps nothing.
|
||||
- Real collection returns once the controller reports no kept archive unheld.
|
||||
|
||||
## Where it stands — 2026-10-05
|
||||
|
||||
Every kept archive is held: the controller's collection command reports 134 of 134 held, none missing. The window the collector needs is no longer reopened by an apply ([issue 224](../224-an-apply-reopens-a-maintenance-window-by-recreating-what-it-held-still/00-report.md)). The collector still runs as a dry run. Turning it to real collection deletes the layers nothing keeps, which is the operator's word to give; this issue resolves when that change lands.
|
||||
+36
@@ -0,0 +1,36 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-controller cmd/mesh-controller, mesh-controller internal/inventory]
|
||||
fixed-by: mesh-controller PR #54
|
||||
amended-design: 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
|
||||
---
|
||||
|
||||
# 254 — Plans for successive merges run over each other, and one was left open
|
||||
|
||||
## What was observed
|
||||
|
||||
Three merges to the catalogue within four minutes made three plans, and all three ran at once:
|
||||
|
||||
- each sent the build agent to every machine;
|
||||
- each asked for the same tier of builds — one module was asked for 32 seconds apart by two plans;
|
||||
- the plan for the middle merge still showed "building" hours later, waiting on nothing.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
A merge's plan is saved without looking at the open plans
|
||||
([ADR 0162](../../02-DECISIONS/0162-a-merge-produces-a-tiered-plan-the-mesh-keeps.md): one plan per merge).
|
||||
Each open plan advances on its own. Nothing ends a plan whose work a newer merge has taken over, and
|
||||
nothing lets a person close a plan that waits on nothing.
|
||||
[Issue 219](../219-an-older-build-that-finishes-later-replaces-a-newer-one/00-report.md) settled only which
|
||||
build's output wins.
|
||||
|
||||
## Fix
|
||||
|
||||
[ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md)
|
||||
§3. A newer merge's plan supersedes the older open plans for the same repository and branch, and takes in
|
||||
the modules they had not built. A person can close a plan by its id.
|
||||
|
||||
## Resolved — 2026-10-05
|
||||
|
||||
A newer merge's plan supersedes the older open plans of its repository and branch, and a person can close a plan by its id. Live: the plan left open since the afternoon was closed by hand, and its note said what it had built and never sent. A merge made while the controller's own plan was open took it over: the older plan reads "superseded at tier 0 by" the newer, which planned its three modules again.
|
||||
@@ -0,0 +1,36 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-catalog modules/systemd]
|
||||
fixed-by: mesh-catalog PR #65
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 255 — The journal verb read nothing for a system service
|
||||
|
||||
## What was observed
|
||||
|
||||
The service manager seat's `journal` verb answered "-- No entries --" for the controller's service on the
|
||||
control node, while the service was logging steadily. Twice in one day a session reading the controller's
|
||||
log fell back to a shell on the machine: the verb that exists for exactly that question answered nothing.
|
||||
On one machine of four the same verb did answer: the only one whose operator account is in the journal's
|
||||
group.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
The module runs as the operator account and escalates the five acts on the system manager with `sudo -n`.
|
||||
It ran `journalctl` unescalated. journalctl shows an account outside the journal's group only that
|
||||
account's own entries, and says "-- No entries --" for everything else. That reads as a quiet service, not
|
||||
as a refusal.
|
||||
|
||||
The mesh grants the operator account passwordless escalation on every machine through its own drop-in, as
|
||||
the sudo module's check confirms.
|
||||
|
||||
## Fix
|
||||
|
||||
A read of the system journal escalates like an act does. The account's own journal, in the user scope,
|
||||
does not. The module was ported to Go as part of the fix, its tests with it.
|
||||
|
||||
## Resolved — 2026-10-05
|
||||
|
||||
The journal verb reads a system unit's journal escalated, in the module ported to Go. Live: the controller's and the host's journals read through the verb on the control node.
|
||||
+31
@@ -0,0 +1,31 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-controller cmd/mesh-controller]
|
||||
fixed-by: mesh-controller PR #55
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 256 — A first machine's report read as stale between the two sends
|
||||
|
||||
## What was observed
|
||||
|
||||
The first plan to roll out under [ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md)
|
||||
sent the build agent to one machine first, and to the other two once that machine reported it applied, 27
|
||||
seconds later. All three applied it within seconds. The plan then said, for seven minutes, that it was
|
||||
"waiting for build-agent" on the first machine to be applied. It went on only when that machine happened to
|
||||
report again for another reason.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
The tier gate asked every machine for a report made after the module's last send. With one machine first,
|
||||
a module is sent twice, and the second send is the later one. The first machine's report came between the
|
||||
two sends, so it read as older than the build. The plan waited for that machine's next report, one report
|
||||
cycle, and never knew why.
|
||||
|
||||
## Fix
|
||||
|
||||
The gate judges each machine from its own send: the first machine from the first send, the rest from the
|
||||
second. A plan's wait is also printed to the second rather than the minute, where it had read "0s", and the
|
||||
first send is printed in the machine's own time rather than in UTC. Checked by the controller's test: a first
|
||||
machine's report between the two sends opens the gate, and a report from before its send does not.
|
||||
+49
@@ -0,0 +1,49 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-host]
|
||||
fixed-by: novox/mesh-host#24 (89a7796)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 257. A plan waited on a declaration its first machine never reported
|
||||
|
||||
## Symptom
|
||||
|
||||
A merge to the controller's repository produced a plan of three tiers, the controller first. The
|
||||
controller was built and sent to its first machine, the anchor, which is also the machine holding the
|
||||
bus. The plan then said, for five minutes and with no end in sight:
|
||||
|
||||
> tier 1 of 3, tier 0 built; waiting for mesh-controller on the anchor, sent first at 20:22 to report
|
||||
> it applied before the rest are sent
|
||||
|
||||
The anchor had applied it. Its host's journal said `applied 470 resource(s)` ten seconds after the
|
||||
send, the new controller was running, and the controller logged the report.
|
||||
|
||||
## Evidence
|
||||
|
||||
The plan waits for a report that names the declaration last sent (ADR 0218, one machine first; the
|
||||
report's `declared` digest against the machine's recorded `sent` digest). Read from the store:
|
||||
|
||||
| | digest | at |
|
||||
|---|---|---|
|
||||
| recorded as sent to the anchor | `09f4d6…` | 20:22:57.699 |
|
||||
| the anchor's report, `declared` | `39b48a…` | 20:23:07.718 |
|
||||
|
||||
The anchor applied and reported a declaration that is not the one the controller recorded sending.
|
||||
So the report never counted, and the plan would have waited until its bound and then failed the
|
||||
rollout at its first machine, for a machine that had applied correctly.
|
||||
|
||||
The other three machines' digests matched their reports at the same moment.
|
||||
|
||||
## What unblocked it
|
||||
|
||||
A push to the anchor, made for another reason (its bus grants), recorded a new send. The anchor's
|
||||
next report named it, the digests matched, and the plan moved on at its next tick and finished.
|
||||
|
||||
## What it costs
|
||||
|
||||
A plan that rolls out the controller itself, to the machine that holds the bus, can stop at its first
|
||||
machine with nothing wrong there. The plan's words are true but useless: they name a machine that
|
||||
already did what was asked. Merges made close together, from several sessions, are the case the
|
||||
plans exist for, and a controller change is among them.
|
||||
+60
@@ -0,0 +1,60 @@
|
||||
# Diagnosis
|
||||
|
||||
## 2026-10-05
|
||||
|
||||
**What is known.** The digest the controller recorded for the anchor (`09f4d6…`, at 20:22:57.699) is
|
||||
not the digest of what the anchor applied (`39b48a…`). The anchor's host logged one apply that
|
||||
updated the controller, starting at 20:23:02 and reporting at 20:23:07. The controller that sent it
|
||||
was replaced at 20:23:04 by that same apply. Its successor logged two reports from the anchor, both at
|
||||
20:23:07.
|
||||
|
||||
**Ruled out.**
|
||||
|
||||
- The catalogue's catch-up: that is the catalogue asking the controller, not a machine asking for its
|
||||
declaration.
|
||||
- The host's five-minute check at 20:22:58: it applied nothing new (only "kept" lines). The apply
|
||||
that moved the controller is the one at 20:23:02.
|
||||
- A digest computed differently by the host and the controller: the other three machines matched,
|
||||
and after the push at 20:28:00 the anchor matched too.
|
||||
|
||||
**Not yet confirmed: two sends to one machine in one step.** The anchor is both the plan's first
|
||||
machine and the machine holding the bus. Issue 249 sends the machine holding the bus before the first
|
||||
machine when its user list must change, and this controller change added bus grants. If the anchor
|
||||
was sent twice in that step, each with its own digest, then the host applying the later one while the
|
||||
earlier one's record landed last would give exactly this picture. The next step is to read
|
||||
`sendToEach` for the order of its sends and of their `recordSent` calls, and to check the node's
|
||||
sequence numbers for two sends at 20:22:57.
|
||||
|
||||
**Whatever the cause, the plan's wait has a second fault.** A first machine that reports `applied` for
|
||||
a declaration the controller cannot place says nothing about the build. The plan should say "it
|
||||
applied something other than what was recorded as sent" instead of "waiting", so a reader acts in
|
||||
minutes rather than at the bound.
|
||||
|
||||
## 2026-10-05, later: found
|
||||
|
||||
**The lead above was wrong.** It was not two sends. The machine's own five-minute reconcile did it.
|
||||
|
||||
The host re-applies what it was last told every five minutes. That reconcile read the kept
|
||||
declaration, then waited for any apply in progress. A declaration the link is applying is kept only
|
||||
once its apply ends. On the anchor:
|
||||
|
||||
- the reconcile was due at 20:22:58;
|
||||
- the controller's send arrived at 20:22:57;
|
||||
- the reconcile read the declaration kept before that send, waited, and applied it after the link's
|
||||
apply;
|
||||
- its report named that older declaration.
|
||||
|
||||
That explains every fact above:
|
||||
|
||||
- the report's digest matched nothing recorded as sent;
|
||||
- the other machines, not due, matched;
|
||||
- the next push matched.
|
||||
|
||||
**Seen a second time, with a visible effect** (issue 261): a module assigned on the laptop was
|
||||
applied and then given back 29 seconds later by the reconcile due during that apply.
|
||||
|
||||
**Fixed** in mesh-host: the reconcile reads the kept declaration once it holds the apply lock. A test
|
||||
reproduces the interleaving: it fails with the old order and passes with the fix.
|
||||
|
||||
The plan's "waiting" wording, raised above as a second fault, is unchanged. With the cause gone, a
|
||||
report naming an unrecorded declaration should no longer occur.
|
||||
@@ -0,0 +1,53 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-controller]
|
||||
fixed-by: novox/mesh-controller#59 (4b382ce)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 258. Every machine bound the mesh's resolver to itself
|
||||
|
||||
## Symptom
|
||||
|
||||
After the mesh moved to one resolver (ADR 0194, 0196), the seat `mesh-dns-resolver` was recorded as
|
||||
held on the anchor, and each other machine was pinned to it for `wildcard-resolution`. Each machine
|
||||
was then pushed. The anchor's resolver configuration named its own private address first. Every
|
||||
other machine's did too: each named **its own** private address, not the anchor's.
|
||||
|
||||
Nothing broke, because each machine still ran a resolver of its own while the move was under way. But
|
||||
the mesh had decided on one resolver, recorded who held it and pinned every machine to it, and no
|
||||
machine used it.
|
||||
|
||||
## Cause
|
||||
|
||||
The controller binds a requirement in one of two branches:
|
||||
|
||||
- **Answered here**, when a module on the same machine provides it.
|
||||
- **Answered elsewhere**, when one on another machine does.
|
||||
|
||||
Only the second branch read the seat's holder (ADR 0110) and a pin. The first took the local provider,
|
||||
and read a pin only to choose between two local ones. Every machine still had its own resolver
|
||||
assigned, so every machine took the first branch.
|
||||
|
||||
The seat's own definition says the requirement "resolves to the holder wherever it is placed". The
|
||||
code did not.
|
||||
|
||||
## Fix
|
||||
|
||||
For a mesh-wide provision, a pin naming another machine, or the seat's holder on another machine,
|
||||
wins over a provider on this one. With neither, the local provider answers as before.
|
||||
|
||||
This is checked by three resolve tests in mesh-controller:
|
||||
|
||||
- the holder elsewhere answers;
|
||||
- a pin elsewhere answers;
|
||||
- a holder here still answers here.
|
||||
|
||||
The tests fail without the fix.
|
||||
|
||||
## Verified
|
||||
|
||||
2026-10-05, after the fix rolled out: all four machines' resolver configuration names the anchor's
|
||||
resolver first and the public one second. Each resolves the mesh's machine names, a wildcard name
|
||||
under a machine, and a public name.
|
||||
@@ -0,0 +1,46 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-controller]
|
||||
fixed-by: novox/mesh-controller#64 (002d5e5)
|
||||
amended-design: 03-DESIGN/01-to-be/30-the-mesh-updates-itself-on-a-push.md
|
||||
---
|
||||
|
||||
# 259. A push to one machine sent every machine
|
||||
|
||||
## Symptom
|
||||
|
||||
A change to the resolver modules was merged with the upgrade policy `record`, so that each machine
|
||||
would get it only when pushed. The plan was to go one machine at a time: the anchor first, then each
|
||||
of the others, each checked before the next.
|
||||
|
||||
`push <anchor>` sent all four machines, each a new declaration carrying the change. A later
|
||||
`push <laptop>` did the same. A further fault in the change (issue 260) was therefore met on every
|
||||
machine at once, not on one.
|
||||
|
||||
## Cause
|
||||
|
||||
A named push ends by flushing every other machine whose declaration differs from what it was last
|
||||
sent (issue 057, ADR 0083). This was meant for consequences of the push, such as a provider's grant
|
||||
list after a consumer was assigned. It cannot tell a consequence from a change the upgrade policy is
|
||||
holding back. Under `record`, every machine running the module differs, so every machine is flushed.
|
||||
|
||||
Issue 249 met the same confusion for the machine holding the bus, and narrowed that check to the user
|
||||
list alone. The flush has no such narrowing.
|
||||
|
||||
## What it costs
|
||||
|
||||
`record` is the policy for a change that must be walked through the mesh by hand. A named push cannot
|
||||
do that, so the policy does not hold the change back from anything but the merge. The command's name
|
||||
says one machine, and the command's output is the only place that says otherwise.
|
||||
|
||||
## Open
|
||||
|
||||
- Should the flush send a machine only what changed as a consequence: grants, user lists and bound
|
||||
facts, not module versions held by a policy?
|
||||
- Or should it list such machines as behind and leave them, as `push --behind` would find them?
|
||||
|
||||
## Verified
|
||||
|
||||
2026-10-05: after the fix, a named push sent only the machine named. Upgrades held by a policy were
|
||||
then sent with `push --behind`, as ADR 0221 has it.
|
||||
@@ -0,0 +1,27 @@
|
||||
# 259 — Diagnosis
|
||||
|
||||
## 2026-10-05
|
||||
|
||||
**Located in the controller's push.** A named push ends with the cascade of
|
||||
[ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md): it sends every other machine
|
||||
whose declaration's digest differs from the digest it was last sent. The comparison is the only one the
|
||||
mesh can make from what it keeps. It keeps a digest of each send and nothing about what the send
|
||||
carried, so a machine behind because a module moved under `record` and a machine behind because of a
|
||||
grant read the same.
|
||||
|
||||
**Not only `record`.** A plan rolling a module out one machine first
|
||||
([ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md))
|
||||
leaves the other machines differing while it waits on the first. Any named push elsewhere in that
|
||||
window sends them the build the plan is holding.
|
||||
|
||||
**And the machine holding the bus.** [Issue 249](../249-a-modules-new-state-is-refused-until-a-push-the-merge-did-not-make/00-report.md)
|
||||
narrowed *whether* that machine is added to a push to its user list alone. Once added, it is sent its
|
||||
whole declaration, held builds included.
|
||||
|
||||
**Ruled out:** composing a held machine with its held modules kept at their last build, so it gets the
|
||||
grant and not the upgrade. Reasons in the decision.
|
||||
|
||||
**Decided** in [ADR 0221](../../02-DECISIONS/0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md):
|
||||
each send records which build of each module it carried, and a push does not send a machine it did not
|
||||
name any build a policy or an open plan holds back. It names the machine and the remedy instead. The
|
||||
fix is a pull request on mesh-controller; `fixed-by:` is filled when it merges.
|
||||
@@ -0,0 +1,50 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-host, mesh-catalog]
|
||||
fixed-by: novox/mesh-host#25 (a566add)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 260. The resolver was restarted before the file it reads existed
|
||||
|
||||
## Symptom
|
||||
|
||||
The first apply of the resolver's new configuration failed on every machine. Each host's journal:
|
||||
|
||||
> failed dnsmasq.service: restarting dnsmasq.service: starting it again: systemctl exited 1
|
||||
|
||||
The resolver's own journal:
|
||||
|
||||
> dnsmasq: cannot read /etc/mesh-resolver/zones.conf: No such file or directory
|
||||
|
||||
The service manager's automatic restart started it again in the same second, and it answered. The
|
||||
host still reported the apply as failed, on every machine.
|
||||
|
||||
## Cause
|
||||
|
||||
The new configuration names a second file, the mesh's zones. The host wrote the configuration and
|
||||
restarted the service on it **before** it created the zones file. The zones file was created next in
|
||||
the same apply. The service's `restart-on` names both files, but nothing orders the restart after
|
||||
every file it reads has been written.
|
||||
|
||||
## What it costs
|
||||
|
||||
A configuration that adds a file it reads fails its first start everywhere. Here the service manager's
|
||||
restart policy hid it within a second. A service without one would have stayed down, holding the
|
||||
machine's name resolution with it.
|
||||
|
||||
## Open
|
||||
|
||||
- Should the host run every restart and reload after all files of the apply are written?
|
||||
- Or should it order each restart after every resource the service names in `restart-on`?
|
||||
|
||||
## Fix
|
||||
|
||||
The host now orders the apply so that a service comes after every resource it names under
|
||||
`restart-on` or `reload-on`. Only services move, and only as far as the last of what they name; every
|
||||
other resource keeps its declared place. The first open question above is answered that way, per
|
||||
service, rather than by moving every restart to the end. This is checked by mesh-host's
|
||||
`restart_order_test.go`: a service restarting on two files, one declared after it, is restarted once
|
||||
both exist, and a later change to the second file still restarts it. The test fails without the
|
||||
ordering.
|
||||
@@ -0,0 +1,40 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-host]
|
||||
fixed-by: novox/mesh-host#24 (89a7796)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 261. A module was applied and given back half a minute later
|
||||
|
||||
## Symptom
|
||||
|
||||
The `hosts` module was assigned to the laptop and pushed. The laptop's host applied it: the module's
|
||||
block was written at the start of `/etc/hosts`, its tools were installed, and the host reported the
|
||||
declaration it had been sent. Twenty-nine seconds later the same host gave the block and the tools
|
||||
back. Nothing on the controller had sent a declaration without the module. Five minutes later the
|
||||
module was back.
|
||||
|
||||
The host's journal on the laptop, in order:
|
||||
|
||||
- `updated hosts.own (/etc/hosts): the mesh's region added at the start`
|
||||
- `applied 342 resource(s)`
|
||||
- `restored hosts.own (/etc/hosts)`
|
||||
- `removed hosts.bundle-tools`
|
||||
|
||||
## Cause
|
||||
|
||||
The host's five-minute reconcile was due during that apply. It read the declaration kept before the
|
||||
push, waited for the push's apply to finish, and then applied the older declaration over it. That
|
||||
older declaration did not have the module, so the module was given back. The next reconcile read the
|
||||
newer declaration, which had been kept by then, and applied the module again.
|
||||
|
||||
This is the same fault as [issue 257](../257-a-plan-waited-on-a-declaration-its-first-machine-never-reported/00-report.md),
|
||||
where it showed only as a report naming a declaration nobody had recorded sending.
|
||||
|
||||
## Fix
|
||||
|
||||
The reconcile reads the kept declaration once it holds the apply lock, so it always applies the
|
||||
latest declaration the mesh sent. This is checked by mesh-host's test
|
||||
`TestAReconcileAppliesWhatWasKeptWhenItsTurnComes`, which fails with the old order.
|
||||
+59
@@ -0,0 +1,59 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-10-05
|
||||
located-in: [mesh-catalog]
|
||||
fixed-by: novox/mesh-catalog#71 (20603b6), novox/mesh-controller#65 (df9231c), novox/mesh-catalog#73 (b4b86c1)
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 262. An Alpine container could not find a machine by its mesh name
|
||||
|
||||
## Symptom
|
||||
|
||||
After the mesh moved to one resolver (ADR 0194, 0196), a workflow container on the home server was
|
||||
restarted so that it would ask the mesh's resolver. It then crash-looped every fourteen seconds:
|
||||
|
||||
> getaddrinfo ENOTFOUND <anchor>.internal
|
||||
|
||||
From the same container, `getent hosts <anchor>.internal` answered correctly. So did the machine
|
||||
itself, every time, in about forty-five milliseconds.
|
||||
|
||||
## Cause
|
||||
|
||||
The mesh's resolver answered a machine's name only through a wildcard rule:
|
||||
|
||||
- asked for the IPv4 address, it gave the address;
|
||||
- asked for the IPv6 address, it answered NXDOMAIN, "no such name", where the correct answer is
|
||||
NODATA, "the name exists and has no such record".
|
||||
|
||||
glibc ignores that. musl, the C library of every Alpine image, asks for both records and takes the
|
||||
NXDOMAIN as final, so the whole lookup failed. The per-machine resolvers this one replaced answered
|
||||
each machine's name from `/etc/hosts`, which gives NODATA, so nothing had depended on the difference
|
||||
before.
|
||||
|
||||
Reproduced with a throwaway resolver of the same version, given only the wildcard rule.
|
||||
|
||||
## Fix
|
||||
|
||||
Each machine's name is now also a host record in the machine list the resolver reads. Asked for an
|
||||
IPv6 address, the resolver then answers that the name exists and has none. A name under a machine,
|
||||
answered only by the wildcard, still answers NXDOMAIN for IPv6, exactly as it did before the move.
|
||||
|
||||
## How it is checked
|
||||
|
||||
Ask the mesh's resolver for the IPv6 address of a machine's name. The answer must be NOERROR with no
|
||||
records. Today that check is done by hand. A controller test that renders the machine list and
|
||||
requires one host record per machine should be added.
|
||||
|
||||
*2026-10-05:* that test is added with [ADR 0223](../../02-DECISIONS/0223-the-mesh-has-two-resolvers-and-a-machine-lists-only-them.md)
|
||||
— mesh-controller's composition test renders the resolver's machine list and requires exactly one
|
||||
host record per machine. The live check stays by hand, on each resolver.
|
||||
|
||||
## Verified, and what else was found
|
||||
|
||||
2026-10-05: the host record fixed the IPv6 answer, but Alpine programs still failed when they read the
|
||||
machine's resolver file directly (host networking, docker's default bridge). They asked the mesh's
|
||||
resolver and the public fallback at once and took the public "no such name". ADR 0223 removes the
|
||||
public fallback: a machine lists only the mesh's resolvers, now two. Afterwards an Alpine container
|
||||
resolved a machine's name ten times out of ten, with host networking and on the default bridge, and
|
||||
the build machines built with no hosts-file line.
|
||||
@@ -0,0 +1,77 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-06
|
||||
located-in: [mesh-controller internal/catalogue, mesh-controller cmd/mesh-controller, mesh-catalog modules]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 263. Every consumer pays for the tightest backend's name limit
|
||||
|
||||
## Symptom
|
||||
|
||||
The operator: "the 20 character limit has bitten us multiple times". Most recently, on 2026-10-06:
|
||||
|
||||
- A change made the network-manager modules require the mesh's resolver provision.
|
||||
- Each machine running them thereby became a consumer of that provision, with an identity
|
||||
`mesh_<machine>_<module>`.
|
||||
- `networkmanager`'s identity is 23 to 26 characters, over the 20 the controller allows.
|
||||
|
||||
The refusal surfaced as the anchor's whole declaration failing to compose, because the anchor carries
|
||||
every consumer's grant as the resolver's holder:
|
||||
|
||||
> networkmanager on <home-server> is identified as "mesh_<home-server>_networkmanager", 23 characters
|
||||
> where a backend (an S3 access key) keeps 20 — give the module a shorter `slug` or shorten the
|
||||
> machine's name
|
||||
|
||||
For as long as that lasted, no push to the anchor could go through. The fix was another slug.
|
||||
|
||||
## Why it keeps happening
|
||||
|
||||
[ADR 0049](../../02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md) bounds **every**
|
||||
consumer's identity by the tightest backend anywhere, an object store's 20-character access key
|
||||
([issue 034](../034-mesh-login-exceeds-s3-access-key-limit/00-report.md)), and makes a module `slug`
|
||||
the remedy. That has two costs:
|
||||
|
||||
- **The limit applies where it means nothing.** The resolver provision mints no credential and keeps
|
||||
no name in any backend, yet its consumers were held to an object store's key length. ADR 0049 named
|
||||
per-provision bounds (its option C) as the refinement for later. It has not been done.
|
||||
- **It is found late, and far from its cause.** The catalogue's module check and the controller's
|
||||
tests passed. The refusal came only when a real machine's name met the module's name, and it showed
|
||||
on the provider's machine, not on the consumer's, as a whole node that could not be pushed. ADR
|
||||
0049 says it is refused at assignment. This one was a new requirement on modules already assigned,
|
||||
so assignment never saw it.
|
||||
|
||||
## What to look into
|
||||
|
||||
1. **Bound each provision by its own backend's limit**, ADR 0049's option C. Only a provision that
|
||||
creates a name in a limited backend carries the limit, and a keyless one carries none.
|
||||
2. **Refuse it before merge.** The catalogue check should judge a module's identity against the
|
||||
longest machine name in the mesh (or the bound per provision) for every provision it requires, so
|
||||
the refusal names the module in its own pull request.
|
||||
3. **Never refuse a provider's whole declaration** for one consumer's identity. Leave that consumer's
|
||||
grant out, say so in `status`, and keep the provider's machine pushable.
|
||||
|
||||
## Where it lives
|
||||
|
||||
- **The one bound** is `identityLimit` in the controller's `internal/catalogue/identity.go`, applied to
|
||||
every consumer whatever it requires. Nothing in a provider's definition could say otherwise.
|
||||
- **The refusal** was raised in `grantsFor` (`cmd/mesh-controller/plan.go`), while composing the
|
||||
*provider's* declaration: an error there fails the whole composition, so the anchor could not be
|
||||
pushed for a module on another machine. No assignment, module check or test judged it before.
|
||||
- **The facts** are in each provider's code in the catalogue: the object store keeps the identity as
|
||||
an access key (20), the databases as roles and database names (63, 128), the identity provider as a
|
||||
client id (255); the resolver and the route providers keep no name of their consumers.
|
||||
|
||||
Decided in [ADR 0225](../../02-DECISIONS/0225-a-consumers-identity-is-bounded-by-the-provision-it-requires.md),
|
||||
which takes ADR 0049's option C and adds the two placements the issue asks for: each offer states its
|
||||
bound, `module check` judges every identity on the longest machine name before merge, and a provider
|
||||
leaves an overflowing consumer out of its grants and says so, in `status` too, instead of refusing.
|
||||
|
||||
## How it is checked (once fixed)
|
||||
|
||||
- A catalogue test: a module requiring a keyless provision composes with a long name.
|
||||
- A module check: a module requiring a limited provision is refused when the longest machine's
|
||||
identity overflows.
|
||||
- A controller test: an overflowing consumer leaves the provider's declaration composable, and is
|
||||
reported.
|
||||
+79
@@ -0,0 +1,79 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-06
|
||||
located-in: [mesh-host]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 264. A self-updating engine lost the report of the apply that delivered it
|
||||
|
||||
## Symptom
|
||||
|
||||
One declaration for the anchor carried both a new controller and a new version of the host. The
|
||||
host applied it and said, in its journal:
|
||||
|
||||
- `host <version> is delivered; standing aside so the launcher runs it`
|
||||
- `applied, and could not tell the mesh: reporting: context canceled`
|
||||
|
||||
The new host started and said `in the mesh, hearing what this node should be`, and nothing more
|
||||
about that declaration. The controller's release plan waits for the first machine to report the
|
||||
exact digest of the declaration it was sent. That report never came, so the plan waited until a
|
||||
person pushed again by hand.
|
||||
|
||||
This is the same family as [issue 257](../257-a-plan-waited-on-a-declaration-its-first-machine-never-reported/00-report.md)
|
||||
and [issue 261](../261-a-module-was-applied-and-given-back-half-a-minute-later/00-report.md): the
|
||||
machine did what it was sent, and the report the plan waits on did not say so.
|
||||
|
||||
## Cause
|
||||
|
||||
Standing aside for a successor ([ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md)) is
|
||||
decided inside the apply, once the apply has finished, and it cancels the link's context. The link
|
||||
then published the apply's report on that same cancelled context, and the bus refused it before it
|
||||
left. The declaration was then acknowledged anyway, so the mesh did not deliver it again, and the
|
||||
new host had nothing that told it a report was owed.
|
||||
|
||||
A crash or power cut between the apply and the report has the same shape. The declaration is
|
||||
delivered again in that case only because it was not yet acknowledged, and nothing on the machine
|
||||
records that a report is owed.
|
||||
|
||||
## Fix
|
||||
|
||||
Two parts, in the host:
|
||||
|
||||
1. The report of an apply, and of a declaration set aside for a newer one, is published on a context
|
||||
that standing aside does not cancel. It is still limited by the host's timeout, and it is sent
|
||||
before the connection is closed. After an apply that stood aside, nothing further is applied.
|
||||
2. The report of the last apply of a declaration from the mesh is kept beside the node's state until
|
||||
the broker has taken it. When a host links, it first sends any report still kept, before it applies
|
||||
anything newly delivered, so the re-sent report cannot arrive after a newer one. What is re-sent
|
||||
is exactly the report the apply made, with the same digest and the same outcome. It is never a new
|
||||
description of the machine, so it cannot claim that something was applied when it was not. A
|
||||
report that names no declaration, such as a refused one, clears what was kept. A node that never
|
||||
applied anything sends nothing.
|
||||
|
||||
## How it is checked
|
||||
|
||||
These tests in the host fail without the fix:
|
||||
|
||||
- `TestAnApplyThatEndsTheLinkIsStillReported`: an apply that cancels the link still has its report
|
||||
delivered, and its declaration is settled after the report.
|
||||
- `TestAReportLostAfterTheApplyIsSaidOnTheNextLink`: a report the broker refused is sent again on the
|
||||
next link, unchanged.
|
||||
|
||||
These tests guard the edges:
|
||||
|
||||
- `TestNothingIsSaidAgainWhenNothingWasLost`: nothing is re-sent when nothing was lost.
|
||||
- `TestTheKeptReportSurvivesTheHost` and `TestNothingKeptWhenNothingWasApplied`: the kept report
|
||||
survives a restart exactly as it was made, and it is cleared only by a report about the same
|
||||
declaration.
|
||||
|
||||
Once the fix is rolled out, the live check is the next declaration that carries a new host version.
|
||||
The first machine's report should reach the plan without anyone pushing by hand.
|
||||
|
||||
## Later
|
||||
|
||||
**2026-10-06.** The next stalled plan after this fix was rolled out looked like this issue coming
|
||||
back, and was suspected of being caused by its kept report. It was not. A reconcile's report about
|
||||
an older declaration reached the controller after the newer apply's report, and replaced it:
|
||||
[issue 267](../267-a-reconciles-report-overtook-the-apply-that-followed-it/00-report.md).
|
||||
@@ -0,0 +1,145 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-06
|
||||
located-in: [mesh-controller internal/link, mesh-controller cmd/mesh-controller, mesh-tools node-tools/internal/console]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 265. A push outlived its caller, and its answer was refused
|
||||
|
||||
## Symptom
|
||||
|
||||
Through the console, a call to the controller's seat — `mesh-controller.push` with a machine, or
|
||||
`mesh-controller.command` with `push <machine>` — ended in:
|
||||
|
||||
> mesh-controller.push did not answer in time. Something is serving it, so this is the tool being slow
|
||||
> rather than absent.
|
||||
|
||||
and the push had happened: the machine applied what it was sent. Over 2026-10-04 and 2026-10-05, 54
|
||||
of 103 pushes asked through the console ended that way. Each one left the operator not knowing whether
|
||||
to push again.
|
||||
|
||||
Some of them also left a line in the controller's journal, written by the bus client library and by
|
||||
nothing of the mesh's own:
|
||||
|
||||
> nats: permissions violation: Permissions Violation for Publish to "_INBOX.<laptop>.node-tools.….…"
|
||||
> on connection [838]
|
||||
|
||||
Thirty such lines in three days. No caller, no `status` and no record said that a call had lost its
|
||||
answer.
|
||||
|
||||
## Evidence
|
||||
|
||||
Three waits govern one call, and they disagreed:
|
||||
|
||||
| Who waits | How long | Where |
|
||||
|---|---|---|
|
||||
| The console, for an answer | 30 s | mesh-tools `node-tools/internal/bus`, `RequestTimeout` |
|
||||
| The bus, for the one answer it lets the controller send | 1 min, one message | the user list the controller composes (`allow_responses`, ADR 0043) |
|
||||
| The controller, for the verb's command | 5 min | mesh-controller `internal/link`, `HandlerTimeout` |
|
||||
|
||||
**A push takes about as long as the console waits.** The calls the console made were paired with
|
||||
their answers from the session transcripts. Of the pushes that were answered, the median took 26.6 s.
|
||||
The longest took up to two minutes, because the console queues calls made side by side. A named push
|
||||
composes the named machine, then the machine holding the bus when its user list changed (issue 249).
|
||||
It then composes every other machine to flush what fell behind (ADR 0083). Composing one machine
|
||||
alone took 9 s for the anchor, measured with a `plan` call. Anything slower than 30 s was reported as
|
||||
"did not answer in time", and its answer went to an inbox nobody read any more. The bus says nothing
|
||||
about that.
|
||||
|
||||
**Every journal line was a timed-out push whose answer the bus refused.** Each violation was set
|
||||
against the console calls before it. Every one of the 30 came 35 to 50 seconds after a push or a
|
||||
`command push …` that had timed out at 30 s. None came after a call that was answered. None came at
|
||||
the 60 s the bus's window would explain. The two from the night of 2026-10-05, in the broker's own
|
||||
log:
|
||||
|
||||
```
|
||||
22:25:00 console asks push of the anchor
|
||||
22:25:15 broker: "the mesh's user list changed; reloading in place" — Reloaded server configuration
|
||||
22:25:30 console gives up (30 s)
|
||||
22:25:44 broker: Publish Violation - User "controller", Subject "_INBOX.<laptop>.node-tools.….LNRyztuX"
|
||||
|
||||
22:41:17 console asks push of the anchor
|
||||
22:41:30 broker: user list changed; reloaded
|
||||
22:41:47 console gives up
|
||||
22:42:00 broker: Publish Violation - User "controller", Subject "_INBOX.<laptop>.node-tools.….BzO7Ep8s"
|
||||
```
|
||||
|
||||
The bus server's source names the cause (`client.go`, `setPermissions`). A reload of the
|
||||
authorization re-registers every connected user and gives each a new, empty table of the replies it
|
||||
may send. The permission to answer a request already received is forgotten. A push sends the machine
|
||||
holding the bus first when its user list changed (issue 249). The broker reloads. When the push
|
||||
finishes, the answer it waited to send is refused.
|
||||
|
||||
## Cause
|
||||
|
||||
**A verb answered only when it finished, however long that took.** The control plane runs each verb's
|
||||
command and answers with what it printed (ADR 0154, design 33). Nothing bounded that by the caller's
|
||||
wait or by the bus's window for an answer. Nothing ruled out the verb destroying that window itself.
|
||||
Three silences followed:
|
||||
|
||||
1. A verb slower than its caller's wait answered nobody. The caller was told it was slow, and read
|
||||
that as failed.
|
||||
2. A verb that reloaded the bus had its answer refused, every time the user list had changed.
|
||||
3. The refusal reached only the client library's default error handler, which prints a line. No
|
||||
record of the mesh's own said which call lost its answer, or what the answer was.
|
||||
|
||||
The two-minute default for `allow_responses` (the hypothesis this was opened under) is not the cause:
|
||||
the mesh sets the window to one minute, and every refusal came well inside it.
|
||||
|
||||
## Fix
|
||||
|
||||
**The controller answers every call within ten seconds and keeps what came of it.** In mesh-controller:
|
||||
|
||||
- `internal/link`: a seat's call is answered once, as before. It gets its whole answer when the
|
||||
command finishes within `AnswerWithin` (10 s). Otherwise it is told the call is running, with the
|
||||
call's id. The command runs on. Its answer is kept in a log of the last hundred calls, and the
|
||||
journal says when it ended.
|
||||
- **A push is answered before it sends anything**, by its verb or through `command`: its own first
|
||||
act can reload the bus. The answer says it is running, with the id.
|
||||
- The connection's error handler is chained, not replaced. A refused answer is matched to its call
|
||||
by the reply subject the bus names. It is recorded on the call and said in the mesh's words, naming
|
||||
the call.
|
||||
- A new verb, `calls`, lists the recent calls: running, answered, finished after its caller was
|
||||
answered, and answer refused by the bus. Given an id, it returns that call's whole answer.
|
||||
Arguments are kept by name, and settings and secrets are never kept.
|
||||
- The bus's window for an answer is now a named constant (`ResponseTTL`). A test holds
|
||||
`AnswerWithin` inside it and inside the console's wait.
|
||||
|
||||
In mesh-tools, the console's timeout no longer reads as a failure. It says the call may still be
|
||||
running and may already have done what was asked, and to check before repeating a call that changes
|
||||
the mesh. For the controller it adds that the answer was lost on the way, and that
|
||||
`mesh-controller.calls` has it.
|
||||
|
||||
Why this shape rather than the others weighed:
|
||||
|
||||
- **Widening the bus's window does nothing for the refusal.** A reload empties the reply table
|
||||
whatever its expiry is.
|
||||
- **Making push faster** narrows the race and keeps it. The flush of every other machine is ADR
|
||||
0083's and stays.
|
||||
- **Answering early with an id** is how `build` already works ("asked, not waited for"). It is the
|
||||
only shape where the answer cannot depend on what the verb does to the bus.
|
||||
|
||||
## How it is checked
|
||||
|
||||
- `internal/link` tests:
|
||||
- a call that finishes in time is answered once, in full;
|
||||
- a call that outlasts the window is answered once, that it is running, with its id, and its final
|
||||
answer is kept under that id;
|
||||
- an acknowledged call is answered before its handler goes on;
|
||||
- the bus's refusal text, with the call's reply subject, is recorded on that call, and other
|
||||
errors are not;
|
||||
- settings and secrets are not kept.
|
||||
- `cmd/mesh-controller` tests:
|
||||
- `push`, and `command push …`, answer first, and `status` and `assign` do not;
|
||||
- `calls` is served, refuses an argument it does not declare, and says plainly when it holds no
|
||||
such call;
|
||||
- `AnswerWithin` is under half the console's wait and under `ResponseTTL`.
|
||||
- mesh-tools console test: a timeout says "not a failure", and for the controller names
|
||||
`mesh-controller.calls`.
|
||||
- Once rolled out:
|
||||
- a push through the console answers at once with an id, and `mesh-controller.calls` with that id
|
||||
shows what it sent;
|
||||
- the controller's journal shows no bare `Permissions Violation for Publish to "_INBOX.…"` without
|
||||
the mesh's own line naming the call beside it.
|
||||
@@ -0,0 +1,104 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-06
|
||||
located-in: [mesh-catalog, mesh-controller]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 266. A merge on the bus was never handed to the controller
|
||||
|
||||
## Symptom
|
||||
|
||||
A pull request was merged into the tools repository, which the runtime and the console are built
|
||||
from. Nothing was built. The runtime was then built by hand.
|
||||
|
||||
The forge's poll said in its journal that it announced the merge, 19 seconds after it was made. The
|
||||
controller's journal has no line for it: no plan, and no "nothing the mesh holds reads it". The
|
||||
controller did log the merges just before and just after it, to other repositories, as it always
|
||||
does. It had restarted three minutes earlier, when a new controller was rolled out.
|
||||
|
||||
## Evidence
|
||||
|
||||
- **The announcement is on the bus.** The events stream holds it, on the merge subject, with the
|
||||
merge commit and the files it changed.
|
||||
- **The controller's consumer is past it and holds nothing.** The controller's durable consumer on
|
||||
the events stream has nothing pending and nothing unacknowledged. Its delivered and acknowledged
|
||||
positions are past the announcement.
|
||||
- **It was never handed over.** Since the controller connected, the stream took 25 messages on the
|
||||
consumer's subjects. The server's own count for the controller's subscription is 24 deliveries.
|
||||
The consumer hands over one message at a time, and the next message, a build outcome, was handed
|
||||
over and acted on 20 seconds after the merge. So the merge was not handed over and then lost by
|
||||
the controller. No other subscriber to the consumer's delivery subject existed, then or since.
|
||||
- **Every way the controller could take it logs.** Acting on a merge, or refusing one, writes a line
|
||||
in every case. An announcement the controller could not read, or one it did not recognise, would
|
||||
be taken without a line, but this one has the subject and the body the controller expects. A
|
||||
merge announced into another repository 8 minutes later was acted on and logged.
|
||||
- **It is not the first.** Across three days, 23 merges the poll announced have no matching line in
|
||||
the controller's journal. Most of them fall in windows when the bus or the controller was being
|
||||
replaced. This one does not.
|
||||
|
||||
## Cause
|
||||
|
||||
The bus server ran **nats 2.10.29**. On that release, a JetStream consumer with **more than one
|
||||
filter subject** sometimes moves its position past a matching message without delivering it. It
|
||||
does not count the message as pending, and it does not redeliver it. The controller's consumer has
|
||||
seven filters: merges, build outcomes, catalogue upgrades, a catalogue's catch-up, and the two
|
||||
provider standings. The controller acts only on what that consumer hands it. Nothing compares what
|
||||
was announced with what was acted on, so a skipped merge is silent.
|
||||
|
||||
Reproduced outside the mesh with the server library. A stream shaped like the events stream, with
|
||||
a consumer configured exactly like the controller's, receives traffic like the mesh's: mostly build
|
||||
logs on new subjects, with the followed events among them.
|
||||
|
||||
- On 2.10.29, a few percent of the followed events are never delivered, and the consumer has
|
||||
nothing pending.
|
||||
- On the same release, a consumer with a single filter received every message.
|
||||
- On 2.11.17, 2.12.15 and 2.14.5, every message was delivered.
|
||||
|
||||
Every module consumer with more than one filter on that bus is exposed in the same way. That
|
||||
includes the catalogue's consumer of build outcomes and every provider's consumer of provisioned and
|
||||
deprovisioned events.
|
||||
|
||||
## Fix
|
||||
|
||||
1. **mesh-catalog: the bus server runs 2.11.17.** The nats module's image is pinned by its
|
||||
multi-architecture index digest. The Dockerfile now also states the release that digest is,
|
||||
because a digest does not state it. 2.11.17 is the smallest step that delivers every message in
|
||||
the reproduction. Moving to 2.11 is a one-way upgrade of the stream store, so plan it for when the
|
||||
bus can be restarted.
|
||||
2. **mesh-controller: missed merges are caught up.** Every five minutes the controller reads the
|
||||
last day of merge announcements back from the events stream. It reads them on an ordered consumer
|
||||
of its own, filtered on the merge subject alone, which acknowledges nothing. For each announcement
|
||||
older than ten minutes, it judges the merge the way acting on it would. A merge that was already
|
||||
acted on reads as history, because acting marks every module the merge moved as looked at since.
|
||||
A merge that would still move a module was never acted on. The controller says so, naming the
|
||||
repository, the commit, how late it is and which modules are behind, and then acts on it. So no
|
||||
record of which merges were handled is needed.
|
||||
|
||||
It only considers modules built from the merged repository. Modules that only package source
|
||||
from that repository have no record that acting touched them, so they would always read as never
|
||||
acted on. A merge that moves both kinds is caught through the first kind, and acting on it also
|
||||
rebuilds the second. A merge older than a day is left to a person, because rebuilding it days
|
||||
later would be a surprise.
|
||||
|
||||
## How it is checked
|
||||
|
||||
- mesh-catalog, `modules/nats/cmd/nats-tools`.
|
||||
`TestAConsumerWithSeveralFiltersIsHandedEveryMessage` runs the reproduction above against the
|
||||
server release the module's tests pin. It fails on 2.10.29 and passes on 2.11.17.
|
||||
`TestTheImageIsTheServerTestedHere` fails if the release named in the Dockerfile differs from the
|
||||
tested server's version. The first test is therefore a statement about the server the mesh runs.
|
||||
- mesh-controller, `cmd/mesh-controller`.
|
||||
`TestAMergeTheBusNeverHandedOverIsActedOnLate` replays this incident: an announcement the
|
||||
controller never acted on, among others that were acted on or that nothing reads. It is left to
|
||||
the controller's own consumer for ten minutes. Then it is said and acted on once, and never
|
||||
again. `TestAMissedMergeThatCouldNotBeActedOnIsTriedAgain` and
|
||||
`TestAMergeOlderThanTheLookBackIsLeftAlone` check the edges. Reading the stream back was also
|
||||
checked against an embedded 2.10.29 server holding exactly the controller's bus permissions.
|
||||
|
||||
## Not covered
|
||||
|
||||
If the forge's poll never announces a merge at all, the catch-up has nothing to read. The poll keeps
|
||||
its own record of what it announced, so a restart of the poll does not lose merges. A poll that is
|
||||
down for longer than a day would still lose them.
|
||||
@@ -0,0 +1,90 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-06
|
||||
located-in: [mesh-host, mesh-controller]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 267. A reconcile's report overtook the apply that followed it
|
||||
|
||||
## Symptom
|
||||
|
||||
A release plan sent the home server a declaration and waited for the home server to report it. The
|
||||
host's journal showed the declaration applied nine seconds later, with one line,
|
||||
`applied 659 resource(s)`. The controller's journal showed that line twice in the same second. The
|
||||
report the controller kept for the home server then named a different, older declaration than the
|
||||
one it had recorded sending. The plan requires those two to match for its first machine, so it
|
||||
waited until a person pushed again by hand.
|
||||
|
||||
The same doubled line appears earlier that night, in a log written before the fix for
|
||||
[issue 264](../264-a-self-updating-engine-lost-the-report-of-the-apply-that-delivered-it/00-report.md)
|
||||
existed.
|
||||
|
||||
## Cause
|
||||
|
||||
Two reports, and the older one arrived last.
|
||||
|
||||
The host's five-minute reconcile was due three seconds before the declaration arrived. It took the
|
||||
apply lock first. It read the declaration kept at that moment, which was the older one, held the
|
||||
machine to it and made its report. A reconcile on an adopted node, or on any node that can say
|
||||
which links face outside, offers that report to the link to send without being asked
|
||||
([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md),
|
||||
[ADR 0140](../../02-DECISIONS/0140-the-filter-constrains-what-arrives-from-outside.md)). The link was busy:
|
||||
the delivery's apply was waiting for the same lock. When the apply finished, the link sent its
|
||||
report about the new declaration, and then went back to its loop and sent the queued reconcile
|
||||
report about the old one.
|
||||
|
||||
The controller keeps one report per machine, and the last one written wins. It already sets aside a
|
||||
report about a declaration it has moved past
|
||||
([design 25](../../03-DESIGN/01-to-be/25-the-bus-on-nats.md) §3), but only a report that is
|
||||
nothing more than an apply's account. A report that also says what the machine is (its links, what
|
||||
it holds, its firewall) is always acted on, because that half is never stale. The reconcile's report
|
||||
said both, so its account of the old declaration replaced the account of the new one.
|
||||
|
||||
[Issue 261](../261-a-module-was-applied-and-given-back-half-a-minute-later/00-report.md) fixed
|
||||
which declaration a reconcile applies when it waits for a delivery. This is the other order: the
|
||||
reconcile goes first, and only its report is late.
|
||||
|
||||
**It was not introduced by the fix for issue 264**, which was suspected first because the stall
|
||||
followed its rollout. That fix re-sends a kept report only when the link is opened, and says so in
|
||||
the journal. The host had linked once, an hour earlier, and its journal has no such line. Its kept
|
||||
report was absent, because the broker had taken the last apply's report. And the same doubled
|
||||
report appears in the controller's log before that fix was written.
|
||||
|
||||
## Fix
|
||||
|
||||
Both ends, so that the order of arrival cannot decide which account is kept:
|
||||
|
||||
1. **The host.** The link remembers the declaration it last applied, or last re-sent from what was
|
||||
kept. A report made without being asked is set aside, and counted as not sent, when it names a
|
||||
different declaration. The apply's own report has already described the machine, and it is more
|
||||
recent. The next reconcile sends anything that is still news. A report that names no declaration
|
||||
is sent as before, and so is one made before this link has applied anything.
|
||||
2. **The controller.** A report about a declaration other than the one the mesh last sent a machine
|
||||
still records what it says about the machine, as before. It no longer replaces the stored account
|
||||
of the apply, or the list of what the machine owns. The controller logs that it kept the report's
|
||||
facts and not its account.
|
||||
|
||||
The report does not carry the declaration's sequence number, so the digest the mesh recorded sending
|
||||
is what decides. This is the same check that already applies to a report that is only an account of
|
||||
an apply.
|
||||
|
||||
## How it is checked
|
||||
|
||||
These tests fail without the fix:
|
||||
|
||||
- mesh-host `TestAReconcileReportOlderThanTheApplyIsNotSaidAfterIt`. A reconcile report about the
|
||||
old declaration is queued while the new one is applied. The mesh hears only the report of the
|
||||
new apply, and the reconcile report is counted as not said.
|
||||
- mesh-controller `TestAnAccountOfAnOlderDeclarationDoesNotReplaceTheNewer`. The mesh sent the new
|
||||
declaration and heard its report, then a report about the old one. The kept account still names
|
||||
the new declaration, and the old report's links are recorded.
|
||||
|
||||
These tests guard the edges:
|
||||
|
||||
- mesh-host `TestAReconcileReportAboutTheAppliedDeclarationIsSaid` and `TestOvertaken`.
|
||||
- mesh-controller `TestAnAccountOfTheSentDeclarationReplacesTheKeptOne`.
|
||||
|
||||
Once the fix is rolled out, the live check is a push that arrives while a reconcile is running. The
|
||||
controller should log the report once, and the plan should move on without anyone pushing by hand.
|
||||
@@ -0,0 +1,138 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-06
|
||||
located-in: [mesh-catalog modules/letta, mesh-catalog modules/docker, mesh-controller cmd/mesh-controller]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 268. letta printed its passwords into its log
|
||||
|
||||
## Symptom
|
||||
|
||||
The `letta` container on the home server writes two secrets into its own log every time it starts:
|
||||
|
||||
- `Creating engine postgresql://<user>:<password>@<host>:<port>/<database>`: the database password
|
||||
the mesh granted it from the `postgres` provision;
|
||||
- `▶ Using secure mode with password: <password>`: the letta server password, one of the module's
|
||||
own secrets.
|
||||
|
||||
The same database URI is printed twice more on each start: by the image's `startup.sh` (`External
|
||||
Postgres configuration detected, using …`) and by its migration step (`Using database: …`).
|
||||
The container had restarted about a hundred times, so each line was in the log about a hundred
|
||||
times. The runtime keeps that log in its own file, up to ten files of 100 MB each. Anyone who reads
|
||||
it can read both passwords: the docker module's `docker_logs` tool, anything with the runtime's
|
||||
socket, and every agent transcript that ever called `docker_logs` on this container.
|
||||
|
||||
No value appears in this record.
|
||||
|
||||
## Cause
|
||||
|
||||
**letta prints what it is given.** letta 0.6.8 prints `LETTA_PG_URI` whole in three places, and
|
||||
prints its server password in `--secure` mode. These are bare `print` calls. No log level or flag
|
||||
turns them off. The newest release (0.16.8) still prints the server password and the migration
|
||||
step's URI, so upgrading does not fix it.
|
||||
|
||||
**The module handed it both.** It composed the database password into `LETTA_PG_URI` in an
|
||||
env-file. letta reads its settings only from the environment
|
||||
([ADR 0086](../../02-DECISIONS/0086-a-secret-reaches-a-process-as-a-file.md)'s exception, declared on
|
||||
the container as `secrets-in-environment`). The declared reason even said that `startup.sh` echoes
|
||||
the URI. The exception was written down to justify the environment, and nobody acted on the log line.
|
||||
|
||||
**Nothing in the mesh looks at what a container prints.**
|
||||
[Issue 041](../041-a-sealed-credential-ends-up-in-the-process-environment/00-report.md) closed the
|
||||
environment path. `docker_inspect` leaves environment values out. But a program's own output carried
|
||||
the same values to more readers, and no check compared a container's log with the secrets it was
|
||||
given.
|
||||
|
||||
The pattern is wider than letta. Ten more modules compose a granted password into a connection URI
|
||||
the same way. Whether each program prints it is the program's business, so a manifest lint cannot
|
||||
decide it. Only reading the log can.
|
||||
|
||||
## Fix
|
||||
|
||||
**letta** (mesh-catalog, `modules/letta`):
|
||||
|
||||
- `LETTA_PG_URI` names no password. libpq, through letta's psycopg2 driver, reads the password from
|
||||
a `pgpass` file the module writes at `0600` and mounts read-only, named by `PGPASSFILE`. All three
|
||||
prints of the URI now carry no secret. The file is mounted directly into the container, so it is
|
||||
part of the container's spec by content
|
||||
([issue 103](../103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md)), and a
|
||||
rotated password recreates the container.
|
||||
- The container starts through a small script the module writes. The script rewrites the one line
|
||||
of the image's code that prints the server password, then executes the image's own `startup.sh`.
|
||||
The image is pinned by digest, so the rewrite is exact. **If the server password is still printed
|
||||
after the rewrite, or a password is back in `LETTA_PG_URI`, the script refuses to start letta and
|
||||
says why.** A letta that does not start is diagnosable. A letta that leaks is silent.
|
||||
- Both own secrets now say they are read at start (`taken: at-start`,
|
||||
[ADR 0114](../../02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md)). Without
|
||||
that, `rotate` refuses to replace the server password.
|
||||
|
||||
Two options were rejected. Filtering the container's output through a pipe would lose the
|
||||
process's signals and exit status, and would have to recognise every secret by a pattern. Building a
|
||||
derived image would put a build of a third-party image into the pipeline to change one line.
|
||||
|
||||
**The safety net** (mesh-catalog, `modules/docker`): a new tool, `docker_secrets_in_logs`. It reads
|
||||
the recent log of every container the mesh holds and compares it with three things: the values in
|
||||
that container's environment whose names say they are secrets, the password in any URI the
|
||||
environment holds, and the shape `scheme://user:password@` anywhere in a line. A finding names the
|
||||
container, the module and the variable, with a count and the first and last time it appeared. It
|
||||
never includes the value or the line. A container whose log cannot be read is listed as not read,
|
||||
never as clean. `docker_logs` redacts the same values before answering, because its answers end up in
|
||||
agent transcripts.
|
||||
|
||||
The tool has a limit. A secret delivered only as a mounted file is not known to it, because the tools
|
||||
run as the operator account, which cannot read the host's `0600` files. Such a secret is caught only
|
||||
when it is printed inside a URI.
|
||||
|
||||
**Rotation** (mesh-controller): before this change, `rotate <provision> --consumer <machine>` was the
|
||||
narrowest way to rotate a pair credential. It replaced the credential of every module on that machine
|
||||
that consumes the provision, and restarted all of them. `--module` (and `module` beside `provision`
|
||||
on the verb) narrows the rotation to one consuming module.
|
||||
|
||||
## Rotation, once the fix runs
|
||||
|
||||
The two secrets have already been exposed, so they are replaced after the fix is live. Each step
|
||||
below needs the operator's approval.
|
||||
|
||||
1. **Roll out the fix.** Merge the catalogue and controller changes and let the pipeline build. Then
|
||||
push the home server. The container's spec changes (its arguments, its mounts, its env-file), so
|
||||
the host removes the `letta` container and creates it again. **The runtime's log file goes with
|
||||
the removed container**, and this purges the retained copies under the runtime's own driver.
|
||||
Check: `docker_logs` on `letta` shows the URI with no password and `Using secure mode (the
|
||||
password is not printed)`, and `docker_secrets_in_logs` names nothing for `letta`.
|
||||
2. **The database password.** Run `mesh-controller.rotate` with `provision: postgres-database`,
|
||||
`consumer: <the home server>` and `module: letta`. The new credential is sent to both ends together.
|
||||
The `pgpass` file changes, so letta is recreated with the new value. Before the controller change
|
||||
is deployed, the only form available is the one without `module`, which also rotates every other
|
||||
module on the machine that consumes `postgres-database`. The command lists them before it acts.
|
||||
3. **The server password.** Run `mesh-controller.rotate` with `node: <the home server>`,
|
||||
`module: letta` and `secret: server-password`. If the mesh made the value, it makes a new one and
|
||||
pushes it. The env-file and the runtime's config file change, and both letta containers start
|
||||
again on the new value. If the value was *given* to the mesh (`secret accept`, when the server
|
||||
already had clients), the controller refuses with the reason
|
||||
([ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md)). In that case, generate a new
|
||||
value into a `0600` file on the controller's machine, run `secret accept <the home server> letta
|
||||
server-password --from <file>`, remove the file, and push. In both cases, every client outside
|
||||
the module that calls letta with that password (a workflow reaches it by its public name) needs
|
||||
the new value too.
|
||||
4. **Copies the runtime did not hold.** The container logs through the runtime's own file driver,
|
||||
not the journal. Confirm with `docker_inspect` (`LogConfig.Type`). Then check whether a backup on
|
||||
the home server includes the runtime's data directory, and if it does, which snapshots predate
|
||||
step 1. Agent transcripts that called `docker_logs` on `letta` hold the old values. Steps 2 and 3
|
||||
make those values useless, and deleting the transcripts is optional.
|
||||
|
||||
## How it is checked
|
||||
|
||||
- The start script was tested against the image's own `app.py`, extracted from the pinned digest. It
|
||||
rewrites the line, and the result compiles. It refuses to start when a print of the password
|
||||
remains or when the URI carries a password. It allows an unchanged rerun.
|
||||
- `mesh-controller module check` passes for `letta` and `docker`. The catalogue tests pass.
|
||||
- The docker module's tests: a scan of a log shaped like this leak names the server password, the
|
||||
password inside an environment URI, and a URI found by its shape, by name only. The answer, marshalled
|
||||
whole, contains no value. A masked `***` is not counted. An unreadable log is reported, not called
|
||||
clean. `docker_logs` returns the same lines redacted.
|
||||
- The controller's tests: the verb passes `module` beside `provision` through as `--module`, and the
|
||||
narrowing touches only that module's credentials.
|
||||
- On the mesh, after step 1: `docker_secrets_in_logs` on the home server. It is run again after any
|
||||
change to a module that composes a secret into a URI.
|
||||
Reference in New Issue
Block a user