diff --git a/02-DECISIONS/0037-where-a-module-lives.md b/02-DECISIONS/0037-where-a-module-lives.md index a697d12..c01e6c6 100644 --- a/02-DECISIONS/0037-where-a-module-lives.md +++ b/02-DECISIONS/0037-where-a-module-lives.md @@ -1,6 +1,6 @@ --- topic: building it -status: proposed +status: accepted date: 2026-09-01 deciders: jochen reconstructed: false @@ -101,3 +101,19 @@ the digest down after building. **Whether kind 4 deserves a module at all.** Thirty-five descriptions that say *install this and write these files* may be better as one module with settings than as thirty-five modules. Left open deliberately; it is a question about the shape of the catalogue, not about whether to have one. + +## Accepted, 2026-09-29, against what was built + +*Marked in a grooming pass: the mesh was built to this record and the record still said `proposed`.* + +The proposal is the arrangement that exists. `mesh-catalog` holds descriptions of software we did +not write and the programs that provision it, and holds neither the mesh's own components nor an +application's own module. The mesh's list of modules is a table in the control plane, filled by +`module add`, and every module records the source it came from with the commit it was read at. + +**One half is not built: `module check` as a command on the control plane's binary.** A manifest is +still validated by a test that reaches into the control plane's internals — which works for this +catalogue and gives nothing at all to somebody describing their own application in their own +repository, which this record says is the case that matters most. That is +[issue 148](../04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md). + diff --git a/02-DECISIONS/0113-the-vault-makes-every-secret.md b/02-DECISIONS/0113-the-vault-makes-every-secret.md index a8ae287..b4b61fa 100644 --- a/02-DECISIONS/0113-the-vault-makes-every-secret.md +++ b/02-DECISIONS/0113-the-vault-makes-every-secret.md @@ -1,6 +1,6 @@ --- topic: what runs on it -status: proposed +status: accepted date: 2026-09-25 deciders: jochen reconstructed: false @@ -275,3 +275,11 @@ On acceptance, each of these is superseded or amended by this record, not edited everything a module needs is a requirement - [Issue 095](../04-ISSUES/095-a-module-assigned-after-genesis-has-no-broker-account/00-report.md), [issue 103](../04-ISSUES/103-a-container-is-not-recreated-when-a-file-it-reads-changes/00-report.md): what fails today + +## Accepted, 2026-09-29, against what was built + +*Marked in a grooming pass.* The vault is a module providing `secret` at mesh scope, and six +modules in the catalogue require it — so a shared secret is a requirement answered by the vault, +which is what this record asks for. Private keys are still made where they are used and never +travel, which is the other half and was never in question. + diff --git a/02-DECISIONS/0115-one-assignment-of-a-module-per-node.md b/02-DECISIONS/0115-one-assignment-of-a-module-per-node.md index 341a3ff..5936529 100644 --- a/02-DECISIONS/0115-one-assignment-of-a-module-per-node.md +++ b/02-DECISIONS/0115-one-assignment-of-a-module-per-node.md @@ -1,6 +1,6 @@ --- topic: what runs on it -status: proposed +status: accepted date: 2026-09-26 deciders: jochen extends: 0112-a-module-definition-names-no-node-mesh-or-path.md @@ -47,3 +47,10 @@ other boundary already is: the module name. node runs one of each (ADR 0115)" — instead of failing on whichever name collides first. - Multi-tenant asks are answered in the catalogue (a second module definition), not in the control plane. + +## Accepted, 2026-09-29, against what was built + +*Marked in a grooming pass.* The rule is enforced where it cannot be forgotten: `assignment`'s +primary key is `(node, module)`, so a second assignment of one module to one machine is not a thing +the mesh can hold. The record read `proposed` while the schema had already settled it. + diff --git a/02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md b/02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md index 317ec90..c5958f1 100644 --- a/02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md +++ b/02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md @@ -12,7 +12,7 @@ extends: 0102-the-mesh-writes-into-a-shared-file-never-over-it.md ## Context When a resource stops being declared — its module unassigned, the node sent a -deliberately-empty declaration ([issue 127](../04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)), +deliberately-empty declaration ([issue 149](../04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)), or a new catalogue version renaming its id — the host undoes it. The host's own code states the rule it means to follow: **it removes what it made and leaves what it merely configured.** For almost every resource it does exactly that: diff --git a/02-DECISIONS/README.md b/02-DECISIONS/README.md index 3e1cf39..cbf086a 100644 --- a/02-DECISIONS/README.md +++ b/02-DECISIONS/README.md @@ -207,9 +207,9 @@ python3 00-META/checks/index.py fail if stale - **0099** — [A step that runs once names what it reads, and runs again when it changed](0099-a-step-that-runs-once-names-what-it-reads.md) - **0110** — [A seat is held by one assignment, from a closed set, and it may deliver a provision](0110-a-seat-is-a-module-assignment-from-a-closed-set.md) - **0112** — [A module definition names no node, no mesh and no path: everything it needs is a requirement the mesh resolves](0112-a-module-definition-names-no-node-mesh-or-path.md) -- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md) *(proposed)* +- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md) - **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md) *(proposed)* -- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md) *(proposed)* +- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md) - **0117** — [A machine's uplink is a seat: the mesh configures the manager, never the link](0117-a-machines-uplink-is-a-seat.md) - **0118** — [Undeclaring removes what the mesh made, and gives a unit back the state it was found in](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md) - **0120** — [A roster fact carries its format as a template: the mesh owns the data, the module owns the format](0120-a-roster-fact-carries-its-format-as-a-template.md) @@ -237,7 +237,7 @@ python3 00-META/checks/index.py fail if stale - **0014** — [No workspace — each module is a standalone package consuming published dependencies](0014-no-npm-workspace.md) - **0015** — [Applications live in their own repository; the monorepo is for the mesh](0015-applications-live-in-their-own-repository.md) - **0016** — [The lab](0016-the-lab.md) -- **0037** — [Where a module lives](0037-where-a-module-lives.md) *(proposed)* +- **0037** — [Where a module lives](0037-where-a-module-lives.md) - **0039** — [What the SDK holds, and what it refuses](0039-what-the-sdk-holds-and-refuses.md) - **0068** — [The lab takes requests, one at a time, and runs each from its own copy](0068-the-lab-takes-requests.md) *(proposed)* - **0069** — [A module is a repository and a path within it](0069-a-module-is-a-repository-and-a-path.md) diff --git a/04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md b/04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md index 7c8280e..2c8c762 100644 --- a/04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md +++ b/04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md @@ -155,3 +155,17 @@ design document here, and get it back. That check fails today by design. **What stands until then** is the signpost, and the honest description of it: reachable, not surfacing. + +## Where this stands, 2026-09-29 + +*Added in a grooming pass.* The knowledge base this record is about is the **predecessor's**, and it +is no longer reachable from anything: the surface that answered `recall_search` speaks the transport +the mesh removed at the cut-over +([issue 147](../147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md)). + +So the sentence in `README.md` that this record catches — *these documents are still indexed into +the knowledge base* — is now wrong twice over: nothing indexed them, and there is nothing to index +them into. The record stays open, and its answer is no longer "index this repository somewhere"; it +is whatever the mesh grows as its own knowledge surface, if it grows one. Until then the honest fix +is the README, which should stop claiming a property nothing provides. + diff --git a/04-ISSUES/088-the-forges-own-address-names-a-port-it-may-not-have/00-report.md b/04-ISSUES/088-the-forges-own-address-names-a-port-it-may-not-have/00-report.md index 4aa2eea..890b6c8 100644 --- a/04-ISSUES/088-the-forges-own-address-names-a-port-it-may-not-have/00-report.md +++ b/04-ISSUES/088-the-forges-own-address-names-a-port-it-may-not-have/00-report.md @@ -1,8 +1,8 @@ --- -status: open +status: resolved opened: 2026-09-22 -located-in: [] -fixed-by: +located-in: [mesh-catalog modules/gitea, mesh-controller internal/catalogue/declaration.go] +fixed-by: mesh-controller 7352c84, merged in #46 — a module is told its port in a container's environment too, as a file already was amended-design: --- @@ -39,3 +39,9 @@ the assignment happens to differ. - Should composition refuse an environment value that names a port the module does not fix, the way it refuses other claims a module cannot make? - Which other modules write their own address, with a port, into their environment? + +## Closed + +*2026-09-29, in a grooming pass rather than by whoever fixed it.* The forge's address follows a moved port the same way every other reader does. Found by +reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was +not re-verified on a machine, and this record says so rather than implying a run that did not happen. diff --git a/04-ISSUES/089-a-contributed-route-names-a-port-the-node-may-have-moved/00-report.md b/04-ISSUES/089-a-contributed-route-names-a-port-the-node-may-have-moved/00-report.md index 80df81c..038b028 100644 --- a/04-ISSUES/089-a-contributed-route-names-a-port-the-node-may-have-moved/00-report.md +++ b/04-ISSUES/089-a-contributed-route-names-a-port-the-node-may-have-moved/00-report.md @@ -1,8 +1,8 @@ --- -status: open +status: resolved opened: 2026-09-22 -located-in: [] -fixed-by: +located-in: [mesh-controller internal/catalogue/declaration.go, mesh-catalog (every routed module)] +fixed-by: mesh-controller bdf965d (a route names the endpoint it serves) with `portOfEndpoint` and `AtPublishedPort` — the contribution carries the endpoint's declared port and the machine-side redirection is applied to it like any other amended-design: --- @@ -45,3 +45,9 @@ precisely because the predecessor holds the usual one. keeping the mapping out of rendered configuration? - What should refuse a declaration whose contributed route names a port nothing on that node listens on? + +## Closed + +*2026-09-29, in a grooming pass rather than by whoever fixed it.* A route names an endpoint rather than a port, and the redirection that turns a declared port into the published one is applied to contributions too. Found by +reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was +not re-verified on a machine, and this record says so rather than implying a run that did not happen. diff --git a/04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md b/04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md index 6b03b96..194fbd7 100644 --- a/04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md +++ b/04-ISSUES/114-should-the-controller-be-a-container-or-a-process/00-report.md @@ -1,8 +1,8 @@ --- -status: open +status: resolved opened: 2026-09-24 located-in: [mesh-controller module.json, mesh-host internal/apply] -fixed-by: +fixed-by: ADR 0142 — the mesh's own components are delivered as binaries on the machine, so the controller is a process; the delivery itself is issue 142 amended-design: --- @@ -84,3 +84,21 @@ restart and run-to-completion semantics — so this would not need host-side wor The two do not collapse into one. The controller is not a code-carrying sidecar, and `network: host` is what makes the asymmetry visible here and nowhere else. + +## Answered + +*2026-09-29, in a grooming pass.* This asked a question rather than reporting a defect, and the +question was taken: [ADR 0142](../../02-DECISIONS/0142-the-mesh-delivers-its-own-components-as-binaries.md) +decides that the mesh's own components — the host, the controller, the catalogue, the builder, the +vault — are **binaries on the machine**, delivered by the mechanism +[ADR 0141](../../02-DECISIONS/0141-the-host-delivers-its-own-successor.md) built, and that +third-party software (the store, the registry, the broker) stays a container because an image is the +right way to carry somebody else's build. + +So the operating experience this record was written from — every mutating command reached through +`docker exec mesh-controller` — is answered, and answered against the container. + +**The delivery is a separate matter and is not this record's.** Step 1 of it is built and no +component travels yet; that is +[issue 142](../142-the-host-is-the-one-thing-the-mesh-does-not-deliver/00-report.md). + diff --git a/04-ISSUES/120-a-provisioner-remembers-what-it-did-not-what-is/00-report.md b/04-ISSUES/120-a-provisioner-remembers-what-it-did-not-what-is/00-report.md index 9b697f3..078d27c 100644 --- a/04-ISSUES/120-a-provisioner-remembers-what-it-did-not-what-is/00-report.md +++ b/04-ISSUES/120-a-provisioner-remembers-what-it-did-not-what-is/00-report.md @@ -1,8 +1,8 @@ --- -status: located +status: resolved opened: 2026-09-26 located-in: [mesh-sdk src/provisioner, mesh-catalog modules/redis] -fixed-by: +fixed-by: mesh-catalog bbda88c, merged in #84 — every credential provider says whether it still holds a consumer amended-design: --- @@ -60,3 +60,9 @@ checks it after the first pass. instance and leaves the gap for the others. - Where does the record of what was applied live, if not in memory? ADR 0114, still proposed, puts rotation state with the vault. The same place may answer this. + +## Closed + +*2026-09-29, in a grooming pass rather than by whoever fixed it.* A provisioner asks the backend what is there rather than trusting what it remembers doing. Found by +reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was +not re-verified on a machine, and this record says so rather than implying a run that did not happen. diff --git a/04-ISSUES/128-the-hosts-file-is-written-whole/00-report.md b/04-ISSUES/128-the-hosts-file-is-written-whole/00-report.md index fa02a00..509ee3f 100644 --- a/04-ISSUES/128-the-hosts-file-is-written-whole/00-report.md +++ b/04-ISSUES/128-the-hosts-file-is-written-whole/00-report.md @@ -1,5 +1,6 @@ --- -status: located +status: resolved +fixed-by: mesh-host 1cb8953 and fdc768c — the mesh writes into a marked block of a text file instead of over it opened: 2026-09-26 located-in: [mesh-controller internal/catalogue/facts.go, mesh-host internal/apply] --- @@ -71,3 +72,9 @@ private network loses that name too. - The host's file resource supports `into: "json"` only; anything else is a whole write. - `node show ` on the adopted workstation: `holds file /etc/hosts mesh-wireguard.fact-node-names`, original kept. + +## Closed + +*2026-09-29, in a grooming pass rather than by whoever fixed it.* A shared hosts file keeps every line that is not the mesh's. Found by +reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was +not re-verified on a machine, and this record says so rather than implying a run that did not happen. diff --git a/04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/01-diagnosis.md b/04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/01-diagnosis.md index 0071b16..e262608 100644 --- a/04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/01-diagnosis.md +++ b/04-ISSUES/129-nothing-makes-a-machine-trust-the-meshs-authority/01-diagnosis.md @@ -30,3 +30,16 @@ network, installs it as a trust anchor, refreshes the machine's bundles, and — unassigned stops its unit, and stopping the unit is what undoes it — takes both away again. The owner is therefore `mesh-catalog`, module `ca-trust`, and nothing in the control plane. + +## The module exists, and this stays open until a machine holds it + +*2026-09-29.* `ca-trust` is in the catalogue and merged +([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md)), and what it renders +is checked in the control plane's own suite: the script fetches from the authority it was bound to, +and the unit runs it both ways. + +**No machine has been assigned it, and nothing has verified a name because of it.** The bed written +for that cannot run ([issue 146](../146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md)), +and the live mesh has not been given the module. So the symptom this record opened on — every +internal name failing verification on every machine — is still true everywhere, and the record stays +`located` until it is not. Closing it on a module that exists would be closing it on an intention. diff --git a/04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md b/04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md index 57fb3de..431a3f8 100644 --- a/04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md +++ b/04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md @@ -1,5 +1,6 @@ --- -status: located +status: resolved +fixed-by: mesh-host 3112c88 — undeclaring gives a unit back the state it was found in, and removes only a process the mesh made opened: 2026-09-27 located-in: [mesh-host internal/apply/apply.go (remove)] amended-design: 02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md @@ -15,7 +16,7 @@ found that the host's `remove` path stops every `service` resource that is no lo delete". `store.Orphans` matches by id alone. So any of these stops the unit: - the module is unassigned — by mistake, or to switch it for another; -- the node is sent a deliberately-empty declaration ([issue 127](../127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)); +- the node is sent a deliberately-empty declaration ([issue 149](../149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)); - a later catalogue version renames the resource's `id`. That is right for a service the mesh brought into being. It is wrong for a unit the mesh @@ -64,3 +65,9 @@ something to settle in passing. The unassign preview is partly answered — the host's plan names each unit it will stop — and the controller's side is left open. + +## Closed + +*2026-09-29, in a grooming pass rather than by whoever fixed it.* Undeclaring no longer stops a unit the mesh only reloaded or only kept running. Found by +reading what the code repositories' commits cite: the fix names this issue and is on `main`. It was +not re-verified on a machine, and this record says so rather than implying a run that did not happen. diff --git a/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md b/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md index bcbd17b..317bb01 100644 --- a/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md +++ b/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md @@ -82,12 +82,100 @@ managing, or genesis carries a user list that includes the first node's enrolmen takes over from there. Both are decisions, not patches, and both belong to the genesis step that was deliberately left until last. +## 5 — the composed user list has to be placed by hand at genesis *(fixed)* + +The account a token is the password of is **not recorded at all**: the composer names an enrolment +user for every machine with a live token, nothing minted a credential for it, and the composition +left it out as a user with no password. The comment above the issuing code already claimed +otherwise — *"the account is created before the token is handed over"* — which is how it went +unnoticed. Issuing a token now records that account, with the token's own secret as its password, +because that is the string the machine will present. + +Placing it is the other half. The list reaches the machine running the bus in that machine's +declaration, which a machine that has not enrolled does not get, so at genesis it cannot arrive +that way. **The control plane composes and says what it composed** — `broker accounts`, to standard +output — and whoever is raising the machine writes it beside the bus's configuration and makes the +server re-read it. Twice, because two accounts come into existence at different moments: the +enrolment when the token is issued, and the machine's own when it enrols. A control plane that +wrote the file itself would have to know where the bus keeps its configuration and how to make it +reload, which is the module's knowledge and is what the module takes over on the first push. + +With that, **a first node enrols against the bus it just raised** — measured, from bare, in the +lab. + +## 6 — and is enrolled twice, keeping a credential the mesh has replaced *(fixed)* + +``` +mesh-controller: enrolled anchor +mesh-controller: enrolled anchor (the same second) +``` + +One `enrol` on the machine, two enrolments in the control plane. Each mints the node a fresh bus +password and returns it; the machine keeps the answer to the first, and the mesh keeps the hash of +the second. The machine then reconnects for ever as a user whose password the mesh rotated out from +under it — *authentication error - User "node.anchor"* on the bus, `Authorization Violation` in the +host's log, and a node that never reports. + +What is ruled out: the host asking twice — it asks again only when the mesh says *try again*, and +a refused attempt is not logged as an enrolment. Redelivery by the consumer — there is one +consumer, its acknowledgement window is thirty seconds, and the handler is quick. + +What is left: the client re-publishing when an acknowledgement is slow, which is what its defaults +do. That was addressed by giving the publish a message id derived from its own bytes, so the stream +discards the copy — **and the duplicate survived it**, so either the id is not reaching the stream +or the second copy is not a copy. This is where the trail stops. + +Worth saying plainly: **the mint is the fragile part, not the delivery.** An enrolment answered +twice is survivable if the answer is the same both times, and it cannot be — the mesh keeps only +the hash, so a second answer is necessarily a different credential. + +**Found, and it is not about enrolment at all** *(fixed)*. The bus's own counters settled it: one +message published, one held in the stream, one delivery, nothing redelivered — and the controller +enrolled the machine twice. So the handler ran twice on one delivery. + +A push consumer delivers onto an ordinary subject, and **everything subscribed to that subject gets +a copy**. The controller holds a consumer called `controller` on CONTROL and another called +`controller` on EVENTS, and the delivery subject was derived from the consumer's name alone — so +both were `_DELIVER.controller`, the one process held both subscriptions, and every message from +either stream was acted on twice. + +Enrolment is where it drew blood, because enrolling twice mints twice and the second credential +replaces the first. But it applied to **every report and every event the controller follows**, and +it is the kind of fault that leaves no trace: nothing is redelivered, no counter is wrong, the work +simply happens twice. The comment in the receiving code about a merge that ran the whole catalogue +five times over on 2026-09-28 is the same shape seen from the other end. + +The stream is in the delivery subject now, because the pair is what identifies a consumer — the +server scopes a durable's name to its stream, and this subject was the one place that scoping was +dropped. A subscriber's permission gains the same shape, keeping the bare name so an existing +consumer keeps working until the controller's next assertion moves it. + +*How it is checked:* the consumers the mesh asks for are asserted to deliver onto distinct subjects, +in the controller's own suite. Against a server it would be invisible, which is the point. + ## Where it belongs `mesh-host` (the bundle and the enrolment path) and `mesh-controller` (the certificate command, and the composition that cannot reach the bus at genesis). Three of the four are fixed on branches; the fourth is the genesis work. +## What made it slow, and what was changed so it is not + +Six faults behind one another, each found by raising a machine and reading what it said. What cost +the most was not the faults: + +- **Every bed's own instructions named the bundle that cannot work**, so the first three attempts + ended in a control plane crash-looping on a missing bus. They name the working one now. +- **A host binary built without its system** refuses everything it is given with *this host was + built for ""*, which reads like a broken bundle. The lab's README says so. +- **`make image` in the control plane had been broken for as long as its base was pinned**: the + Dockerfile's fallback is a Go older than the module asks for, and the pipeline never saw it + because the pipeline passes the declared base in. It reads the base from the manifest now. +- **Leaving the machine standing is what answers the question.** Every finding above came from + shelling in afterwards — the host's log, the bus's log, the file the bus was actually handed — + and none from the test's own output, which says only that nothing converged. The bed takes + `MESH_LAB_KEEP`, and the README says to reach for it first. + ## What it cost, for the next person Every lab bed still names `foundation-first-node.lock` in its own instructions, and that bundle diff --git a/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md b/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md new file mode 100644 index 0000000..b056a21 --- /dev/null +++ b/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md @@ -0,0 +1,75 @@ +--- +status: located +opened: 2026-09-29 +located-in: [nothing in the mesh — the surface in use is the predecessor's, installed on the workstation] +--- + +# 147 — the operator's tools still dial the bus that was removed + +## What was observed + +Every tool call an operator makes against the mesh fails, on every machine, with the same answer: + +``` +AMQP not connected — cannot reach hal/mesh@novox +AMQP not connected — cannot reach hal/mesh@shanks +``` + +The mesh moved to one bus and the previous transport was deleted +([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md), cut over +2026-09-28). The surface an operator drives the mesh through — the tool bridge their assistant +speaks to — still opens an AMQP connection, so it cannot reach anything. Not one node answers, +including the machine the operator is sitting at. + +**What that leaves.** The mesh itself is healthy: nodes are current, modules run, the bus carries +the mesh's own traffic. What is gone is the way a person asks it anything. Every question — what a +node runs, what is assigned, what a module's settings are — has to be asked by opening a shell on +the machine and running the control plane's binary inside its container, which is precisely the +path the tool surface exists to remove, and which nothing checks, records or permits. + +It also silently changes how work gets done: an assistant told to use the mesh's tools finds them +dead, falls back to `ssh` and `docker exec`, and the operator discovers the fallback rather than +the fault. + +## Why this is here and not a note in the knowledge base + +The rule is that everything happens over the mesh's bus. A surface that cannot reach that bus is +that rule enforced by nothing — and the mesh reports itself healthy throughout, because what broke +is not a module, a node or a provision but the thing standing outside asking them questions. + +The same cut-over removed the assistant's long-term memory (`recall`, `memorize`) for the same +reason, which is why lessons from the last two days were written into this repository by hand. + +## What would have prevented it + +- **The tool surface as a consumer of the bus like any other**, so moving the bus moves it — rather + than a separate bridge with its own connection settings that nothing resolves. +- **A check that a tool call reaches a node**, run where the mesh's other checks run. Every check + the mesh makes today is about the relationship between the mesh and a machine; none asks whether + a person can ask it anything ([issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md), + the same shape one level out). + +## Diagnosed at once, because the answer was in the configuration + +**The surface is not the mesh's.** The tool server the operator's assistant speaks to is the +predecessor's brain, installed on the workstation and started as a local process, with the +predecessor's broker URL — `amqp://…@` — written into the assistant's own +configuration. Nothing about it is a module: it has no manifest, no assignment, no seat, no account +on the mesh's bus, and the mesh has never known it exists. + +So nothing regressed. The mesh removed a transport that this program still dials, and the program +was never part of the mesh to be moved. **The mesh has a tool model** — `tools` in a manifest, the +calls a seat accepts, the subjects a module answers on, and a control-plane command that asks one — +and **nothing publishes an operator-facing surface onto it.** What an operator uses is the thing +that came before, kept alive by a URL in a file. + +That is the issue, and it is larger than a broken connection: the way a person drives this mesh is +outside the mesh. + +## Evidence to carry into diagnosis + +- The failure text names `hal/mesh@`, so the bridge is resolving a node and then dialling the + old transport. +- It fails identically for the local machine, which rules out reachability and points at the + transport alone. +- The mesh's own traffic over the new bus is unaffected: nodes report, declarations apply. diff --git a/04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md b/04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md new file mode 100644 index 0000000..1bd3f00 --- /dev/null +++ b/04-ISSUES/148-a-manifest-outside-this-catalogue-has-no-check/00-report.md @@ -0,0 +1,37 @@ +--- +status: open +opened: 2026-09-29 +located-in: [mesh-controller cmd/mesh-controller] +--- + +# 148 — a manifest outside this catalogue has no check + +## What was observed + +A module's manifest is validated by a **test** — `internal/catalogue`'s suite parses every manifest +in the catalogue checkout beside it and fails on one it cannot resolve. That works, and it is how +several real faults were caught before a machine saw them. + +It is available to exactly one repository: this one. Somebody describing their own application in +their own repository — the case +[ADR 0037](../../02-DECISIONS/0037-where-a-module-lives.md) calls *the case that matters most* — +has no check at all. They write a manifest, register it with a running mesh, and find out whether +it is valid when the mesh refuses it, or later, when a machine applies something that resolved and +should not have. + +The same record asks for the answer: **a `module check` command on the control plane's binary**, so +a manifest is checked by the tool rather than by a test that imports the tool's internals. + +## What would have prevented it + +Nothing prevents this; it was noticed and left. ADR 0037 named it on 2026-09-01 and the record sat +`proposed` until 2026-09-29, so the missing half was never anybody's task. + +## Evidence to carry into diagnosis + +- `mesh-controller/internal/catalogue` — `ParseManifest` and `CatalogueProblems` are the check, and + both are internal. +- The catalogue-wide test is `TestEveryCatalogueManifestDeclaresWhatItMounts` and its siblings; they + take a path from `MESH_CATALOG`, so the mechanism is already path-driven and not repository-bound. +- `mesh-controller module add` refuses a bad manifest at registration, which is the same check far + too late: by then it is in a running mesh's records. diff --git a/04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md b/04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md similarity index 57% rename from 04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md rename to 04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md index 9a57cb6..b5d797b 100644 --- a/04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md +++ b/04-ISSUES/149-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md @@ -1,10 +1,15 @@ --- -status: located +status: resolved opened: 2026-09-27 -located-in: [mesh-controller cmd/mesh-controller/push.go] +located-in: [mesh-controller cmd/mesh-controller/push.go, mesh-controller cmd/mesh-controller/sendable.go] +fixed-by: mesh-controller sendable.go and push.go — an empty declaration is sent carrying `owns_nothing`, and the host refuses an empty body that does not carry it --- -# A declaration that shrinks to empty is skipped, so the node keeps what it should drop +# 149 — a declaration that shrinks to empty is skipped, so the node keeps what it should drop + +*Opened as 127 and renumbered on 2026-09-29: two records were given that number on the same day.* +*The other kept it, because three documents and three source files cite it by number and nothing +cited this one but a decision and a sibling issue, both corrected with this move.* ## What was observed @@ -37,3 +42,17 @@ mean "own nothing", which the host already applies correctly when it receives on On ace, one command drops it permanently (the corrected controller never re-composes it): `sudo ufw delete allow 5671`. At ace's converge it would clear on its own. + +## Closed + +*2026-09-29, in a grooming pass.* Both halves are on `main` and both name this issue. + +- The control plane **sends** it: a declaration that composes to no resources goes out with + `owns_nothing`, and `push` says *sent, not skipped*. +- The host **refuses an empty body that does not carry it**, so a truncated or mis-composed + declaration can never be read as "own nothing" — which is the failure the fix had to avoid while + making the empty case expressible. + +Closed by reading the code rather than by watching a machine let go of a stray resource; the record +says so rather than implying a run. + diff --git a/04-ISSUES/147-a-route-is-contributed-before-its-module-is-taken/00-report.md b/04-ISSUES/150-a-route-is-contributed-before-its-module-is-taken/00-report.md similarity index 97% rename from 04-ISSUES/147-a-route-is-contributed-before-its-module-is-taken/00-report.md rename to 04-ISSUES/150-a-route-is-contributed-before-its-module-is-taken/00-report.md index f9afe62..77220e6 100644 --- a/04-ISSUES/147-a-route-is-contributed-before-its-module-is-taken/00-report.md +++ b/04-ISSUES/150-a-route-is-contributed-before-its-module-is-taken/00-report.md @@ -6,7 +6,7 @@ located-in: fixed-by: --- -# 147 — A route is contributed before its module is taken +# 150 — A route is contributed before its module is taken ## What was observed diff --git a/04-ISSUES/148-a-new-name-recreates-every-container-in-the-mesh/00-report.md b/04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md similarity index 96% rename from 04-ISSUES/148-a-new-name-recreates-every-container-in-the-mesh/00-report.md rename to 04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md index 8185c6a..bf44ae6 100644 --- a/04-ISSUES/148-a-new-name-recreates-every-container-in-the-mesh/00-report.md +++ b/04-ISSUES/151-a-new-name-recreates-every-container-in-the-mesh/00-report.md @@ -7,13 +7,13 @@ located-in: fixed-by: --- -# 148 — A new name recreates every container in the mesh +# 151 — A new name recreates every container in the mesh ## What was observed Migrating one small module on ace (searxng) took four routine controller actions: `node public-domain ace zurag.be`, `assign ace searxng` + push, `unassign ace searxng` + push (a rollback, see -[issue 147](../147-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take +[issue 150](../150-a-route-is-contributed-before-its-module-is-taken/00-report.md)), and assign + take + push again. Each push to ace also pushed novox ("this push left g14, novox, shanks behind … sending it too"). novox's host then **replaced every container it runs, twice**: diff --git a/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md new file mode 100644 index 0000000..6dcd31f --- /dev/null +++ b/04-ISSUES/152-a-nodes-plan-failure-silently-drops-its-routed-names/00-report.md @@ -0,0 +1,124 @@ +--- +status: resolved +opened: 2026-09-29 +located-in: + - mesh-controller cmd/mesh-controller/plan.go (routeNamesInTheMesh skips a node whose plan will not compose) +fixed-by: mesh-control fix/152-a-lookup-failure-is-not-an-absence +amended-design: +--- + +# 152 — A node whose plan will not compose silently removes its names from every machine + +## What was observed + +For twenty-five minutes, and for seventeen of them after the last operator action, the control-node's +host applied all 327 of its resources every ~6.5 minutes without pause, replacing every container on +the machine each time — the control plane's own store and registry, the edge proxy, the forge, the directory, mail, +and the bus the mesh runs on. The forge's web surface answered `502` throughout; load on the machine +sat near 8. Nothing was converging: each pass ended and the next began four seconds later. + +The two machines carrying no containers were not churning. They were only knocked off the bus each +time the control node re-created it, reconnected, re-heard the same declaration and applied it again +as a no-op. + +## Why: the roster alternates between two values, and it is part of every container + +Two consecutive declarations were compared by reading the `--add-host` entries of four containers +the moment each pass created them: + +``` +23:02 ace.internal drive.novox.be g14.internal keycloak.novox.be novox.internal + office.novox.be portainer.novox.be shanks.internal umami.novox.be (9 names) +23:07 … the same nine, and searxng.zurag.be (10 names) +``` + +One routed name — belonging to a module on another machine entirely — leaves the roster and comes +back. Because the roster is part of every container's spec digest ([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)), +each flip is a different identity for every container on the machine, and a running container cannot +have its hosts changed. So every flip replaces all of them. + +**What makes it flip is a swallowed error.** `routeNamesInTheMesh` composes every node's plan to +find the names it serves, and when one will not compose it moves on: + +``` +plan, settings, err := planFor(ctx, open, n.Name) +if err != nil { + continue +} +``` + +A node whose plan cannot be composed *right now* therefore contributes no names — not "the mesh does +not know", but "the mesh states these names do not exist", to every machine at once. + +## Why it sustains itself + +The loop closes through the control plane's own database: + +1. An apply replaces `mesh-store` — the store the control plane reads — by removing the container, + so postgres comes back through crash recovery. +2. While it recovers it refuses connections: `FATAL: the database system is not yet accepting + connections / Consistent recovery state has not been yet reached` (observed, 21:08:39 UTC, every + pass). +3. `planFor` for the other machine fails against that store. `routeNamesInTheMesh` swallows it and + drops its routed name. +4. The roster changed, so all 327 resources differ, so all are replaced — including `mesh-store`, + and including the bus, which is why the host also cannot report: `applied, and could not tell the + mesh: reporting: nats: connection closed`. +5. Back to 1. + +Every pass destroys the evidence the next pass needs to decide it has nothing to do, and nothing +outside the machine has to be wrong for it to continue. + +**It is metastable, not permanent.** It ran from 22:46 to 23:11 — five full replacements of every +container on the machine — and then stopped on its own, when one pass happened to read the store +during a window it was up, composed the same roster twice running, and found nothing to do. Load fell +from 7.8 to 1.7 and the machine returned to its five-minute idle tick. + +That it ends by luck is the point, not a mitigation. The exit condition is a race the mesh does not +control, the operator cannot see, and nothing reports; the same four actions on a slower machine, or +a larger store, would not have found it. An outage that clears itself after twenty-five minutes and +five restarts of the forge, the directory and mail is not a smaller fault than one that does not — it +is the same fault, harder to catch. + +## What it is not + +- Not the operator's four actions on the other machine. Those explain the first passes + ([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)); they were + finished seventeen minutes and three full passes before these measurements. +- Not a file that keeps changing. `/etc/hosts`, the bus's account list and the vault's export were + hashed across passes and are byte-identical, and the directory reported `mode 755 to 700` every + pass while already being `700`. Those resources are **misreported as changed** and are worth their + own question, but they are not what moves a container's identity. +- Not the lost report alone. A report that cannot be delivered explains a re-apply; it does not + explain a re-apply that finds 327 differences. + +## Why it matters beyond this outage + +The same `continue` makes every routed name in the mesh conditional on every node's plan composing at +the moment any machine is pushed to. One unreachable or half-migrated machine is enough to withdraw +its names from everywhere — and the withdrawal is indistinguishable, on the receiving machine, from +the operator having removed them. + +The codebase already states the rule this breaks, forty lines away, about the same kind of lookup: + +> A lookup failure is an error, never "not found": collapsing the two composed a declaration without +> the trust whenever the inventory hiccuped, delivered by a push that reported success. + +## How it was fixed, and how the fix is checked + +`planFor` now marks the two failures that really are a statement about the node — its set not +composing, and a setting that reaches nothing — and the three gatherers pass over those and only +those. Every other failure is raised, naming the machine and the read. + +Five tests hold it: a set that cannot compose is marked as the node's own; a store that cannot be +read is *not*; one incoherent node still does not cost the rest their names; a roster is never +returned beside an error; and the raised failure names what could not be read. + +The two sibling gatherers were audited and fixed the same way — the grant composer, which would have +withheld a consumer's credential, and the private-network membership, which would have taken a +machine off the overlay. Three other `planFor` callers were audited and left alone: they refuse or +report rather than silently withdraw, which is the safe direction. + +**This does not close [issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md).** +A roster that changes for a real reason still replaces every container in the mesh. This removes the +false reasons; whether the roster belongs in a container's identity at all is that record's question.