diff --git a/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md b/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md index bcbd17b..317bb01 100644 --- a/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md +++ b/04-ISSUES/146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/01-diagnosis.md @@ -82,12 +82,100 @@ managing, or genesis carries a user list that includes the first node's enrolmen takes over from there. Both are decisions, not patches, and both belong to the genesis step that was deliberately left until last. +## 5 — the composed user list has to be placed by hand at genesis *(fixed)* + +The account a token is the password of is **not recorded at all**: the composer names an enrolment +user for every machine with a live token, nothing minted a credential for it, and the composition +left it out as a user with no password. The comment above the issuing code already claimed +otherwise — *"the account is created before the token is handed over"* — which is how it went +unnoticed. Issuing a token now records that account, with the token's own secret as its password, +because that is the string the machine will present. + +Placing it is the other half. The list reaches the machine running the bus in that machine's +declaration, which a machine that has not enrolled does not get, so at genesis it cannot arrive +that way. **The control plane composes and says what it composed** — `broker accounts`, to standard +output — and whoever is raising the machine writes it beside the bus's configuration and makes the +server re-read it. Twice, because two accounts come into existence at different moments: the +enrolment when the token is issued, and the machine's own when it enrols. A control plane that +wrote the file itself would have to know where the bus keeps its configuration and how to make it +reload, which is the module's knowledge and is what the module takes over on the first push. + +With that, **a first node enrols against the bus it just raised** — measured, from bare, in the +lab. + +## 6 — and is enrolled twice, keeping a credential the mesh has replaced *(fixed)* + +``` +mesh-controller: enrolled anchor +mesh-controller: enrolled anchor (the same second) +``` + +One `enrol` on the machine, two enrolments in the control plane. Each mints the node a fresh bus +password and returns it; the machine keeps the answer to the first, and the mesh keeps the hash of +the second. The machine then reconnects for ever as a user whose password the mesh rotated out from +under it — *authentication error - User "node.anchor"* on the bus, `Authorization Violation` in the +host's log, and a node that never reports. + +What is ruled out: the host asking twice — it asks again only when the mesh says *try again*, and +a refused attempt is not logged as an enrolment. Redelivery by the consumer — there is one +consumer, its acknowledgement window is thirty seconds, and the handler is quick. + +What is left: the client re-publishing when an acknowledgement is slow, which is what its defaults +do. That was addressed by giving the publish a message id derived from its own bytes, so the stream +discards the copy — **and the duplicate survived it**, so either the id is not reaching the stream +or the second copy is not a copy. This is where the trail stops. + +Worth saying plainly: **the mint is the fragile part, not the delivery.** An enrolment answered +twice is survivable if the answer is the same both times, and it cannot be — the mesh keeps only +the hash, so a second answer is necessarily a different credential. + +**Found, and it is not about enrolment at all** *(fixed)*. The bus's own counters settled it: one +message published, one held in the stream, one delivery, nothing redelivered — and the controller +enrolled the machine twice. So the handler ran twice on one delivery. + +A push consumer delivers onto an ordinary subject, and **everything subscribed to that subject gets +a copy**. The controller holds a consumer called `controller` on CONTROL and another called +`controller` on EVENTS, and the delivery subject was derived from the consumer's name alone — so +both were `_DELIVER.controller`, the one process held both subscriptions, and every message from +either stream was acted on twice. + +Enrolment is where it drew blood, because enrolling twice mints twice and the second credential +replaces the first. But it applied to **every report and every event the controller follows**, and +it is the kind of fault that leaves no trace: nothing is redelivered, no counter is wrong, the work +simply happens twice. The comment in the receiving code about a merge that ran the whole catalogue +five times over on 2026-09-28 is the same shape seen from the other end. + +The stream is in the delivery subject now, because the pair is what identifies a consumer — the +server scopes a durable's name to its stream, and this subject was the one place that scoping was +dropped. A subscriber's permission gains the same shape, keeping the bare name so an existing +consumer keeps working until the controller's next assertion moves it. + +*How it is checked:* the consumers the mesh asks for are asserted to deliver onto distinct subjects, +in the controller's own suite. Against a server it would be invisible, which is the point. + ## Where it belongs `mesh-host` (the bundle and the enrolment path) and `mesh-controller` (the certificate command, and the composition that cannot reach the bus at genesis). Three of the four are fixed on branches; the fourth is the genesis work. +## What made it slow, and what was changed so it is not + +Six faults behind one another, each found by raising a machine and reading what it said. What cost +the most was not the faults: + +- **Every bed's own instructions named the bundle that cannot work**, so the first three attempts + ended in a control plane crash-looping on a missing bus. They name the working one now. +- **A host binary built without its system** refuses everything it is given with *this host was + built for ""*, which reads like a broken bundle. The lab's README says so. +- **`make image` in the control plane had been broken for as long as its base was pinned**: the + Dockerfile's fallback is a Go older than the module asks for, and the pipeline never saw it + because the pipeline passes the declared base in. It reads the base from the manifest now. +- **Leaving the machine standing is what answers the question.** Every finding above came from + shelling in afterwards — the host's log, the bus's log, the file the bus was actually handed — + and none from the test's own output, which says only that nothing converged. The bed takes + `MESH_LAB_KEEP`, and the README says to reach for it first. + ## What it cost, for the next person Every lab bed still names `foundation-first-node.lock` in its own instructions, and that bundle diff --git a/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md b/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md new file mode 100644 index 0000000..b056a21 --- /dev/null +++ b/04-ISSUES/147-the-operators-tools-still-dial-the-bus-that-was-removed/00-report.md @@ -0,0 +1,75 @@ +--- +status: located +opened: 2026-09-29 +located-in: [nothing in the mesh — the surface in use is the predecessor's, installed on the workstation] +--- + +# 147 — the operator's tools still dial the bus that was removed + +## What was observed + +Every tool call an operator makes against the mesh fails, on every machine, with the same answer: + +``` +AMQP not connected — cannot reach hal/mesh@novox +AMQP not connected — cannot reach hal/mesh@shanks +``` + +The mesh moved to one bus and the previous transport was deleted +([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md), cut over +2026-09-28). The surface an operator drives the mesh through — the tool bridge their assistant +speaks to — still opens an AMQP connection, so it cannot reach anything. Not one node answers, +including the machine the operator is sitting at. + +**What that leaves.** The mesh itself is healthy: nodes are current, modules run, the bus carries +the mesh's own traffic. What is gone is the way a person asks it anything. Every question — what a +node runs, what is assigned, what a module's settings are — has to be asked by opening a shell on +the machine and running the control plane's binary inside its container, which is precisely the +path the tool surface exists to remove, and which nothing checks, records or permits. + +It also silently changes how work gets done: an assistant told to use the mesh's tools finds them +dead, falls back to `ssh` and `docker exec`, and the operator discovers the fallback rather than +the fault. + +## Why this is here and not a note in the knowledge base + +The rule is that everything happens over the mesh's bus. A surface that cannot reach that bus is +that rule enforced by nothing — and the mesh reports itself healthy throughout, because what broke +is not a module, a node or a provision but the thing standing outside asking them questions. + +The same cut-over removed the assistant's long-term memory (`recall`, `memorize`) for the same +reason, which is why lessons from the last two days were written into this repository by hand. + +## What would have prevented it + +- **The tool surface as a consumer of the bus like any other**, so moving the bus moves it — rather + than a separate bridge with its own connection settings that nothing resolves. +- **A check that a tool call reaches a node**, run where the mesh's other checks run. Every check + the mesh makes today is about the relationship between the mesh and a machine; none asks whether + a person can ask it anything ([issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md), + the same shape one level out). + +## Diagnosed at once, because the answer was in the configuration + +**The surface is not the mesh's.** The tool server the operator's assistant speaks to is the +predecessor's brain, installed on the workstation and started as a local process, with the +predecessor's broker URL — `amqp://…@` — written into the assistant's own +configuration. Nothing about it is a module: it has no manifest, no assignment, no seat, no account +on the mesh's bus, and the mesh has never known it exists. + +So nothing regressed. The mesh removed a transport that this program still dials, and the program +was never part of the mesh to be moved. **The mesh has a tool model** — `tools` in a manifest, the +calls a seat accepts, the subjects a module answers on, and a control-plane command that asks one — +and **nothing publishes an operator-facing surface onto it.** What an operator uses is the thing +that came before, kept alive by a URL in a file. + +That is the issue, and it is larger than a broken connection: the way a person drives this mesh is +outside the mesh. + +## Evidence to carry into diagnosis + +- The failure text names `hal/mesh@`, so the bridge is resolving a node and then dialling the + old transport. +- It fails identically for the local machine, which rules out reachability and points at the + transport alone. +- The mesh's own traffic over the new bus is unaffected: nodes report, declarations apply.