Merge pull request 'Issue 146: a first node now enrols, and is enrolled twice' (#185) from issue/146-diagnosis into main
This commit was merged in pull request #185.
This commit is contained in:
+88
@@ -82,12 +82,100 @@ managing, or genesis carries a user list that includes the first node's enrolmen
|
||||
takes over from there. Both are decisions, not patches, and both belong to the genesis step that was
|
||||
deliberately left until last.
|
||||
|
||||
## 5 — the composed user list has to be placed by hand at genesis *(fixed)*
|
||||
|
||||
The account a token is the password of is **not recorded at all**: the composer names an enrolment
|
||||
user for every machine with a live token, nothing minted a credential for it, and the composition
|
||||
left it out as a user with no password. The comment above the issuing code already claimed
|
||||
otherwise — *"the account is created before the token is handed over"* — which is how it went
|
||||
unnoticed. Issuing a token now records that account, with the token's own secret as its password,
|
||||
because that is the string the machine will present.
|
||||
|
||||
Placing it is the other half. The list reaches the machine running the bus in that machine's
|
||||
declaration, which a machine that has not enrolled does not get, so at genesis it cannot arrive
|
||||
that way. **The control plane composes and says what it composed** — `broker accounts`, to standard
|
||||
output — and whoever is raising the machine writes it beside the bus's configuration and makes the
|
||||
server re-read it. Twice, because two accounts come into existence at different moments: the
|
||||
enrolment when the token is issued, and the machine's own when it enrols. A control plane that
|
||||
wrote the file itself would have to know where the bus keeps its configuration and how to make it
|
||||
reload, which is the module's knowledge and is what the module takes over on the first push.
|
||||
|
||||
With that, **a first node enrols against the bus it just raised** — measured, from bare, in the
|
||||
lab.
|
||||
|
||||
## 6 — and is enrolled twice, keeping a credential the mesh has replaced *(fixed)*
|
||||
|
||||
```
|
||||
mesh-controller: enrolled anchor
|
||||
mesh-controller: enrolled anchor (the same second)
|
||||
```
|
||||
|
||||
One `enrol` on the machine, two enrolments in the control plane. Each mints the node a fresh bus
|
||||
password and returns it; the machine keeps the answer to the first, and the mesh keeps the hash of
|
||||
the second. The machine then reconnects for ever as a user whose password the mesh rotated out from
|
||||
under it — *authentication error - User "node.anchor"* on the bus, `Authorization Violation` in the
|
||||
host's log, and a node that never reports.
|
||||
|
||||
What is ruled out: the host asking twice — it asks again only when the mesh says *try again*, and
|
||||
a refused attempt is not logged as an enrolment. Redelivery by the consumer — there is one
|
||||
consumer, its acknowledgement window is thirty seconds, and the handler is quick.
|
||||
|
||||
What is left: the client re-publishing when an acknowledgement is slow, which is what its defaults
|
||||
do. That was addressed by giving the publish a message id derived from its own bytes, so the stream
|
||||
discards the copy — **and the duplicate survived it**, so either the id is not reaching the stream
|
||||
or the second copy is not a copy. This is where the trail stops.
|
||||
|
||||
Worth saying plainly: **the mint is the fragile part, not the delivery.** An enrolment answered
|
||||
twice is survivable if the answer is the same both times, and it cannot be — the mesh keeps only
|
||||
the hash, so a second answer is necessarily a different credential.
|
||||
|
||||
**Found, and it is not about enrolment at all** *(fixed)*. The bus's own counters settled it: one
|
||||
message published, one held in the stream, one delivery, nothing redelivered — and the controller
|
||||
enrolled the machine twice. So the handler ran twice on one delivery.
|
||||
|
||||
A push consumer delivers onto an ordinary subject, and **everything subscribed to that subject gets
|
||||
a copy**. The controller holds a consumer called `controller` on CONTROL and another called
|
||||
`controller` on EVENTS, and the delivery subject was derived from the consumer's name alone — so
|
||||
both were `_DELIVER.controller`, the one process held both subscriptions, and every message from
|
||||
either stream was acted on twice.
|
||||
|
||||
Enrolment is where it drew blood, because enrolling twice mints twice and the second credential
|
||||
replaces the first. But it applied to **every report and every event the controller follows**, and
|
||||
it is the kind of fault that leaves no trace: nothing is redelivered, no counter is wrong, the work
|
||||
simply happens twice. The comment in the receiving code about a merge that ran the whole catalogue
|
||||
five times over on 2026-09-28 is the same shape seen from the other end.
|
||||
|
||||
The stream is in the delivery subject now, because the pair is what identifies a consumer — the
|
||||
server scopes a durable's name to its stream, and this subject was the one place that scoping was
|
||||
dropped. A subscriber's permission gains the same shape, keeping the bare name so an existing
|
||||
consumer keeps working until the controller's next assertion moves it.
|
||||
|
||||
*How it is checked:* the consumers the mesh asks for are asserted to deliver onto distinct subjects,
|
||||
in the controller's own suite. Against a server it would be invisible, which is the point.
|
||||
|
||||
## Where it belongs
|
||||
|
||||
`mesh-host` (the bundle and the enrolment path) and `mesh-controller` (the certificate command, and
|
||||
the composition that cannot reach the bus at genesis). Three of the four are fixed on branches; the
|
||||
fourth is the genesis work.
|
||||
|
||||
## What made it slow, and what was changed so it is not
|
||||
|
||||
Six faults behind one another, each found by raising a machine and reading what it said. What cost
|
||||
the most was not the faults:
|
||||
|
||||
- **Every bed's own instructions named the bundle that cannot work**, so the first three attempts
|
||||
ended in a control plane crash-looping on a missing bus. They name the working one now.
|
||||
- **A host binary built without its system** refuses everything it is given with *this host was
|
||||
built for ""*, which reads like a broken bundle. The lab's README says so.
|
||||
- **`make image` in the control plane had been broken for as long as its base was pinned**: the
|
||||
Dockerfile's fallback is a Go older than the module asks for, and the pipeline never saw it
|
||||
because the pipeline passes the declared base in. It reads the base from the manifest now.
|
||||
- **Leaving the machine standing is what answers the question.** Every finding above came from
|
||||
shelling in afterwards — the host's log, the bus's log, the file the bus was actually handed —
|
||||
and none from the test's own output, which says only that nothing converged. The bed takes
|
||||
`MESH_LAB_KEEP`, and the README says to reach for it first.
|
||||
|
||||
## What it cost, for the next person
|
||||
|
||||
Every lab bed still names `foundation-first-node.lock` in its own instructions, and that bundle
|
||||
|
||||
@@ -0,0 +1,75 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-29
|
||||
located-in: [nothing in the mesh — the surface in use is the predecessor's, installed on the workstation]
|
||||
---
|
||||
|
||||
# 147 — the operator's tools still dial the bus that was removed
|
||||
|
||||
## What was observed
|
||||
|
||||
Every tool call an operator makes against the mesh fails, on every machine, with the same answer:
|
||||
|
||||
```
|
||||
AMQP not connected — cannot reach hal/mesh@novox
|
||||
AMQP not connected — cannot reach hal/mesh@shanks
|
||||
```
|
||||
|
||||
The mesh moved to one bus and the previous transport was deleted
|
||||
([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md), cut over
|
||||
2026-09-28). The surface an operator drives the mesh through — the tool bridge their assistant
|
||||
speaks to — still opens an AMQP connection, so it cannot reach anything. Not one node answers,
|
||||
including the machine the operator is sitting at.
|
||||
|
||||
**What that leaves.** The mesh itself is healthy: nodes are current, modules run, the bus carries
|
||||
the mesh's own traffic. What is gone is the way a person asks it anything. Every question — what a
|
||||
node runs, what is assigned, what a module's settings are — has to be asked by opening a shell on
|
||||
the machine and running the control plane's binary inside its container, which is precisely the
|
||||
path the tool surface exists to remove, and which nothing checks, records or permits.
|
||||
|
||||
It also silently changes how work gets done: an assistant told to use the mesh's tools finds them
|
||||
dead, falls back to `ssh` and `docker exec`, and the operator discovers the fallback rather than
|
||||
the fault.
|
||||
|
||||
## Why this is here and not a note in the knowledge base
|
||||
|
||||
The rule is that everything happens over the mesh's bus. A surface that cannot reach that bus is
|
||||
that rule enforced by nothing — and the mesh reports itself healthy throughout, because what broke
|
||||
is not a module, a node or a provision but the thing standing outside asking them questions.
|
||||
|
||||
The same cut-over removed the assistant's long-term memory (`recall`, `memorize`) for the same
|
||||
reason, which is why lessons from the last two days were written into this repository by hand.
|
||||
|
||||
## What would have prevented it
|
||||
|
||||
- **The tool surface as a consumer of the bus like any other**, so moving the bus moves it — rather
|
||||
than a separate bridge with its own connection settings that nothing resolves.
|
||||
- **A check that a tool call reaches a node**, run where the mesh's other checks run. Every check
|
||||
the mesh makes today is about the relationship between the mesh and a machine; none asks whether
|
||||
a person can ask it anything ([issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
|
||||
the same shape one level out).
|
||||
|
||||
## Diagnosed at once, because the answer was in the configuration
|
||||
|
||||
**The surface is not the mesh's.** The tool server the operator's assistant speaks to is the
|
||||
predecessor's brain, installed on the workstation and started as a local process, with the
|
||||
predecessor's broker URL — `amqp://…@<the control node>` — written into the assistant's own
|
||||
configuration. Nothing about it is a module: it has no manifest, no assignment, no seat, no account
|
||||
on the mesh's bus, and the mesh has never known it exists.
|
||||
|
||||
So nothing regressed. The mesh removed a transport that this program still dials, and the program
|
||||
was never part of the mesh to be moved. **The mesh has a tool model** — `tools` in a manifest, the
|
||||
calls a seat accepts, the subjects a module answers on, and a control-plane command that asks one —
|
||||
and **nothing publishes an operator-facing surface onto it.** What an operator uses is the thing
|
||||
that came before, kept alive by a URL in a file.
|
||||
|
||||
That is the issue, and it is larger than a broken connection: the way a person drives this mesh is
|
||||
outside the mesh.
|
||||
|
||||
## Evidence to carry into diagnosis
|
||||
|
||||
- The failure text names `hal/mesh@<node>`, so the bridge is resolving a node and then dialling the
|
||||
old transport.
|
||||
- It fails identically for the local machine, which rules out reachability and points at the
|
||||
transport alone.
|
||||
- The mesh's own traffic over the new bus is unaffected: nodes report, declarations apply.
|
||||
Reference in New Issue
Block a user