Compare commits

..
5 Commits
Author SHA1 Message Date
jschoubben 3c0f7082e6 Issue 147: the tool surface is not the mesh's, it is the predecessor's
Diagnosed from the configuration: the tool server is HAL's brain, a local
process on the workstation with the predecessor's broker URL in the assistant's
own config. No manifest, no assignment, no seat, no account. Nothing regressed
— the mesh removed a transport this program still dials, and the program was
never part of the mesh. The mesh has a tool model and nothing publishes an
operator-facing surface onto it.
2026-09-29 21:50:50 +02:00
jschoubben 0dd00e88b6 Issue 147: the operator's tools still dial the bus that was removed
Every tool call fails with 'AMQP not connected', on every node including the
local one, because the tool surface still opens an AMQP connection and that
transport was deleted at the cut-over. The mesh reports healthy throughout —
what broke is the thing standing outside asking it questions, so nothing the
mesh checks is about it.
2026-09-29 21:39:55 +02:00
jschoubben eef54917ec Issue 146: the double enrolment was two consumers sharing a delivery subject
Not about enrolment. A push consumer delivers onto an ordinary subject and
everything subscribed to it gets a copy; the controller's two consumers were
both named after it, so both were given the same subject and the one process
acted on every message twice. Enrolment is where it drew blood because a second
enrolment mints a second credential.
2026-09-29 21:32:52 +02:00
jschoubben e9b1010bc0 Issue 146: what made it slow, and what was changed so it is not 2026-09-29 17:45:25 +02:00
jschoubben f6ed3545b7 Issue 146: a first node now enrols, and is enrolled twice
Two more faults behind the three already fixed. The account a token is the
password of was never recorded, and the comment above the issuing code said
it was; issuing now records it. Placing the composed list at genesis is the
other half — the control plane says what it composed and whoever raises the
machine writes it beside the bus, because no declaration can reach a machine
that has not enrolled.

With that a first node enrols. It is then enrolled twice from one attempt,
each minting a credential, and it keeps the answer to the first while the mesh
keeps the second. The trail for that one stops at a duplicate that survived
message-id deduplication.
2026-09-29 17:36:40 +02:00
2 changed files with 163 additions and 0 deletions
@@ -82,12 +82,100 @@ managing, or genesis carries a user list that includes the first node's enrolmen
takes over from there. Both are decisions, not patches, and both belong to the genesis step that was takes over from there. Both are decisions, not patches, and both belong to the genesis step that was
deliberately left until last. deliberately left until last.
## 5 — the composed user list has to be placed by hand at genesis *(fixed)*
The account a token is the password of is **not recorded at all**: the composer names an enrolment
user for every machine with a live token, nothing minted a credential for it, and the composition
left it out as a user with no password. The comment above the issuing code already claimed
otherwise — *"the account is created before the token is handed over"* — which is how it went
unnoticed. Issuing a token now records that account, with the token's own secret as its password,
because that is the string the machine will present.
Placing it is the other half. The list reaches the machine running the bus in that machine's
declaration, which a machine that has not enrolled does not get, so at genesis it cannot arrive
that way. **The control plane composes and says what it composed** — `broker accounts`, to standard
output — and whoever is raising the machine writes it beside the bus's configuration and makes the
server re-read it. Twice, because two accounts come into existence at different moments: the
enrolment when the token is issued, and the machine's own when it enrols. A control plane that
wrote the file itself would have to know where the bus keeps its configuration and how to make it
reload, which is the module's knowledge and is what the module takes over on the first push.
With that, **a first node enrols against the bus it just raised** — measured, from bare, in the
lab.
## 6 — and is enrolled twice, keeping a credential the mesh has replaced *(fixed)*
```
mesh-controller: enrolled anchor
mesh-controller: enrolled anchor (the same second)
```
One `enrol` on the machine, two enrolments in the control plane. Each mints the node a fresh bus
password and returns it; the machine keeps the answer to the first, and the mesh keeps the hash of
the second. The machine then reconnects for ever as a user whose password the mesh rotated out from
under it — *authentication error - User "node.anchor"* on the bus, `Authorization Violation` in the
host's log, and a node that never reports.
What is ruled out: the host asking twice — it asks again only when the mesh says *try again*, and
a refused attempt is not logged as an enrolment. Redelivery by the consumer — there is one
consumer, its acknowledgement window is thirty seconds, and the handler is quick.
What is left: the client re-publishing when an acknowledgement is slow, which is what its defaults
do. That was addressed by giving the publish a message id derived from its own bytes, so the stream
discards the copy — **and the duplicate survived it**, so either the id is not reaching the stream
or the second copy is not a copy. This is where the trail stops.
Worth saying plainly: **the mint is the fragile part, not the delivery.** An enrolment answered
twice is survivable if the answer is the same both times, and it cannot be — the mesh keeps only
the hash, so a second answer is necessarily a different credential.
**Found, and it is not about enrolment at all** *(fixed)*. The bus's own counters settled it: one
message published, one held in the stream, one delivery, nothing redelivered — and the controller
enrolled the machine twice. So the handler ran twice on one delivery.
A push consumer delivers onto an ordinary subject, and **everything subscribed to that subject gets
a copy**. The controller holds a consumer called `controller` on CONTROL and another called
`controller` on EVENTS, and the delivery subject was derived from the consumer's name alone — so
both were `_DELIVER.controller`, the one process held both subscriptions, and every message from
either stream was acted on twice.
Enrolment is where it drew blood, because enrolling twice mints twice and the second credential
replaces the first. But it applied to **every report and every event the controller follows**, and
it is the kind of fault that leaves no trace: nothing is redelivered, no counter is wrong, the work
simply happens twice. The comment in the receiving code about a merge that ran the whole catalogue
five times over on 2026-09-28 is the same shape seen from the other end.
The stream is in the delivery subject now, because the pair is what identifies a consumer — the
server scopes a durable's name to its stream, and this subject was the one place that scoping was
dropped. A subscriber's permission gains the same shape, keeping the bare name so an existing
consumer keeps working until the controller's next assertion moves it.
*How it is checked:* the consumers the mesh asks for are asserted to deliver onto distinct subjects,
in the controller's own suite. Against a server it would be invisible, which is the point.
## Where it belongs ## Where it belongs
`mesh-host` (the bundle and the enrolment path) and `mesh-controller` (the certificate command, and `mesh-host` (the bundle and the enrolment path) and `mesh-controller` (the certificate command, and
the composition that cannot reach the bus at genesis). Three of the four are fixed on branches; the the composition that cannot reach the bus at genesis). Three of the four are fixed on branches; the
fourth is the genesis work. fourth is the genesis work.
## What made it slow, and what was changed so it is not
Six faults behind one another, each found by raising a machine and reading what it said. What cost
the most was not the faults:
- **Every bed's own instructions named the bundle that cannot work**, so the first three attempts
ended in a control plane crash-looping on a missing bus. They name the working one now.
- **A host binary built without its system** refuses everything it is given with *this host was
built for ""*, which reads like a broken bundle. The lab's README says so.
- **`make image` in the control plane had been broken for as long as its base was pinned**: the
Dockerfile's fallback is a Go older than the module asks for, and the pipeline never saw it
because the pipeline passes the declared base in. It reads the base from the manifest now.
- **Leaving the machine standing is what answers the question.** Every finding above came from
shelling in afterwards — the host's log, the bus's log, the file the bus was actually handed —
and none from the test's own output, which says only that nothing converged. The bed takes
`MESH_LAB_KEEP`, and the README says to reach for it first.
## What it cost, for the next person ## What it cost, for the next person
Every lab bed still names `foundation-first-node.lock` in its own instructions, and that bundle Every lab bed still names `foundation-first-node.lock` in its own instructions, and that bundle
@@ -0,0 +1,75 @@
---
status: located
opened: 2026-09-29
located-in: [nothing in the mesh — the surface in use is the predecessor's, installed on the workstation]
---
# 147 — the operator's tools still dial the bus that was removed
## What was observed
Every tool call an operator makes against the mesh fails, on every machine, with the same answer:
```
AMQP not connected — cannot reach hal/mesh@novox
AMQP not connected — cannot reach hal/mesh@shanks
```
The mesh moved to one bus and the previous transport was deleted
([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md), cut over
2026-09-28). The surface an operator drives the mesh through — the tool bridge their assistant
speaks to — still opens an AMQP connection, so it cannot reach anything. Not one node answers,
including the machine the operator is sitting at.
**What that leaves.** The mesh itself is healthy: nodes are current, modules run, the bus carries
the mesh's own traffic. What is gone is the way a person asks it anything. Every question — what a
node runs, what is assigned, what a module's settings are — has to be asked by opening a shell on
the machine and running the control plane's binary inside its container, which is precisely the
path the tool surface exists to remove, and which nothing checks, records or permits.
It also silently changes how work gets done: an assistant told to use the mesh's tools finds them
dead, falls back to `ssh` and `docker exec`, and the operator discovers the fallback rather than
the fault.
## Why this is here and not a note in the knowledge base
The rule is that everything happens over the mesh's bus. A surface that cannot reach that bus is
that rule enforced by nothing — and the mesh reports itself healthy throughout, because what broke
is not a module, a node or a provision but the thing standing outside asking them questions.
The same cut-over removed the assistant's long-term memory (`recall`, `memorize`) for the same
reason, which is why lessons from the last two days were written into this repository by hand.
## What would have prevented it
- **The tool surface as a consumer of the bus like any other**, so moving the bus moves it — rather
than a separate bridge with its own connection settings that nothing resolves.
- **A check that a tool call reaches a node**, run where the mesh's other checks run. Every check
the mesh makes today is about the relationship between the mesh and a machine; none asks whether
a person can ask it anything ([issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
the same shape one level out).
## Diagnosed at once, because the answer was in the configuration
**The surface is not the mesh's.** The tool server the operator's assistant speaks to is the
predecessor's brain, installed on the workstation and started as a local process, with the
predecessor's broker URL — `amqp://…@<the control node>` — written into the assistant's own
configuration. Nothing about it is a module: it has no manifest, no assignment, no seat, no account
on the mesh's bus, and the mesh has never known it exists.
So nothing regressed. The mesh removed a transport that this program still dials, and the program
was never part of the mesh to be moved. **The mesh has a tool model** — `tools` in a manifest, the
calls a seat accepts, the subjects a module answers on, and a control-plane command that asks one —
and **nothing publishes an operator-facing surface onto it.** What an operator uses is the thing
that came before, kept alive by a URL in a file.
That is the issue, and it is larger than a broken connection: the way a person drives this mesh is
outside the mesh.
## Evidence to carry into diagnosis
- The failure text names `hal/mesh@<node>`, so the bridge is resolving a node and then dialling the
old transport.
- It fails identically for the local machine, which rules out reachability and points at the
transport alone.
- The mesh's own traffic over the new bus is unaffected: nodes report, declarations apply.