Compare commits
5
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
3c0f7082e6 | ||
|
|
0dd00e88b6 | ||
|
|
eef54917ec | ||
|
|
e9b1010bc0 | ||
|
|
f6ed3545b7 |
+88
@@ -82,12 +82,100 @@ managing, or genesis carries a user list that includes the first node's enrolmen
|
|||||||
takes over from there. Both are decisions, not patches, and both belong to the genesis step that was
|
takes over from there. Both are decisions, not patches, and both belong to the genesis step that was
|
||||||
deliberately left until last.
|
deliberately left until last.
|
||||||
|
|
||||||
|
## 5 — the composed user list has to be placed by hand at genesis *(fixed)*
|
||||||
|
|
||||||
|
The account a token is the password of is **not recorded at all**: the composer names an enrolment
|
||||||
|
user for every machine with a live token, nothing minted a credential for it, and the composition
|
||||||
|
left it out as a user with no password. The comment above the issuing code already claimed
|
||||||
|
otherwise — *"the account is created before the token is handed over"* — which is how it went
|
||||||
|
unnoticed. Issuing a token now records that account, with the token's own secret as its password,
|
||||||
|
because that is the string the machine will present.
|
||||||
|
|
||||||
|
Placing it is the other half. The list reaches the machine running the bus in that machine's
|
||||||
|
declaration, which a machine that has not enrolled does not get, so at genesis it cannot arrive
|
||||||
|
that way. **The control plane composes and says what it composed** — `broker accounts`, to standard
|
||||||
|
output — and whoever is raising the machine writes it beside the bus's configuration and makes the
|
||||||
|
server re-read it. Twice, because two accounts come into existence at different moments: the
|
||||||
|
enrolment when the token is issued, and the machine's own when it enrols. A control plane that
|
||||||
|
wrote the file itself would have to know where the bus keeps its configuration and how to make it
|
||||||
|
reload, which is the module's knowledge and is what the module takes over on the first push.
|
||||||
|
|
||||||
|
With that, **a first node enrols against the bus it just raised** — measured, from bare, in the
|
||||||
|
lab.
|
||||||
|
|
||||||
|
## 6 — and is enrolled twice, keeping a credential the mesh has replaced *(fixed)*
|
||||||
|
|
||||||
|
```
|
||||||
|
mesh-controller: enrolled anchor
|
||||||
|
mesh-controller: enrolled anchor (the same second)
|
||||||
|
```
|
||||||
|
|
||||||
|
One `enrol` on the machine, two enrolments in the control plane. Each mints the node a fresh bus
|
||||||
|
password and returns it; the machine keeps the answer to the first, and the mesh keeps the hash of
|
||||||
|
the second. The machine then reconnects for ever as a user whose password the mesh rotated out from
|
||||||
|
under it — *authentication error - User "node.anchor"* on the bus, `Authorization Violation` in the
|
||||||
|
host's log, and a node that never reports.
|
||||||
|
|
||||||
|
What is ruled out: the host asking twice — it asks again only when the mesh says *try again*, and
|
||||||
|
a refused attempt is not logged as an enrolment. Redelivery by the consumer — there is one
|
||||||
|
consumer, its acknowledgement window is thirty seconds, and the handler is quick.
|
||||||
|
|
||||||
|
What is left: the client re-publishing when an acknowledgement is slow, which is what its defaults
|
||||||
|
do. That was addressed by giving the publish a message id derived from its own bytes, so the stream
|
||||||
|
discards the copy — **and the duplicate survived it**, so either the id is not reaching the stream
|
||||||
|
or the second copy is not a copy. This is where the trail stops.
|
||||||
|
|
||||||
|
Worth saying plainly: **the mint is the fragile part, not the delivery.** An enrolment answered
|
||||||
|
twice is survivable if the answer is the same both times, and it cannot be — the mesh keeps only
|
||||||
|
the hash, so a second answer is necessarily a different credential.
|
||||||
|
|
||||||
|
**Found, and it is not about enrolment at all** *(fixed)*. The bus's own counters settled it: one
|
||||||
|
message published, one held in the stream, one delivery, nothing redelivered — and the controller
|
||||||
|
enrolled the machine twice. So the handler ran twice on one delivery.
|
||||||
|
|
||||||
|
A push consumer delivers onto an ordinary subject, and **everything subscribed to that subject gets
|
||||||
|
a copy**. The controller holds a consumer called `controller` on CONTROL and another called
|
||||||
|
`controller` on EVENTS, and the delivery subject was derived from the consumer's name alone — so
|
||||||
|
both were `_DELIVER.controller`, the one process held both subscriptions, and every message from
|
||||||
|
either stream was acted on twice.
|
||||||
|
|
||||||
|
Enrolment is where it drew blood, because enrolling twice mints twice and the second credential
|
||||||
|
replaces the first. But it applied to **every report and every event the controller follows**, and
|
||||||
|
it is the kind of fault that leaves no trace: nothing is redelivered, no counter is wrong, the work
|
||||||
|
simply happens twice. The comment in the receiving code about a merge that ran the whole catalogue
|
||||||
|
five times over on 2026-09-28 is the same shape seen from the other end.
|
||||||
|
|
||||||
|
The stream is in the delivery subject now, because the pair is what identifies a consumer — the
|
||||||
|
server scopes a durable's name to its stream, and this subject was the one place that scoping was
|
||||||
|
dropped. A subscriber's permission gains the same shape, keeping the bare name so an existing
|
||||||
|
consumer keeps working until the controller's next assertion moves it.
|
||||||
|
|
||||||
|
*How it is checked:* the consumers the mesh asks for are asserted to deliver onto distinct subjects,
|
||||||
|
in the controller's own suite. Against a server it would be invisible, which is the point.
|
||||||
|
|
||||||
## Where it belongs
|
## Where it belongs
|
||||||
|
|
||||||
`mesh-host` (the bundle and the enrolment path) and `mesh-controller` (the certificate command, and
|
`mesh-host` (the bundle and the enrolment path) and `mesh-controller` (the certificate command, and
|
||||||
the composition that cannot reach the bus at genesis). Three of the four are fixed on branches; the
|
the composition that cannot reach the bus at genesis). Three of the four are fixed on branches; the
|
||||||
fourth is the genesis work.
|
fourth is the genesis work.
|
||||||
|
|
||||||
|
## What made it slow, and what was changed so it is not
|
||||||
|
|
||||||
|
Six faults behind one another, each found by raising a machine and reading what it said. What cost
|
||||||
|
the most was not the faults:
|
||||||
|
|
||||||
|
- **Every bed's own instructions named the bundle that cannot work**, so the first three attempts
|
||||||
|
ended in a control plane crash-looping on a missing bus. They name the working one now.
|
||||||
|
- **A host binary built without its system** refuses everything it is given with *this host was
|
||||||
|
built for ""*, which reads like a broken bundle. The lab's README says so.
|
||||||
|
- **`make image` in the control plane had been broken for as long as its base was pinned**: the
|
||||||
|
Dockerfile's fallback is a Go older than the module asks for, and the pipeline never saw it
|
||||||
|
because the pipeline passes the declared base in. It reads the base from the manifest now.
|
||||||
|
- **Leaving the machine standing is what answers the question.** Every finding above came from
|
||||||
|
shelling in afterwards — the host's log, the bus's log, the file the bus was actually handed —
|
||||||
|
and none from the test's own output, which says only that nothing converged. The bed takes
|
||||||
|
`MESH_LAB_KEEP`, and the README says to reach for it first.
|
||||||
|
|
||||||
## What it cost, for the next person
|
## What it cost, for the next person
|
||||||
|
|
||||||
Every lab bed still names `foundation-first-node.lock` in its own instructions, and that bundle
|
Every lab bed still names `foundation-first-node.lock` in its own instructions, and that bundle
|
||||||
|
|||||||
@@ -0,0 +1,75 @@
|
|||||||
|
---
|
||||||
|
status: located
|
||||||
|
opened: 2026-09-29
|
||||||
|
located-in: [nothing in the mesh — the surface in use is the predecessor's, installed on the workstation]
|
||||||
|
---
|
||||||
|
|
||||||
|
# 147 — the operator's tools still dial the bus that was removed
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
Every tool call an operator makes against the mesh fails, on every machine, with the same answer:
|
||||||
|
|
||||||
|
```
|
||||||
|
AMQP not connected — cannot reach hal/mesh@novox
|
||||||
|
AMQP not connected — cannot reach hal/mesh@shanks
|
||||||
|
```
|
||||||
|
|
||||||
|
The mesh moved to one bus and the previous transport was deleted
|
||||||
|
([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md), cut over
|
||||||
|
2026-09-28). The surface an operator drives the mesh through — the tool bridge their assistant
|
||||||
|
speaks to — still opens an AMQP connection, so it cannot reach anything. Not one node answers,
|
||||||
|
including the machine the operator is sitting at.
|
||||||
|
|
||||||
|
**What that leaves.** The mesh itself is healthy: nodes are current, modules run, the bus carries
|
||||||
|
the mesh's own traffic. What is gone is the way a person asks it anything. Every question — what a
|
||||||
|
node runs, what is assigned, what a module's settings are — has to be asked by opening a shell on
|
||||||
|
the machine and running the control plane's binary inside its container, which is precisely the
|
||||||
|
path the tool surface exists to remove, and which nothing checks, records or permits.
|
||||||
|
|
||||||
|
It also silently changes how work gets done: an assistant told to use the mesh's tools finds them
|
||||||
|
dead, falls back to `ssh` and `docker exec`, and the operator discovers the fallback rather than
|
||||||
|
the fault.
|
||||||
|
|
||||||
|
## Why this is here and not a note in the knowledge base
|
||||||
|
|
||||||
|
The rule is that everything happens over the mesh's bus. A surface that cannot reach that bus is
|
||||||
|
that rule enforced by nothing — and the mesh reports itself healthy throughout, because what broke
|
||||||
|
is not a module, a node or a provision but the thing standing outside asking them questions.
|
||||||
|
|
||||||
|
The same cut-over removed the assistant's long-term memory (`recall`, `memorize`) for the same
|
||||||
|
reason, which is why lessons from the last two days were written into this repository by hand.
|
||||||
|
|
||||||
|
## What would have prevented it
|
||||||
|
|
||||||
|
- **The tool surface as a consumer of the bus like any other**, so moving the bus moves it — rather
|
||||||
|
than a separate bridge with its own connection settings that nothing resolves.
|
||||||
|
- **A check that a tool call reaches a node**, run where the mesh's other checks run. Every check
|
||||||
|
the mesh makes today is about the relationship between the mesh and a machine; none asks whether
|
||||||
|
a person can ask it anything ([issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
|
||||||
|
the same shape one level out).
|
||||||
|
|
||||||
|
## Diagnosed at once, because the answer was in the configuration
|
||||||
|
|
||||||
|
**The surface is not the mesh's.** The tool server the operator's assistant speaks to is the
|
||||||
|
predecessor's brain, installed on the workstation and started as a local process, with the
|
||||||
|
predecessor's broker URL — `amqp://…@<the control node>` — written into the assistant's own
|
||||||
|
configuration. Nothing about it is a module: it has no manifest, no assignment, no seat, no account
|
||||||
|
on the mesh's bus, and the mesh has never known it exists.
|
||||||
|
|
||||||
|
So nothing regressed. The mesh removed a transport that this program still dials, and the program
|
||||||
|
was never part of the mesh to be moved. **The mesh has a tool model** — `tools` in a manifest, the
|
||||||
|
calls a seat accepts, the subjects a module answers on, and a control-plane command that asks one —
|
||||||
|
and **nothing publishes an operator-facing surface onto it.** What an operator uses is the thing
|
||||||
|
that came before, kept alive by a URL in a file.
|
||||||
|
|
||||||
|
That is the issue, and it is larger than a broken connection: the way a person drives this mesh is
|
||||||
|
outside the mesh.
|
||||||
|
|
||||||
|
## Evidence to carry into diagnosis
|
||||||
|
|
||||||
|
- The failure text names `hal/mesh@<node>`, so the bridge is resolving a node and then dialling the
|
||||||
|
old transport.
|
||||||
|
- It fails identically for the local machine, which rules out reachability and points at the
|
||||||
|
transport alone.
|
||||||
|
- The mesh's own traffic over the new bus is unaffected: nodes report, declarations apply.
|
||||||
Reference in New Issue
Block a user