Compare commits
2
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
e9b1010bc0 | ||
|
|
f6ed3545b7 |
+65
@@ -82,12 +82,77 @@ managing, or genesis carries a user list that includes the first node's enrolmen
|
||||
takes over from there. Both are decisions, not patches, and both belong to the genesis step that was
|
||||
deliberately left until last.
|
||||
|
||||
## 5 — the composed user list has to be placed by hand at genesis *(fixed)*
|
||||
|
||||
The account a token is the password of is **not recorded at all**: the composer names an enrolment
|
||||
user for every machine with a live token, nothing minted a credential for it, and the composition
|
||||
left it out as a user with no password. The comment above the issuing code already claimed
|
||||
otherwise — *"the account is created before the token is handed over"* — which is how it went
|
||||
unnoticed. Issuing a token now records that account, with the token's own secret as its password,
|
||||
because that is the string the machine will present.
|
||||
|
||||
Placing it is the other half. The list reaches the machine running the bus in that machine's
|
||||
declaration, which a machine that has not enrolled does not get, so at genesis it cannot arrive
|
||||
that way. **The control plane composes and says what it composed** — `broker accounts`, to standard
|
||||
output — and whoever is raising the machine writes it beside the bus's configuration and makes the
|
||||
server re-read it. Twice, because two accounts come into existence at different moments: the
|
||||
enrolment when the token is issued, and the machine's own when it enrols. A control plane that
|
||||
wrote the file itself would have to know where the bus keeps its configuration and how to make it
|
||||
reload, which is the module's knowledge and is what the module takes over on the first push.
|
||||
|
||||
With that, **a first node enrols against the bus it just raised** — measured, from bare, in the
|
||||
lab.
|
||||
|
||||
## 6 — and is enrolled twice, keeping a credential the mesh has replaced *(open)*
|
||||
|
||||
```
|
||||
mesh-controller: enrolled anchor
|
||||
mesh-controller: enrolled anchor (the same second)
|
||||
```
|
||||
|
||||
One `enrol` on the machine, two enrolments in the control plane. Each mints the node a fresh bus
|
||||
password and returns it; the machine keeps the answer to the first, and the mesh keeps the hash of
|
||||
the second. The machine then reconnects for ever as a user whose password the mesh rotated out from
|
||||
under it — *authentication error - User "node.anchor"* on the bus, `Authorization Violation` in the
|
||||
host's log, and a node that never reports.
|
||||
|
||||
What is ruled out: the host asking twice — it asks again only when the mesh says *try again*, and
|
||||
a refused attempt is not logged as an enrolment. Redelivery by the consumer — there is one
|
||||
consumer, its acknowledgement window is thirty seconds, and the handler is quick.
|
||||
|
||||
What is left: the client re-publishing when an acknowledgement is slow, which is what its defaults
|
||||
do. That was addressed by giving the publish a message id derived from its own bytes, so the stream
|
||||
discards the copy — **and the duplicate survived it**, so either the id is not reaching the stream
|
||||
or the second copy is not a copy. This is where the trail stops.
|
||||
|
||||
Worth saying plainly: **the mint is the fragile part, not the delivery.** An enrolment answered
|
||||
twice is survivable if the answer is the same both times, and it cannot be — the mesh keeps only
|
||||
the hash, so a second answer is necessarily a different credential. Whatever closes this either
|
||||
makes the enrolment arrive once, or stops the second arrival from rotating anything.
|
||||
|
||||
## Where it belongs
|
||||
|
||||
`mesh-host` (the bundle and the enrolment path) and `mesh-controller` (the certificate command, and
|
||||
the composition that cannot reach the bus at genesis). Three of the four are fixed on branches; the
|
||||
fourth is the genesis work.
|
||||
|
||||
## What made it slow, and what was changed so it is not
|
||||
|
||||
Six faults behind one another, each found by raising a machine and reading what it said. What cost
|
||||
the most was not the faults:
|
||||
|
||||
- **Every bed's own instructions named the bundle that cannot work**, so the first three attempts
|
||||
ended in a control plane crash-looping on a missing bus. They name the working one now.
|
||||
- **A host binary built without its system** refuses everything it is given with *this host was
|
||||
built for ""*, which reads like a broken bundle. The lab's README says so.
|
||||
- **`make image` in the control plane had been broken for as long as its base was pinned**: the
|
||||
Dockerfile's fallback is a Go older than the module asks for, and the pipeline never saw it
|
||||
because the pipeline passes the declared base in. It reads the base from the manifest now.
|
||||
- **Leaving the machine standing is what answers the question.** Every finding above came from
|
||||
shelling in afterwards — the host's log, the bus's log, the file the bus was actually handed —
|
||||
and none from the test's own output, which says only that nothing converged. The bed takes
|
||||
`MESH_LAB_KEEP`, and the README says to reach for it first.
|
||||
|
||||
## What it cost, for the next person
|
||||
|
||||
Every lab bed still names `foundation-first-node.lock` in its own instructions, and that bundle
|
||||
|
||||
Reference in New Issue
Block a user