Issues 203–206 diagnosed and located #310

Merged
mesh-admin merged 2 commits from issues/203-206-diagnosed into main 2026-10-03 09:43:00 +00:00
8 changed files with 104 additions and 4 deletions
@@ -1,5 +1,5 @@
--- ---
status: open status: located
opened: 2026-10-02 opened: 2026-10-02
located-in: located-in:
- mesh-controller - mesh-controller
@@ -0,0 +1,18 @@
# Diagnosis — 203
**2026-10-03.** Two acts, one effect. `assign` records the assignment and, at composition, every
`own-secrets` entry a module declares gets a sealed value from the controller's `Needed` map — a
value minted so the file exists, which for `broker` is a random secret, not a credential. The bus
credential is composed only by `module issue <module> --node <node>` (cmd/mesh-controller/modules.go,
`issueOnTheNewBus` → `issueWith`): it mints the bus user, records its hash, and seals the credential
JSON into the same `broker` need. Nothing joins the two: `push` composes and sends whatever the need
holds, and the only warning is the standing line listing every bus user without a minted credential,
printed on every push regardless of what was just assigned.
Ruled out: the host (it wrote the file it was given, owned as asked); the runtime (it refused a file
that is not JSON, correctly, and said so); the manifest (`own-secrets.broker` is the shape every
module uses).
**Owner:** mesh-controller — the assign path. **Fix direction:** assigning a module that declares
`own-secrets.broker` issues its credential in the same act, idempotently; a push of a module whose bus
user is unminted is refused by name rather than sent with a placeholder.
@@ -1,5 +1,5 @@
--- ---
status: open status: located
opened: 2026-10-02 opened: 2026-10-02
located-in: located-in:
- mesh-controller - mesh-controller
@@ -0,0 +1,37 @@
# Diagnosis — 204
**2026-10-03, in the controller's code.**
Ruled out: a start-up re-send (a starting controller sends nothing; it asserts the bus, resumes plans,
follows events); a plan sending a recorded declaration (a plan records artifact digests and module
states, and its rollout composes fresh at send time); a cache (every compose reads the store); a path
that does not record its send (push, the cascade and the rollout all record after sending; only the raw
`declare <node> <file>` command did not); the host applying an older sequence (it already refuses a
declaration numbered below the one it kept).
Found, two faults that together produce the evidence:
1. **The sequence number went on at send time, after composing.** Every sending path composed first and
numbered each declaration as it was sent; a multi-machine send composes every machine before sending
any. So a declaration composed *before* an assignment changed and sent *after* a fresher one carried
the higher number — and the host, refusing only lower numbers, applied the older content as the mesh's
newest word. The stale declaration was accepted, so its number was higher, so it was composed earlier
and sent later.
2. **The send record was written on the sender's own context, after the send.** A controller being
replaced in that second has its context cancelled between telling the machine and writing the record;
the machine was told, the record never written, and status showed only the person's earlier send.
The likely sender, consistent with both and with the timing: the outgoing controller's reaction to a
catalogue registration during the build round, which re-sends the machines running the registered
module (one ran on exactly the two machines affected), composed under its hold before the person's
assignments, numbered and sent at 21:29:19–20 as the controller was being replaced. The old container's
log is gone, so the sender is inferred from code and timing; the mechanism is not.
**Owner:** mesh-controller. **Fix direction:** number a declaration before composing it, in every path,
so what was composed earlier is numbered lower whatever order the sends happen in and the host's
existing refusal does its job; record a send on a context that outlives the sender; the raw `declare`
command records too.
Two questions left for HQ: whether status should show the sequence a machine was last sent beside the
digest, and whether a sender's hold should also cover the assignment verbs, which today run between a
hold's compose and its send without waiting.
@@ -1,8 +1,9 @@
--- ---
status: open status: located
opened: 2026-10-02 opened: 2026-10-02
located-in: located-in:
- mesh-host - mesh-host
- 00-META/how-we-build.md
fixed-by: fixed-by:
amended-design: amended-design:
--- ---
@@ -0,0 +1,18 @@
# Diagnosis — 205
**2026-10-03.** The host's package step on an Arch machine installs with the package manager against
the database the machine has (mesh-host internal/system/arch.go); it neither refreshes it nor can
safely, since a refresh plus one install is the partial upgrade the distribution warns against. On
the control node the database and keyring were from 24 July; the mirrors no longer served the version
it named, so every mirror answered 404 and the one cached file failed its signature. The host reported
the package manager's output whole, which reads as a mirror outage.
Two owners. The **narrow** half is the host's: classify that failure and say what it is — the database
is stale, the operator must upgrade — rather than relaying forty mirror lines. The **wide** half is a
rule nobody has written: who keeps a machine current enough for its own declarations to apply, and
how that is checked. ADR 0173 makes the machine the mesh's; `00-META/how-we-build.md` says nothing
about its package database. That is a decision (playbook 02), not a code fix: a `package-manager` seat
holder with a schedule, the host, or the operator by rule.
Ruled out: the manifest (`package: nodejs` is correct for the distribution and installed on three
machines the same hour); the network (the mirrors answered, with 404s).
@@ -1,5 +1,5 @@
--- ---
status: open status: located
opened: 2026-10-03 opened: 2026-10-03
located-in: located-in:
- mesh-controller - mesh-controller
@@ -0,0 +1,26 @@
# Diagnosis — 206
**2026-10-03.** Three faults in one handover, all the controller's.
1. **The worker's shape is asserted, not reconciled.** The controller creates a seat's worker if
absent (internal/broker, the consumer assertion on start) and leaves an existing one as it is. A
change of shape — here push to pull — therefore never reaches a bus that already has the worker
until somebody deletes it. The new build machine bound a worker whose type its code no longer
speaks.
2. **The plan rolled the holder before the definer.** The merge's plan tiered by artifacts (ADR 0162):
the build machine's image stands on nothing of the controller's, so it came first. For every other
module the order is indifferent; for the holder of the build seat, the controller that defines its
worker must run first, or the build that would bring the controller cannot be taken.
3. **A credential older than its shape.** The build machine's credential was sealed on 2026-09-28,
before credentials carried `claims`; the new binary read none and fell back to the new seat, for
which it had no grant. Re-issuing the credential fixed it; nothing had said it was stale.
Also seen: the build machine's container restarts on its environment file and not on its credential,
so a re-issued credential reaches it only by chance (shared with issue 203's fix direction).
Ruled out: the bus (it refused exactly what the grants and the consumer type said to refuse); the
build machine's new code (it did what its credential told it).
**Fix direction:** the controller reconciles every consumer it owns to the shape it derives, recreating
one whose type changed and saying so; a plan rolls the controller before any holder of the build seat;
a credential whose shape predates what the binary reads is listed and re-issued.