Compare commits

...
1 Commits
3 changed files with 140 additions and 0 deletions
@@ -0,0 +1,45 @@
---
status: open
opened: 2026-10-02
located-in:
- mesh-controller
fixed-by:
amended-design:
---
# 203 — A fresh assignment is pushed before its credential exists, and the runtime crash-loops
## What was observed
2026-10-02, the first live assignment of the node tools runtime (to-be 38 WP3). `assign` put the
module on a machine and `push` sent the declaration. The host applied everything: the bundle unpacked,
the unit written and started, the module's `broker` secret file written and owned by the operator
account. The runtime then restarted thirteen times in a minute:
```
mesh-tools: cannot read the broker credential at …/broker: SyntaxError: Unexpected token 'O',
"Oj6j2Ssa-v"... is not valid JSON
```
The file held a 40-byte random secret, not a bus credential. The push's own output had said why,
one line among forty: *the bus's user list leaves out … `<node>.node-tools`. Each is a user that cannot
connect until one is issued.* The credential exists only after `module issue <module> --node <node>`,
a separate act; a second push then carried the real credential and the runtime came up. The same
sequence repeated on the next two machines, by hand, in the right order.
## Why it matters beyond this instance
Every module that speaks on the bus declares an `own-secrets.broker`; the mesh seals *something*
there on assignment and the real credential only on issue. So the first push of any fresh assignment
delivers a process that cannot authenticate and will crash-loop until a person runs a second verb and
a second push. Nothing refuses the first push, and the warning is a line in a long list that is
printed on every push regardless. The design says the mesh issues an assignment's subjects and the
runtime serves what it is issued ([ADR 0160](../../02-DECISIONS/0160-the-mesh-issues-an-assignments-subjects-and-a-runtime-serves-what-it-is-issued.md));
an assignment whose credential is not issued is half an assignment, and the mesh lets it through.
## Questions HQ must answer
- Is issuing the credential part of assigning, so `assign` mints it, or must a push refuse a module
whose bus user is unminted, naming the verb?
- Is a placeholder sealed where a credential belongs ever right, or should the resource be absent
until the credential exists, so the host never writes a file the process cannot read?
@@ -0,0 +1,49 @@
---
status: open
opened: 2026-10-02
located-in:
- mesh-controller
fixed-by:
amended-design:
---
# 204 — A controller handover re-sent every node a declaration composed from a stale view
## What was observed
2026-10-02 21:29 UTC. Two machines were assigned the node tools runtime and had the console taken off
them, and were pushed; the host on each applied it (the console removed, the runtime created and
running). Two seconds later, on each, a second declaration arrived that undid it — the host's log:
```
23:29:18 created node-tools.runtime (node-tools): 239 file(s), running as node-tools.service
23:29:18 applied 569 resource(s)
23:29:20 removed node-tools.runtime (node-tools)
23:29:20 forgotten node-tools.interpreter (nodejs)
23:29:24 created mesh-console.needs-broker, created mesh-console.server (mesh-console)
```
The second declaration had the console assigned and the runtime absent: the assignments as they were
a minute earlier. The controller's status knew of one send per machine, the person's. At that minute a
plan from an unrelated merge was rolling a new controller build onto the control node, so an instance
was starting while another was stopping. A third push by hand, two minutes later, restored both
machines and nothing undid it again.
Which instance sent the stale declaration, and from what, is not established: the outgoing one on its
way down, the incoming one at start-up before its view was current, or the rolling plan sending what it
had composed when it was made — the same family as [issue 201](../201-a-push-recreated-the-controller-behind-the-row-its-successor-wrote/00-report.md),
where a plan sent a digest older than the one a successor had written.
## Why it matters beyond this instance
A declaration the mesh sends is the mesh's word on what a machine should be; a machine applies it in
full, including removing what it no longer names. A stale one is therefore not a no-op: it tears down
whatever was assigned since, and the controller's own record does not show the send, so the next
person reads "applied, current" over a machine that is wrong. Two people merging within a minute is
ordinary, and a controller handover happens on every controller merge.
## Questions HQ must answer
- Where does the stale view come from, and is every send recorded so status can show it?
- Should a declaration carry the assignment generation it was composed from, so a host refuses one
older than the last it applied, as it already refuses a declaration it cannot verify?
@@ -0,0 +1,46 @@
---
status: open
opened: 2026-10-02
located-in:
- mesh-host
fixed-by:
amended-design:
---
# 205 — A package resource fails against a stale package database, and nothing keeps it fresh
## What was observed
2026-10-02. The node tools runtime's manifest declares the interpreter as a package. On three machines
the package installed. On the control node the host failed the declaration three times and reported the
machine wrong, stuck:
```
applying "node-tools.interpreter": installing nodejs: pacman exited 1:
error: failed retrieving file 'nodejs-26.5.0-1-x86_64.pkg.tar.zst' from <mirror>: 404
… (every mirror)
error: nodejs: signature from "<packager>" is invalid
```
The machine's package database was from 24 July, ten weeks earlier; the mirrors had long moved on from
the version it asked for, and its keyring was as old. The host asks the package manager to install from
whatever database the machine has and does not refresh it; refreshing on the host's own initiative is
not safe either, because on a rolling distribution a refreshed database plus a single install is a
partial upgrade, which the distribution warns against. The way out was a full system upgrade by the
operator, outside the mesh.
## Why it matters beyond this instance
A `package` resource is one of the host's shapes and every environment module leans on it (design 37).
Its success depends on a machine fact the mesh neither records nor keeps: how old the package database
is. A machine that has not been upgraded in months fails every new package the mesh declares, with an
error that reads as a mirror outage. The mesh says the operator's machine is the mesh's
([ADR 0173](../../02-DECISIONS/0173-the-operators-machine-is-the-meshs-and-a-module-is-what-it-declares.md));
nothing in it says who keeps the machine current enough for its own declarations to apply, or checks.
## Questions HQ must answer
- Is keeping a machine's package database and keyring current a module's job (a `package-manager`
seat holder with a schedule), the host's, or the operator's — and how is it checked?
- Should a `package` resource's failure distinguish "the database is stale" from "the mirror is down",
so the report says what to do?