Compare commits

..
5 Commits
2 changed files with 148 additions and 0 deletions
@@ -0,0 +1,61 @@
---
status: open
opened: 2026-10-04
located-in: []
fixed-by:
amended-design:
---
# 229 — A rollout cannot be followed through the mesh's tools, so an agent goes round them
## What was observed
2026-10-04, rolling out to-be 41. An agent drove the rollout through the mesh's MCP tools: the
controller seat's `status`, `plans`, `command`, and the forge's merge. Four times it left those tools
and posted JSON-RPC by hand to the node console's HTTP endpoint with `curl`:
1. **To wait for a plan.** `plans` answers once, with prose. Nothing waits for a plan to reach a tier,
finish or fail. An agent's tools cannot be called from a shell loop, so the only way to be told
when a plan moved was a background `curl` loop polling the console every twenty seconds and
matching the plan's line with `grep`.
2. **To read `status`.** `status` answers a paragraph of prose (the bus's user list), then a JSON
document, both inside one string. Picking out `behind`, `waiting` and `reported` took a script
that cut the string at the first brace and parsed the rest.
3. **To read one module out of `module list`,** whose output was too long to read whole for one line.
4. **To call a tool that arrived after the agent's session began.** The modules rolled out in that same
session added `node-login-shell.execute` and `zsh.zsh_config` to one machine. The agent's MCP
connection had been opened before the console moved to discovery ([ADR 0195](../../02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md)).
It still held the flat catalogue the console announced then, which lacks both the new verbs and the five discovery tools (`mesh_call` among
them) the console announces now. Clearing a session does not reconnect its MCP servers, and the
console never sends a list-changed notice, so nothing told the client its list was stale. The agent
posted `mesh_machine` and `mesh_call` by hand. Reconnecting the server would have given it the
discovery tools, which reach any tool by address the moment it exists.
The calls were authorised, because the console is the operator's own surface. But each is a raw call
the mesh's tools were meant to make unnecessary ([ADR 0154](../../02-DECISIONS/0154-the-meshs-own-verbs-are-the-controller-seats-tools.md)
puts the controller's verbs behind the seat). Each is also a script that breaks silently when a
sentence in the prose changes.
## Why it matters beyond this instance
Every rollout an agent drives has the same shape: merge, wait for a plan, push, wait for reports,
check `status`. When the tools answer only once and only in prose, every agent writes its own poller
and its own parser. Those are invisible to review, different each time, and wrong the first time the
wording moves. An agent that cannot wait also tends to act early, which is the opposite of what a
rollout needs.
## What a fix has to settle
- A way to **wait** on the mesh's own progress. For example, `plans` and `status` could take a plan
or node and a bound, and answer when it moves or the bound passes. Or a verb could follow one plan
to its end.
- **Structured answers** from the controller's verbs, with the prose as a field beside the data, not
around it.
- Whether `command`'s generic answer should take a filter, or whether the verbs it is used for most
(`module list`, `node show`) deserve verbs of their own.
- **A client is told when the console's own surface changes.** The console announces `listChanged`
and sends the notice when what it lists changes, for example after an upgrade that changes its
tools. A long-running session then never keeps a list the console no longer serves. Discovery
already makes every module's tools reachable without the list changing.
How each is checked belongs to the record that settles it.
@@ -0,0 +1,87 @@
---
status: open
opened: 2026-10-04
located-in:
- mesh-host
- mesh-controller
fixed-by:
amended-design:
---
# 230 — A host that hands over to a newer one loses its report, and a plan waits for it for ever without saying so
## What was observed
2026-10-04, rolling out to-be 41. A new host build and a new controller were merged together. The
controller's plan built build-agent and sent every machine a fresh declaration, which also delivered
the new host. Three of the four machines logged, within the same second:
```
host 4bd7df099757 is delivered; standing aside so the launcher runs it
applied 546 resource(s)
applied, and could not tell the mesh: reporting: context canceled
nox-mesh-host-launch: running /usr/lib/nox-mesh-host/versions/4bd7df099757/nox-mesh-host
```
The new host came up and waited for its next declaration. The mesh never heard that the old one had
applied.
The plan then sat at "tier 1 built; waiting for build-agent on [three machines] to be applied", and
everything the mesh said about it read as healthy:
- `status` showed it as `rolling` with `"late": false`;
- `plans` printed "for 0s" on every look, so the wait never appeared to grow;
- the three machines' reports showed `current: false`, which reads like a machine that is merely slow.
Nothing logged, alerted or counted the wait. It was found because a person asked twice for the plan's
state, and the cause was found by reading a machine's own journal. A push to each of the three machines
released it: each new host applied and reported, and the plan moved on.
The same day, a second way to lose a report showed up. Assigning modules with tools to a workstation
changed the bus's user list, which the control machine's declaration carries. Applying it replaced the
bus's container, which cut every machine off for about fifteen seconds. The control machine itself
then logged `applied, and could not tell the mesh: reporting: nats: connection closed`. The report was
lost because the bus restarted under the apply that restarted it.
## Why it matters beyond this instance
**Every genuine host upgrade loses one report.**
[Issue 163](../163-a-delivered-host-stood-aside-on-every-push-and-reported-nothing/00-report.md) fixed
the host that stood aside on every push for the version it already ran. It named the mechanism, that
standing aside cancels the context the report is published with. That fix made standing aside
happen only for a real new version, but left the mechanism in place. So whenever a host build reaches
a machine, that apply's report is lost.
**And the mesh cannot tell a stuck wait from a slow one.** A plan that waits on a report that will
never come waits for ever, and nothing about it changes:
- its age does not grow ("for 0s");
- `late` stays false;
- nothing logs, emits an event or alerts.
This is [issue 187](../187-the-mesh-tells-nobody-when-it-stops-working/00-report.md)'s class of fault
again, *the mesh tells nobody when it stops working*, now in the rollout machinery that every merge
goes through. The operator's rule from issue 163 applies: if an answer has not come in the time an
answer takes, something is wrong, and the mesh must say so itself.
## What a fix has to settle
1. **A report survives whatever its own apply restarts.** The host publishes its report and has it
acknowledged before it stands aside. It retries a report the bus dropped once the link is back.
Failing that, the new host should report the declaration it took over,
naming the apply its predecessor finished. A lost report must be impossible, not merely unlikely.
2. **A plan's wait has an age and a bound.**
- "for 0s" must be the real time since the wait began.
- A wait past a bound, set by how long an apply takes rather than by a guess, makes the plan
`late`.
3. **Late is said where people and agents look.**
- It is said in `status` and in `plans`.
- It is logged as a warning by the controller.
- It is emitted as an event under the controller seat, so something can alert on it.
4. **A plan waiting on a machine the mesh has stopped hearing from** says that, by name, instead of
waiting. The machine's heartbeat already tells the controller it is alive. A live machine with an
unacknowledged declaration is the stuck case itself.
How each is checked belongs to the fix. For the host: a delivered upgrade, applied, is reported. For
the controller: a plan whose machine never reports turns `late` within its bound, and says so in
`status`, the log and an event.