Merge pull request 'Issues 229 and 230: a rollout cannot be followed through the mesh's tools; a host hand-over loses its report and a plan waits for ever' (#352) from issues/229-a-rollout-cannot-be-followed-through-the-meshs-tools into main
This commit was merged in pull request #352.
This commit is contained in:
@@ -0,0 +1,61 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-04
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 229 — A rollout cannot be followed through the mesh's tools, so an agent goes round them
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-04, rolling out to-be 41. An agent drove the rollout through the mesh's MCP tools: the
|
||||
controller seat's `status`, `plans`, `command`, and the forge's merge. Four times it left those tools
|
||||
and posted JSON-RPC by hand to the node console's HTTP endpoint with `curl`:
|
||||
|
||||
1. **To wait for a plan.** `plans` answers once, with prose. Nothing waits for a plan to reach a tier,
|
||||
finish or fail. An agent's tools cannot be called from a shell loop, so the only way to be told
|
||||
when a plan moved was a background `curl` loop polling the console every twenty seconds and
|
||||
matching the plan's line with `grep`.
|
||||
2. **To read `status`.** `status` answers a paragraph of prose (the bus's user list), then a JSON
|
||||
document, both inside one string. Picking out `behind`, `waiting` and `reported` took a script
|
||||
that cut the string at the first brace and parsed the rest.
|
||||
3. **To read one module out of `module list`,** whose output was too long to read whole for one line.
|
||||
4. **To call a tool that arrived after the agent's session began.** The modules rolled out in that same
|
||||
session added `node-login-shell.execute` and `zsh.zsh_config` to one machine. The agent's MCP
|
||||
connection had been opened before the console moved to discovery ([ADR 0195](../../02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md)).
|
||||
It still held the flat catalogue the console announced then, which lacks both the new verbs and the five discovery tools (`mesh_call` among
|
||||
them) the console announces now. Clearing a session does not reconnect its MCP servers, and the
|
||||
console never sends a list-changed notice, so nothing told the client its list was stale. The agent
|
||||
posted `mesh_machine` and `mesh_call` by hand. Reconnecting the server would have given it the
|
||||
discovery tools, which reach any tool by address the moment it exists.
|
||||
|
||||
The calls were authorised, because the console is the operator's own surface. But each is a raw call
|
||||
the mesh's tools were meant to make unnecessary ([ADR 0154](../../02-DECISIONS/0154-the-meshs-own-verbs-are-the-controller-seats-tools.md)
|
||||
puts the controller's verbs behind the seat). Each is also a script that breaks silently when a
|
||||
sentence in the prose changes.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
Every rollout an agent drives has the same shape: merge, wait for a plan, push, wait for reports,
|
||||
check `status`. When the tools answer only once and only in prose, every agent writes its own poller
|
||||
and its own parser. Those are invisible to review, different each time, and wrong the first time the
|
||||
wording moves. An agent that cannot wait also tends to act early, which is the opposite of what a
|
||||
rollout needs.
|
||||
|
||||
## What a fix has to settle
|
||||
|
||||
- A way to **wait** on the mesh's own progress. For example, `plans` and `status` could take a plan
|
||||
or node and a bound, and answer when it moves or the bound passes. Or a verb could follow one plan
|
||||
to its end.
|
||||
- **Structured answers** from the controller's verbs, with the prose as a field beside the data, not
|
||||
around it.
|
||||
- Whether `command`'s generic answer should take a filter, or whether the verbs it is used for most
|
||||
(`module list`, `node show`) deserve verbs of their own.
|
||||
- **A client is told when the console's own surface changes.** The console announces `listChanged`
|
||||
and sends the notice when what it lists changes, for example after an upgrade that changes its
|
||||
tools. A long-running session then never keeps a list the console no longer serves. Discovery
|
||||
already makes every module's tools reachable without the list changing.
|
||||
|
||||
How each is checked belongs to the record that settles it.
|
||||
+87
@@ -0,0 +1,87 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-10-04
|
||||
located-in:
|
||||
- mesh-host
|
||||
- mesh-controller
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 230 — A host that hands over to a newer one loses its report, and a plan waits for it for ever without saying so
|
||||
|
||||
## What was observed
|
||||
|
||||
2026-10-04, rolling out to-be 41. A new host build and a new controller were merged together. The
|
||||
controller's plan built build-agent and sent every machine a fresh declaration, which also delivered
|
||||
the new host. Three of the four machines logged, within the same second:
|
||||
|
||||
```
|
||||
host 4bd7df099757 is delivered; standing aside so the launcher runs it
|
||||
applied 546 resource(s)
|
||||
applied, and could not tell the mesh: reporting: context canceled
|
||||
nox-mesh-host-launch: running /usr/lib/nox-mesh-host/versions/4bd7df099757/nox-mesh-host
|
||||
```
|
||||
|
||||
The new host came up and waited for its next declaration. The mesh never heard that the old one had
|
||||
applied.
|
||||
|
||||
The plan then sat at "tier 1 built; waiting for build-agent on [three machines] to be applied", and
|
||||
everything the mesh said about it read as healthy:
|
||||
|
||||
- `status` showed it as `rolling` with `"late": false`;
|
||||
- `plans` printed "for 0s" on every look, so the wait never appeared to grow;
|
||||
- the three machines' reports showed `current: false`, which reads like a machine that is merely slow.
|
||||
|
||||
Nothing logged, alerted or counted the wait. It was found because a person asked twice for the plan's
|
||||
state, and the cause was found by reading a machine's own journal. A push to each of the three machines
|
||||
released it: each new host applied and reported, and the plan moved on.
|
||||
|
||||
The same day, a second way to lose a report showed up. Assigning modules with tools to a workstation
|
||||
changed the bus's user list, which the control machine's declaration carries. Applying it replaced the
|
||||
bus's container, which cut every machine off for about fifteen seconds. The control machine itself
|
||||
then logged `applied, and could not tell the mesh: reporting: nats: connection closed`. The report was
|
||||
lost because the bus restarted under the apply that restarted it.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
**Every genuine host upgrade loses one report.**
|
||||
[Issue 163](../163-a-delivered-host-stood-aside-on-every-push-and-reported-nothing/00-report.md) fixed
|
||||
the host that stood aside on every push for the version it already ran. It named the mechanism, that
|
||||
standing aside cancels the context the report is published with. That fix made standing aside
|
||||
happen only for a real new version, but left the mechanism in place. So whenever a host build reaches
|
||||
a machine, that apply's report is lost.
|
||||
|
||||
**And the mesh cannot tell a stuck wait from a slow one.** A plan that waits on a report that will
|
||||
never come waits for ever, and nothing about it changes:
|
||||
|
||||
- its age does not grow ("for 0s");
|
||||
- `late` stays false;
|
||||
- nothing logs, emits an event or alerts.
|
||||
|
||||
This is [issue 187](../187-the-mesh-tells-nobody-when-it-stops-working/00-report.md)'s class of fault
|
||||
again, *the mesh tells nobody when it stops working*, now in the rollout machinery that every merge
|
||||
goes through. The operator's rule from issue 163 applies: if an answer has not come in the time an
|
||||
answer takes, something is wrong, and the mesh must say so itself.
|
||||
|
||||
## What a fix has to settle
|
||||
|
||||
1. **A report survives whatever its own apply restarts.** The host publishes its report and has it
|
||||
acknowledged before it stands aside. It retries a report the bus dropped once the link is back.
|
||||
Failing that, the new host should report the declaration it took over,
|
||||
naming the apply its predecessor finished. A lost report must be impossible, not merely unlikely.
|
||||
2. **A plan's wait has an age and a bound.**
|
||||
- "for 0s" must be the real time since the wait began.
|
||||
- A wait past a bound, set by how long an apply takes rather than by a guess, makes the plan
|
||||
`late`.
|
||||
3. **Late is said where people and agents look.**
|
||||
- It is said in `status` and in `plans`.
|
||||
- It is logged as a warning by the controller.
|
||||
- It is emitted as an event under the controller seat, so something can alert on it.
|
||||
4. **A plan waiting on a machine the mesh has stopped hearing from** says that, by name, instead of
|
||||
waiting. The machine's heartbeat already tells the controller it is alive. A live machine with an
|
||||
unacknowledged declaration is the stuck case itself.
|
||||
|
||||
How each is checked belongs to the fix. For the host: a delivered upgrade, applied, is reported. For
|
||||
the controller: a plan whose machine never reports turns `late` within its bound, and says so in
|
||||
`status`, the log and an event.
|
||||
Reference in New Issue
Block a user