diff --git a/04-ISSUES/229-a-rollout-cannot-be-followed-through-the-meshs-tools/00-report.md b/04-ISSUES/229-a-rollout-cannot-be-followed-through-the-meshs-tools/00-report.md new file mode 100644 index 0000000..2b6a662 --- /dev/null +++ b/04-ISSUES/229-a-rollout-cannot-be-followed-through-the-meshs-tools/00-report.md @@ -0,0 +1,61 @@ +--- +status: open +opened: 2026-10-04 +located-in: [] +fixed-by: +amended-design: +--- + +# 229 — A rollout cannot be followed through the mesh's tools, so an agent goes round them + +## What was observed + +2026-10-04, rolling out to-be 41. An agent drove the rollout through the mesh's MCP tools: the +controller seat's `status`, `plans`, `command`, and the forge's merge. Four times it left those tools +and posted JSON-RPC by hand to the node console's HTTP endpoint with `curl`: + +1. **To wait for a plan.** `plans` answers once, with prose. Nothing waits for a plan to reach a tier, + finish or fail. An agent's tools cannot be called from a shell loop, so the only way to be told + when a plan moved was a background `curl` loop polling the console every twenty seconds and + matching the plan's line with `grep`. +2. **To read `status`.** `status` answers a paragraph of prose (the bus's user list), then a JSON + document, both inside one string. Picking out `behind`, `waiting` and `reported` took a script + that cut the string at the first brace and parsed the rest. +3. **To read one module out of `module list`,** whose output was too long to read whole for one line. +4. **To call a tool that arrived after the agent's session began.** The modules rolled out in that same + session added `node-login-shell.execute` and `zsh.zsh_config` to one machine. The agent's MCP + connection had been opened before the console moved to discovery ([ADR 0195](../../02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md)). + It still held the flat catalogue the console announced then, which lacks both the new verbs and the five discovery tools (`mesh_call` among + them) the console announces now. Clearing a session does not reconnect its MCP servers, and the + console never sends a list-changed notice, so nothing told the client its list was stale. The agent + posted `mesh_machine` and `mesh_call` by hand. Reconnecting the server would have given it the + discovery tools, which reach any tool by address the moment it exists. + +The calls were authorised, because the console is the operator's own surface. But each is a raw call +the mesh's tools were meant to make unnecessary ([ADR 0154](../../02-DECISIONS/0154-the-meshs-own-verbs-are-the-controller-seats-tools.md) +puts the controller's verbs behind the seat). Each is also a script that breaks silently when a +sentence in the prose changes. + +## Why it matters beyond this instance + +Every rollout an agent drives has the same shape: merge, wait for a plan, push, wait for reports, +check `status`. When the tools answer only once and only in prose, every agent writes its own poller +and its own parser. Those are invisible to review, different each time, and wrong the first time the +wording moves. An agent that cannot wait also tends to act early, which is the opposite of what a +rollout needs. + +## What a fix has to settle + +- A way to **wait** on the mesh's own progress. For example, `plans` and `status` could take a plan + or node and a bound, and answer when it moves or the bound passes. Or a verb could follow one plan + to its end. +- **Structured answers** from the controller's verbs, with the prose as a field beside the data, not + around it. +- Whether `command`'s generic answer should take a filter, or whether the verbs it is used for most + (`module list`, `node show`) deserve verbs of their own. +- **A client is told when the console's own surface changes.** The console announces `listChanged` + and sends the notice when what it lists changes, for example after an upgrade that changes its + tools. A long-running session then never keeps a list the console no longer serves. Discovery + already makes every module's tools reachable without the list changing. + +How each is checked belongs to the record that settles it. diff --git a/04-ISSUES/230-a-host-that-hands-over-to-a-newer-one-loses-its-report-and-a-plan-waits-for-ever/00-report.md b/04-ISSUES/230-a-host-that-hands-over-to-a-newer-one-loses-its-report-and-a-plan-waits-for-ever/00-report.md new file mode 100644 index 0000000..247a258 --- /dev/null +++ b/04-ISSUES/230-a-host-that-hands-over-to-a-newer-one-loses-its-report-and-a-plan-waits-for-ever/00-report.md @@ -0,0 +1,87 @@ +--- +status: open +opened: 2026-10-04 +located-in: + - mesh-host + - mesh-controller +fixed-by: +amended-design: +--- + +# 230 — A host that hands over to a newer one loses its report, and a plan waits for it for ever without saying so + +## What was observed + +2026-10-04, rolling out to-be 41. A new host build and a new controller were merged together. The +controller's plan built build-agent and sent every machine a fresh declaration, which also delivered +the new host. Three of the four machines logged, within the same second: + +``` +host 4bd7df099757 is delivered; standing aside so the launcher runs it +applied 546 resource(s) +applied, and could not tell the mesh: reporting: context canceled +nox-mesh-host-launch: running /usr/lib/nox-mesh-host/versions/4bd7df099757/nox-mesh-host +``` + +The new host came up and waited for its next declaration. The mesh never heard that the old one had +applied. + +The plan then sat at "tier 1 built; waiting for build-agent on [three machines] to be applied", and +everything the mesh said about it read as healthy: + +- `status` showed it as `rolling` with `"late": false`; +- `plans` printed "for 0s" on every look, so the wait never appeared to grow; +- the three machines' reports showed `current: false`, which reads like a machine that is merely slow. + +Nothing logged, alerted or counted the wait. It was found because a person asked twice for the plan's +state, and the cause was found by reading a machine's own journal. A push to each of the three machines +released it: each new host applied and reported, and the plan moved on. + +The same day, a second way to lose a report showed up. Assigning modules with tools to a workstation +changed the bus's user list, which the control machine's declaration carries. Applying it replaced the +bus's container, which cut every machine off for about fifteen seconds. The control machine itself +then logged `applied, and could not tell the mesh: reporting: nats: connection closed`. The report was +lost because the bus restarted under the apply that restarted it. + +## Why it matters beyond this instance + +**Every genuine host upgrade loses one report.** +[Issue 163](../163-a-delivered-host-stood-aside-on-every-push-and-reported-nothing/00-report.md) fixed +the host that stood aside on every push for the version it already ran. It named the mechanism, that +standing aside cancels the context the report is published with. That fix made standing aside +happen only for a real new version, but left the mechanism in place. So whenever a host build reaches +a machine, that apply's report is lost. + +**And the mesh cannot tell a stuck wait from a slow one.** A plan that waits on a report that will +never come waits for ever, and nothing about it changes: + +- its age does not grow ("for 0s"); +- `late` stays false; +- nothing logs, emits an event or alerts. + +This is [issue 187](../187-the-mesh-tells-nobody-when-it-stops-working/00-report.md)'s class of fault +again, *the mesh tells nobody when it stops working*, now in the rollout machinery that every merge +goes through. The operator's rule from issue 163 applies: if an answer has not come in the time an +answer takes, something is wrong, and the mesh must say so itself. + +## What a fix has to settle + +1. **A report survives whatever its own apply restarts.** The host publishes its report and has it + acknowledged before it stands aside. It retries a report the bus dropped once the link is back. + Failing that, the new host should report the declaration it took over, + naming the apply its predecessor finished. A lost report must be impossible, not merely unlikely. +2. **A plan's wait has an age and a bound.** + - "for 0s" must be the real time since the wait began. + - A wait past a bound, set by how long an apply takes rather than by a guess, makes the plan + `late`. +3. **Late is said where people and agents look.** + - It is said in `status` and in `plans`. + - It is logged as a warning by the controller. + - It is emitted as an event under the controller seat, so something can alert on it. +4. **A plan waiting on a machine the mesh has stopped hearing from** says that, by name, instead of + waiting. The machine's heartbeat already tells the controller it is alive. A live machine with an + unacknowledged declaration is the stuck case itself. + +How each is checked belongs to the fix. For the host: a delivered upgrade, applied, is reported. For +the controller: a plan whose machine never reports turns `late` within its bound, and says so in +`status`, the log and an event.