From 8d83d946592aeb7aab0a6ca2d265193e06b4e7d5 Mon Sep 17 00:00:00 2001 From: jochen Date: Sun, 4 Oct 2026 10:51:45 +0200 Subject: [PATCH 1/5] Issue 229: a rollout cannot be followed through the mesh's tools --- .../00-report.md | 49 +++++++++++++++++++ 1 file changed, 49 insertions(+) create mode 100644 04-ISSUES/229-a-rollout-cannot-be-followed-through-the-meshs-tools/00-report.md diff --git a/04-ISSUES/229-a-rollout-cannot-be-followed-through-the-meshs-tools/00-report.md b/04-ISSUES/229-a-rollout-cannot-be-followed-through-the-meshs-tools/00-report.md new file mode 100644 index 0000000..88b243d --- /dev/null +++ b/04-ISSUES/229-a-rollout-cannot-be-followed-through-the-meshs-tools/00-report.md @@ -0,0 +1,49 @@ +--- +status: open +opened: 2026-10-04 +located-in: [] +fixed-by: +amended-design: +--- + +# 229 — A rollout cannot be followed through the mesh's tools, so an agent goes round them + +## What was observed + +2026-10-04, rolling out to-be 41. An agent drove the rollout through the mesh's MCP tools: the +controller seat's `status`, `plans`, `command`, and the forge's merge. Three times it left those tools +and posted JSON-RPC by hand to the node console's HTTP endpoint with `curl`: + +1. **To wait for a plan.** `plans` answers once, with prose. Nothing waits for a plan to reach a tier, + finish or fail. An agent's tools cannot be called from a shell loop, so the only way to be told + when a plan moved was a background `curl` loop polling the console every twenty seconds and + matching the plan's line with `grep`. +2. **To read `status`.** `status` answers a paragraph of prose (the bus's user list), then a JSON + document, both inside one string. Picking out `behind`, `waiting` and `reported` took a script + that cut the string at the first brace and parsed the rest. +3. **To read one module out of `module list`,** whose output was too long to read whole for one line. + +The calls were authorised, because the console is the operator's own surface. But each is a raw call +the mesh's tools were meant to make unnecessary ([ADR 0154](../../02-DECISIONS/0154-the-meshs-own-verbs-are-the-controller-seats-tools.md) +puts the controller's verbs behind the seat). Each is also a script that breaks silently when a +sentence in the prose changes. + +## Why it matters beyond this instance + +Every rollout an agent drives has the same shape: merge, wait for a plan, push, wait for reports, +check `status`. When the tools answer only once and only in prose, every agent writes its own poller +and its own parser. Those are invisible to review, different each time, and wrong the first time the +wording moves. An agent that cannot wait also tends to act early, which is the opposite of what a +rollout needs. + +## What a fix has to settle + +- A way to **wait** on the mesh's own progress. For example, `plans` and `status` could take a plan + or node and a bound, and answer when it moves or the bound passes. Or a verb could follow one plan + to its end. +- **Structured answers** from the controller's verbs, with the prose as a field beside the data, not + around it. +- Whether `command`'s generic answer should take a filter, or whether the verbs it is used for most + (`module list`, `node show`) deserve verbs of their own. + +How each is checked belongs to the record that settles it. From 2db0ea268df50cec1d9e22179b49b07b7ab01abe Mon Sep 17 00:00:00 2001 From: jochen Date: Sun, 4 Oct 2026 10:54:35 +0200 Subject: [PATCH 2/5] Issue 230: a host that hands over loses its report, and a plan waits for it for ever without saying so --- .../00-report.md | 80 +++++++++++++++++++ 1 file changed, 80 insertions(+) create mode 100644 04-ISSUES/230-a-host-that-hands-over-to-a-newer-one-loses-its-report-and-a-plan-waits-for-ever/00-report.md diff --git a/04-ISSUES/230-a-host-that-hands-over-to-a-newer-one-loses-its-report-and-a-plan-waits-for-ever/00-report.md b/04-ISSUES/230-a-host-that-hands-over-to-a-newer-one-loses-its-report-and-a-plan-waits-for-ever/00-report.md new file mode 100644 index 0000000..7fce3f3 --- /dev/null +++ b/04-ISSUES/230-a-host-that-hands-over-to-a-newer-one-loses-its-report-and-a-plan-waits-for-ever/00-report.md @@ -0,0 +1,80 @@ +--- +status: open +opened: 2026-10-04 +located-in: + - mesh-host + - mesh-controller +fixed-by: +amended-design: +--- + +# 230 — A host that hands over to a newer one loses its report, and a plan waits for it for ever without saying so + +## What was observed + +2026-10-04, rolling out to-be 41. A new host build and a new controller were merged together. The +controller's plan built build-agent and sent every machine a fresh declaration, which also delivered +the new host. Three of the four machines logged, within the same second: + +``` +host 4bd7df099757 is delivered; standing aside so the launcher runs it +applied 546 resource(s) +applied, and could not tell the mesh: reporting: context canceled +nox-mesh-host-launch: running /usr/lib/nox-mesh-host/versions/4bd7df099757/nox-mesh-host +``` + +The new host came up and waited for its next declaration. The mesh never heard that the old one had +applied. + +The plan then sat at "tier 1 built; waiting for build-agent on [three machines] to be applied", and +everything the mesh said about it read as healthy: + +- `status` showed it as `rolling` with `"late": false`; +- `plans` printed "for 0s" on every look, so the wait never appeared to grow; +- the three machines' reports showed `current: false`, which reads like a machine that is merely slow. + +Nothing logged, alerted or counted the wait. It was found because a person asked twice for the plan's +state, and the cause was found by reading a machine's own journal. A push to each of the three machines +released it: each new host applied and reported, and the plan moved on. + +## Why it matters beyond this instance + +**Every genuine host upgrade loses one report.** +[Issue 163](../163-a-delivered-host-stood-aside-on-every-push-and-reported-nothing/00-report.md) fixed +the host that stood aside on every push for the version it already ran. It named the mechanism, that +standing aside cancels the context the report is published with. That fix made standing aside +happen only for a real new version, but left the mechanism in place. So whenever a host build reaches +a machine, that apply's report is lost. + +**And the mesh cannot tell a stuck wait from a slow one.** A plan that waits on a report that will +never come waits for ever, and nothing about it changes: + +- its age does not grow ("for 0s"); +- `late` stays false; +- nothing logs, emits an event or alerts. + +This is [issue 187](../187-the-mesh-tells-nobody-when-it-stops-working/00-report.md)'s class of fault +again, *the mesh tells nobody when it stops working*, now in the rollout machinery that every merge +goes through. The operator's rule from issue 163 applies: if an answer has not come in the time an +answer takes, something is wrong, and the mesh must say so itself. + +## What a fix has to settle + +1. **The host reports before it stands aside.** The report should be published and acknowledged + before the old host exits. Failing that, the new host should report the declaration it took over, + naming the apply its predecessor finished. A lost report must be impossible, not merely unlikely. +2. **A plan's wait has an age and a bound.** + - "for 0s" must be the real time since the wait began. + - A wait past a bound, set by how long an apply takes rather than by a guess, makes the plan + `late`. +3. **Late is said where people and agents look.** + - It is said in `status` and in `plans`. + - It is logged as a warning by the controller. + - It is emitted as an event under the controller seat, so something can alert on it. +4. **A plan waiting on a machine the mesh has stopped hearing from** says that, by name, instead of + waiting. The machine's heartbeat already tells the controller it is alive. A live machine with an + unacknowledged declaration is the stuck case itself. + +How each is checked belongs to the fix. For the host: a delivered upgrade, applied, is reported. For +the controller: a plan whose machine never reports turns `late` within its bound, and says so in +`status`, the log and an event. From 9c13c89fa3893a79803552ab0a06cb26cf33363a Mon Sep 17 00:00:00 2001 From: jochen Date: Sun, 4 Oct 2026 11:00:13 +0200 Subject: [PATCH 3/5] Issue 229: new tools never reach an agent's connection, and the console's generic tools are not on it --- .../00-report.md | 14 +++++++++++++- 1 file changed, 13 insertions(+), 1 deletion(-) diff --git a/04-ISSUES/229-a-rollout-cannot-be-followed-through-the-meshs-tools/00-report.md b/04-ISSUES/229-a-rollout-cannot-be-followed-through-the-meshs-tools/00-report.md index 88b243d..ec71d3c 100644 --- a/04-ISSUES/229-a-rollout-cannot-be-followed-through-the-meshs-tools/00-report.md +++ b/04-ISSUES/229-a-rollout-cannot-be-followed-through-the-meshs-tools/00-report.md @@ -11,7 +11,7 @@ amended-design: ## What was observed 2026-10-04, rolling out to-be 41. An agent drove the rollout through the mesh's MCP tools: the -controller seat's `status`, `plans`, `command`, and the forge's merge. Three times it left those tools +controller seat's `status`, `plans`, `command`, and the forge's merge. Five times it left those tools and posted JSON-RPC by hand to the node console's HTTP endpoint with `curl`: 1. **To wait for a plan.** `plans` answers once, with prose. Nothing waits for a plan to reach a tier, @@ -22,6 +22,15 @@ and posted JSON-RPC by hand to the node console's HTTP endpoint with `curl`: document, both inside one string. Picking out `behind`, `waiting` and `reported` took a script that cut the string at the first brace and parsed the rest. 3. **To read one module out of `module list`,** whose output was too long to read whole for one line. +4. **To call a tool that arrived after the agent's session began.** The modules rolled out in that same + session added `node-login-shell.execute` and `zsh.zsh_config` to one machine. The agent's MCP + connection to the mesh had listed its tools once, when the session started, and never listed them + again: no list-changed notice reached it, and a search of its tools found neither verb. So the only + way to prove the new verb through the mesh was the console's `mesh_call`, posted by hand. +5. **To reach the console's own generic tools at all.** The console serves `mesh_overview`, + `mesh_machine`, `mesh_search`, `mesh_describe` and `mesh_call`. The MCP connection the agent had + flattens every module's tools into one list, and exposes none of those five. The one surface that + can reach any tool by address is therefore missing from the connection agents use. The calls were authorised, because the console is the operator's own surface. But each is a raw call the mesh's tools were meant to make unnecessary ([ADR 0154](../../02-DECISIONS/0154-the-meshs-own-verbs-are-the-controller-seats-tools.md) @@ -45,5 +54,8 @@ rollout needs. around it. - Whether `command`'s generic answer should take a filter, or whether the verbs it is used for most (`module list`, `node show`) deserve verbs of their own. +- **A tool list that follows the mesh.** When a module's tools arrive or leave, the connection says so + (MCP's list-changed notification). Failing that, it always carries the console's addressable + `mesh_call` and `mesh_describe`, so a new verb is reachable without leaving the agent's tools. How each is checked belongs to the record that settles it. From f23a71e0d7f13c1a09e2e844711f8a133d5ebab8 Mon Sep 17 00:00:00 2001 From: jochen Date: Sun, 4 Oct 2026 11:00:27 +0200 Subject: [PATCH 4/5] Issue 230: a report is also lost when the apply restarts the bus --- .../00-report.md | 11 +++++++++-- 1 file changed, 9 insertions(+), 2 deletions(-) diff --git a/04-ISSUES/230-a-host-that-hands-over-to-a-newer-one-loses-its-report-and-a-plan-waits-for-ever/00-report.md b/04-ISSUES/230-a-host-that-hands-over-to-a-newer-one-loses-its-report-and-a-plan-waits-for-ever/00-report.md index 7fce3f3..247a258 100644 --- a/04-ISSUES/230-a-host-that-hands-over-to-a-newer-one-loses-its-report-and-a-plan-waits-for-ever/00-report.md +++ b/04-ISSUES/230-a-host-that-hands-over-to-a-newer-one-loses-its-report-and-a-plan-waits-for-ever/00-report.md @@ -37,6 +37,12 @@ Nothing logged, alerted or counted the wait. It was found because a person asked state, and the cause was found by reading a machine's own journal. A push to each of the three machines released it: each new host applied and reported, and the plan moved on. +The same day, a second way to lose a report showed up. Assigning modules with tools to a workstation +changed the bus's user list, which the control machine's declaration carries. Applying it replaced the +bus's container, which cut every machine off for about fifteen seconds. The control machine itself +then logged `applied, and could not tell the mesh: reporting: nats: connection closed`. The report was +lost because the bus restarted under the apply that restarted it. + ## Why it matters beyond this instance **Every genuine host upgrade loses one report.** @@ -60,8 +66,9 @@ answer takes, something is wrong, and the mesh must say so itself. ## What a fix has to settle -1. **The host reports before it stands aside.** The report should be published and acknowledged - before the old host exits. Failing that, the new host should report the declaration it took over, +1. **A report survives whatever its own apply restarts.** The host publishes its report and has it + acknowledged before it stands aside. It retries a report the bus dropped once the link is back. + Failing that, the new host should report the declaration it took over, naming the apply its predecessor finished. A lost report must be impossible, not merely unlikely. 2. **A plan's wait has an age and a bound.** - "for 0s" must be the real time since the wait began. From 0c2eae07c54c61fe438a1c224bfc849f68328c36 Mon Sep 17 00:00:00 2001 From: jochen Date: Sun, 4 Oct 2026 11:01:50 +0200 Subject: [PATCH 5/5] Issue 229: the stale tool list is a connection opened before discovery; the fix is list-changed --- .../00-report.md | 22 +++++++++---------- 1 file changed, 11 insertions(+), 11 deletions(-) diff --git a/04-ISSUES/229-a-rollout-cannot-be-followed-through-the-meshs-tools/00-report.md b/04-ISSUES/229-a-rollout-cannot-be-followed-through-the-meshs-tools/00-report.md index ec71d3c..2b6a662 100644 --- a/04-ISSUES/229-a-rollout-cannot-be-followed-through-the-meshs-tools/00-report.md +++ b/04-ISSUES/229-a-rollout-cannot-be-followed-through-the-meshs-tools/00-report.md @@ -11,7 +11,7 @@ amended-design: ## What was observed 2026-10-04, rolling out to-be 41. An agent drove the rollout through the mesh's MCP tools: the -controller seat's `status`, `plans`, `command`, and the forge's merge. Five times it left those tools +controller seat's `status`, `plans`, `command`, and the forge's merge. Four times it left those tools and posted JSON-RPC by hand to the node console's HTTP endpoint with `curl`: 1. **To wait for a plan.** `plans` answers once, with prose. Nothing waits for a plan to reach a tier, @@ -24,13 +24,12 @@ and posted JSON-RPC by hand to the node console's HTTP endpoint with `curl`: 3. **To read one module out of `module list`,** whose output was too long to read whole for one line. 4. **To call a tool that arrived after the agent's session began.** The modules rolled out in that same session added `node-login-shell.execute` and `zsh.zsh_config` to one machine. The agent's MCP - connection to the mesh had listed its tools once, when the session started, and never listed them - again: no list-changed notice reached it, and a search of its tools found neither verb. So the only - way to prove the new verb through the mesh was the console's `mesh_call`, posted by hand. -5. **To reach the console's own generic tools at all.** The console serves `mesh_overview`, - `mesh_machine`, `mesh_search`, `mesh_describe` and `mesh_call`. The MCP connection the agent had - flattens every module's tools into one list, and exposes none of those five. The one surface that - can reach any tool by address is therefore missing from the connection agents use. + connection had been opened before the console moved to discovery ([ADR 0195](../../02-DECISIONS/0195-the-meshs-tools-are-found-by-address-not-announced-whole.md)). + It still held the flat catalogue the console announced then, which lacks both the new verbs and the five discovery tools (`mesh_call` among + them) the console announces now. Clearing a session does not reconnect its MCP servers, and the + console never sends a list-changed notice, so nothing told the client its list was stale. The agent + posted `mesh_machine` and `mesh_call` by hand. Reconnecting the server would have given it the + discovery tools, which reach any tool by address the moment it exists. The calls were authorised, because the console is the operator's own surface. But each is a raw call the mesh's tools were meant to make unnecessary ([ADR 0154](../../02-DECISIONS/0154-the-meshs-own-verbs-are-the-controller-seats-tools.md) @@ -54,8 +53,9 @@ rollout needs. around it. - Whether `command`'s generic answer should take a filter, or whether the verbs it is used for most (`module list`, `node show`) deserve verbs of their own. -- **A tool list that follows the mesh.** When a module's tools arrive or leave, the connection says so - (MCP's list-changed notification). Failing that, it always carries the console's addressable - `mesh_call` and `mesh_describe`, so a new verb is reachable without leaving the agent's tools. +- **A client is told when the console's own surface changes.** The console announces `listChanged` + and sends the notice when what it lists changes, for example after an upgrade that changes its + tools. A long-running session then never keeps a list the console no longer serves. Discovery + already makes every module's tools reachable without the list changing. How each is checked belongs to the record that settles it.