diff --git a/04-ISSUES/230-a-host-that-hands-over-to-a-newer-one-loses-its-report-and-a-plan-waits-for-ever/00-report.md b/04-ISSUES/230-a-host-that-hands-over-to-a-newer-one-loses-its-report-and-a-plan-waits-for-ever/00-report.md new file mode 100644 index 0000000..7fce3f3 --- /dev/null +++ b/04-ISSUES/230-a-host-that-hands-over-to-a-newer-one-loses-its-report-and-a-plan-waits-for-ever/00-report.md @@ -0,0 +1,80 @@ +--- +status: open +opened: 2026-10-04 +located-in: + - mesh-host + - mesh-controller +fixed-by: +amended-design: +--- + +# 230 — A host that hands over to a newer one loses its report, and a plan waits for it for ever without saying so + +## What was observed + +2026-10-04, rolling out to-be 41. A new host build and a new controller were merged together. The +controller's plan built build-agent and sent every machine a fresh declaration, which also delivered +the new host. Three of the four machines logged, within the same second: + +``` +host 4bd7df099757 is delivered; standing aside so the launcher runs it +applied 546 resource(s) +applied, and could not tell the mesh: reporting: context canceled +nox-mesh-host-launch: running /usr/lib/nox-mesh-host/versions/4bd7df099757/nox-mesh-host +``` + +The new host came up and waited for its next declaration. The mesh never heard that the old one had +applied. + +The plan then sat at "tier 1 built; waiting for build-agent on [three machines] to be applied", and +everything the mesh said about it read as healthy: + +- `status` showed it as `rolling` with `"late": false`; +- `plans` printed "for 0s" on every look, so the wait never appeared to grow; +- the three machines' reports showed `current: false`, which reads like a machine that is merely slow. + +Nothing logged, alerted or counted the wait. It was found because a person asked twice for the plan's +state, and the cause was found by reading a machine's own journal. A push to each of the three machines +released it: each new host applied and reported, and the plan moved on. + +## Why it matters beyond this instance + +**Every genuine host upgrade loses one report.** +[Issue 163](../163-a-delivered-host-stood-aside-on-every-push-and-reported-nothing/00-report.md) fixed +the host that stood aside on every push for the version it already ran. It named the mechanism, that +standing aside cancels the context the report is published with. That fix made standing aside +happen only for a real new version, but left the mechanism in place. So whenever a host build reaches +a machine, that apply's report is lost. + +**And the mesh cannot tell a stuck wait from a slow one.** A plan that waits on a report that will +never come waits for ever, and nothing about it changes: + +- its age does not grow ("for 0s"); +- `late` stays false; +- nothing logs, emits an event or alerts. + +This is [issue 187](../187-the-mesh-tells-nobody-when-it-stops-working/00-report.md)'s class of fault +again, *the mesh tells nobody when it stops working*, now in the rollout machinery that every merge +goes through. The operator's rule from issue 163 applies: if an answer has not come in the time an +answer takes, something is wrong, and the mesh must say so itself. + +## What a fix has to settle + +1. **The host reports before it stands aside.** The report should be published and acknowledged + before the old host exits. Failing that, the new host should report the declaration it took over, + naming the apply its predecessor finished. A lost report must be impossible, not merely unlikely. +2. **A plan's wait has an age and a bound.** + - "for 0s" must be the real time since the wait began. + - A wait past a bound, set by how long an apply takes rather than by a guess, makes the plan + `late`. +3. **Late is said where people and agents look.** + - It is said in `status` and in `plans`. + - It is logged as a warning by the controller. + - It is emitted as an event under the controller seat, so something can alert on it. +4. **A plan waiting on a machine the mesh has stopped hearing from** says that, by name, instead of + waiting. The machine's heartbeat already tells the controller it is alive. A live machine with an + unacknowledged declaration is the stuck case itself. + +How each is checked belongs to the fix. For the host: a delivered upgrade, applied, is reported. For +the controller: a plan whose machine never reports turns `late` within its bound, and says so in +`status`, the log and an event.