diff --git a/04-ISSUES/281-a-tier-sent-one-module-at-a-time-blamed-a-module-for-its-machine/00-report.md b/04-ISSUES/281-a-tier-sent-one-module-at-a-time-blamed-a-module-for-its-machine/00-report.md new file mode 100644 index 00000000..02b75e8e --- /dev/null +++ b/04-ISSUES/281-a-tier-sent-one-module-at-a-time-blamed-a-module-for-its-machine/00-report.md @@ -0,0 +1,90 @@ +--- +status: located +opened: 2026-10-06 +located-in: [mesh-controller cmd/mesh-controller] +fixed-by: mesh-controller PR #100 +amended-design: +--- + +# 281. A tier sent one module at a time blamed a module for its machine + +## Symptom + +On 2026-10-06 a catalogue merge rebuilt about seventy modules in one tier. The plan sent the laptop, +the first machine for most of them, **one module at a time**: a separate send every few seconds, each +logged as "tier 1 built; sent X to the laptop first, the rest once it reports it applied". The +laptop's node-engine applied a new declaration every ten to twenty seconds and set several aside ("a +newer one arrived with it"). The controller's version-split probe raised `core-behind` for the laptop +("runs the node tools it was last sent, not yet reported applied"), and the workstation showed it too. + +Thirteen minutes after its send, the gate failed one desktop module on the laptop with "raised since it +was sent: machine.laptop.core-behind". That condition was about the machine and had nothing to do with +the module. The gate then could not put the module back: "NOT put back: no build of it from is still kept". It raised `rollback-failed`, urgent. The plan stopped at its last tier with +about sixty built modules never sent. + +## Diagnosis + +**A send per module.** A send always carries the machine's whole declaration +([ADR 0221](../../02-DECISIONS/0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md)). The plan's tier loop still made one gated send per module to +its first machine, then returned and took the next module on the next step. Seventy modules made +seventy whole declarations for one machine within a few minutes. The machine's reports never caught up +with the newest one, which is exactly what the probe's `core-behind` says. The churn caused the +condition. + +**A machine-level condition read as the module's.** The gate's health check counted every +machine-scope condition raised on the judged machine since the send against the module it was judging. +It did not ask whether the condition named the module, or whether something else in the send caused +it. Each module had a gate of its own, so the first one to reach the ten-minute bound took the blame. + +**Why there was "no build kept".** The plan reads the build the first machine ran before from that +machine's record of what it was last sent, at the moment of the module's own send. An earlier send of +the same tier had already carried this module's new build to the laptop, because a send carries +everything that waits there. So the plan recorded the move as "new build to new build" and searched for +an earlier build *of the new commit* to put back. The build the laptop really ran before was kept all +along. Its artifact was byte-identical to the new one, so nothing of the module had changed on the +laptop at any point. A gate failed a build that no send had moved, and then looked for its predecessor +under the wrong commit. + +Ruled out: build retention. The previous build was within the kept builds, and this was not the +module's first build on the machine. + +## Fix + +- **One send per machine per tier.** Every module of a tier whose first machine is the same goes there + in one gated send. What the machine ran before is read once, before that send. One gate judges the + whole send. The rest of the tier's machines are sent once each, together, after the gate passes. + What a failed send puts back goes back in one send per machine. A retried rollout sends each machine + once. +- **A machine-level condition is the machine's.** The gate sorts conditions about the machine itself: + - a condition naming a module the send moved belongs to that module; + - `core-behind` belongs to the node-engine or the node tools when the send moved them, and they + get the gate's bound as their settle window. When the send did not move them, it is no evidence + about what was sent, because the reports already say whether the send was applied; + - anything else holds the machine back as a whole. Everything the send moved there waits on it and, + at the bound, fails together, once, with " as a whole: …" as the reason. + + The module a gate is kept on is put back only when it is among those found wanting. +- **A send that changed nothing of a module is no verdict on its build.** This covers the same commit, + the same source, or the same artifacts. Such a module is left as it was, is not marked failed, and + raises no rollback condition. A first build on a machine says so plainly: there is nothing earlier + to put back. `plans retry` takes a plan stopped at such a gate, asks the module again under a new + id, and sends the tier's unfinished first sends again as one. + +## Left open + +- The version-split probe still raises `core-behind` for the moment between a send and its report. + The gate no longer reads that as evidence, but `status` still shows it for that moment. +- Choosing first machines: each module still takes the first machine heard from that runs it. A tier + whose modules spread over several first machines sends each of those machines once, one after the + other. + +## How it is checked + +| Rule | Checked by | +|---|---| +| a tier is sent to each machine once | the controller's test: a tier of twelve modules on two machines makes exactly two sends, one per machine, judged by one gate, with every build's pass kept | +| the node tools settling fail no other module | the controller's test: `core-behind` raised while the node tools, sent in the same tier, settle; at the bound only the node tools are put back, and the other modules stay at the new build, unmarked | +| a machine unhealthy as a whole fails its send together | the controller's test: any other machine-level condition fails every module of the send with "as a whole", put back in one send | +| a send that changed nothing is no verdict | the controller's test: a module already carried to its first machine is left as it was, unmarked, with no rollback condition, and `plans retry` asks it again | +| live | the next merge that rebuilds a wide tier logs one "sent N module(s) to first in one send" line per first machine, and no gate fails for `core-behind` |