Merge pull request 'Issue 281: a tier sent one module at a time blamed a module for its machine' (#154) from issues/281-a-tier-sent-one-module-at-a-time-blamed-a-module-for-its-machine into main

This commit was merged in pull request #154.
This commit is contained in:
2026-10-06 20:26:05 +00:00
@@ -0,0 +1,90 @@
---
status: located
opened: 2026-10-06
located-in: [mesh-controller cmd/mesh-controller]
fixed-by: mesh-controller PR #100
amended-design:
---
# 281. A tier sent one module at a time blamed a module for its machine
## Symptom
On 2026-10-06 a catalogue merge rebuilt about seventy modules in one tier. The plan sent the laptop,
the first machine for most of them, **one module at a time**: a separate send every few seconds, each
logged as "tier 1 built; sent X to the laptop first, the rest once it reports it applied". The
laptop's node-engine applied a new declaration every ten to twenty seconds and set several aside ("a
newer one arrived with it"). The controller's version-split probe raised `core-behind` for the laptop
("runs the node tools it was last sent, not yet reported applied"), and the workstation showed it too.
Thirteen minutes after its send, the gate failed one desktop module on the laptop with "raised since it
was sent: machine.laptop.core-behind". That condition was about the machine and had nothing to do with
the module. The gate then could not put the module back: "NOT put back: no build of it from <the new
commit> is still kept". It raised `rollback-failed`, urgent. The plan stopped at its last tier with
about sixty built modules never sent.
## Diagnosis
**A send per module.** A send always carries the machine's whole declaration
([ADR 0221](../../02-DECISIONS/0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md)). The plan's tier loop still made one gated send per module to
its first machine, then returned and took the next module on the next step. Seventy modules made
seventy whole declarations for one machine within a few minutes. The machine's reports never caught up
with the newest one, which is exactly what the probe's `core-behind` says. The churn caused the
condition.
**A machine-level condition read as the module's.** The gate's health check counted every
machine-scope condition raised on the judged machine since the send against the module it was judging.
It did not ask whether the condition named the module, or whether something else in the send caused
it. Each module had a gate of its own, so the first one to reach the ten-minute bound took the blame.
**Why there was "no build kept".** The plan reads the build the first machine ran before from that
machine's record of what it was last sent, at the moment of the module's own send. An earlier send of
the same tier had already carried this module's new build to the laptop, because a send carries
everything that waits there. So the plan recorded the move as "new build to new build" and searched for
an earlier build *of the new commit* to put back. The build the laptop really ran before was kept all
along. Its artifact was byte-identical to the new one, so nothing of the module had changed on the
laptop at any point. A gate failed a build that no send had moved, and then looked for its predecessor
under the wrong commit.
Ruled out: build retention. The previous build was within the kept builds, and this was not the
module's first build on the machine.
## Fix
- **One send per machine per tier.** Every module of a tier whose first machine is the same goes there
in one gated send. What the machine ran before is read once, before that send. One gate judges the
whole send. The rest of the tier's machines are sent once each, together, after the gate passes.
What a failed send puts back goes back in one send per machine. A retried rollout sends each machine
once.
- **A machine-level condition is the machine's.** The gate sorts conditions about the machine itself:
- a condition naming a module the send moved belongs to that module;
- `core-behind` belongs to the node-engine or the node tools when the send moved them, and they
get the gate's bound as their settle window. When the send did not move them, it is no evidence
about what was sent, because the reports already say whether the send was applied;
- anything else holds the machine back as a whole. Everything the send moved there waits on it and,
at the bound, fails together, once, with "<machine> as a whole: …" as the reason.
The module a gate is kept on is put back only when it is among those found wanting.
- **A send that changed nothing of a module is no verdict on its build.** This covers the same commit,
the same source, or the same artifacts. Such a module is left as it was, is not marked failed, and
raises no rollback condition. A first build on a machine says so plainly: there is nothing earlier
to put back. `plans retry` takes a plan stopped at such a gate, asks the module again under a new
id, and sends the tier's unfinished first sends again as one.
## Left open
- The version-split probe still raises `core-behind` for the moment between a send and its report.
The gate no longer reads that as evidence, but `status` still shows it for that moment.
- Choosing first machines: each module still takes the first machine heard from that runs it. A tier
whose modules spread over several first machines sends each of those machines once, one after the
other.
## How it is checked
| Rule | Checked by |
|---|---|
| a tier is sent to each machine once | the controller's test: a tier of twelve modules on two machines makes exactly two sends, one per machine, judged by one gate, with every build's pass kept |
| the node tools settling fail no other module | the controller's test: `core-behind` raised while the node tools, sent in the same tier, settle; at the bound only the node tools are put back, and the other modules stay at the new build, unmarked |
| a machine unhealthy as a whole fails its send together | the controller's test: any other machine-level condition fails every module of the send with "as a whole", put back in one send |
| a send that changed nothing is no verdict | the controller's test: a module already carried to its first machine is left as it was, unmarked, with no rollback condition, and `plans retry` asks it again |
| live | the next merge that rebuilds a wide tier logs one "sent N module(s) to <machine> first in one send" line per first machine, and no gate fails for `core-behind` |