Issue 281: a tier sent one module at a time blamed a module for its machine
mesh/merge-gate the mesh is checking this head against every machine
mesh/delivery delivered
mesh/merge-gate the mesh is checking this head against every machine
mesh/delivery delivered
This commit is contained in:
+90
@@ -0,0 +1,90 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-06
|
||||
located-in: [mesh-controller cmd/mesh-controller]
|
||||
fixed-by: mesh-controller PR #100
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 281. A tier sent one module at a time blamed a module for its machine
|
||||
|
||||
## Symptom
|
||||
|
||||
On 2026-10-06 a catalogue merge rebuilt about seventy modules in one tier. The plan sent the laptop,
|
||||
the first machine for most of them, **one module at a time**: a separate send every few seconds, each
|
||||
logged as "tier 1 built; sent X to the laptop first, the rest once it reports it applied". The
|
||||
laptop's node-engine applied a new declaration every ten to twenty seconds and set several aside ("a
|
||||
newer one arrived with it"). The controller's version-split probe raised `core-behind` for the laptop
|
||||
("runs the node tools it was last sent, not yet reported applied"), and the workstation showed it too.
|
||||
|
||||
Thirteen minutes after its send, the gate failed one desktop module on the laptop with "raised since it
|
||||
was sent: machine.laptop.core-behind". That condition was about the machine and had nothing to do with
|
||||
the module. The gate then could not put the module back: "NOT put back: no build of it from <the new
|
||||
commit> is still kept". It raised `rollback-failed`, urgent. The plan stopped at its last tier with
|
||||
about sixty built modules never sent.
|
||||
|
||||
## Diagnosis
|
||||
|
||||
**A send per module.** A send always carries the machine's whole declaration
|
||||
([ADR 0221](../../02-DECISIONS/0221-a-push-sends-no-build-a-policy-or-a-plan-holds-back-except-to-the-machine-it-names.md)). The plan's tier loop still made one gated send per module to
|
||||
its first machine, then returned and took the next module on the next step. Seventy modules made
|
||||
seventy whole declarations for one machine within a few minutes. The machine's reports never caught up
|
||||
with the newest one, which is exactly what the probe's `core-behind` says. The churn caused the
|
||||
condition.
|
||||
|
||||
**A machine-level condition read as the module's.** The gate's health check counted every
|
||||
machine-scope condition raised on the judged machine since the send against the module it was judging.
|
||||
It did not ask whether the condition named the module, or whether something else in the send caused
|
||||
it. Each module had a gate of its own, so the first one to reach the ten-minute bound took the blame.
|
||||
|
||||
**Why there was "no build kept".** The plan reads the build the first machine ran before from that
|
||||
machine's record of what it was last sent, at the moment of the module's own send. An earlier send of
|
||||
the same tier had already carried this module's new build to the laptop, because a send carries
|
||||
everything that waits there. So the plan recorded the move as "new build to new build" and searched for
|
||||
an earlier build *of the new commit* to put back. The build the laptop really ran before was kept all
|
||||
along. Its artifact was byte-identical to the new one, so nothing of the module had changed on the
|
||||
laptop at any point. A gate failed a build that no send had moved, and then looked for its predecessor
|
||||
under the wrong commit.
|
||||
|
||||
Ruled out: build retention. The previous build was within the kept builds, and this was not the
|
||||
module's first build on the machine.
|
||||
|
||||
## Fix
|
||||
|
||||
- **One send per machine per tier.** Every module of a tier whose first machine is the same goes there
|
||||
in one gated send. What the machine ran before is read once, before that send. One gate judges the
|
||||
whole send. The rest of the tier's machines are sent once each, together, after the gate passes.
|
||||
What a failed send puts back goes back in one send per machine. A retried rollout sends each machine
|
||||
once.
|
||||
- **A machine-level condition is the machine's.** The gate sorts conditions about the machine itself:
|
||||
- a condition naming a module the send moved belongs to that module;
|
||||
- `core-behind` belongs to the node-engine or the node tools when the send moved them, and they
|
||||
get the gate's bound as their settle window. When the send did not move them, it is no evidence
|
||||
about what was sent, because the reports already say whether the send was applied;
|
||||
- anything else holds the machine back as a whole. Everything the send moved there waits on it and,
|
||||
at the bound, fails together, once, with "<machine> as a whole: …" as the reason.
|
||||
|
||||
The module a gate is kept on is put back only when it is among those found wanting.
|
||||
- **A send that changed nothing of a module is no verdict on its build.** This covers the same commit,
|
||||
the same source, or the same artifacts. Such a module is left as it was, is not marked failed, and
|
||||
raises no rollback condition. A first build on a machine says so plainly: there is nothing earlier
|
||||
to put back. `plans retry` takes a plan stopped at such a gate, asks the module again under a new
|
||||
id, and sends the tier's unfinished first sends again as one.
|
||||
|
||||
## Left open
|
||||
|
||||
- The version-split probe still raises `core-behind` for the moment between a send and its report.
|
||||
The gate no longer reads that as evidence, but `status` still shows it for that moment.
|
||||
- Choosing first machines: each module still takes the first machine heard from that runs it. A tier
|
||||
whose modules spread over several first machines sends each of those machines once, one after the
|
||||
other.
|
||||
|
||||
## How it is checked
|
||||
|
||||
| Rule | Checked by |
|
||||
|---|---|
|
||||
| a tier is sent to each machine once | the controller's test: a tier of twelve modules on two machines makes exactly two sends, one per machine, judged by one gate, with every build's pass kept |
|
||||
| the node tools settling fail no other module | the controller's test: `core-behind` raised while the node tools, sent in the same tier, settle; at the bound only the node tools are put back, and the other modules stay at the new build, unmarked |
|
||||
| a machine unhealthy as a whole fails its send together | the controller's test: any other machine-level condition fails every module of the send with "as a whole", put back in one send |
|
||||
| a send that changed nothing is no verdict | the controller's test: a module already carried to its first machine is left as it was, unmarked, with no rollback condition, and `plans retry` asks it again |
|
||||
| live | the next merge that rebuilds a wide tier logs one "sent N module(s) to <machine> first in one send" line per first machine, and no gate fails for `core-behind` |
|
||||
Reference in New Issue
Block a user