diff --git a/04-ISSUES/257-a-plan-waited-on-a-declaration-its-first-machine-never-reported/00-report.md b/04-ISSUES/257-a-plan-waited-on-a-declaration-its-first-machine-never-reported/00-report.md new file mode 100644 index 0000000..e9b01f1 --- /dev/null +++ b/04-ISSUES/257-a-plan-waited-on-a-declaration-its-first-machine-never-reported/00-report.md @@ -0,0 +1,49 @@ +--- +status: diagnosing +opened: 2026-10-05 +located-in: [] +fixed-by: +amended-design: +--- + +# 257. A plan waited on a declaration its first machine never reported + +## Symptom + +A merge to the controller's repository produced a plan of three tiers, the controller first. The +controller was built and sent to its first machine, the anchor, which is also the machine holding the +bus. The plan then said, for five minutes and with no end in sight: + +> tier 1 of 3, tier 0 built; waiting for mesh-controller on the anchor, sent first at 20:22 to report +> it applied before the rest are sent + +The anchor had applied it. Its host's journal said `applied 470 resource(s)` ten seconds after the +send, the new controller was running, and the controller logged the report. + +## Evidence + +The plan waits for a report that names the declaration last sent (ADR 0218, one machine first; the +report's `declared` digest against the machine's recorded `sent` digest). Read from the store: + +| | digest | at | +|---|---|---| +| recorded as sent to the anchor | `09f4d6…` | 20:22:57.699 | +| the anchor's report, `declared` | `39b48a…` | 20:23:07.718 | + +The anchor applied and reported a declaration that is not the one the controller recorded sending. +So the report never counted, and the plan would have waited until its bound and then failed the +rollout at its first machine, for a machine that had applied correctly. + +The other three machines' digests matched their reports at the same moment. + +## What unblocked it + +A push to the anchor, made for another reason (its bus grants), recorded a new send. The anchor's +next report named it, the digests matched, and the plan moved on at its next tick and finished. + +## What it costs + +A plan that rolls out the controller itself, to the machine that holds the bus, can stop at its first +machine with nothing wrong there. The plan's words are true but useless: they name a machine that +already did what was asked. Merges made close together, from several sessions, are the case the +plans exist for, and a controller change is among them. diff --git a/04-ISSUES/257-a-plan-waited-on-a-declaration-its-first-machine-never-reported/01-diagnosis.md b/04-ISSUES/257-a-plan-waited-on-a-declaration-its-first-machine-never-reported/01-diagnosis.md new file mode 100644 index 0000000..870bce6 --- /dev/null +++ b/04-ISSUES/257-a-plan-waited-on-a-declaration-its-first-machine-never-reported/01-diagnosis.md @@ -0,0 +1,31 @@ +# Diagnosis + +## 2026-10-05 + +**What is known.** The digest the controller recorded for the anchor (`09f4d6…`, at 20:22:57.699) is +not the digest of what the anchor applied (`39b48a…`). The anchor's host logged one apply that +updated the controller, starting at 20:23:02 and reporting at 20:23:07. The controller that sent it +was replaced at 20:23:04 by that same apply. Its successor logged two reports from the anchor, both at +20:23:07. + +**Ruled out.** + +- The catalogue's catch-up: that is the catalogue asking the controller, not a machine asking for its + declaration. +- The host's five-minute check at 20:22:58: it applied nothing new (only "kept" lines). The apply + that moved the controller is the one at 20:23:02. +- A digest computed differently by the host and the controller: the other three machines matched, + and after the push at 20:28:00 the anchor matched too. + +**Not yet confirmed: two sends to one machine in one step.** The anchor is both the plan's first +machine and the machine holding the bus. Issue 249 sends the machine holding the bus before the first +machine when its user list must change, and this controller change added bus grants. If the anchor +was sent twice in that step, each with its own digest, then the host applying the later one while the +earlier one's record landed last would give exactly this picture. The next step is to read +`sendToEach` for the order of its sends and of their `recordSent` calls, and to check the node's +sequence numbers for two sends at 20:22:57. + +**Whatever the cause, the plan's wait has a second fault.** A first machine that reports `applied` for +a declaration the controller cannot place says nothing about the build. The plan should say "it +applied something other than what was recorded as sent" instead of "waiting", so a reader acts in +minutes rather than at the bound.