From aa221abfee2f1f043526ed1a2428e6db3d7559f6 Mon Sep 17 00:00:00 2001 From: jochen Date: Thu, 8 Oct 2026 13:23:49 +0200 Subject: [PATCH] Issues 316-319: holds that never lift, a hold a later merge walks past, and one module's wait stalling every walk Four delivery defects seen on 2026-10-08, each located, so their decisions can be made before code moves. --- .../00-report.md | 62 +++++++++++++++ .../01-diagnosis.md | 49 ++++++++++++ .../00-report.md | 54 +++++++++++++ .../01-diagnosis.md | 42 ++++++++++ .../00-report.md | 69 ++++++++++++++++ .../01-diagnosis.md | 55 +++++++++++++ .../00-report.md | 78 +++++++++++++++++++ .../01-diagnosis.md | 48 ++++++++++++ 8 files changed, 457 insertions(+) create mode 100644 04-ISSUES/316-a-group-member-held-before-its-composed-check-answered-stayed-held-after-it-passed/00-report.md create mode 100644 04-ISSUES/316-a-group-member-held-before-its-composed-check-answered-stayed-held-after-it-passed/01-diagnosis.md create mode 100644 04-ISSUES/317-a-held-deliverys-module-was-walked-by-a-later-merges-plan/00-report.md create mode 100644 04-ISSUES/317-a-held-deliverys-module-was-walked-by-a-later-merges-plan/01-diagnosis.md create mode 100644 04-ISSUES/318-one-modules-wait-for-a-new-login-held-every-other-modules-walk/00-report.md create mode 100644 04-ISSUES/318-one-modules-wait-for-a-new-login-held-every-other-modules-walk/01-diagnosis.md create mode 100644 04-ISSUES/319-a-held-delivery-a-later-commit-superseded-has-no-end-but-release/00-report.md create mode 100644 04-ISSUES/319-a-held-delivery-a-later-commit-superseded-has-no-end-but-release/01-diagnosis.md diff --git a/04-ISSUES/316-a-group-member-held-before-its-composed-check-answered-stayed-held-after-it-passed/00-report.md b/04-ISSUES/316-a-group-member-held-before-its-composed-check-answered-stayed-held-after-it-passed/00-report.md new file mode 100644 index 00000000..1078438c --- /dev/null +++ b/04-ISSUES/316-a-group-member-held-before-its-composed-check-answered-stayed-held-after-it-passed/00-report.md @@ -0,0 +1,62 @@ +--- +status: located +opened: 2026-10-08 +located-in: [mesh-catalog modules/mesh-delivery (holder.go, groupTurns and Checked; table.go, the rows out of held)] +fixed-by: +amended-design: +--- + +# 316. A group member held before its composed check answered stayed held after it passed + +## Symptom + +On 2026-10-08 the delivery group `feat/module-groups` held two pull requests: the node-engine's +(`mesh-host` #54) and the catalogue's (mesh-catalog #125, the lighting module's account group of +[ADR 0252](../../02-DECISIONS/0252-a-module-puts-an-account-in-a-group-and-the-mesh-says-when-a-new-login-is-needed.md)). +Each passed its own merge check and was merged as soon as it did. The group's composed check had been +asked, but had not answered yet. + +What `mesh-delivery.show` and the holder's journal on the control node said, in UTC: + +| time | what happened | +|---|---| +| 10:06:07 | the catalogue's member `ready` | +| 10:08:11 | the node-engine's member `ready` | +| 10:08:19 | the composed check asked of the controller, for both heads | +| 10:09:23 | the node-engine's pull request merged | +| 10:09:24 | the catalogue's pull request merged | +| 10:09:45 | both `published`; the node-engine's member went straight on to `delivering` (the controller's own path) | +| 10:09:48 | the catalogue's member `held`: "its group feat/module-groups's composed check did not pass for the heads that merged; a person decides that it goes on" | +| 10:10:21 | the holder logged "group feat/module-groups composed: pass — every machine composes with the change as it did without (4 of 4 compose)" | + +An hour later the group's record said `composed: pass` for exactly the two heads that merged, and the +catalogue's member still said `held`, waiting for a person, for a reason that was no longer true. Its +`mesh/delivery` status on the pull request said the same. `mesh-delivery.checks` for mesh-catalog #125 +showed `mesh/delivery-group` with the verdict "composed check pass" beside `mesh/delivery` saying "did not +pass". + +The forge let both pull requests merge: the trunk requires `mesh/merge-gate` and `mesh/repo-check`, not +`mesh/delivery-group`. + +## Why it is a design issue + +A hold is a question to a person. Here the question was asked on a fact that was not yet known ("did not +pass" read from "not answered"), and the answer the mesh received two minutes later never reached the +question. The person is asked to decide something the mesh already decided, and nothing says the hold is +stale: the failure is quiet in the direction that wastes a person's attention and delays a change. + +The two checks a merge reads and the composed check a group's turn reads are also not the same set. So +two members can merge correctly, each by its own rules, and still meet a hold that only a person can +lift. + +## Open questions + +1. Is a composed check that was asked and has not answered a reason to hold, or a reason to wait in + `published` until it answers? +2. When a composed check answers after a member was held for it, does the hold lift by observation, as + every other state in the table moves on its facts? +3. Should `mesh/delivery-group` be required on the trunks, so a group cannot merge before its composed + check passed? To-be 47 leaves this to the operator; this issue is the case that setting decides. +4. The group's order put the catalogue's member first, and the node-engine's member walked at once on the + controller's own path. Is a group's order meant to hold for a member whose walk never waits for + mesh-delivery? diff --git a/04-ISSUES/316-a-group-member-held-before-its-composed-check-answered-stayed-held-after-it-passed/01-diagnosis.md b/04-ISSUES/316-a-group-member-held-before-its-composed-check-answered-stayed-held-after-it-passed/01-diagnosis.md new file mode 100644 index 00000000..b0540d6d --- /dev/null +++ b/04-ISSUES/316-a-group-member-held-before-its-composed-check-answered-stayed-held-after-it-passed/01-diagnosis.md @@ -0,0 +1,49 @@ +# 316 — diagnosis + +**2026-10-08.** Read from mesh-delivery's code on the catalogue's main (`modules/mesh-delivery/cmd/mesh-delivery`), +from `mesh-delivery.show` of both members, and from the holder's journal on the control node. + +**Where the hold is set.** `groupTurns` (holder.go) walks a closed group's members in order. For a member +in `published` it computes `composed`: the group's check exists, its verdict is `pass` or `warning`, and it +names the members' current heads. When `composed` is false it writes the member's `HeldWhy` ("its group's +composed check did not pass for the heads that merged") and settles it, and the table's `hold` row takes it +to `held`. `composed` is false in three different cases, and the code does not tell them apart: + +- the composed check failed; +- it was for other heads; +- it was asked and has not answered (`Verdict` empty). This is the case here: asked at 10:08:19, answered + at 10:10:21, and the catalogue's member reached `published` at 10:09:45, in between. + +**Why nothing lifts it.** Three places would have to act, and none does: + +- `Checked`, for a group's verdict, stores the verdict and summary on the group and returns. It does not + look at the members. +- `groupTurns` skips a member in `held` (`default: return nil`), so the tick after the verdict sees the + passing check and does nothing with it. +- The table has two rows out of `held`: `go` (the walk started on a word that was not this owner's) and + `release` (a person's act). There is no observed row that takes a held delivery back when the reason it + was held is gone, and `HeldWhy` is never cleared once set. + +**What the design says.** ADR 0239 decision 2's row reads "merged, and its group's composed check did not +pass for the heads that merged". To-be 47 says the composed check is asked when the group forms or a +member's head moves. Neither says what a member merged while its group's check is still out should do. +The code read "did not pass" as "has not passed yet". + +**Why the merge was possible.** The trunk's protection requires `mesh/merge-gate` and `mesh/repo-check` +(`mesh-delivery.checks`, mesh-catalog #125: `required`). `mesh/delivery-group` is posted on every member's +head and is not required, so the forge merged each member on its own checks, 64 seconds after the composed +check was asked. To-be 47 lists whether to require it as not decided. + +**Ruled out.** + +- The controller's composed check. It answered `pass` for the very heads that merged (the group's + `check.commits`), in about two minutes, an ordinary time for a check. +- A lost verdict. The holder logged it at 10:10:21 and the group's record holds it. +- The order. The group's order was asked and kept before the check; it plays no part in the hold. + +**Not this issue, noted.** The node-engine's member walked at once (`published` → `delivering` at the +same instant), because the controller's own path does not wait for mesh-delivery (ADR 0239 decision 8). +The group's order put it second. Open question 4 in the report asks whether that is intended. + +**Located** in mesh-delivery: `groupTurns` holds on a composed check that has not answered, and neither +`Checked` nor the table brings a held member back when the check later passes. diff --git a/04-ISSUES/317-a-held-deliverys-module-was-walked-by-a-later-merges-plan/00-report.md b/04-ISSUES/317-a-held-deliverys-module-was-walked-by-a-later-merges-plan/00-report.md new file mode 100644 index 00000000..129de146 --- /dev/null +++ b/04-ISSUES/317-a-held-deliverys-module-was-walked-by-a-later-merges-plan/00-report.md @@ -0,0 +1,54 @@ +--- +status: located +opened: 2026-10-08 +located-in: [mesh-controller cmd/mesh-controller (release_plan.go, supersededBy), mesh-catalog modules/mesh-delivery (table.go, no row out of held for a superseded walk)] +fixed-by: +amended-design: +--- + +# 317. A held delivery's module was walked by a later merge's plan + +## Symptom + +On 2026-10-08 the catalogue's delivery of mesh-catalog #125 (the lighting module `openrazer`, its account +group) was `held` from 10:09:48 UTC, waiting for a person ([issue 316](../316-a-group-member-held-before-its-composed-check-answered-stayed-held-after-it-passed/00-report.md)). +Nobody released it. `mesh-delivery.show` still said `held` at 11:20 UTC, and `mesh-delivery.checks` for +mesh-catalog #125 still said `mesh/delivery` pending, "held for a person". + +Its module reached a machine anyway. Three later merges to the catalogue's main touched other modules: +`docker` (#122), the artifact store's tools (#123) and `gitea` (#124). Each made a plan, and each plan +took `openrazer` along. `mesh-controller.plans`: + +| plan (catalogue commit) | made (UTC) | says | +|---|---|---| +| a6385479 (#125's merge) | 10:09:45 | superseded at tier 0 by the plan at 3f552274; openrazer planned there again | +| 3f552274 (#122's merge) | 11:06:07 | superseded at tier 0 by the plan at c78b5fc9; docker, openrazer planned there again | +| c78b5fc9 (#124's merge) | 11:07:02 | superseded at tier 0 by the plan at 86077780; docker, gitea, openrazer planned there again | +| 86077780 (#123's merge) | 11:07:03 | tier 0: distribution, docker, gitea and openrazer built from c78b5fc9; openrazer's gate on the laptop judging | + +The controller's walk of the builds waiting for a gate then sent the laptop `openrazer` 88135ad0 → +c78b5fc9, beside `docker` and `node-tools`. From 11:10:58 UTC the laptop's condition for openrazer said +its account was *relogin needed*. That verdict exists only in the build #125 added, so the held change was +running on the laptop while its delivery said it waited for a person. + +## Why it is a design issue + +A hold is the mesh's promise that a change goes no further until a person says so +([ADR 0239](../../02-DECISIONS/0239-a-delivery-is-owned-by-the-mesh-delivery-module-and-runs-from-commit-to-delivered.md)). +Any later merge to the same repository and branch broke that promise without a person's word and without +saying so, through a different rule: a newer merge takes over an older plan's unfinished modules +([ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md) +decision 3). The two decisions were made apart, and nothing composes them. + +The delivery's record is now false in both directions: it says the change waits, and the change has +gone out. A person who releases it later releases nothing; a person who stops it stops nothing. + +## Open questions + +1. When a newer merge supersedes an older plan, does a module the older plan's walk was holding for its + delivery's word come along, or stay held with that delivery? +2. If it comes along, what does the held delivery become: superseded by the newer delivery, with the hold + carried over to it, or something else? Today the table has no row from `held` for a superseded walk. +3. A trunk commit holds every earlier commit, so a later build of a module carries the held change in + any case. Is a hold on one delivery meaningful once a later merge of the same repository is walked, + or does a hold have to hold the module on that trunk until it is released? diff --git a/04-ISSUES/317-a-held-deliverys-module-was-walked-by-a-later-merges-plan/01-diagnosis.md b/04-ISSUES/317-a-held-deliverys-module-was-walked-by-a-later-merges-plan/01-diagnosis.md new file mode 100644 index 00000000..d50ed1ed --- /dev/null +++ b/04-ISSUES/317-a-held-deliverys-module-was-walked-by-a-later-merges-plan/01-diagnosis.md @@ -0,0 +1,42 @@ +# 317 — diagnosis + +**2026-10-08.** Read from `mesh-controller.plans` (the list, and each plan by id), `mesh-delivery.show` +and `deliveries`, and the code on the main branches of the controller and the catalogue. + +**The held delivery's walk.** #125's merge made the plan at a6385479. Its walk waited for mesh-delivery's +word (`waited_for: mesh-delivery`, `waits: true`) and was never let go, because its delivery was held. Its +one module, `openrazer`, was "not yet asked": nothing was built. + +**How a later merge took it.** `supersededBy` (the controller, `release_plan.go`) runs when a merge makes a +plan. For every open plan of the same repository and branch made before it, it takes each module the +older plan had not built (not asked, asked and not answered, or built and not yet sent), adds it to the +newer plan, and closes the older plan as superseded with the note "… planned there again". It does not +ask why the older plan had not built them. A walk waiting for a held delivery's word is open, so it is +superseded like any slow plan. This happened three times over (3f552274, c78b5fc9, 86077780), each newer +plan folding in what the one before it held. + +The newest plan's own delivery (#123) was not held, so mesh-delivery let its walk go, and `openrazer` was +built from c78b5fc9. The controller's walk of the builds waiting for a gate then carried it to the laptop, +since a send carries every move waiting on its machine (ADR 0236). + +**Why the held delivery did not notice.** mesh-delivery's table has `superseded` rows only from +`published` and `delivering` ("a newer delivery to the same trunk took over its walk"). From `held` there +are only `go` and `release`. When the walk record said `superseded`, no row applied, so the delivery stayed +`held` with `held_why` unchanged, and its walk shows "superseded … openrazer planned there again". Two of +the three later deliveries (#122, #124) went to `superseded` correctly, through the `published` row. + +**What the design says.** ADR 0218 decision 3 makes a newer merge take over an older plan's unfinished +work, and was written before deliveries could be held. ADR 0239 makes a held delivery wait for a person, +and to-be 47 says that nothing of a published delivery is registered before its walk starts, so that a +send cannot carry it early. Neither says what a supersession does to a module a hold is keeping back. + +**Ruled out.** + +- A person's act. #125's transitions end at `hold`, with no `release`, and its own walk was superseded + while still waiting (`let_go` empty), so no `plans go` started it. +- Issue 295's path, a recorded build carried by a send: here the build was asked by a plan, not left + recorded. + +**Located** in the controller's `supersededBy`, which folds a held walk's modules into a newer plan, and in +mesh-delivery's table, which has no way out of `held` when that happens, so the delivery's state stops +being true. diff --git a/04-ISSUES/318-one-modules-wait-for-a-new-login-held-every-other-modules-walk/00-report.md b/04-ISSUES/318-one-modules-wait-for-a-new-login-held-every-other-modules-walk/00-report.md new file mode 100644 index 00000000..0c15dc5b --- /dev/null +++ b/04-ISSUES/318-one-modules-wait-for-a-new-login-held-every-other-modules-walk/00-report.md @@ -0,0 +1,69 @@ +--- +status: located +opened: 2026-10-08 +located-in: [mesh-controller cmd/mesh-controller (release.go, gatedSend and ungatedIn; gate.go, judgeMoves; the walk of the builds waiting for a gate)] +fixed-by: +amended-design: +--- + +# 318. One module's wait for a new login held every other module's walk + +## Symptom + +On 2026-10-08 at 11:08 UTC the controller opened its walk of the builds waiting for a gate: the laptop +first, then the workstation, then the control node. On the laptop it sent three moves in one send and +judged them together: + +| module | move | +|---|---| +| `docker` | ba6058ab → c78b5fc9 | +| `node-tools` | 9a4140b4 → d4845f7c | +| `openrazer` | 88135ad0 → c78b5fc9 (the held change of [issue 317](../317-a-held-deliverys-module-was-walked-by-a-later-merges-plan/00-report.md)) | + +The laptop's gate said, for its whole ten minutes, "judging (0 of 3 passes; wanting: its unit +openrazer.daemon on the laptop is unhealthy: failed in the account's own service manager (exit-code))". +The module's condition on the laptop said why: the daemon's unit had failed since 10:13 UTC, and the +account was *relogin needed* from 11:10:58 UTC, which +[ADR 0252](../../02-DECISIONS/0252-a-module-puts-an-account-in-a-group-and-the-mesh-says-when-a-new-login-is-needed.md) +expects until the operator logs out and in again. + +Meanwhile three other walks could not send to the rest of the mesh. `mesh-controller.plans` said of the +tool runner's (`novox/mesh-tools` d4845f7c, `node-tools`), the catalogue's (86077780, `docker`) and the +controller's (950562af, `build-agent`): "a build waits there for its gate: … the workstation (…, openrazer +88135ad0 → c78b5fc9) — a walk sends them, one machine at a time, each judged". + +At 11:21:06 UTC the laptop's gate failed: "not healthy within 10m0s of its apply: its unit +openrazer.daemon … is unhealthy", and `openrazer` was put back to 88135ad0 on the laptop +(`build.openrazer..rolled-back`). The walk ended `failed`. After it, all three walks still said +the same: `openrazer` c78b5fc9 waits on the workstation for a gate, so nothing is sent there or to the +machines after it. + +## Why it is a design issue + +The wait was known, and it was a person's: ADR 0252 says a new login holds a delivery *of that module* on +that machine. It held every module on every machine after it instead, for changes that had nothing to do +with it: the tool runner, the container engine and the build agent. The end was worse than the wait. The +gate's bound turned "waiting for a person" into "failed", the build that added the account group was put +back, and the walk that carries the other modules failed with it. A failed walk of the waiting builds +does not open again on its own: a person releases it. + +Three things meet here, each sound alone: + +- a send carries every move waiting on its machine, and the gate judges them all at once + ([ADR 0236](../../02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)); +- no other send may go where a move waits for a gate, so every walk waits for the one that judges it; +- a module's health that waits for a person reads as *not yet healthy*, which fails at the gate's bound. + +## Open questions + +1. Should *relogin needed* count against a gate at all? It says the build did what it should and a person + has one step left. It is neither a pass nor a fault of the build. +2. If it counts, should it hold only that module's move, and let the moves judged beside it pass or fail on + their own? +3. A wait for a person has no bound a machine can meet. Should it end in *failed* and a rollback after ten + minutes? The build put back declares no account, and by ADR 0252 decision 3 the node-engine then gives + back the account group it added, so the login the gate waited for can no longer help. +4. The daemon's unit had failed since 10:13 UTC, almost an hour before the new build reached the laptop. + Should a gate fail a build for a failure that was there before it was sent? +5. While a move waits for a gate on one machine, should walks of unrelated modules still be refused there, + or only those whose send would carry it? diff --git a/04-ISSUES/318-one-modules-wait-for-a-new-login-held-every-other-modules-walk/01-diagnosis.md b/04-ISSUES/318-one-modules-wait-for-a-new-login-held-every-other-modules-walk/01-diagnosis.md new file mode 100644 index 00000000..b51b570b --- /dev/null +++ b/04-ISSUES/318-one-modules-wait-for-a-new-login-held-every-other-modules-walk/01-diagnosis.md @@ -0,0 +1,55 @@ +# 318 — diagnosis + +**2026-10-08.** Read from `mesh-controller.plans` (the list, and the walk of the waiting builds by id) between +11:17 and 11:22 UTC, `mesh-controller.conditions` and `status`, and the controller's code on its main +branch (`cmd/mesh-controller`). + +**One send, one verdict.** `gatedSend` (release.go) sends a machine every move waiting there and hands the +gate all of them as carried moves. `judgeMoves` (gate.go) reads each module on each machine and keeps the +worst reading. Any module *not yet healthy* resets the passes to zero for the whole send: `docker` and +`node-tools` were healthy, and the gate still said "0 of 3 passes" because `openrazer` was not. Past +`gateBound` (ten minutes after the apply) a *not yet healthy* gate fails. `g.Failing` names only the +modules found wanting, so only `openrazer` was put back, but the walk itself ended `failed`. + +**Everything else waits for that verdict.** `ungatedIn` (release.go) refuses any send that is not a person's +and would carry a move no gate has seen, outside the send's own scope. A plan's sends to the rest of its +machines carry the machine's whole declaration, so while `openrazer` c78b5fc9 waited for a gate on the +workstation, the tool runner's, the catalogue's and the controller's walks were each refused there, and +they say so on every retry. They do not fail; they wait for the walk of the waiting builds, which goes one +machine at a time and had stopped on the laptop. After it failed, the release code's rule applies: a +failed walk of the waiting builds stops the next one opening on its own until a person releases it. + +**How *relogin needed* reached the gate.** ADR 0252 decision 5 says the node-engine states the account's +health as the module's resource of kind `account`, unhealthy with *relogin needed*, and that "the +first-node gate reads it like any other". The module's health reading (`judgeHealth`) turns an unhealthy +resource into *not yet healthy*. There is no reading for "waits for a person". The nearest one, +*waiting* (a module waiting on an unhealthy provider), neither passes nor fails at the bound, but it is +kept for providers (ADR 0240 rule 5). + +Here the gate's words named the unit, not the account: the daemon's unit had failed since 10:13 UTC on +both the laptop and the workstation, before the new build reached either (issue 315 records the same daemon +failing at every login). I think the unit fails for the same reason the account is *relogin needed*: the +running session lacks the group the daemon checks for. That is a guess, not checked. Either way, both +readings existed before or because of the person's step, and neither is a fault the build brought. + +**What the design says.** ADR 0252's consequences say "a new login holds a delivery of that module on that +machine" and accept that a walk that starts on a workstation waits for the person. They do not consider the +other modules carried in the same send, the walks refused behind it, or the gate's ten-minute bound turning +the wait into a rollback. ADR 0236 makes one send judge everything waiting on a machine, to stop unjudged +builds from moving; it does not distinguish a module waiting for a person from one that is broken. + +**What the rollback does to the account group.** By ADR 0252 decision 3, an account resource that is no +longer declared gives back the groups the node-engine added. The build put back (88135ad0) declares no +account, and after the rollback the laptop's condition for `openrazer` named only the unit, not the account. +Whether the account actually left the group was not checked. + +**Ruled out.** + +- A fault in `docker` or `node-tools`: both read healthy at every judging; they were held only by the + verdict they shared. +- A machine-wide fault on the laptop: its node-engine reported `applied` and current, and the gate's + wanting named one module's unit. + +**Located** in the controller's walk and first-node gate: one verdict per send (`judgeMoves`), sends refused +on any machine where a move waits (`ungatedIn`), and no reading for a health that waits for a person, so a +known wait fails at the bound and stops every walk behind it. diff --git a/04-ISSUES/319-a-held-delivery-a-later-commit-superseded-has-no-end-but-release/00-report.md b/04-ISSUES/319-a-held-delivery-a-later-commit-superseded-has-no-end-but-release/00-report.md new file mode 100644 index 00000000..968d347f --- /dev/null +++ b/04-ISSUES/319-a-held-delivery-a-later-commit-superseded-has-no-end-but-release/00-report.md @@ -0,0 +1,78 @@ +--- +status: located +opened: 2026-10-08 +located-in: [mesh-catalog modules/mesh-delivery (table.go, the rows out of held and its bound; verbs.go, retire-history, close and stop), 03-DESIGN/01-to-be/47-delivery-from-commit-to-delivered.md (the state table)] +fixed-by: +amended-design: +--- + +# 319. A held delivery that a later commit superseded has no end but release + +## Symptom + +On 2026-10-08 at 11:20 UTC sixteen deliveries were `held` for a person. `mesh-delivery.retire-history`, +run as a dry run, retired none of them and gave the same reason for each: "retire from held is refused — +its guard does not hold: this owner heard its head: it is not history made from the forge's word". + +What they are, read with `mesh-delivery.deliveries` and `show`: + +| how many | repository | why held | what it builds | +|---|---|---|---| +| 4 | hq | merged without a passing check | nothing: "the change touches no module of the mesh's graph" | +| 2 | catalogue | no walk opened within ten minutes of the merge | a new module, sent nowhere | +| 8 | controller | no walk opened within ten minutes of the merge | the controller, the build agent, the route proxy | +| 2 | catalogue | its group's composed check did not pass for the heads that merged ([issue 316](../316-a-group-member-held-before-its-composed-check-answered-stayed-held-after-it-passed/00-report.md)) | one module each | + +Two examples: + +- `novox/hq@055550802096`, the pull request that opened issue 281. Merged 2026-10-06 20:26 UTC; its head was + announced to mesh-delivery at 23:05 UTC, at the owner's first start, and it went straight from + `proposed` to `held`: "merged without a passing check". Its plan builds nothing. +- `novox/mesh-controller@ec3769a8a669`, controller #105. Ready at 2026-10-07 00:26 UTC, merged at 00:27, + held at 00:37: "merged, and the controller opened no walk for it within 10m0s". The controller's modules + have been walked since from later commits of the same main, for example the group + `feat/delivery-checks-verb`, whose controller member is `delivered`. + +Five of them (the four hq merges and controller #105) are past the held bound of one day. The self-check's +probe D14 raises each as `delivery..stalled`, "healer H2 may none: the state is the operator's". They +were five of the ten conditions open that hour. + +## Why it is a design issue + +A held delivery asks a person "does this go on?". For these sixteen the question has no meaningful answer: +a later commit of the same trunk has already delivered what they held, or they hold nothing to deliver. The +state table offers a person four ways out of `held`, and none fits: + +- `release` starts a walk of the old commit: wrong for a commit a later one superseded, and nothing for one + that builds nothing. +- `retire-history` is for history made from the forge's word ([ADR 0250](../../02-DECISIONS/0250-a-merge-heard-only-from-the-buss-history-is-retired-and-owed-work-the-forge-cannot-do-ends.md)), + and refuses every one of these because their heads were heard. +- `close`, healer H2's verb, has no transition from `held`: the bound names none. +- `stop` is allowed from `held`, but it says a person stopped a delivery and sets `mesh/delivery` to error + on the commit: the wrong record for a change that did reach the machines through a later commit + (issue 313 ruled it out for the same reason). + +So the condition stays open for ever, and each new case adds one. A condition nobody can clear correctly +teaches people to ignore the list it is in. + +## What is needed + +The issue states the need; the decision is for graduation. + +- **An end for a held delivery that was superseded**, as `published` and `delivering` already have + (`superseded`, "a newer delivery to the same trunk took over its walk"): either because a later commit's + delivered walk carried the same modules on the same trunk, or because it builds nothing at all. +- **Automatic where nothing is lost**: a held delivery whose plan builds nothing has nothing for a person to + decide. +- **A person's act, with why, otherwise**: a held delivery whose modules a later commit delivered is ended + by a person who says so, and recorded as superseded, not stopped. + +## Open questions + +1. Is "a later commit delivered every module this one moves" something mesh-delivery can read from its own + records, or does it need the controller to say which build each machine runs? +2. The four hq merges were heard at the owner's first start from the bus's history, as announcements rather + than merges. Are they history in the sense of ADR 0250, missed by its rule because their head was + replayed too? +3. Why did the controller open no walk for eight of its own merges within ten minutes? The controller's own + path never waits, so its walk may have been missed rather than absent. Not diagnosed here. diff --git a/04-ISSUES/319-a-held-delivery-a-later-commit-superseded-has-no-end-but-release/01-diagnosis.md b/04-ISSUES/319-a-held-delivery-a-later-commit-superseded-has-no-end-but-release/01-diagnosis.md new file mode 100644 index 00000000..9a4a2de4 --- /dev/null +++ b/04-ISSUES/319-a-held-delivery-a-later-commit-superseded-has-no-end-but-release/01-diagnosis.md @@ -0,0 +1,48 @@ +# 319 — diagnosis + +**2026-10-08.** Read from mesh-delivery's code on the catalogue's main (`modules/mesh-delivery/cmd/mesh-delivery`), +`mesh-delivery.retire-history` (dry), `deliveries`, `show`, and `mesh-controller.conditions`. + +**The rows out of `held`** (table.go), each with what it needs: + +| event | kind | needs | +|---|---|---| +| `go` | observed | its walk started on a word not this owner's (the controller's own path, or a person's `plans go`) | +| `release` | a person's act | why and who; it walks the delivery's own commit | +| `retire` | a person's act | `isHistory`: made from the forge's word, merge heard 15 minutes or more after it happened, no walk open | +| `stop` | a person's act | why and who; from any state not final | + +There is no `superseded` row from `held`. `published` and `delivering` have one, guarded by the walk's +record saying `superseded`. A held delivery whose walk never opened has no walk record to read, and one +whose walk was superseded ([issue 317](../317-a-held-deliverys-module-was-walked-by-a-later-merges-plan/00-report.md)) +does not match it either. + +**Why `retire-history` refuses them.** `isHistory` requires `FromForgesWord()`: the delivery was made from +the forge's `pull.merged` because its head was never heard. All sixteen had their heads announced. The four +hq merges were announced at 23:05:28 UTC on 2026-10-06, the owner's first start, hours after their +merge, so their heads came from the bus's history as much as the 688 merges issue 313 retired. The rule +does not see that, because it reads how the delivery was made, not when. + +**Why `close` refuses them.** `Close` reads the state's bound. `held`'s bound is one day with no H2 +events ("it waits for the operator"), so it answers "the state is the operator's". + +**Why a delivery that builds nothing is held.** `merged-unchecked` (proposed → held) is tried before +anything about the plan. The row `published` → `delivered` "nothing for a walk to move" exists, but a +delivery merged unchecked never reaches `published`. Its plan was known (`builds nothing`), and nothing +reads it once held. + +**The stalled conditions.** D14 lists every delivery past its state's bound. For `held`, H2 may do nothing, +so the condition's resolver is `operator`, and it clears only when the delivery leaves `held`. None of the +person's acts is a correct way out, so the condition has no correct end. + +**What the design says.** To-be 47's state table gives `held` a bound of one day and "nothing: it is the +operator's" after it. ADR 0239 decision 2's table has `superseded` only from `published` and `delivering`. +Neither considers a held delivery whose content a later commit delivered. + +**Ruled out.** + +- A missing release. Releasing any of these would walk an old commit over a newer one, or walk nothing. +- D14 itself. It reports the state truthfully; the state has no way to change. + +**Located** in mesh-delivery's state table (no end for a superseded or empty held delivery, and history +read by how a delivery was made) and in to-be 47, whose state table carries the same gap.