Issues 316-319: holds that never lift, a hold a later merge walks past, and one module's wait stalling every walk
Four delivery defects seen on 2026-10-08, each located, so their decisions can be made before code moves.
This commit is contained in:
+62
@@ -0,0 +1,62 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-08
|
||||
located-in: [mesh-catalog modules/mesh-delivery (holder.go, groupTurns and Checked; table.go, the rows out of held)]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 316. A group member held before its composed check answered stayed held after it passed
|
||||
|
||||
## Symptom
|
||||
|
||||
On 2026-10-08 the delivery group `feat/module-groups` held two pull requests: the node-engine's
|
||||
(`mesh-host` #54) and the catalogue's (mesh-catalog #125, the lighting module's account group of
|
||||
[ADR 0252](../../02-DECISIONS/0252-a-module-puts-an-account-in-a-group-and-the-mesh-says-when-a-new-login-is-needed.md)).
|
||||
Each passed its own merge check and was merged as soon as it did. The group's composed check had been
|
||||
asked, but had not answered yet.
|
||||
|
||||
What `mesh-delivery.show` and the holder's journal on the control node said, in UTC:
|
||||
|
||||
| time | what happened |
|
||||
|---|---|
|
||||
| 10:06:07 | the catalogue's member `ready` |
|
||||
| 10:08:11 | the node-engine's member `ready` |
|
||||
| 10:08:19 | the composed check asked of the controller, for both heads |
|
||||
| 10:09:23 | the node-engine's pull request merged |
|
||||
| 10:09:24 | the catalogue's pull request merged |
|
||||
| 10:09:45 | both `published`; the node-engine's member went straight on to `delivering` (the controller's own path) |
|
||||
| 10:09:48 | the catalogue's member `held`: "its group feat/module-groups's composed check did not pass for the heads that merged; a person decides that it goes on" |
|
||||
| 10:10:21 | the holder logged "group feat/module-groups composed: pass — every machine composes with the change as it did without (4 of 4 compose)" |
|
||||
|
||||
An hour later the group's record said `composed: pass` for exactly the two heads that merged, and the
|
||||
catalogue's member still said `held`, waiting for a person, for a reason that was no longer true. Its
|
||||
`mesh/delivery` status on the pull request said the same. `mesh-delivery.checks` for mesh-catalog #125
|
||||
showed `mesh/delivery-group` with the verdict "composed check pass" beside `mesh/delivery` saying "did not
|
||||
pass".
|
||||
|
||||
The forge let both pull requests merge: the trunk requires `mesh/merge-gate` and `mesh/repo-check`, not
|
||||
`mesh/delivery-group`.
|
||||
|
||||
## Why it is a design issue
|
||||
|
||||
A hold is a question to a person. Here the question was asked on a fact that was not yet known ("did not
|
||||
pass" read from "not answered"), and the answer the mesh received two minutes later never reached the
|
||||
question. The person is asked to decide something the mesh already decided, and nothing says the hold is
|
||||
stale: the failure is quiet in the direction that wastes a person's attention and delays a change.
|
||||
|
||||
The two checks a merge reads and the composed check a group's turn reads are also not the same set. So
|
||||
two members can merge correctly, each by its own rules, and still meet a hold that only a person can
|
||||
lift.
|
||||
|
||||
## Open questions
|
||||
|
||||
1. Is a composed check that was asked and has not answered a reason to hold, or a reason to wait in
|
||||
`published` until it answers?
|
||||
2. When a composed check answers after a member was held for it, does the hold lift by observation, as
|
||||
every other state in the table moves on its facts?
|
||||
3. Should `mesh/delivery-group` be required on the trunks, so a group cannot merge before its composed
|
||||
check passed? To-be 47 leaves this to the operator; this issue is the case that setting decides.
|
||||
4. The group's order put the catalogue's member first, and the node-engine's member walked at once on the
|
||||
controller's own path. Is a group's order meant to hold for a member whose walk never waits for
|
||||
mesh-delivery?
|
||||
+49
@@ -0,0 +1,49 @@
|
||||
# 316 — diagnosis
|
||||
|
||||
**2026-10-08.** Read from mesh-delivery's code on the catalogue's main (`modules/mesh-delivery/cmd/mesh-delivery`),
|
||||
from `mesh-delivery.show` of both members, and from the holder's journal on the control node.
|
||||
|
||||
**Where the hold is set.** `groupTurns` (holder.go) walks a closed group's members in order. For a member
|
||||
in `published` it computes `composed`: the group's check exists, its verdict is `pass` or `warning`, and it
|
||||
names the members' current heads. When `composed` is false it writes the member's `HeldWhy` ("its group's
|
||||
composed check did not pass for the heads that merged") and settles it, and the table's `hold` row takes it
|
||||
to `held`. `composed` is false in three different cases, and the code does not tell them apart:
|
||||
|
||||
- the composed check failed;
|
||||
- it was for other heads;
|
||||
- it was asked and has not answered (`Verdict` empty). This is the case here: asked at 10:08:19, answered
|
||||
at 10:10:21, and the catalogue's member reached `published` at 10:09:45, in between.
|
||||
|
||||
**Why nothing lifts it.** Three places would have to act, and none does:
|
||||
|
||||
- `Checked`, for a group's verdict, stores the verdict and summary on the group and returns. It does not
|
||||
look at the members.
|
||||
- `groupTurns` skips a member in `held` (`default: return nil`), so the tick after the verdict sees the
|
||||
passing check and does nothing with it.
|
||||
- The table has two rows out of `held`: `go` (the walk started on a word that was not this owner's) and
|
||||
`release` (a person's act). There is no observed row that takes a held delivery back when the reason it
|
||||
was held is gone, and `HeldWhy` is never cleared once set.
|
||||
|
||||
**What the design says.** ADR 0239 decision 2's row reads "merged, and its group's composed check did not
|
||||
pass for the heads that merged". To-be 47 says the composed check is asked when the group forms or a
|
||||
member's head moves. Neither says what a member merged while its group's check is still out should do.
|
||||
The code read "did not pass" as "has not passed yet".
|
||||
|
||||
**Why the merge was possible.** The trunk's protection requires `mesh/merge-gate` and `mesh/repo-check`
|
||||
(`mesh-delivery.checks`, mesh-catalog #125: `required`). `mesh/delivery-group` is posted on every member's
|
||||
head and is not required, so the forge merged each member on its own checks, 64 seconds after the composed
|
||||
check was asked. To-be 47 lists whether to require it as not decided.
|
||||
|
||||
**Ruled out.**
|
||||
|
||||
- The controller's composed check. It answered `pass` for the very heads that merged (the group's
|
||||
`check.commits`), in about two minutes, an ordinary time for a check.
|
||||
- A lost verdict. The holder logged it at 10:10:21 and the group's record holds it.
|
||||
- The order. The group's order was asked and kept before the check; it plays no part in the hold.
|
||||
|
||||
**Not this issue, noted.** The node-engine's member walked at once (`published` → `delivering` at the
|
||||
same instant), because the controller's own path does not wait for mesh-delivery (ADR 0239 decision 8).
|
||||
The group's order put it second. Open question 4 in the report asks whether that is intended.
|
||||
|
||||
**Located** in mesh-delivery: `groupTurns` holds on a composed check that has not answered, and neither
|
||||
`Checked` nor the table brings a held member back when the check later passes.
|
||||
@@ -0,0 +1,54 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-08
|
||||
located-in: [mesh-controller cmd/mesh-controller (release_plan.go, supersededBy), mesh-catalog modules/mesh-delivery (table.go, no row out of held for a superseded walk)]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 317. A held delivery's module was walked by a later merge's plan
|
||||
|
||||
## Symptom
|
||||
|
||||
On 2026-10-08 the catalogue's delivery of mesh-catalog #125 (the lighting module `openrazer`, its account
|
||||
group) was `held` from 10:09:48 UTC, waiting for a person ([issue 316](../316-a-group-member-held-before-its-composed-check-answered-stayed-held-after-it-passed/00-report.md)).
|
||||
Nobody released it. `mesh-delivery.show` still said `held` at 11:20 UTC, and `mesh-delivery.checks` for
|
||||
mesh-catalog #125 still said `mesh/delivery` pending, "held for a person".
|
||||
|
||||
Its module reached a machine anyway. Three later merges to the catalogue's main touched other modules:
|
||||
`docker` (#122), the artifact store's tools (#123) and `gitea` (#124). Each made a plan, and each plan
|
||||
took `openrazer` along. `mesh-controller.plans`:
|
||||
|
||||
| plan (catalogue commit) | made (UTC) | says |
|
||||
|---|---|---|
|
||||
| a6385479 (#125's merge) | 10:09:45 | superseded at tier 0 by the plan at 3f552274; openrazer planned there again |
|
||||
| 3f552274 (#122's merge) | 11:06:07 | superseded at tier 0 by the plan at c78b5fc9; docker, openrazer planned there again |
|
||||
| c78b5fc9 (#124's merge) | 11:07:02 | superseded at tier 0 by the plan at 86077780; docker, gitea, openrazer planned there again |
|
||||
| 86077780 (#123's merge) | 11:07:03 | tier 0: distribution, docker, gitea and openrazer built from c78b5fc9; openrazer's gate on the laptop judging |
|
||||
|
||||
The controller's walk of the builds waiting for a gate then sent the laptop `openrazer` 88135ad0 →
|
||||
c78b5fc9, beside `docker` and `node-tools`. From 11:10:58 UTC the laptop's condition for openrazer said
|
||||
its account was *relogin needed*. That verdict exists only in the build #125 added, so the held change was
|
||||
running on the laptop while its delivery said it waited for a person.
|
||||
|
||||
## Why it is a design issue
|
||||
|
||||
A hold is the mesh's promise that a change goes no further until a person says so
|
||||
([ADR 0239](../../02-DECISIONS/0239-a-delivery-is-owned-by-the-mesh-delivery-module-and-runs-from-commit-to-delivered.md)).
|
||||
Any later merge to the same repository and branch broke that promise without a person's word and without
|
||||
saying so, through a different rule: a newer merge takes over an older plan's unfinished modules
|
||||
([ADR 0218](../../02-DECISIONS/0218-a-plan-sends-grants-before-code-rolls-out-one-machine-first-and-a-newer-merge-takes-over-an-older-plan.md)
|
||||
decision 3). The two decisions were made apart, and nothing composes them.
|
||||
|
||||
The delivery's record is now false in both directions: it says the change waits, and the change has
|
||||
gone out. A person who releases it later releases nothing; a person who stops it stops nothing.
|
||||
|
||||
## Open questions
|
||||
|
||||
1. When a newer merge supersedes an older plan, does a module the older plan's walk was holding for its
|
||||
delivery's word come along, or stay held with that delivery?
|
||||
2. If it comes along, what does the held delivery become: superseded by the newer delivery, with the hold
|
||||
carried over to it, or something else? Today the table has no row from `held` for a superseded walk.
|
||||
3. A trunk commit holds every earlier commit, so a later build of a module carries the held change in
|
||||
any case. Is a hold on one delivery meaningful once a later merge of the same repository is walked,
|
||||
or does a hold have to hold the module on that trunk until it is released?
|
||||
+42
@@ -0,0 +1,42 @@
|
||||
# 317 — diagnosis
|
||||
|
||||
**2026-10-08.** Read from `mesh-controller.plans` (the list, and each plan by id), `mesh-delivery.show`
|
||||
and `deliveries`, and the code on the main branches of the controller and the catalogue.
|
||||
|
||||
**The held delivery's walk.** #125's merge made the plan at a6385479. Its walk waited for mesh-delivery's
|
||||
word (`waited_for: mesh-delivery`, `waits: true`) and was never let go, because its delivery was held. Its
|
||||
one module, `openrazer`, was "not yet asked": nothing was built.
|
||||
|
||||
**How a later merge took it.** `supersededBy` (the controller, `release_plan.go`) runs when a merge makes a
|
||||
plan. For every open plan of the same repository and branch made before it, it takes each module the
|
||||
older plan had not built (not asked, asked and not answered, or built and not yet sent), adds it to the
|
||||
newer plan, and closes the older plan as superseded with the note "… planned there again". It does not
|
||||
ask why the older plan had not built them. A walk waiting for a held delivery's word is open, so it is
|
||||
superseded like any slow plan. This happened three times over (3f552274, c78b5fc9, 86077780), each newer
|
||||
plan folding in what the one before it held.
|
||||
|
||||
The newest plan's own delivery (#123) was not held, so mesh-delivery let its walk go, and `openrazer` was
|
||||
built from c78b5fc9. The controller's walk of the builds waiting for a gate then carried it to the laptop,
|
||||
since a send carries every move waiting on its machine (ADR 0236).
|
||||
|
||||
**Why the held delivery did not notice.** mesh-delivery's table has `superseded` rows only from
|
||||
`published` and `delivering` ("a newer delivery to the same trunk took over its walk"). From `held` there
|
||||
are only `go` and `release`. When the walk record said `superseded`, no row applied, so the delivery stayed
|
||||
`held` with `held_why` unchanged, and its walk shows "superseded … openrazer planned there again". Two of
|
||||
the three later deliveries (#122, #124) went to `superseded` correctly, through the `published` row.
|
||||
|
||||
**What the design says.** ADR 0218 decision 3 makes a newer merge take over an older plan's unfinished
|
||||
work, and was written before deliveries could be held. ADR 0239 makes a held delivery wait for a person,
|
||||
and to-be 47 says that nothing of a published delivery is registered before its walk starts, so that a
|
||||
send cannot carry it early. Neither says what a supersession does to a module a hold is keeping back.
|
||||
|
||||
**Ruled out.**
|
||||
|
||||
- A person's act. #125's transitions end at `hold`, with no `release`, and its own walk was superseded
|
||||
while still waiting (`let_go` empty), so no `plans go` started it.
|
||||
- Issue 295's path, a recorded build carried by a send: here the build was asked by a plan, not left
|
||||
recorded.
|
||||
|
||||
**Located** in the controller's `supersededBy`, which folds a held walk's modules into a newer plan, and in
|
||||
mesh-delivery's table, which has no way out of `held` when that happens, so the delivery's state stops
|
||||
being true.
|
||||
+69
@@ -0,0 +1,69 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-08
|
||||
located-in: [mesh-controller cmd/mesh-controller (release.go, gatedSend and ungatedIn; gate.go, judgeMoves; the walk of the builds waiting for a gate)]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 318. One module's wait for a new login held every other module's walk
|
||||
|
||||
## Symptom
|
||||
|
||||
On 2026-10-08 at 11:08 UTC the controller opened its walk of the builds waiting for a gate: the laptop
|
||||
first, then the workstation, then the control node. On the laptop it sent three moves in one send and
|
||||
judged them together:
|
||||
|
||||
| module | move |
|
||||
|---|---|
|
||||
| `docker` | ba6058ab → c78b5fc9 |
|
||||
| `node-tools` | 9a4140b4 → d4845f7c |
|
||||
| `openrazer` | 88135ad0 → c78b5fc9 (the held change of [issue 317](../317-a-held-deliverys-module-was-walked-by-a-later-merges-plan/00-report.md)) |
|
||||
|
||||
The laptop's gate said, for its whole ten minutes, "judging (0 of 3 passes; wanting: its unit
|
||||
openrazer.daemon on the laptop is unhealthy: failed in the account's own service manager (exit-code))".
|
||||
The module's condition on the laptop said why: the daemon's unit had failed since 10:13 UTC, and the
|
||||
account was *relogin needed* from 11:10:58 UTC, which
|
||||
[ADR 0252](../../02-DECISIONS/0252-a-module-puts-an-account-in-a-group-and-the-mesh-says-when-a-new-login-is-needed.md)
|
||||
expects until the operator logs out and in again.
|
||||
|
||||
Meanwhile three other walks could not send to the rest of the mesh. `mesh-controller.plans` said of the
|
||||
tool runner's (`novox/mesh-tools` d4845f7c, `node-tools`), the catalogue's (86077780, `docker`) and the
|
||||
controller's (950562af, `build-agent`): "a build waits there for its gate: … the workstation (…, openrazer
|
||||
88135ad0 → c78b5fc9) — a walk sends them, one machine at a time, each judged".
|
||||
|
||||
At 11:21:06 UTC the laptop's gate failed: "not healthy within 10m0s of its apply: its unit
|
||||
openrazer.daemon … is unhealthy", and `openrazer` was put back to 88135ad0 on the laptop
|
||||
(`build.openrazer.<laptop>.rolled-back`). The walk ended `failed`. After it, all three walks still said
|
||||
the same: `openrazer` c78b5fc9 waits on the workstation for a gate, so nothing is sent there or to the
|
||||
machines after it.
|
||||
|
||||
## Why it is a design issue
|
||||
|
||||
The wait was known, and it was a person's: ADR 0252 says a new login holds a delivery *of that module* on
|
||||
that machine. It held every module on every machine after it instead, for changes that had nothing to do
|
||||
with it: the tool runner, the container engine and the build agent. The end was worse than the wait. The
|
||||
gate's bound turned "waiting for a person" into "failed", the build that added the account group was put
|
||||
back, and the walk that carries the other modules failed with it. A failed walk of the waiting builds
|
||||
does not open again on its own: a person releases it.
|
||||
|
||||
Three things meet here, each sound alone:
|
||||
|
||||
- a send carries every move waiting on its machine, and the gate judges them all at once
|
||||
([ADR 0236](../../02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md));
|
||||
- no other send may go where a move waits for a gate, so every walk waits for the one that judges it;
|
||||
- a module's health that waits for a person reads as *not yet healthy*, which fails at the gate's bound.
|
||||
|
||||
## Open questions
|
||||
|
||||
1. Should *relogin needed* count against a gate at all? It says the build did what it should and a person
|
||||
has one step left. It is neither a pass nor a fault of the build.
|
||||
2. If it counts, should it hold only that module's move, and let the moves judged beside it pass or fail on
|
||||
their own?
|
||||
3. A wait for a person has no bound a machine can meet. Should it end in *failed* and a rollback after ten
|
||||
minutes? The build put back declares no account, and by ADR 0252 decision 3 the node-engine then gives
|
||||
back the account group it added, so the login the gate waited for can no longer help.
|
||||
4. The daemon's unit had failed since 10:13 UTC, almost an hour before the new build reached the laptop.
|
||||
Should a gate fail a build for a failure that was there before it was sent?
|
||||
5. While a move waits for a gate on one machine, should walks of unrelated modules still be refused there,
|
||||
or only those whose send would carry it?
|
||||
+55
@@ -0,0 +1,55 @@
|
||||
# 318 — diagnosis
|
||||
|
||||
**2026-10-08.** Read from `mesh-controller.plans` (the list, and the walk of the waiting builds by id) between
|
||||
11:17 and 11:22 UTC, `mesh-controller.conditions` and `status`, and the controller's code on its main
|
||||
branch (`cmd/mesh-controller`).
|
||||
|
||||
**One send, one verdict.** `gatedSend` (release.go) sends a machine every move waiting there and hands the
|
||||
gate all of them as carried moves. `judgeMoves` (gate.go) reads each module on each machine and keeps the
|
||||
worst reading. Any module *not yet healthy* resets the passes to zero for the whole send: `docker` and
|
||||
`node-tools` were healthy, and the gate still said "0 of 3 passes" because `openrazer` was not. Past
|
||||
`gateBound` (ten minutes after the apply) a *not yet healthy* gate fails. `g.Failing` names only the
|
||||
modules found wanting, so only `openrazer` was put back, but the walk itself ended `failed`.
|
||||
|
||||
**Everything else waits for that verdict.** `ungatedIn` (release.go) refuses any send that is not a person's
|
||||
and would carry a move no gate has seen, outside the send's own scope. A plan's sends to the rest of its
|
||||
machines carry the machine's whole declaration, so while `openrazer` c78b5fc9 waited for a gate on the
|
||||
workstation, the tool runner's, the catalogue's and the controller's walks were each refused there, and
|
||||
they say so on every retry. They do not fail; they wait for the walk of the waiting builds, which goes one
|
||||
machine at a time and had stopped on the laptop. After it failed, the release code's rule applies: a
|
||||
failed walk of the waiting builds stops the next one opening on its own until a person releases it.
|
||||
|
||||
**How *relogin needed* reached the gate.** ADR 0252 decision 5 says the node-engine states the account's
|
||||
health as the module's resource of kind `account`, unhealthy with *relogin needed*, and that "the
|
||||
first-node gate reads it like any other". The module's health reading (`judgeHealth`) turns an unhealthy
|
||||
resource into *not yet healthy*. There is no reading for "waits for a person". The nearest one,
|
||||
*waiting* (a module waiting on an unhealthy provider), neither passes nor fails at the bound, but it is
|
||||
kept for providers (ADR 0240 rule 5).
|
||||
|
||||
Here the gate's words named the unit, not the account: the daemon's unit had failed since 10:13 UTC on
|
||||
both the laptop and the workstation, before the new build reached either (issue 315 records the same daemon
|
||||
failing at every login). I think the unit fails for the same reason the account is *relogin needed*: the
|
||||
running session lacks the group the daemon checks for. That is a guess, not checked. Either way, both
|
||||
readings existed before or because of the person's step, and neither is a fault the build brought.
|
||||
|
||||
**What the design says.** ADR 0252's consequences say "a new login holds a delivery of that module on that
|
||||
machine" and accept that a walk that starts on a workstation waits for the person. They do not consider the
|
||||
other modules carried in the same send, the walks refused behind it, or the gate's ten-minute bound turning
|
||||
the wait into a rollback. ADR 0236 makes one send judge everything waiting on a machine, to stop unjudged
|
||||
builds from moving; it does not distinguish a module waiting for a person from one that is broken.
|
||||
|
||||
**What the rollback does to the account group.** By ADR 0252 decision 3, an account resource that is no
|
||||
longer declared gives back the groups the node-engine added. The build put back (88135ad0) declares no
|
||||
account, and after the rollback the laptop's condition for `openrazer` named only the unit, not the account.
|
||||
Whether the account actually left the group was not checked.
|
||||
|
||||
**Ruled out.**
|
||||
|
||||
- A fault in `docker` or `node-tools`: both read healthy at every judging; they were held only by the
|
||||
verdict they shared.
|
||||
- A machine-wide fault on the laptop: its node-engine reported `applied` and current, and the gate's
|
||||
wanting named one module's unit.
|
||||
|
||||
**Located** in the controller's walk and first-node gate: one verdict per send (`judgeMoves`), sends refused
|
||||
on any machine where a move waits (`ungatedIn`), and no reading for a health that waits for a person, so a
|
||||
known wait fails at the bound and stops every walk behind it.
|
||||
+78
@@ -0,0 +1,78 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-08
|
||||
located-in: [mesh-catalog modules/mesh-delivery (table.go, the rows out of held and its bound; verbs.go, retire-history, close and stop), 03-DESIGN/01-to-be/47-delivery-from-commit-to-delivered.md (the state table)]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 319. A held delivery that a later commit superseded has no end but release
|
||||
|
||||
## Symptom
|
||||
|
||||
On 2026-10-08 at 11:20 UTC sixteen deliveries were `held` for a person. `mesh-delivery.retire-history`,
|
||||
run as a dry run, retired none of them and gave the same reason for each: "retire from held is refused —
|
||||
its guard does not hold: this owner heard its head: it is not history made from the forge's word".
|
||||
|
||||
What they are, read with `mesh-delivery.deliveries` and `show`:
|
||||
|
||||
| how many | repository | why held | what it builds |
|
||||
|---|---|---|---|
|
||||
| 4 | hq | merged without a passing check | nothing: "the change touches no module of the mesh's graph" |
|
||||
| 2 | catalogue | no walk opened within ten minutes of the merge | a new module, sent nowhere |
|
||||
| 8 | controller | no walk opened within ten minutes of the merge | the controller, the build agent, the route proxy |
|
||||
| 2 | catalogue | its group's composed check did not pass for the heads that merged ([issue 316](../316-a-group-member-held-before-its-composed-check-answered-stayed-held-after-it-passed/00-report.md)) | one module each |
|
||||
|
||||
Two examples:
|
||||
|
||||
- `novox/hq@055550802096`, the pull request that opened issue 281. Merged 2026-10-06 20:26 UTC; its head was
|
||||
announced to mesh-delivery at 23:05 UTC, at the owner's first start, and it went straight from
|
||||
`proposed` to `held`: "merged without a passing check". Its plan builds nothing.
|
||||
- `novox/mesh-controller@ec3769a8a669`, controller #105. Ready at 2026-10-07 00:26 UTC, merged at 00:27,
|
||||
held at 00:37: "merged, and the controller opened no walk for it within 10m0s". The controller's modules
|
||||
have been walked since from later commits of the same main, for example the group
|
||||
`feat/delivery-checks-verb`, whose controller member is `delivered`.
|
||||
|
||||
Five of them (the four hq merges and controller #105) are past the held bound of one day. The self-check's
|
||||
probe D14 raises each as `delivery.<id>.stalled`, "healer H2 may none: the state is the operator's". They
|
||||
were five of the ten conditions open that hour.
|
||||
|
||||
## Why it is a design issue
|
||||
|
||||
A held delivery asks a person "does this go on?". For these sixteen the question has no meaningful answer:
|
||||
a later commit of the same trunk has already delivered what they held, or they hold nothing to deliver. The
|
||||
state table offers a person four ways out of `held`, and none fits:
|
||||
|
||||
- `release` starts a walk of the old commit: wrong for a commit a later one superseded, and nothing for one
|
||||
that builds nothing.
|
||||
- `retire-history` is for history made from the forge's word ([ADR 0250](../../02-DECISIONS/0250-a-merge-heard-only-from-the-buss-history-is-retired-and-owed-work-the-forge-cannot-do-ends.md)),
|
||||
and refuses every one of these because their heads were heard.
|
||||
- `close`, healer H2's verb, has no transition from `held`: the bound names none.
|
||||
- `stop` is allowed from `held`, but it says a person stopped a delivery and sets `mesh/delivery` to error
|
||||
on the commit: the wrong record for a change that did reach the machines through a later commit
|
||||
(issue 313 ruled it out for the same reason).
|
||||
|
||||
So the condition stays open for ever, and each new case adds one. A condition nobody can clear correctly
|
||||
teaches people to ignore the list it is in.
|
||||
|
||||
## What is needed
|
||||
|
||||
The issue states the need; the decision is for graduation.
|
||||
|
||||
- **An end for a held delivery that was superseded**, as `published` and `delivering` already have
|
||||
(`superseded`, "a newer delivery to the same trunk took over its walk"): either because a later commit's
|
||||
delivered walk carried the same modules on the same trunk, or because it builds nothing at all.
|
||||
- **Automatic where nothing is lost**: a held delivery whose plan builds nothing has nothing for a person to
|
||||
decide.
|
||||
- **A person's act, with why, otherwise**: a held delivery whose modules a later commit delivered is ended
|
||||
by a person who says so, and recorded as superseded, not stopped.
|
||||
|
||||
## Open questions
|
||||
|
||||
1. Is "a later commit delivered every module this one moves" something mesh-delivery can read from its own
|
||||
records, or does it need the controller to say which build each machine runs?
|
||||
2. The four hq merges were heard at the owner's first start from the bus's history, as announcements rather
|
||||
than merges. Are they history in the sense of ADR 0250, missed by its rule because their head was
|
||||
replayed too?
|
||||
3. Why did the controller open no walk for eight of its own merges within ten minutes? The controller's own
|
||||
path never waits, so its walk may have been missed rather than absent. Not diagnosed here.
|
||||
+48
@@ -0,0 +1,48 @@
|
||||
# 319 — diagnosis
|
||||
|
||||
**2026-10-08.** Read from mesh-delivery's code on the catalogue's main (`modules/mesh-delivery/cmd/mesh-delivery`),
|
||||
`mesh-delivery.retire-history` (dry), `deliveries`, `show`, and `mesh-controller.conditions`.
|
||||
|
||||
**The rows out of `held`** (table.go), each with what it needs:
|
||||
|
||||
| event | kind | needs |
|
||||
|---|---|---|
|
||||
| `go` | observed | its walk started on a word not this owner's (the controller's own path, or a person's `plans go`) |
|
||||
| `release` | a person's act | why and who; it walks the delivery's own commit |
|
||||
| `retire` | a person's act | `isHistory`: made from the forge's word, merge heard 15 minutes or more after it happened, no walk open |
|
||||
| `stop` | a person's act | why and who; from any state not final |
|
||||
|
||||
There is no `superseded` row from `held`. `published` and `delivering` have one, guarded by the walk's
|
||||
record saying `superseded`. A held delivery whose walk never opened has no walk record to read, and one
|
||||
whose walk was superseded ([issue 317](../317-a-held-deliverys-module-was-walked-by-a-later-merges-plan/00-report.md))
|
||||
does not match it either.
|
||||
|
||||
**Why `retire-history` refuses them.** `isHistory` requires `FromForgesWord()`: the delivery was made from
|
||||
the forge's `pull.merged` because its head was never heard. All sixteen had their heads announced. The four
|
||||
hq merges were announced at 23:05:28 UTC on 2026-10-06, the owner's first start, hours after their
|
||||
merge, so their heads came from the bus's history as much as the 688 merges issue 313 retired. The rule
|
||||
does not see that, because it reads how the delivery was made, not when.
|
||||
|
||||
**Why `close` refuses them.** `Close` reads the state's bound. `held`'s bound is one day with no H2
|
||||
events ("it waits for the operator"), so it answers "the state is the operator's".
|
||||
|
||||
**Why a delivery that builds nothing is held.** `merged-unchecked` (proposed → held) is tried before
|
||||
anything about the plan. The row `published` → `delivered` "nothing for a walk to move" exists, but a
|
||||
delivery merged unchecked never reaches `published`. Its plan was known (`builds nothing`), and nothing
|
||||
reads it once held.
|
||||
|
||||
**The stalled conditions.** D14 lists every delivery past its state's bound. For `held`, H2 may do nothing,
|
||||
so the condition's resolver is `operator`, and it clears only when the delivery leaves `held`. None of the
|
||||
person's acts is a correct way out, so the condition has no correct end.
|
||||
|
||||
**What the design says.** To-be 47's state table gives `held` a bound of one day and "nothing: it is the
|
||||
operator's" after it. ADR 0239 decision 2's table has `superseded` only from `published` and `delivering`.
|
||||
Neither considers a held delivery whose content a later commit delivered.
|
||||
|
||||
**Ruled out.**
|
||||
|
||||
- A missing release. Releasing any of these would walk an old commit over a newer one, or walk nothing.
|
||||
- D14 itself. It reports the state truthfully; the state has no way to change.
|
||||
|
||||
**Located** in mesh-delivery's state table (no end for a superseded or empty held delivery, and history
|
||||
read by how a delivery was made) and in to-be 47, whose state table carries the same gap.
|
||||
Reference in New Issue
Block a user