ADR 0254: a wait for a person's new login passes the gate, so one module's wait stalls no walk (issue 318)
mesh/merge-gate pass: the change touches no module of the mesh's graph
mesh/repo-check pass: its merge-check.sh passed
mesh/delivery-group group fix/318-a-wait-for-a-person-is-not-a-failure checking: 3 of 4 member(s) ready
mesh/delivery superseded: a newer head of the same pull request

This commit is contained in:
jochen
2026-10-08 13:43:08 +02:00
parent 73fa97efde
commit 5c91bbd4f2
7 changed files with 255 additions and 3 deletions
@@ -99,6 +99,13 @@ gated: it is the module saying it must change everywhere at once.
> is the condition `module.<module>.<machine>.unhealthy` that *no condition raised since the send* already
> holds the module on. Designed in [to-be 48](../03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md).
> **The mechanism changed — 2026-10-08, by [ADR 0254](0254-a-wait-for-a-persons-new-login-is-not-a-gates-failure-and-each-module-of-a-send-keeps-its-own-pass.md).**
> What stands: the three judgings, their spacing and bound, one verdict per send, and what was found wanting
> put back. What moved: a module whose only fault is a wait for a person's new login reads *waits for a
> person*, which counts as a pass and is carried in the verdict; and each module of a send counts its own
> passes, so that when the send fails, a module healthy on its own for the passes asked keeps a pass and is
> not put back (issue 318).
**3. A build that fails its gate is put back there, once, and never sent again on its own.** The plan
stops; the build is marked failed at its gate; the module's registered build goes back to the build the
first machine ran before — the newest successful build of that commit still kept ([ADR 0189](0189-the-store-keeps-what-the-records-name.md)
@@ -9,6 +9,15 @@ extends: 02-DECISIONS/0176-the-login-shell-is-a-node-seat-and-execute-is-its-con
# 252. A module puts an account in a group, and the mesh says when a new login is needed
> **Superseded in part — 2026-10-08, by [ADR 0254](0254-a-wait-for-a-persons-new-login-is-not-a-gates-failure-and-each-module-of-a-send-keeps-its-own-pass.md).** Decision 5's
> last sentence, "the first-node gate reads it like any other: a module whose account needs a new login is
> not yet good on that machine", and the consequence "a new login holds a delivery of that module on that
> machine" no longer hold. The gate reads *relogin needed*, and a unit that fails in that account's own
> service manager because of it, as a wait for a person: it passes, the wait carried in its verdict and
> said as the module's condition `relogin-needed`. On 2026-10-08 the old reading failed the gate at its
> bound, put the build back and so took the group back, and held every other module's walk behind it
> (issue 318). Decisions 1 to 4, and what decision 5 says the node-engine reads and states, stand.
## Context
The module of a peripheral-lighting daemon installs the daemon, its kernel driver and its client
@@ -0,0 +1,202 @@
---
topic: the mesh
status: accepted
date: 2026-10-08
deciders: jochen
reconstructed: false
supersedes-in-part:
- 0252-a-module-puts-an-account-in-a-group-and-the-mesh-says-when-a-new-login-is-needed.md
extends: 02-DECISIONS/0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md
---
# 254. A wait for a person's new login is not a gate's failure, and each module of a send keeps its own pass
## Context
On 2026-10-08 the controller's walk of the builds waiting for a gate sent the laptop three moves in one
send: the container engine's module, the tool runner, and the lighting module, whose new build put the
operator's account in the group its daemon needs ([ADR 0252](0252-a-module-puts-an-account-in-a-group-and-the-mesh-says-when-a-new-login-is-needed.md)).
What followed is [issue 318](../04-ISSUES/318-one-modules-wait-for-a-new-login-held-every-other-modules-walk/00-report.md):
- **The laptop said what ADR 0252 expects.** The node-engine put the account in the group and stated the
module's account resource *relogin needed*: the session running began before the group. The daemon's
unit, started by the account's own service manager, kept failing for the same reason.
- **The gate read that as *not yet healthy* for ten minutes and failed at its bound.** ADR 0252 decision 5
said so on purpose: "a module whose account needs a new login is not yet good on that machine". The gate
had two outcomes for a reading that is not a pass: wait (a provider is unhealthy, ADR 0240 rule 5) or
fail at the bound. A wait for a person is neither.
- **The rollback made the login useless.** The build put back declares no account, so by ADR 0252
decision 3 the node-engine takes back the group it added. A new login after that gives the daemon
nothing.
- **One verdict held every module of the send.** `judgeMoves` keeps the worst reading of every module the
send carried. The other two modules read healthy at every judging, and the gate still said "0 of 3
passes". When it failed, only the lighting module was put back, but the other two got no verdict at all.
- **A move without a verdict stops every other walk there.** `ungatedIn` refuses any send, other than a
person's, that would carry a move no gate has seen (ADR 0236 §4a). The tool runner's, the catalogue's and
the controller's walks were each refused on the workstation for as long as the walk of the waiting
builds was judging the laptop, and after it failed, until a person released it.
What the controller and the node-engine had, on the main branches of 2026-10-08:
- The node-engine states a module's failed unit as a resource of kind `unit` whose reason says "in the
account's own service manager", and knows which account that manager is (the unit judge keeps it as
evidence). It does not send the account to the controller.
- It states the account a module put in a group as a resource of kind `account`, its target the account,
its reason starting *relogin needed* when only a new login is missing.
- A gate's record (`PlanGate`) counts passes for the send as a whole and names the modules the last
judging found wanting (`Failing`). A failure puts back the modules found wanting, and leaves the rest of
the send with no verdict.
**Checked against GENESIS.** *Failure must be loud*, and a wait is not a failure: a gate that turns "the
operator has one step left" into "the build failed" and undoes the build says something false, and loudly.
*Anything requiring a human to notice it will be noticed late*: a person's step must be said to that
person, once, in words they act on, not discovered in a failed walk. Nothing here conflicts with GENESIS.
## Considered Options
**What the gate does with a wait for a person.**
1. **Keep it *not yet healthy*** (ADR 0252 as decided). It is what failed on 2026-10-08, and the rollback
removes the very group the login would bring. Rejected.
2. **Read it as *waiting*, as a provider being down is read**: no pass, no failure, past the bound too.
The walk then waits for the person's login, on a laptop that may be asleep or away for days, and every
walk refused behind it waits with it. That is issue 318's stall without its rollback. Rejected.
3. **Leave the account out of the gate's reading.** The daemon's unit still fails, and the gate still
fails on it; and the condition that tells the operator what to do would disappear with it. Rejected.
4. **A reading of its own, *waits for a person*, that passes, carried along in the verdict, and is said
to the operator as the module's condition.** The build did what it should; what is left is a person's.
Chosen.
**Which failures the wait covers.** Excusing too much hides a fault; excusing too little fails the build
anyway, since the unit that cannot run is what the gate names first.
5. **Every unhealthy resource of a module whose account says *relogin needed*.** A container that crashes
for a reason of its own would pass. Rejected.
6. **The account, and only units in the own service manager of that same account, of the same module.**
Those are the processes that began before the group, and nothing else can be. Chosen. The node-engine
says the account on such a unit; a unit whose engine does not say it is not covered.
**How one module's reading reaches the others of its send.**
7. **Split a send per module.** A send carries the machine's whole declaration (ADR 0221, issue 281). Not
possible.
8. **Let other walks' sends go to a machine whose only pending gate waits for a person.** With the wait a
pass, that case no longer arises. In its general form, a send that carries a move still being judged,
it would end ADR 0236 §4a's rule that no build moves without a gate. Rejected.
9. **One verdict per send, as now, but each module's passes counted apart**: when the send fails, a
module that was healthy on its own for the passes the gate asks keeps a pass, is not put back, and
so is not one more move every other walk waits for. Chosen. The send's verdict, its bound, its
rollback of what was found wanting and the walk stopping after a failure stay as ADR 0236 decided.
## Decision
1. **A wait for a person is a gate reading of its own: *waits for a person*.** Today it has one source: a
module's resource of kind `account` is unhealthy with a reason starting *relogin needed* (ADR 0252
decision 5), and every other unhealthy resource of that module on that machine runs in the own
service manager of that same account. Anything else unhealthy of the module (a container, a unit of
the machine's own manager, a unit in another account's manager, a unit whose manager is not said, an
account *not in the group* or not read) is judged as before, and the module is not yet healthy. A
resource still starting is not yet healthy, as before.
2. **The gate passes a module that waits for a person, and carries the wait.** The reading counts as a
pass. The verdict says the wait ("… and it waits for a person: relogin needed on the laptop: …"), in the
plan, in the gate's record and in each carried build's verdict. Nothing is put back for it, so the group
the build added stays.
3. **The node-engine says whose manager a unit runs in.** A module's resource that runs in an account's
own service manager (a failed unit, or a service liveness judges) names that account, and so does a
resource of kind `account`. An engine older than this says none, and its unit failures are judged as
before.
4. **The wait is said to the operator as the module's condition**, `module.<module>.<machine>.relogin-needed`,
instead of its `unhealthy` condition: a warning, the operator's, never escalated by age, raised on two
statements in a row as every module condition is, and cleared on the first statement that no longer
says it. Its summary is one sentence: "relogin needed on `<machine>`: `<module>` waits for a new login
of `<account>` … log out of every session and in again, or reboot". The gate does not read this
condition as one raised since the send. When the login comes and a unit still fails, the wait clears
and the module's own `unhealthy` condition is raised in its place, with no wait to cover it.
5. **Each module of a send keeps its own count of passes.** The gate's record counts, per module, the
judgings in a row that found it healthy or waiting for a person on every machine judged. The send's
verdict is still one. When it fails, at the bound or on a fault, a module whose count reached the
passes the gate asks, after the settle time, **keeps a pass**: its verdict is recorded as passed "on its
own", it is not put back, and the walk's note names it. What was found wanting is put back as ADR
0236 §3 says, and the walk stops as ADR 0236 says.
## Consequences
- **Issue 318's walk passes.** The laptop's send passes after its three judgings, the lighting module's
wait carried along; the walk goes on to the workstation and the control node, and the walks refused
behind it send once the moves they waited for have their passes.
- **The operator reads one sentence per machine**, "relogin needed on the laptop: …", in place of an
urgent rollback, and it clears when they have logged in again.
- **ADR 0252's consequence "a new login holds a delivery of that module on that machine" no longer holds**,
and decision 5's last sentence is superseded: the gate passes such a module. The false success ADR 0252
feared is not a pass on "applied": the wait is said in the verdict and as the module's condition, and
the daemon is judged again once the login has come.
- **What the gate cannot tell.** A unit in the account's own manager that would fail after the login for a
reason of its own is excused until the login. It is said then, as the module's own condition. A build
that passed its gate is not put back by a later condition; a newer build or a person's push is the way.
- **One failure of a send no longer strands the healthy modules of that send** without a verdict. A walk
that fails still stops, and the next one still waits for a person's release (ADR 0236). What remains
for it to carry is the module that failed, which is put back.
- **The gate's record grows three fields**: the per-module counts, the waits, and the modules that kept a
pass. A plan written before them reads as having none.
- **Rollout order**: the node-engine first, so that units name their account; then the controller.
A controller before the node-engine is harmless: no unit names an account, so no wait covers a unit,
and the gate fails as it did.
- **A new condition kind.** The plain words per condition kind (a headline, an explanation and a resolved
line), decided on a branch beside this record the same day for issue 320, will need words for
`relogin-needed`: whichever of the two lands second adds them.
**How it meets issues 316, 317 and 319**, which the same delivery of the lighting module met that day.
None of them is fixed here.
- **[Issue 317](../04-ISSUES/317-a-held-deliverys-module-was-walked-by-a-later-merges-plan/00-report.md)**
is how the lighting module's build reached the laptop at all: a later merge's plan folded a held
delivery's module into its own walk. This record changes what the gate does with the build once it is
sent, not whether it should have been sent. With 317 fixed, the same wait would be met by the lighting
module's own walk, and would pass the same way.
- **[Issue 316](../04-ISSUES/316-a-group-member-held-before-its-composed-check-answered-stayed-held-after-it-passed/00-report.md)
and [issue 319](../04-ISSUES/319-a-held-delivery-a-later-commit-superseded-has-no-end-but-release/00-report.md)**
are mesh-delivery's: a hold that does not lift, and a hold with no end but release. A gate that fails on a
wait for a person no longer adds a failed walk to them. A held delivery whose module is later walked by
another commit stays in 319's open question, and a pass carried with a wait is not a reason to hold.
## How it is checked
- **The controller** (`mesh-controller`):
- `cmd/mesh-controller/person_wait_test.go`:
- only the account with *relogin needed* and units in that same account's manager make a wait; an
account not in its group, a unit with no account said, a unit in another account's manager and a
container beside the wait do not;
- the gate reads a wait as its own reading, and the same module with a container down as not yet;
- the condition is `relogin-needed`, a warning, the operator's, its summary starting "relogin needed
on"; after the login a unit still failing is the module's own `unhealthy`, and after the unit runs
nothing is open;
- a send that fails at its bound keeps the pass of the module that was healthy on its own, marks the
one that failed and puts back only that one.
- `cmd/mesh-controller/replay318_test.go`: issue 318 replayed, below.
- **The node-engine** (`mesh-host`), `cmd/mesh-host/units_test.go`: a failed unit and a liveness-judged
service in an account's own manager name the account, the account resource names itself, and what the
machine's own manager runs names none.
- **The replay** `R318` in mesh-lab's register: the walk of the waiting builds sends the laptop the three
moves of 2026-10-08, the laptop states the account *relogin needed* and the daemon's unit failed in
that account's manager, and the gate's bound is reached at the first judging that is not a pass. Before
the fix the walk fails and the lighting module is put back; on it the walk is done on both machines,
every move has its pass, the lighting module's carries the wait, and its condition says *relogin
needed*. Proved with `go run ./cmd/prove R318`.
- **Live**, after rollout: on the next send of a module that puts the operator's account in a group,
`plans <id>` shows the gate passed with "it waits for a person", `conditions` shows
`module.<module>.<machine>.relogin-needed` until the operator logs in again, and no
`build.<module>.<machine>.rolled-back` is raised for it.
## References
- [Issue 318](../04-ISSUES/318-one-modules-wait-for-a-new-login-held-every-other-modules-walk/00-report.md),
and issues 316, 317 and 319 beside it.
- [ADR 0236](0236-a-build-is-judged-on-its-first-machine-and-put-back-by-something-other-than-itself-and-so-it-rolls-out-unattended.md)
(the gate, a send judged as one, the walk of the waiting builds),
[ADR 0240](0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md) (a module's stated
health, the *waiting* reading for a provider down),
[ADR 0252](0252-a-module-puts-an-account-in-a-group-and-the-mesh-says-when-a-new-login-is-needed.md)
(the account group and *relogin needed*; decision 5's gate sentence superseded here).
- [To-be 48](../03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md) §5, and
[to-be 45](../03-DESIGN/01-to-be/45-a-core-that-cannot-fail-silently.md) §8: the gate.
- mesh-controller, mesh-host and mesh-lab: the branches `fix/318-a-wait-for-a-person-is-not-a-failure`.
@@ -4,6 +4,7 @@ status: in-progress
code: [mesh-controller, mesh-host, mesh-tools, mesh-catalog, mesh-sdk, mesh-lab]
updated: 2026-10-08
decisions:
- 02-DECISIONS/0254-a-wait-for-a-persons-new-login-is-not-a-gates-failure-and-each-module-of-a-send-keeps-its-own-pass.md
- 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md
- 02-DECISIONS/0239-a-delivery-is-owned-by-the-mesh-delivery-module-and-runs-from-commit-to-delivered.md
- 02-DECISIONS/0238-a-commit-is-the-build-at-hand-one-commit-one-change-plan-checked-off-the-trunk-and-published-only-on-it.md
@@ -497,6 +498,11 @@ the other machines follow. "Reported applied" is not enough.
or leaves the machine while a build no gate has seen waits there; a rebuild with the same artifacts and
manifest is no move. What waits is moved by a **walk**, one machine at a time, the control
node last, each judged; one that fails holds the next until `upgrade release-backlog --why`.
- **A wait for a person passes, and each module of a send keeps its own pass**
([ADR 0254](../../02-DECISIONS/0254-a-wait-for-a-persons-new-login-is-not-a-gates-failure-and-each-module-of-a-send-keeps-its-own-pass.md)): a module whose only fault is a new login the operator owes
is passed with the wait carried and said as its condition `relogin-needed`; when a send fails, a module
healthy on its own for the passes asked keeps a pass and is not put back. Designed in
[to-be 48](48-a-module-says-how-it-is-healthy.md) §5.
- **A new controller that passes its gate sends the bus's machine the user list it composes**, when that
changed and nothing held back would go with it.
- **The default policy is to roll** (ADR 0236 §4); `record` stays where a module says why, keeps
@@ -2,8 +2,9 @@
layer: to-be
status: in-progress
code: [mesh-host, mesh-controller, mesh-lab, mesh-catalog]
updated: 2026-10-07
updated: 2026-10-08
decisions:
- 02-DECISIONS/0254-a-wait-for-a-persons-new-login-is-not-a-gates-failure-and-each-module-of-a-send-keeps-its-own-pass.md
- 02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md
- 02-DECISIONS/0241-a-machine-says-how-its-network-is-and-an-outside-writer-of-a-mesh-file-is-a-finding.md
- 02-DECISIONS/0247-a-machine-with-a-vpn-client-routes-names-by-domain-through-a-resolver-of-its-own.md
@@ -139,6 +140,11 @@ that says no resource is unhealthy clears it. `module` joins the scopes of the c
restarts are evidence. Held to the content rule (ADR 0234 §6), as every condition.
- **Resolver:** `self` — it clears on observation. A healer that acts on it is a later record.
**A wait for a person's new login is its own condition**, `module.<module>.<machine>.relogin-needed`, in
place of `unhealthy` ([ADR 0254](../../02-DECISIONS/0254-a-wait-for-a-persons-new-login-is-not-a-gates-failure-and-each-module-of-a-send-keeps-its-own-pass.md)): a warning, the operator's, never escalated by age, raised and
cleared as above, its summary one sentence that says to log out and in again. When the login has come and
a unit still fails, it clears and `unhealthy` is raised.
**`status`** lists it with the other conditions; **`node show`** shows each module's resources with their
state and since when, so "is it working" has an answer without opening a terminal on the machine.
@@ -156,6 +162,19 @@ send and within ten minutes of it. Two of its points now read this design:
At the bound, a module not healthy is put back, as ADR 0236 §3 says.
**A wait for a person is not a fault** ([ADR 0254](../../02-DECISIONS/0254-a-wait-for-a-persons-new-login-is-not-a-gates-failure-and-each-module-of-a-send-keeps-its-own-pass.md), issue 318). A module's
account that is *relogin needed* (ADR 0252), and units of that module failing in the own service manager of
that same account, read **waits for a person**: the build did what it should, and the operator has one
step left. The reading counts as a pass and is carried in the verdict; nothing is put back for it, so the
account group the build added stays. Anything else unhealthy of the module beside it is judged as above.
The node-engine names the account on every resource that runs in an account's own manager, which is how
the controller tells such a unit from one that is broken.
**Each module of a send keeps its own pass** (ADR 0254). A send is still judged as one, with one verdict.
The gate also counts, per module, the judgings in a row that found it healthy or waiting for a person;
when the send fails, a module whose count reached the passes asked keeps a pass, said "on its own", and is
neither put back nor left as a move no gate has seen for every other walk to wait on.
## 6. A provider down is said once
A check that names, in `needs`, the provision it exercises is tied to that provision's provider for this
@@ -2,8 +2,9 @@
status: located
opened: 2026-10-08
located-in: [mesh-controller cmd/mesh-controller (release.go, gatedSend and ungatedIn; gate.go, judgeMoves; the walk of the builds waiting for a gate)]
fixed-by:
amended-design:
fixed-by: novox/mesh-host PR #55, novox/mesh-controller PR #142, novox/mesh-lab PR #67
replay: R318
amended-design: 03-DESIGN/01-to-be/48-a-module-says-how-it-is-healthy.md
---
# 318. One module's wait for a new login held every other module's walk
@@ -53,3 +53,11 @@ Whether the account actually left the group was not checked.
**Located** in the controller's walk and first-node gate: one verdict per send (`judgeMoves`), sends refused
on any machine where a move waits (`ungatedIn`), and no reading for a health that waits for a person, so a
known wait fails at the bound and stops every walk behind it.
**Decided and fixed, 2026-10-08**, by [ADR 0254](../../02-DECISIONS/0254-a-wait-for-a-persons-new-login-is-not-a-gates-failure-and-each-module-of-a-send-keeps-its-own-pass.md).
Open questions 1 to 3: *relogin needed*, and a unit failing in that same account's own manager, read *waits
for a person*, which passes with the wait carried and is said as the module's condition `relogin-needed`;
nothing is put back for it. Question 2's other half: each module of a send counts its own passes, and keeps
a pass when another fails beside it. Question 5: other walks are still refused while a move waits for a
gate (ADR 0236 §4a stands); with the wait a pass and every module of a send given a verdict, nothing is left
waiting for them. Question 4, a unit that failed before the send, is not decided here.