Issue 083: the diagnosis describes held messages, the enrolment's order, and what is not closed

This commit is contained in:
2026-09-22 14:24:31 +02:00
parent 2d45aa5f42
commit 266ee34b5c
@@ -2,31 +2,56 @@
1. **Build results, upgrade announcements and catch-up requests can simply be asked again.**
Recording a build result ignores one already kept; acting on an upgrade pushes declarations,
which converge; replaying builds re-announces what is recorded. So each is handed back to the
broker while the store cannot be reached, exactly as a report is (issue 082): after a pause, for
at most two minutes, then let go with a line saying it was lost. Anything the store answered no
to is settled as before, so none of them can stop the queue behind it. A catch-up request used to
be acknowledged before the work; it is acknowledged after.
2. **An enrolment cannot be asked again as it was.** It spent the token first and then wrote the
node's key, its broker account and its other keys — in two databases, the token in the
inventory and the key in the identity store, so no one transaction could cover both. A store
gone between the two left a spent token and a node with no key, and the host, which generates
new keys on every attempt, could not recover without a person issuing a new token.
3. **So the token is claimed, not spent, until the enrolment is complete.** The presenting key
which converge; replaying builds re-announces what is recorded. A catch-up request used to be
acknowledged before the work; it is acknowledged after.
2. **But asking again must not stop the queue.** Issue 082's first shape slept on a message and
handed it back to the head of the queue. The control plane takes one message at a time, so
while one report waited for the store, every enrolment waited behind it — and a host waiting
on its answer gave up after ten seconds, while its request stayed queued and would later spend
the token for keys nobody held. So a message the store cannot take is now **held**: kept
unacknowledged, set aside, and tried again on a timer, while the loop goes on to the next
message. It is held for at most two minutes, then let go with a line saying it was lost; a
control plane that stops returns everything it held to the broker, since none of it was
acknowledged. Each message is held under its subject — a node's report, a module's upgrade,
the catch-up, each build on its own — and a newer message for the same subject sets the held
one aside, so an old report recorded late can never overwrite what the node is doing now. An
upgrade is held only when reading the store failed, not when a push did, so a push that timed
out is not repeated every few seconds.
3. **An enrolment cannot simply be asked again.** It spent the token first and then wrote the
node's key and its other keys — in two databases, the token in the inventory and the key in the
identity store, so no one transaction could cover both. A store gone between the two left a
spent token and a node with no key, and the host, which generates new keys on every attempt,
could not recover without a person issuing a new token.
4. **So the token is claimed, not spent, until the enrolment is complete.** The presenting key
claims it for a short lease; every write that follows overwrites, so an attempt interrupted
part-way can be made again; and the token is spent last, only by the key holding the claim. A
second presenter is held off while the claim is live, and may take it once the lease lapses. A
store that cannot be asked, or a token held by another presenter, is answered "try again"
rather than refused, with nothing spent; the host asks again with the same request — the same
keys — for as long as the claim lasts, and then says the token was not spent.
part-way can be made again by the same presenter; the token is spent once everything is in the
store; and only then does the token's secret stop being the node's broker password — after the
spend, because a password replaced by an attempt that then failed would be held by nobody. A
second presenter is held off while the claim is live and may take it once the lease lapses. The
presenter that spent a token may claim it again, so an answer lost after the spend does not lock
out a machine the mesh holds as enrolled. A store that cannot be asked, or a token held by
another presenter, is answered "try again" rather than refused; the host asks again with the
same request — the same keys — for as long as the claim lasts, and then says to wait that long
before running enrol again.
**How it is checked:** unit tests hold each of the three handlers to handing back a message while
the store restarts and settling it otherwise; tests against a real store hold the claim to one
presenter, its lapse, and the spend to the claimant, and an enrolment met by a held token to "try
again"; a unit test holds the host to asking again while told to and stopping when its patience
runs out. The store-window bed stops the store, has a second machine enrol into the gap, starts the
store again, and asserts the enrolment completes on its own with one identity, and that the
machine then applies what it is pushed and is heard from.
**What is not closed.** An enrolment request sent to a control plane that is not running at all is
still queued; if the host gives up and the control plane starts later, the request completes for
keys nobody holds. That predates this issue and is not the store's window.
**How it is checked:** unit tests hold a report to being held while the store is away and recorded
when it is back, to being set aside by a newer report from the same node, and to being let go past
the bound; they hold build results, catch-ups and upgrades to being held on the store and settled
otherwise, and an upgrade to not being held on a push that timed out. Tests against a real store
hold the claim to one presenter, its lapse, the spend to the claimant, and a claim again by the
presenter that spent it; an enrolment met by a held token is answered "try again". A unit test holds
the host to asking again while told to and stopping when its patience runs out. The store-window
bed pushes a slow step to the control-node, stops the store while its report is on the way, and
has a second machine enrol into the gap: the report is held, the enrolment is answered "not now"
meanwhile, and when the store is back the report is recorded and the enrolment completes on its
own — with the mesh holding exactly the keys the machine generated — after which the machine applies
what it is pushed and is heard from. Reviewed twice independently; the first review's findings —
the queue stopped behind a retry, the password replaced too early, a lost answer after the spend —
are what points 2 and 4 describe.
**Located in:** the control plane's consumer and enrolment, the inventory's tokens, and the host's
enrolment. Not a decision: the token stays "useless once used" (ADR 0004); what changed is when it