Issue 083 resolved: nothing the control queue carries is lost while the store restarts #73

Merged
jschoubben merged 3 commits from multiple-fixes into main 2026-09-22 12:43:00 +00:00
Showing only changes of commit 4de74880cb - Show all commits
@@ -27,31 +27,52 @@
part-way can be made again by the same presenter; the token is spent once everything is in the
store; and only then does the token's secret stop being the node's broker password — after the
spend, because a password replaced by an attempt that then failed would be held by nobody. A
second presenter is held off while the claim is live and may take it once the lease lapses. The
presenter that spent a token may claim it again, so an answer lost after the spend does not lock
out a machine the mesh holds as enrolled. A store that cannot be asked, or a token held by
second presenter is held off while the claim is live and may take it once the lease lapses.
5. **Finishing after the spend takes proof.** An answer lost after the spend would lock out a
machine the mesh holds as enrolled, so the presenter that spent a token may finish — but a
presenter is named by its public key, and a public key is no secret. The first version let
anyone holding a leaked token and a node's public key replay the spent token, overwrite the
node's keys and take its broker password; review found it. The host now signs its request with
the identity key it generated, over the token and every key it presents; the mesh lets a
spent token be finished only with a signature that verifies, while the claim's lease is live,
and never on a request the broker handed over a second time, which may already have been
answered. Without the proof a spent token is spent to everyone: "useless once used" (ADR 0004)
holds, with one narrow exception — the node's own signed request, replayed word for word
within the two minutes after its spend, finishes again. That takes reading the request off the
control plane's queue, where the token's secret travels anyway, so only the broker's operator
or the control plane itself could do it. A store that cannot be asked, or a token held by
another presenter, is answered "try again" rather than refused; the host asks again with the
same request — the same keys — for as long as the claim lasts, and then says to wait that long
before running enrol again.
**What is not closed.** An enrolment request sent to a control plane that is not running at all is
still queued; if the host gives up and the control plane starts later, the request completes for
keys nobody holds. That predates this issue and is not the store's window.
keys nobody holds. That predates this issue and is not the store's window. The host still gives up
on its first unanswered ask rather than asking again with the same keys. A control plane that stops
after spending a token and before answering refuses the redelivered request, since it may already
have answered; the node is then told its token cannot be used, and needs a new one — rare, and said
so in the control plane's log. A mesh past about fifty
nodes whose reports all arrive while the store is away holds what the prefetch allows and lets the
rest go, loudly.
**How it is checked:** unit tests hold a report to being held while the store is away and recorded
when it is back, to being set aside by a newer report from the same node, and to being let go past
the bound; they hold build results, catch-ups and upgrades to being held on the store and settled
otherwise, and an upgrade to not being held on a push that timed out. Tests against a real store
hold the claim to one presenter, its lapse, the spend to the claimant, and a claim again by the
presenter that spent it; an enrolment met by a held token is answered "try again". A unit test holds
presenter that spent it only with proof and inside its lease; an enrolment met by a held token is
answered "try again", and a spent token presented with a public key alone, with a proof made by
another key, with a proof made for another request, or redelivered, is refused. A known answer for
the signed bytes is repeated in both repositories' tests. A unit test holds
the host to asking again while told to and stopping when its patience runs out. The store-window
bed pushes a slow step to the control-node, stops the store while its report is on the way, and
has a second machine enrol into the gap: the report is held, the enrolment is answered "not now"
meanwhile, and when the store is back the report is recorded and the enrolment completes on its
own — with the mesh holding exactly the keys the machine generated — after which the machine applies
what it is pushed and is heard from. Reviewed twice independently; the first review's findings —
the queue stopped behind a retry, the password replaced too early, a lost answer after the spend —
are what points 2 and 4 describe.
what it is pushed and is heard from. Two independent reviews shaped this: the first found the queue
stopped behind a retry and the password replaced too early (points 2 and 4); the second found the
replay (point 5), messages lost on shutdown, and identical build results stranding a delivery; a
third confirmed the replay closed and found one handler still settling on shutdown, fixed.
**Located in:** the control plane's consumer and enrolment, the inventory's tokens, and the host's
enrolment. Not a decision: the token stays "useless once used" (ADR 0004); what changed is when it