Merge pull request 'Issue 083 resolved: nothing the control queue carries is lost while the store restarts' (#73) from multiple-fixes into main

This commit was merged in pull request #73.
This commit is contained in:
2026-09-22 14:43:00 +02:00
2 changed files with 82 additions and 2 deletions
@@ -1,7 +1,8 @@
--- ---
status: open status: resolved
opened: 2026-09-22 opened: 2026-09-22
located-in: [mesh-controller internal/link (serve: enrolment, build results, upgrade and catch-up announcements)] located-in: [mesh-controller internal/link (serve, enrolment), mesh-controller internal/inventory (tokens), mesh-host internal/link (enrol)]
fixed-by: mesh-controller multiple-fixes (build results, upgrades and catch-ups handed back while the store is away, bounded; an enrolment claims its token and spends it last, and is answered "try again"); mesh-host multiple-fixes (the host asks again with the same request); mesh-lab multiple-fixes (the store-window bed)
--- ---
# 083 — Other control messages are lost while the store restarts # 083 — Other control messages are lost while the store restarts
@@ -0,0 +1,79 @@
# Diagnosis — 2026-09-22
1. **Build results, upgrade announcements and catch-up requests can simply be asked again.**
Recording a build result ignores one already kept; acting on an upgrade pushes declarations,
which converge; replaying builds re-announces what is recorded. A catch-up request used to be
acknowledged before the work; it is acknowledged after.
2. **But asking again must not stop the queue.** Issue 082's first shape slept on a message and
handed it back to the head of the queue. The control plane takes one message at a time, so
while one report waited for the store, every enrolment waited behind it — and a host waiting
on its answer gave up after ten seconds, while its request stayed queued and would later spend
the token for keys nobody held. So a message the store cannot take is now **held**: kept
unacknowledged, set aside, and tried again on a timer, while the loop goes on to the next
message. It is held for at most two minutes, then let go with a line saying it was lost; a
control plane that stops returns everything it held to the broker, since none of it was
acknowledged. Each message is held under its subject — a node's report, a module's upgrade,
the catch-up, each build on its own — and a newer message for the same subject sets the held
one aside, so an old report recorded late can never overwrite what the node is doing now. An
upgrade is held only when reading the store failed, not when a push did, so a push that timed
out is not repeated every few seconds.
3. **An enrolment cannot simply be asked again.** It spent the token first and then wrote the
node's key and its other keys — in two databases, the token in the inventory and the key in the
identity store, so no one transaction could cover both. A store gone between the two left a
spent token and a node with no key, and the host, which generates new keys on every attempt,
could not recover without a person issuing a new token.
4. **So the token is claimed, not spent, until the enrolment is complete.** The presenting key
claims it for a short lease; every write that follows overwrites, so an attempt interrupted
part-way can be made again by the same presenter; the token is spent once everything is in the
store; and only then does the token's secret stop being the node's broker password — after the
spend, because a password replaced by an attempt that then failed would be held by nobody. A
second presenter is held off while the claim is live and may take it once the lease lapses.
5. **Finishing after the spend takes proof.** An answer lost after the spend would lock out a
machine the mesh holds as enrolled, so the presenter that spent a token may finish — but a
presenter is named by its public key, and a public key is no secret. The first version let
anyone holding a leaked token and a node's public key replay the spent token, overwrite the
node's keys and take its broker password; review found it. The host now signs its request with
the identity key it generated, over the token and every key it presents; the mesh lets a
spent token be finished only with a signature that verifies, while the claim's lease is live,
and never on a request the broker handed over a second time, which may already have been
answered. Without the proof a spent token is spent to everyone: "useless once used" (ADR 0004)
holds, with one narrow exception — the node's own signed request, replayed word for word
within the two minutes after its spend, finishes again. That takes reading the request off the
control plane's queue, where the token's secret travels anyway, so only the broker's operator
or the control plane itself could do it. A store that cannot be asked, or a token held by
another presenter, is answered "try again" rather than refused; the host asks again with the
same request — the same keys — for as long as the claim lasts, and then says to wait that long
before running enrol again.
**What is not closed.** An enrolment request sent to a control plane that is not running at all is
still queued; if the host gives up and the control plane starts later, the request completes for
keys nobody holds. That predates this issue and is not the store's window. The host still gives up
on its first unanswered ask rather than asking again with the same keys. A control plane that stops
after spending a token and before answering refuses the redelivered request, since it may already
have answered; the node is then told its token cannot be used, and needs a new one — rare, and said
so in the control plane's log. A mesh past about fifty
nodes whose reports all arrive while the store is away holds what the prefetch allows and lets the
rest go, loudly.
**How it is checked:** unit tests hold a report to being held while the store is away and recorded
when it is back, to being set aside by a newer report from the same node, and to being let go past
the bound; they hold build results, catch-ups and upgrades to being held on the store and settled
otherwise, and an upgrade to not being held on a push that timed out. Tests against a real store
hold the claim to one presenter, its lapse, the spend to the claimant, and a claim again by the
presenter that spent it only with proof and inside its lease; an enrolment met by a held token is
answered "try again", and a spent token presented with a public key alone, with a proof made by
another key, with a proof made for another request, or redelivered, is refused. A known answer for
the signed bytes is repeated in both repositories' tests. A unit test holds
the host to asking again while told to and stopping when its patience runs out. The store-window
bed pushes a slow step to the control-node, stops the store while its report is on the way, and
has a second machine enrol into the gap: the report is held, the enrolment is answered "not now"
meanwhile, and when the store is back the report is recorded and the enrolment completes on its
own — with the mesh holding exactly the keys the machine generated — after which the machine applies
what it is pushed and is heard from. Two independent reviews shaped this: the first found the queue
stopped behind a retry and the password replaced too early (points 2 and 4); the second found the
replay (point 5), messages lost on shutdown, and identical build results stranding a delivery; a
third confirmed the replay closed and found one handler still settling on shutdown, fixed.
**Located in:** the control plane's consumer and enrolment, the inventory's tokens, and the host's
enrolment. Not a decision: the token stays "useless once used" (ADR 0004); what changed is when it
is used.