Issue 083 resolved: nothing the control queue carries is lost while the store restarts; an enrolment claims its token and spends it last

This commit is contained in:
2026-09-22 14:11:44 +02:00
parent 4e74cf29e4
commit 2d45aa5f42
2 changed files with 36 additions and 2 deletions
@@ -1,7 +1,8 @@
---
status: open
status: resolved
opened: 2026-09-22
located-in: [mesh-controller internal/link (serve: enrolment, build results, upgrade and catch-up announcements)]
located-in: [mesh-controller internal/link (serve, enrolment), mesh-controller internal/inventory (tokens), mesh-host internal/link (enrol)]
fixed-by: mesh-controller multiple-fixes (build results, upgrades and catch-ups handed back while the store is away, bounded; an enrolment claims its token and spends it last, and is answered "try again"); mesh-host multiple-fixes (the host asks again with the same request); mesh-lab multiple-fixes (the store-window bed)
---
# 083 — Other control messages are lost while the store restarts
@@ -0,0 +1,33 @@
# Diagnosis — 2026-09-22
1. **Build results, upgrade announcements and catch-up requests can simply be asked again.**
Recording a build result ignores one already kept; acting on an upgrade pushes declarations,
which converge; replaying builds re-announces what is recorded. So each is handed back to the
broker while the store cannot be reached, exactly as a report is (issue 082): after a pause, for
at most two minutes, then let go with a line saying it was lost. Anything the store answered no
to is settled as before, so none of them can stop the queue behind it. A catch-up request used to
be acknowledged before the work; it is acknowledged after.
2. **An enrolment cannot be asked again as it was.** It spent the token first and then wrote the
node's key, its broker account and its other keys — in two databases, the token in the
inventory and the key in the identity store, so no one transaction could cover both. A store
gone between the two left a spent token and a node with no key, and the host, which generates
new keys on every attempt, could not recover without a person issuing a new token.
3. **So the token is claimed, not spent, until the enrolment is complete.** The presenting key
claims it for a short lease; every write that follows overwrites, so an attempt interrupted
part-way can be made again; and the token is spent last, only by the key holding the claim. A
second presenter is held off while the claim is live, and may take it once the lease lapses. A
store that cannot be asked, or a token held by another presenter, is answered "try again"
rather than refused, with nothing spent; the host asks again with the same request — the same
keys — for as long as the claim lasts, and then says the token was not spent.
**How it is checked:** unit tests hold each of the three handlers to handing back a message while
the store restarts and settling it otherwise; tests against a real store hold the claim to one
presenter, its lapse, and the spend to the claimant, and an enrolment met by a held token to "try
again"; a unit test holds the host to asking again while told to and stopping when its patience
runs out. The store-window bed stops the store, has a second machine enrol into the gap, starts the
store again, and asserts the enrolment completes on its own with one identity, and that the
machine then applies what it is pushed and is heard from.
**Located in:** the control plane's consumer and enrolment, the inventory's tokens, and the host's
enrolment. Not a decision: the token stays "useless once used" (ADR 0004); what changed is when it
is used.