Files
hq/04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/01-diagnosis.md

6.6 KiB

Diagnosis — 2026-09-22

  1. Build results, upgrade announcements and catch-up requests can simply be asked again. Recording a build result ignores one already kept; acting on an upgrade pushes declarations, which converge; replaying builds re-announces what is recorded. A catch-up request used to be acknowledged before the work; it is acknowledged after.
  2. But asking again must not stop the queue. Issue 082's first shape slept on a message and handed it back to the head of the queue. The control plane takes one message at a time, so while one report waited for the store, every enrolment waited behind it — and a host waiting on its answer gave up after ten seconds, while its request stayed queued and would later spend the token for keys nobody held. So a message the store cannot take is now held: kept unacknowledged, set aside, and tried again on a timer, while the loop goes on to the next message. It is held for at most two minutes, then let go with a line saying it was lost; a control plane that stops returns everything it held to the broker, since none of it was acknowledged. Each message is held under its subject — a node's report, a module's upgrade, the catch-up, each build on its own — and a newer message for the same subject sets the held one aside, so an old report recorded late can never overwrite what the node is doing now. An upgrade is held only when reading the store failed, not when a push did, so a push that timed out is not repeated every few seconds.
  3. An enrolment cannot simply be asked again. It spent the token first and then wrote the node's key and its other keys — in two databases, the token in the inventory and the key in the identity store, so no one transaction could cover both. A store gone between the two left a spent token and a node with no key, and the host, which generates new keys on every attempt, could not recover without a person issuing a new token.
  4. So the token is claimed, not spent, until the enrolment is complete. The presenting key claims it for a short lease; every write that follows overwrites, so an attempt interrupted part-way can be made again by the same presenter; the token is spent once everything is in the store; and only then does the token's secret stop being the node's broker password — after the spend, because a password replaced by an attempt that then failed would be held by nobody. A second presenter is held off while the claim is live and may take it once the lease lapses.
  5. Finishing after the spend takes proof. An answer lost after the spend would lock out a machine the mesh holds as enrolled, so the presenter that spent a token may finish — but a presenter is named by its public key, and a public key is no secret. The first version let anyone holding a leaked token and a node's public key replay the spent token, overwrite the node's keys and take its broker password; review found it. The host now signs its request with the identity key it generated, over the token and every key it presents; the mesh lets a spent token be finished only with a signature that verifies, while the claim's lease is live, and never on a request the broker handed over a second time, which may already have been answered. Without the proof a spent token is spent to everyone: "useless once used" (ADR 0004) holds, with one narrow exception — the node's own signed request, replayed word for word within the two minutes after its spend, finishes again. That takes reading the request off the control plane's queue, where the token's secret travels anyway, so only the broker's operator or the control plane itself could do it. A store that cannot be asked, or a token held by another presenter, is answered "try again" rather than refused; the host asks again with the same request — the same keys — for as long as the claim lasts, and then says to wait that long before running enrol again.

What is not closed. An enrolment request sent to a control plane that is not running at all is still queued; if the host gives up and the control plane starts later, the request completes for keys nobody holds. That predates this issue and is not the store's window. The host still gives up on its first unanswered ask rather than asking again with the same keys. A control plane that stops after spending a token and before answering refuses the redelivered request, since it may already have answered; the node is then told its token cannot be used, and needs a new one — rare, and said so in the control plane's log. A mesh past about fifty nodes whose reports all arrive while the store is away holds what the prefetch allows and lets the rest go, loudly.

How it is checked: unit tests hold a report to being held while the store is away and recorded when it is back, to being set aside by a newer report from the same node, and to being let go past the bound; they hold build results, catch-ups and upgrades to being held on the store and settled otherwise, and an upgrade to not being held on a push that timed out. Tests against a real store hold the claim to one presenter, its lapse, the spend to the claimant, and a claim again by the presenter that spent it only with proof and inside its lease; an enrolment met by a held token is answered "try again", and a spent token presented with a public key alone, with a proof made by another key, with a proof made for another request, or redelivered, is refused. A known answer for the signed bytes is repeated in both repositories' tests. A unit test holds the host to asking again while told to and stopping when its patience runs out. The store-window bed pushes a slow step to the control-node, stops the store while its report is on the way, and has a second machine enrol into the gap: the report is held, the enrolment is answered "not now" meanwhile, and when the store is back the report is recorded and the enrolment completes on its own — with the mesh holding exactly the keys the machine generated — after which the machine applies what it is pushed and is heard from. Two independent reviews shaped this: the first found the queue stopped behind a retry and the password replaced too early (points 2 and 4); the second found the replay (point 5), messages lost on shutdown, and identical build results stranding a delivery; a third confirmed the replay closed and found one handler still settling on shutdown, fixed.

Located in: the control plane's consumer and enrolment, the inventory's tokens, and the host's enrolment. Not a decision: the token stays "useless once used" (ADR 0004); what changed is when it is used.