diff --git a/04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/01-diagnosis.md b/04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/01-diagnosis.md index 283c5da..e60b6f5 100644 --- a/04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/01-diagnosis.md +++ b/04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/01-diagnosis.md @@ -2,31 +2,56 @@ 1. **Build results, upgrade announcements and catch-up requests can simply be asked again.** Recording a build result ignores one already kept; acting on an upgrade pushes declarations, - which converge; replaying builds re-announces what is recorded. So each is handed back to the - broker while the store cannot be reached, exactly as a report is (issue 082): after a pause, for - at most two minutes, then let go with a line saying it was lost. Anything the store answered no - to is settled as before, so none of them can stop the queue behind it. A catch-up request used to - be acknowledged before the work; it is acknowledged after. -2. **An enrolment cannot be asked again as it was.** It spent the token first and then wrote the - node's key, its broker account and its other keys — in two databases, the token in the - inventory and the key in the identity store, so no one transaction could cover both. A store - gone between the two left a spent token and a node with no key, and the host, which generates - new keys on every attempt, could not recover without a person issuing a new token. -3. **So the token is claimed, not spent, until the enrolment is complete.** The presenting key + which converge; replaying builds re-announces what is recorded. A catch-up request used to be + acknowledged before the work; it is acknowledged after. +2. **But asking again must not stop the queue.** Issue 082's first shape slept on a message and + handed it back to the head of the queue. The control plane takes one message at a time, so + while one report waited for the store, every enrolment waited behind it — and a host waiting + on its answer gave up after ten seconds, while its request stayed queued and would later spend + the token for keys nobody held. So a message the store cannot take is now **held**: kept + unacknowledged, set aside, and tried again on a timer, while the loop goes on to the next + message. It is held for at most two minutes, then let go with a line saying it was lost; a + control plane that stops returns everything it held to the broker, since none of it was + acknowledged. Each message is held under its subject — a node's report, a module's upgrade, + the catch-up, each build on its own — and a newer message for the same subject sets the held + one aside, so an old report recorded late can never overwrite what the node is doing now. An + upgrade is held only when reading the store failed, not when a push did, so a push that timed + out is not repeated every few seconds. +3. **An enrolment cannot simply be asked again.** It spent the token first and then wrote the + node's key and its other keys — in two databases, the token in the inventory and the key in the + identity store, so no one transaction could cover both. A store gone between the two left a + spent token and a node with no key, and the host, which generates new keys on every attempt, + could not recover without a person issuing a new token. +4. **So the token is claimed, not spent, until the enrolment is complete.** The presenting key claims it for a short lease; every write that follows overwrites, so an attempt interrupted - part-way can be made again; and the token is spent last, only by the key holding the claim. A - second presenter is held off while the claim is live, and may take it once the lease lapses. A - store that cannot be asked, or a token held by another presenter, is answered "try again" - rather than refused, with nothing spent; the host asks again with the same request — the same - keys — for as long as the claim lasts, and then says the token was not spent. + part-way can be made again by the same presenter; the token is spent once everything is in the + store; and only then does the token's secret stop being the node's broker password — after the + spend, because a password replaced by an attempt that then failed would be held by nobody. A + second presenter is held off while the claim is live and may take it once the lease lapses. The + presenter that spent a token may claim it again, so an answer lost after the spend does not lock + out a machine the mesh holds as enrolled. A store that cannot be asked, or a token held by + another presenter, is answered "try again" rather than refused; the host asks again with the + same request — the same keys — for as long as the claim lasts, and then says to wait that long + before running enrol again. -**How it is checked:** unit tests hold each of the three handlers to handing back a message while -the store restarts and settling it otherwise; tests against a real store hold the claim to one -presenter, its lapse, and the spend to the claimant, and an enrolment met by a held token to "try -again"; a unit test holds the host to asking again while told to and stopping when its patience -runs out. The store-window bed stops the store, has a second machine enrol into the gap, starts the -store again, and asserts the enrolment completes on its own with one identity, and that the -machine then applies what it is pushed and is heard from. +**What is not closed.** An enrolment request sent to a control plane that is not running at all is +still queued; if the host gives up and the control plane starts later, the request completes for +keys nobody holds. That predates this issue and is not the store's window. + +**How it is checked:** unit tests hold a report to being held while the store is away and recorded +when it is back, to being set aside by a newer report from the same node, and to being let go past +the bound; they hold build results, catch-ups and upgrades to being held on the store and settled +otherwise, and an upgrade to not being held on a push that timed out. Tests against a real store +hold the claim to one presenter, its lapse, the spend to the claimant, and a claim again by the +presenter that spent it; an enrolment met by a held token is answered "try again". A unit test holds +the host to asking again while told to and stopping when its patience runs out. The store-window +bed pushes a slow step to the control-node, stops the store while its report is on the way, and +has a second machine enrol into the gap: the report is held, the enrolment is answered "not now" +meanwhile, and when the store is back the report is recorded and the enrolment completes on its +own — with the mesh holding exactly the keys the machine generated — after which the machine applies +what it is pushed and is heard from. Reviewed twice independently; the first review's findings — +the queue stopped behind a retry, the password replaced too early, a lost answer after the spend — +are what points 2 and 4 describe. **Located in:** the control plane's consumer and enrolment, the inventory's tokens, and the host's enrolment. Not a decision: the token stays "useless once used" (ADR 0004); what changed is when it