Issue 083: the diagnosis describes held messages, the enrolment's order, and what is not closed
This commit is contained in:
+48
-23
@@ -2,31 +2,56 @@
|
||||
|
||||
1. **Build results, upgrade announcements and catch-up requests can simply be asked again.**
|
||||
Recording a build result ignores one already kept; acting on an upgrade pushes declarations,
|
||||
which converge; replaying builds re-announces what is recorded. So each is handed back to the
|
||||
broker while the store cannot be reached, exactly as a report is (issue 082): after a pause, for
|
||||
at most two minutes, then let go with a line saying it was lost. Anything the store answered no
|
||||
to is settled as before, so none of them can stop the queue behind it. A catch-up request used to
|
||||
be acknowledged before the work; it is acknowledged after.
|
||||
2. **An enrolment cannot be asked again as it was.** It spent the token first and then wrote the
|
||||
node's key, its broker account and its other keys — in two databases, the token in the
|
||||
inventory and the key in the identity store, so no one transaction could cover both. A store
|
||||
gone between the two left a spent token and a node with no key, and the host, which generates
|
||||
new keys on every attempt, could not recover without a person issuing a new token.
|
||||
3. **So the token is claimed, not spent, until the enrolment is complete.** The presenting key
|
||||
which converge; replaying builds re-announces what is recorded. A catch-up request used to be
|
||||
acknowledged before the work; it is acknowledged after.
|
||||
2. **But asking again must not stop the queue.** Issue 082's first shape slept on a message and
|
||||
handed it back to the head of the queue. The control plane takes one message at a time, so
|
||||
while one report waited for the store, every enrolment waited behind it — and a host waiting
|
||||
on its answer gave up after ten seconds, while its request stayed queued and would later spend
|
||||
the token for keys nobody held. So a message the store cannot take is now **held**: kept
|
||||
unacknowledged, set aside, and tried again on a timer, while the loop goes on to the next
|
||||
message. It is held for at most two minutes, then let go with a line saying it was lost; a
|
||||
control plane that stops returns everything it held to the broker, since none of it was
|
||||
acknowledged. Each message is held under its subject — a node's report, a module's upgrade,
|
||||
the catch-up, each build on its own — and a newer message for the same subject sets the held
|
||||
one aside, so an old report recorded late can never overwrite what the node is doing now. An
|
||||
upgrade is held only when reading the store failed, not when a push did, so a push that timed
|
||||
out is not repeated every few seconds.
|
||||
3. **An enrolment cannot simply be asked again.** It spent the token first and then wrote the
|
||||
node's key and its other keys — in two databases, the token in the inventory and the key in the
|
||||
identity store, so no one transaction could cover both. A store gone between the two left a
|
||||
spent token and a node with no key, and the host, which generates new keys on every attempt,
|
||||
could not recover without a person issuing a new token.
|
||||
4. **So the token is claimed, not spent, until the enrolment is complete.** The presenting key
|
||||
claims it for a short lease; every write that follows overwrites, so an attempt interrupted
|
||||
part-way can be made again; and the token is spent last, only by the key holding the claim. A
|
||||
second presenter is held off while the claim is live, and may take it once the lease lapses. A
|
||||
store that cannot be asked, or a token held by another presenter, is answered "try again"
|
||||
rather than refused, with nothing spent; the host asks again with the same request — the same
|
||||
keys — for as long as the claim lasts, and then says the token was not spent.
|
||||
part-way can be made again by the same presenter; the token is spent once everything is in the
|
||||
store; and only then does the token's secret stop being the node's broker password — after the
|
||||
spend, because a password replaced by an attempt that then failed would be held by nobody. A
|
||||
second presenter is held off while the claim is live and may take it once the lease lapses. The
|
||||
presenter that spent a token may claim it again, so an answer lost after the spend does not lock
|
||||
out a machine the mesh holds as enrolled. A store that cannot be asked, or a token held by
|
||||
another presenter, is answered "try again" rather than refused; the host asks again with the
|
||||
same request — the same keys — for as long as the claim lasts, and then says to wait that long
|
||||
before running enrol again.
|
||||
|
||||
**How it is checked:** unit tests hold each of the three handlers to handing back a message while
|
||||
the store restarts and settling it otherwise; tests against a real store hold the claim to one
|
||||
presenter, its lapse, and the spend to the claimant, and an enrolment met by a held token to "try
|
||||
again"; a unit test holds the host to asking again while told to and stopping when its patience
|
||||
runs out. The store-window bed stops the store, has a second machine enrol into the gap, starts the
|
||||
store again, and asserts the enrolment completes on its own with one identity, and that the
|
||||
machine then applies what it is pushed and is heard from.
|
||||
**What is not closed.** An enrolment request sent to a control plane that is not running at all is
|
||||
still queued; if the host gives up and the control plane starts later, the request completes for
|
||||
keys nobody holds. That predates this issue and is not the store's window.
|
||||
|
||||
**How it is checked:** unit tests hold a report to being held while the store is away and recorded
|
||||
when it is back, to being set aside by a newer report from the same node, and to being let go past
|
||||
the bound; they hold build results, catch-ups and upgrades to being held on the store and settled
|
||||
otherwise, and an upgrade to not being held on a push that timed out. Tests against a real store
|
||||
hold the claim to one presenter, its lapse, the spend to the claimant, and a claim again by the
|
||||
presenter that spent it; an enrolment met by a held token is answered "try again". A unit test holds
|
||||
the host to asking again while told to and stopping when its patience runs out. The store-window
|
||||
bed pushes a slow step to the control-node, stops the store while its report is on the way, and
|
||||
has a second machine enrol into the gap: the report is held, the enrolment is answered "not now"
|
||||
meanwhile, and when the store is back the report is recorded and the enrolment completes on its
|
||||
own — with the mesh holding exactly the keys the machine generated — after which the machine applies
|
||||
what it is pushed and is heard from. Reviewed twice independently; the first review's findings —
|
||||
the queue stopped behind a retry, the password replaced too early, a lost answer after the spend —
|
||||
are what points 2 and 4 describe.
|
||||
|
||||
**Located in:** the control plane's consumer and enrolment, the inventory's tokens, and the host's
|
||||
enrolment. Not a decision: the token stays "useless once used" (ADR 0004); what changed is when it
|
||||
|
||||
Reference in New Issue
Block a user