Issue 083 opened (other control messages lost while the store restarts); 082's diagnosis carries its review
This commit is contained in:
@@ -19,3 +19,16 @@ report marked "not now" and acknowledging one refused or recorded. The two-node
|
||||
it, runs the adoption push.
|
||||
|
||||
**Located in:** the control plane's report consumer and its recorder. Not a decision.
|
||||
|
||||
*On review.* An independent review stopped, killed and restarted a real store under the recorder
|
||||
and found two gaps. A store killed rather than stopped surfaces as a bare "unexpected EOF", which
|
||||
was not recognised as an outage, so the report was still lost; and a failed connection was always
|
||||
taken for an outage, so a wrong password would have been asked again every two seconds for ever,
|
||||
holding every enrolment and report behind it. The classification now checks the server's answer
|
||||
first — a wrong password, a cancelled statement, a dropped database are answers — and counts a
|
||||
network error, a connection that ended mid-conversation, a timeout, or anything the driver marks
|
||||
safe to retry as an outage. One report holds the queue at most two minutes, then is let go with a
|
||||
line saying it was lost; shutting down does not wait out the pause. What remains: a failed report
|
||||
retried after a partial write counts its failure twice (the recorder's writes are not one
|
||||
transaction), and the other control messages have the same loss
|
||||
([issue 083](../083-other-control-messages-are-lost-while-the-store-restarts/00-report.md)).
|
||||
|
||||
@@ -0,0 +1,39 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-22
|
||||
located-in: [mesh-controller internal/link (serve: enrolment, build results, upgrade and catch-up announcements)]
|
||||
---
|
||||
|
||||
# 083 — Other control messages are lost while the store restarts
|
||||
|
||||
## Symptom
|
||||
|
||||
[Issue 082](../082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md) kept a
|
||||
node's report for another attempt while the store is restarting. The same control queue carries
|
||||
four other kinds of message, and each is still settled on a failed write as if it had been
|
||||
handled:
|
||||
|
||||
- **An enrolment** is answered with a refusal when the store cannot be asked, so a valid token is
|
||||
refused. Worse, if the store goes after the token is spent and before the node's key is kept,
|
||||
the node holds no key and needs a new token — and trying the enrolment again is not safe,
|
||||
because spending a token cannot be done twice.
|
||||
- **A build result** is rejected on any failure to record it, and nothing announces it again.
|
||||
- **An upgrade announcement** is acknowledged whatever happens, so a module that moved while the
|
||||
store was down is never sent to the machines that run it.
|
||||
- **A catch-up announcement** is acknowledged before it is acted on, so it is lost until the
|
||||
catalogue restarts.
|
||||
|
||||
Found by review of 082, by reading the consumer. Not reproduced.
|
||||
|
||||
## Why it matters beyond the instance
|
||||
|
||||
A migration adopts the store while nodes enrol, builds run and modules move — the window 082
|
||||
found for reports, for every other message the control plane cannot afford to drop.
|
||||
|
||||
## What would close it
|
||||
|
||||
Each kind decided on its own: which can be handed back and asked again (build results and
|
||||
announcements, once recording them is shown to be safe to repeat), and which needs a different
|
||||
shape (enrolment, whose token must be spent once — spent and keyed in one write, or answered "not
|
||||
now" so the node asks again with the same token). A test for each, and a bed that restarts the
|
||||
store while a node enrols and a build lands.
|
||||
Reference in New Issue
Block a user