Issue 083 opened (other control messages lost while the store restarts); 082's diagnosis carries its review

This commit is contained in:
2026-09-22 13:34:05 +02:00
parent 9a75f2b6a9
commit f6ded3102d
2 changed files with 52 additions and 0 deletions
@@ -19,3 +19,16 @@ report marked "not now" and acknowledging one refused or recorded. The two-node
it, runs the adoption push.
**Located in:** the control plane's report consumer and its recorder. Not a decision.
*On review.* An independent review stopped, killed and restarted a real store under the recorder
and found two gaps. A store killed rather than stopped surfaces as a bare "unexpected EOF", which
was not recognised as an outage, so the report was still lost; and a failed connection was always
taken for an outage, so a wrong password would have been asked again every two seconds for ever,
holding every enrolment and report behind it. The classification now checks the server's answer
first — a wrong password, a cancelled statement, a dropped database are answers — and counts a
network error, a connection that ended mid-conversation, a timeout, or anything the driver marks
safe to retry as an outage. One report holds the queue at most two minutes, then is let go with a
line saying it was lost; shutting down does not wait out the pause. What remains: a failed report
retried after a partial write counts its failure twice (the recorder's writes are not one
transaction), and the other control messages have the same loss
([issue 083](../083-other-control-messages-are-lost-while-the-store-restarts/00-report.md)).
@@ -0,0 +1,39 @@
---
status: open
opened: 2026-09-22
located-in: [mesh-controller internal/link (serve: enrolment, build results, upgrade and catch-up announcements)]
---
# 083 — Other control messages are lost while the store restarts
## Symptom
[Issue 082](../082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md) kept a
node's report for another attempt while the store is restarting. The same control queue carries
four other kinds of message, and each is still settled on a failed write as if it had been
handled:
- **An enrolment** is answered with a refusal when the store cannot be asked, so a valid token is
refused. Worse, if the store goes after the token is spent and before the node's key is kept,
the node holds no key and needs a new token — and trying the enrolment again is not safe,
because spending a token cannot be done twice.
- **A build result** is rejected on any failure to record it, and nothing announces it again.
- **An upgrade announcement** is acknowledged whatever happens, so a module that moved while the
store was down is never sent to the machines that run it.
- **A catch-up announcement** is acknowledged before it is acted on, so it is lost until the
catalogue restarts.
Found by review of 082, by reading the consumer. Not reproduced.
## Why it matters beyond the instance
A migration adopts the store while nodes enrol, builds run and modules move — the window 082
found for reports, for every other message the control plane cannot afford to drop.
## What would close it
Each kind decided on its own: which can be handed back and asked again (build results and
announcements, once recording them is shown to be safe to repeat), and which needs a different
shape (enrolment, whose token must be spent once — spent and keyed in one write, or answered "not
now" so the node asks again with the same token). A test for each, and a bed that restarts the
store while a node enrols and a build lands.