From f6ded3102dab31bd54f4cc9f3bef067d00f9e2b3 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 22 Sep 2026 13:34:05 +0200 Subject: [PATCH] Issue 083 opened (other control messages lost while the store restarts); 082's diagnosis carries its review --- .../01-diagnosis.md | 13 +++++++ .../00-report.md | 39 +++++++++++++++++++ 2 files changed, 52 insertions(+) create mode 100644 04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/00-report.md diff --git a/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md b/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md index 2f20b8d..64ec06a 100644 --- a/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md +++ b/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md @@ -19,3 +19,16 @@ report marked "not now" and acknowledging one refused or recorded. The two-node it, runs the adoption push. **Located in:** the control plane's report consumer and its recorder. Not a decision. + +*On review.* An independent review stopped, killed and restarted a real store under the recorder +and found two gaps. A store killed rather than stopped surfaces as a bare "unexpected EOF", which +was not recognised as an outage, so the report was still lost; and a failed connection was always +taken for an outage, so a wrong password would have been asked again every two seconds for ever, +holding every enrolment and report behind it. The classification now checks the server's answer +first — a wrong password, a cancelled statement, a dropped database are answers — and counts a +network error, a connection that ended mid-conversation, a timeout, or anything the driver marks +safe to retry as an outage. One report holds the queue at most two minutes, then is let go with a +line saying it was lost; shutting down does not wait out the pause. What remains: a failed report +retried after a partial write counts its failure twice (the recorder's writes are not one +transaction), and the other control messages have the same loss +([issue 083](../083-other-control-messages-are-lost-while-the-store-restarts/00-report.md)). diff --git a/04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/00-report.md b/04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/00-report.md new file mode 100644 index 0000000..f8fbe68 --- /dev/null +++ b/04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/00-report.md @@ -0,0 +1,39 @@ +--- +status: open +opened: 2026-09-22 +located-in: [mesh-controller internal/link (serve: enrolment, build results, upgrade and catch-up announcements)] +--- + +# 083 — Other control messages are lost while the store restarts + +## Symptom + +[Issue 082](../082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md) kept a +node's report for another attempt while the store is restarting. The same control queue carries +four other kinds of message, and each is still settled on a failed write as if it had been +handled: + +- **An enrolment** is answered with a refusal when the store cannot be asked, so a valid token is + refused. Worse, if the store goes after the token is spent and before the node's key is kept, + the node holds no key and needs a new token — and trying the enrolment again is not safe, + because spending a token cannot be done twice. +- **A build result** is rejected on any failure to record it, and nothing announces it again. +- **An upgrade announcement** is acknowledged whatever happens, so a module that moved while the + store was down is never sent to the machines that run it. +- **A catch-up announcement** is acknowledged before it is acted on, so it is lost until the + catalogue restarts. + +Found by review of 082, by reading the consumer. Not reproduced. + +## Why it matters beyond the instance + +A migration adopts the store while nodes enrol, builds run and modules move — the window 082 +found for reports, for every other message the control plane cannot afford to drop. + +## What would close it + +Each kind decided on its own: which can be handed back and asked again (build results and +announcements, once recording them is shown to be safe to repeat), and which needs a different +shape (enrolment, whose token must be spent once — spent and keyed in one write, or answered "not +now" so the node asks again with the same token). A test for each, and a bed that restarts the +store while a node enrols and a build lands.