2.5 KiB
Diagnosis — 2026-09-22
- The host publishes a report once per apply and acknowledges the declaration after publishing; that half is sound, and the host logged no failure. The control plane's consumer takes one message at a time and acknowledged every report, recorded or not.
- The store it records into is the foundation store, which the adoption push had just recreated. The control plane's log carried the proof: the report was received and could not be written because the store was starting up.
- Fixed in the control plane. The inventory says whether an error means the store could not be asked right now — a connection that failed, or the server's "connection exception" and "operator intervention" classes, which include starting and shutting down. The recorder marks such a failure as worth another attempt; the consumer waits a moment and hands the report back to the broker, which delivers it again. Any other failure is acknowledged as before, so a report for a node the mesh does not know cannot spin.
How it is checked: a unit test classifies a starting store and a lost connection as "not now" and a constraint the store enforced as an answer; a unit test holds the consumer to handing back a report marked "not now" and acknowledging one refused or recorded. The two-node bed, which found it, runs the adoption push.
Located in: the control plane's report consumer and its recorder. Not a decision.
On review. An independent review stopped, killed and restarted a real store under the recorder and found two gaps. A store killed rather than stopped surfaces as a bare "unexpected EOF", which was not recognised as an outage, so the report was still lost; and a failed connection was always taken for an outage, so a wrong password would have been asked again every two seconds for ever, holding every enrolment and report behind it. The classification now checks the server's answer first — a wrong password, a cancelled statement, a dropped database are answers — and counts a network error, a connection that ended mid-conversation, a timeout, or anything the driver marks safe to retry as an outage. One report holds the queue at most two minutes, then is let go with a line saying it was lost; shutting down does not wait out the pause. What remains: a failed report retried after a partial write counts its failure twice (the recorder's writes are not one transaction), and the other control messages have the same loss (issue 083).