Files
hq/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md
T

2.5 KiB

Diagnosis — 2026-09-22

  1. The host publishes a report once per apply and acknowledges the declaration after publishing; that half is sound, and the host logged no failure. The control plane's consumer takes one message at a time and acknowledged every report, recorded or not.
  2. The store it records into is the foundation store, which the adoption push had just recreated. The control plane's log carried the proof: the report was received and could not be written because the store was starting up.
  3. Fixed in the control plane. The inventory says whether an error means the store could not be asked right now — a connection that failed, or the server's "connection exception" and "operator intervention" classes, which include starting and shutting down. The recorder marks such a failure as worth another attempt; the consumer waits a moment and hands the report back to the broker, which delivers it again. Any other failure is acknowledged as before, so a report for a node the mesh does not know cannot spin.

How it is checked: a unit test classifies a starting store and a lost connection as "not now" and a constraint the store enforced as an answer; a unit test holds the consumer to handing back a report marked "not now" and acknowledging one refused or recorded. The two-node bed, which found it, runs the adoption push.

Located in: the control plane's report consumer and its recorder. Not a decision.

On review. An independent review stopped, killed and restarted a real store under the recorder and found two gaps. A store killed rather than stopped surfaces as a bare "unexpected EOF", which was not recognised as an outage, so the report was still lost; and a failed connection was always taken for an outage, so a wrong password would have been asked again every two seconds for ever, holding every enrolment and report behind it. The classification now checks the server's answer first — a wrong password, a cancelled statement, a dropped database are answers — and counts a network error, a connection that ended mid-conversation, a timeout, or anything the driver marks safe to retry as an outage. One report holds the queue at most two minutes, then is let go with a line saying it was lost; shutting down does not wait out the pause. What remains: a failed report retried after a partial write counts its failure twice (the recorder's writes are not one transaction), and the other control messages have the same loss (issue 083).