35 lines
2.5 KiB
Markdown
35 lines
2.5 KiB
Markdown
# Diagnosis — 2026-09-22
|
|
|
|
1. The host publishes a report once per apply and acknowledges the declaration after publishing;
|
|
that half is sound, and the host logged no failure. The control plane's consumer takes one
|
|
message at a time and acknowledged every report, recorded or not.
|
|
2. The store it records into is the foundation store, which the adoption push had just recreated.
|
|
The control plane's log carried the proof: the report was received and could not be written
|
|
because the store was starting up.
|
|
3. Fixed in the control plane. The inventory says whether an error means the store could not be
|
|
asked right now — a connection that failed, or the server's "connection exception" and
|
|
"operator intervention" classes, which include starting and shutting down. The recorder marks
|
|
such a failure as worth another attempt; the consumer waits a moment and hands the report back
|
|
to the broker, which delivers it again. Any other failure is acknowledged as before, so a
|
|
report for a node the mesh does not know cannot spin.
|
|
|
|
**How it is checked:** a unit test classifies a starting store and a lost connection as "not now"
|
|
and a constraint the store enforced as an answer; a unit test holds the consumer to handing back a
|
|
report marked "not now" and acknowledging one refused or recorded. The two-node bed, which found
|
|
it, runs the adoption push.
|
|
|
|
**Located in:** the control plane's report consumer and its recorder. Not a decision.
|
|
|
|
*On review.* An independent review stopped, killed and restarted a real store under the recorder
|
|
and found two gaps. A store killed rather than stopped surfaces as a bare "unexpected EOF", which
|
|
was not recognised as an outage, so the report was still lost; and a failed connection was always
|
|
taken for an outage, so a wrong password would have been asked again every two seconds for ever,
|
|
holding every enrolment and report behind it. The classification now checks the server's answer
|
|
first — a wrong password, a cancelled statement, a dropped database are answers — and counts a
|
|
network error, a connection that ended mid-conversation, a timeout, or anything the driver marks
|
|
safe to retry as an outage. One report holds the queue at most two minutes, then is let go with a
|
|
line saying it was lost; shutting down does not wait out the pause. What remains: a failed report
|
|
retried after a partial write counts its failure twice (the recorder's writes are not one
|
|
transaction), and the other control messages have the same loss
|
|
([issue 083](../083-other-control-messages-are-lost-while-the-store-restarts/00-report.md)).
|