Files
hq/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md
T

34 lines
1.8 KiB
Markdown

---
status: resolved
opened: 2026-09-22
located-in: [mesh-controller internal/link (serve, enrolment), mesh-controller internal/inventory]
fixed-by: mesh-controller multiple-fixes (a report the store cannot take right now is handed back to the broker after a short pause; any other failure is still acknowledged); unit tests on both halves; the two-node bed
---
# 082 — A report that arrives while the store restarts is lost, and the node is never heard from
## Symptom
A push that adopts the foundation's store on the control-node recreates the store. The node
applies the declaration and publishes its report once. If the report reaches the control plane
while the store is still starting, recording it fails — "the database system is starting up" —
and the control plane logs the failure and acknowledges the report anyway. The node never
reports that apply again: its periodic re-apply does not report. The mesh shows the node as sent
and never answered, while every container on it is up and doing what it was told.
Found as an intermittent timeout in the two-node lab bed — twice in three runs — read off the
control plane's own log on machines kept for the purpose.
## Why it matters beyond the instance
Adopting the store and broker is the first thing a control-node does at genesis and the first
thing a migration does. Anything that restarts the store the control plane writes into — an
upgrade of it, a host reboot, a failover — opens the same window. A report is the only way the
mesh learns a node did what it was told, and it is sent exactly once.
## What would close it
A failure that means "not now" — the store unreachable or starting — leaves the report with the
broker to be asked again, while a failure that is an answer is still acknowledged so it cannot
come back for ever. A test on each half, and the bed that found it passing.