34 lines
1.8 KiB
Markdown
34 lines
1.8 KiB
Markdown
---
|
|
status: resolved
|
|
opened: 2026-09-22
|
|
located-in: [mesh-controller internal/link (serve, enrolment), mesh-controller internal/inventory]
|
|
fixed-by: mesh-controller multiple-fixes (a report the store cannot take right now is handed back to the broker after a short pause; any other failure is still acknowledged); unit tests on both halves; the two-node bed
|
|
---
|
|
|
|
# 082 — A report that arrives while the store restarts is lost, and the node is never heard from
|
|
|
|
## Symptom
|
|
|
|
A push that adopts the foundation's store on the control-node recreates the store. The node
|
|
applies the declaration and publishes its report once. If the report reaches the control plane
|
|
while the store is still starting, recording it fails — "the database system is starting up" —
|
|
and the control plane logs the failure and acknowledges the report anyway. The node never
|
|
reports that apply again: its periodic re-apply does not report. The mesh shows the node as sent
|
|
and never answered, while every container on it is up and doing what it was told.
|
|
|
|
Found as an intermittent timeout in the two-node lab bed — twice in three runs — read off the
|
|
control plane's own log on machines kept for the purpose.
|
|
|
|
## Why it matters beyond the instance
|
|
|
|
Adopting the store and broker is the first thing a control-node does at genesis and the first
|
|
thing a migration does. Anything that restarts the store the control plane writes into — an
|
|
upgrade of it, a host reboot, a failover — opens the same window. A report is the only way the
|
|
mesh learns a node did what it was told, and it is sent exactly once.
|
|
|
|
## What would close it
|
|
|
|
A failure that means "not now" — the store unreachable or starting — leaves the report with the
|
|
broker to be asked again, while a failure that is an answer is still acknowledged so it cannot
|
|
come back for ever. A test on each half, and the bed that found it passing.
|