Issue 082: a report that arrives while the store restarts was lost; 081's diagnosis corrected on review

This commit is contained in:
2026-09-22 13:25:33 +02:00
parent 5d1a0372cb
commit 9a75f2b6a9
3 changed files with 69 additions and 6 deletions
@@ -12,12 +12,14 @@ Read from the two pinned images rather than from their documentation.
are fixed names in its code, and a per-consumer grant allows no channels. Those channels are
used only in queue mode. The catalogue runs n8n in its regular mode, where it opens no connection
to the cache at all: the requirement bought a login nothing used.
3. **baserow cannot.** It has a username setting, and its caches, task queue, results, scheduler and
channel layer each take a prefix a settings module could set. But its own code also writes
through raw connections under fixed names: rate-limit keys, singleton locks, a queue-size health
check, an async client. Those cannot be moved without changing baserow. Its container, run
without a cache address, starts a cache of its own inside the container, under a password it
makes itself; that is its standard single-container deployment.
3. **baserow cannot.** It has a username setting, and its caches, task queue, results, scheduler,
channel layer and task locks each take a prefix a settings module could set — none of them from
the environment. But its own code also writes under names fixed in the code: presence keys, a
recently-used-workspaces key that is updated by default, rate-limit keys, and a health check that
reads the task queue by its unprefixed name. Those cannot be moved without changing baserow. Its
container, run without a cache address, starts a cache of its own inside the container, under a
password it generates on first start and keeps on its data volume; that is its standard
single-container deployment.
4. **So neither module takes the shared cache.** baserow uses its own; n8n needs none until
someone runs it in queue mode, which would need a cache of its own for the same pub/sub reason.
The shared cache remains for software that keeps its keys under the login it is given, which the
@@ -32,3 +34,10 @@ passed only because it never looked at the cache.
**Located in:** the two modules. Not a decision: a module choosing its own cache over a shared one it
cannot share is the module's.
*On review.* An independent review re-read both images and confirmed the chain above step by step,
corrected which baserow keys are movable (the task-lock prefix is, through a settings module), and
tightened the bed: it now asserts baserow said it chose its own cache and that the cache challenges
for its password, rather than looking for the absence of errors, which a startup race could have
produced. Upgrading a running node drops any task still queued in the shared cache; the provider
removes the consumer's login when the grant goes, a path this change does not exercise.
@@ -0,0 +1,33 @@
---
status: resolved
opened: 2026-09-22
located-in: [mesh-controller internal/link (serve, enrolment), mesh-controller internal/inventory]
fixed-by: mesh-controller multiple-fixes (a report the store cannot take right now is handed back to the broker after a short pause; any other failure is still acknowledged); unit tests on both halves; the two-node bed
---
# 082 — A report that arrives while the store restarts is lost, and the node is never heard from
## Symptom
A push that adopts the foundation's store on the control-node recreates the store. The node
applies the declaration and publishes its report once. If the report reaches the control plane
while the store is still starting, recording it fails — "the database system is starting up" —
and the control plane logs the failure and acknowledges the report anyway. The node never
reports that apply again: its periodic re-apply does not report. The mesh shows the node as sent
and never answered, while every container on it is up and doing what it was told.
Found as an intermittent timeout in the two-node lab bed — twice in three runs — read off the
control plane's own log on machines kept for the purpose.
## Why it matters beyond the instance
Adopting the store and broker is the first thing a control-node does at genesis and the first
thing a migration does. Anything that restarts the store the control plane writes into — an
upgrade of it, a host reboot, a failover — opens the same window. A report is the only way the
mesh learns a node did what it was told, and it is sent exactly once.
## What would close it
A failure that means "not now" — the store unreachable or starting — leaves the report with the
broker to be asked again, while a failure that is an answer is still acknowledged so it cannot
come back for ever. A test on each half, and the bed that found it passing.
@@ -0,0 +1,21 @@
# Diagnosis — 2026-09-22
1. The host publishes a report once per apply and acknowledges the declaration after publishing;
that half is sound, and the host logged no failure. The control plane's consumer takes one
message at a time and acknowledged every report, recorded or not.
2. The store it records into is the foundation store, which the adoption push had just recreated.
The control plane's log carried the proof: the report was received and could not be written
because the store was starting up.
3. Fixed in the control plane. The inventory says whether an error means the store could not be
asked right now — a connection that failed, or the server's "connection exception" and
"operator intervention" classes, which include starting and shutting down. The recorder marks
such a failure as worth another attempt; the consumer waits a moment and hands the report back
to the broker, which delivers it again. Any other failure is acknowledged as before, so a
report for a node the mesh does not know cannot spin.
**How it is checked:** a unit test classifies a starting store and a lost connection as "not now"
and a constraint the store enforced as an answer; a unit test holds the consumer to handing back a
report marked "not now" and acknowledging one refused or recorded. The two-node bed, which found
it, runs the adoption push.
**Located in:** the control plane's report consumer and its recorder. Not a decision.