Issue 082: a report that arrives while the store restarts was lost; 081's diagnosis corrected on review
This commit is contained in:
+15
-6
@@ -12,12 +12,14 @@ Read from the two pinned images rather than from their documentation.
|
|||||||
are fixed names in its code, and a per-consumer grant allows no channels. Those channels are
|
are fixed names in its code, and a per-consumer grant allows no channels. Those channels are
|
||||||
used only in queue mode. The catalogue runs n8n in its regular mode, where it opens no connection
|
used only in queue mode. The catalogue runs n8n in its regular mode, where it opens no connection
|
||||||
to the cache at all: the requirement bought a login nothing used.
|
to the cache at all: the requirement bought a login nothing used.
|
||||||
3. **baserow cannot.** It has a username setting, and its caches, task queue, results, scheduler and
|
3. **baserow cannot.** It has a username setting, and its caches, task queue, results, scheduler,
|
||||||
channel layer each take a prefix a settings module could set. But its own code also writes
|
channel layer and task locks each take a prefix a settings module could set — none of them from
|
||||||
through raw connections under fixed names: rate-limit keys, singleton locks, a queue-size health
|
the environment. But its own code also writes under names fixed in the code: presence keys, a
|
||||||
check, an async client. Those cannot be moved without changing baserow. Its container, run
|
recently-used-workspaces key that is updated by default, rate-limit keys, and a health check that
|
||||||
without a cache address, starts a cache of its own inside the container, under a password it
|
reads the task queue by its unprefixed name. Those cannot be moved without changing baserow. Its
|
||||||
makes itself; that is its standard single-container deployment.
|
container, run without a cache address, starts a cache of its own inside the container, under a
|
||||||
|
password it generates on first start and keeps on its data volume; that is its standard
|
||||||
|
single-container deployment.
|
||||||
4. **So neither module takes the shared cache.** baserow uses its own; n8n needs none until
|
4. **So neither module takes the shared cache.** baserow uses its own; n8n needs none until
|
||||||
someone runs it in queue mode, which would need a cache of its own for the same pub/sub reason.
|
someone runs it in queue mode, which would need a cache of its own for the same pub/sub reason.
|
||||||
The shared cache remains for software that keeps its keys under the login it is given, which the
|
The shared cache remains for software that keeps its keys under the login it is given, which the
|
||||||
@@ -32,3 +34,10 @@ passed only because it never looked at the cache.
|
|||||||
|
|
||||||
**Located in:** the two modules. Not a decision: a module choosing its own cache over a shared one it
|
**Located in:** the two modules. Not a decision: a module choosing its own cache over a shared one it
|
||||||
cannot share is the module's.
|
cannot share is the module's.
|
||||||
|
|
||||||
|
*On review.* An independent review re-read both images and confirmed the chain above step by step,
|
||||||
|
corrected which baserow keys are movable (the task-lock prefix is, through a settings module), and
|
||||||
|
tightened the bed: it now asserts baserow said it chose its own cache and that the cache challenges
|
||||||
|
for its password, rather than looking for the absence of errors, which a startup race could have
|
||||||
|
produced. Upgrading a running node drops any task still queued in the shared cache; the provider
|
||||||
|
removes the consumer's login when the grant goes, a path this change does not exercise.
|
||||||
|
|||||||
@@ -0,0 +1,33 @@
|
|||||||
|
---
|
||||||
|
status: resolved
|
||||||
|
opened: 2026-09-22
|
||||||
|
located-in: [mesh-controller internal/link (serve, enrolment), mesh-controller internal/inventory]
|
||||||
|
fixed-by: mesh-controller multiple-fixes (a report the store cannot take right now is handed back to the broker after a short pause; any other failure is still acknowledged); unit tests on both halves; the two-node bed
|
||||||
|
---
|
||||||
|
|
||||||
|
# 082 — A report that arrives while the store restarts is lost, and the node is never heard from
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
A push that adopts the foundation's store on the control-node recreates the store. The node
|
||||||
|
applies the declaration and publishes its report once. If the report reaches the control plane
|
||||||
|
while the store is still starting, recording it fails — "the database system is starting up" —
|
||||||
|
and the control plane logs the failure and acknowledges the report anyway. The node never
|
||||||
|
reports that apply again: its periodic re-apply does not report. The mesh shows the node as sent
|
||||||
|
and never answered, while every container on it is up and doing what it was told.
|
||||||
|
|
||||||
|
Found as an intermittent timeout in the two-node lab bed — twice in three runs — read off the
|
||||||
|
control plane's own log on machines kept for the purpose.
|
||||||
|
|
||||||
|
## Why it matters beyond the instance
|
||||||
|
|
||||||
|
Adopting the store and broker is the first thing a control-node does at genesis and the first
|
||||||
|
thing a migration does. Anything that restarts the store the control plane writes into — an
|
||||||
|
upgrade of it, a host reboot, a failover — opens the same window. A report is the only way the
|
||||||
|
mesh learns a node did what it was told, and it is sent exactly once.
|
||||||
|
|
||||||
|
## What would close it
|
||||||
|
|
||||||
|
A failure that means "not now" — the store unreachable or starting — leaves the report with the
|
||||||
|
broker to be asked again, while a failure that is an answer is still acknowledged so it cannot
|
||||||
|
come back for ever. A test on each half, and the bed that found it passing.
|
||||||
@@ -0,0 +1,21 @@
|
|||||||
|
# Diagnosis — 2026-09-22
|
||||||
|
|
||||||
|
1. The host publishes a report once per apply and acknowledges the declaration after publishing;
|
||||||
|
that half is sound, and the host logged no failure. The control plane's consumer takes one
|
||||||
|
message at a time and acknowledged every report, recorded or not.
|
||||||
|
2. The store it records into is the foundation store, which the adoption push had just recreated.
|
||||||
|
The control plane's log carried the proof: the report was received and could not be written
|
||||||
|
because the store was starting up.
|
||||||
|
3. Fixed in the control plane. The inventory says whether an error means the store could not be
|
||||||
|
asked right now — a connection that failed, or the server's "connection exception" and
|
||||||
|
"operator intervention" classes, which include starting and shutting down. The recorder marks
|
||||||
|
such a failure as worth another attempt; the consumer waits a moment and hands the report back
|
||||||
|
to the broker, which delivers it again. Any other failure is acknowledged as before, so a
|
||||||
|
report for a node the mesh does not know cannot spin.
|
||||||
|
|
||||||
|
**How it is checked:** a unit test classifies a starting store and a lost connection as "not now"
|
||||||
|
and a constraint the store enforced as an answer; a unit test holds the consumer to handing back a
|
||||||
|
report marked "not now" and acknowledging one refused or recorded. The two-node bed, which found
|
||||||
|
it, runs the adoption push.
|
||||||
|
|
||||||
|
**Located in:** the control plane's report consumer and its recorder. Not a decision.
|
||||||
Reference in New Issue
Block a user