Issues 081 and 082 resolved; 083 opened #72
+15
-6
@@ -12,12 +12,14 @@ Read from the two pinned images rather than from their documentation.
|
||||
are fixed names in its code, and a per-consumer grant allows no channels. Those channels are
|
||||
used only in queue mode. The catalogue runs n8n in its regular mode, where it opens no connection
|
||||
to the cache at all: the requirement bought a login nothing used.
|
||||
3. **baserow cannot.** It has a username setting, and its caches, task queue, results, scheduler and
|
||||
channel layer each take a prefix a settings module could set. But its own code also writes
|
||||
through raw connections under fixed names: rate-limit keys, singleton locks, a queue-size health
|
||||
check, an async client. Those cannot be moved without changing baserow. Its container, run
|
||||
without a cache address, starts a cache of its own inside the container, under a password it
|
||||
makes itself; that is its standard single-container deployment.
|
||||
3. **baserow cannot.** It has a username setting, and its caches, task queue, results, scheduler,
|
||||
channel layer and task locks each take a prefix a settings module could set — none of them from
|
||||
the environment. But its own code also writes under names fixed in the code: presence keys, a
|
||||
recently-used-workspaces key that is updated by default, rate-limit keys, and a health check that
|
||||
reads the task queue by its unprefixed name. Those cannot be moved without changing baserow. Its
|
||||
container, run without a cache address, starts a cache of its own inside the container, under a
|
||||
password it generates on first start and keeps on its data volume; that is its standard
|
||||
single-container deployment.
|
||||
4. **So neither module takes the shared cache.** baserow uses its own; n8n needs none until
|
||||
someone runs it in queue mode, which would need a cache of its own for the same pub/sub reason.
|
||||
The shared cache remains for software that keeps its keys under the login it is given, which the
|
||||
@@ -32,3 +34,10 @@ passed only because it never looked at the cache.
|
||||
|
||||
**Located in:** the two modules. Not a decision: a module choosing its own cache over a shared one it
|
||||
cannot share is the module's.
|
||||
|
||||
*On review.* An independent review re-read both images and confirmed the chain above step by step,
|
||||
corrected which baserow keys are movable (the task-lock prefix is, through a settings module), and
|
||||
tightened the bed: it now asserts baserow said it chose its own cache and that the cache challenges
|
||||
for its password, rather than looking for the absence of errors, which a startup race could have
|
||||
produced. Upgrading a running node drops any task still queued in the shared cache; the provider
|
||||
removes the consumer's login when the grant goes, a path this change does not exercise.
|
||||
|
||||
@@ -0,0 +1,33 @@
|
||||
---
|
||||
status: resolved
|
||||
opened: 2026-09-22
|
||||
located-in: [mesh-controller internal/link (serve, enrolment), mesh-controller internal/inventory]
|
||||
fixed-by: mesh-controller multiple-fixes (a report the store cannot take right now is handed back to the broker after a short pause; any other failure is still acknowledged); unit tests on both halves; the two-node bed
|
||||
---
|
||||
|
||||
# 082 — A report that arrives while the store restarts is lost, and the node is never heard from
|
||||
|
||||
## Symptom
|
||||
|
||||
A push that adopts the foundation's store on the control-node recreates the store. The node
|
||||
applies the declaration and publishes its report once. If the report reaches the control plane
|
||||
while the store is still starting, recording it fails — "the database system is starting up" —
|
||||
and the control plane logs the failure and acknowledges the report anyway. The node never
|
||||
reports that apply again: its periodic re-apply does not report. The mesh shows the node as sent
|
||||
and never answered, while every container on it is up and doing what it was told.
|
||||
|
||||
Found as an intermittent timeout in the two-node lab bed — twice in three runs — read off the
|
||||
control plane's own log on machines kept for the purpose.
|
||||
|
||||
## Why it matters beyond the instance
|
||||
|
||||
Adopting the store and broker is the first thing a control-node does at genesis and the first
|
||||
thing a migration does. Anything that restarts the store the control plane writes into — an
|
||||
upgrade of it, a host reboot, a failover — opens the same window. A report is the only way the
|
||||
mesh learns a node did what it was told, and it is sent exactly once.
|
||||
|
||||
## What would close it
|
||||
|
||||
A failure that means "not now" — the store unreachable or starting — leaves the report with the
|
||||
broker to be asked again, while a failure that is an answer is still acknowledged so it cannot
|
||||
come back for ever. A test on each half, and the bed that found it passing.
|
||||
@@ -0,0 +1,21 @@
|
||||
# Diagnosis — 2026-09-22
|
||||
|
||||
1. The host publishes a report once per apply and acknowledges the declaration after publishing;
|
||||
that half is sound, and the host logged no failure. The control plane's consumer takes one
|
||||
message at a time and acknowledged every report, recorded or not.
|
||||
2. The store it records into is the foundation store, which the adoption push had just recreated.
|
||||
The control plane's log carried the proof: the report was received and could not be written
|
||||
because the store was starting up.
|
||||
3. Fixed in the control plane. The inventory says whether an error means the store could not be
|
||||
asked right now — a connection that failed, or the server's "connection exception" and
|
||||
"operator intervention" classes, which include starting and shutting down. The recorder marks
|
||||
such a failure as worth another attempt; the consumer waits a moment and hands the report back
|
||||
to the broker, which delivers it again. Any other failure is acknowledged as before, so a
|
||||
report for a node the mesh does not know cannot spin.
|
||||
|
||||
**How it is checked:** a unit test classifies a starting store and a lost connection as "not now"
|
||||
and a constraint the store enforced as an answer; a unit test holds the consumer to handing back a
|
||||
report marked "not now" and acknowledging one refused or recorded. The two-node bed, which found
|
||||
it, runs the adoption push.
|
||||
|
||||
**Located in:** the control plane's report consumer and its recorder. Not a decision.
|
||||
Reference in New Issue
Block a user