diff --git a/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md b/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md index b7c8c97..114aa7c 100644 --- a/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md +++ b/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md @@ -12,12 +12,14 @@ Read from the two pinned images rather than from their documentation. are fixed names in its code, and a per-consumer grant allows no channels. Those channels are used only in queue mode. The catalogue runs n8n in its regular mode, where it opens no connection to the cache at all: the requirement bought a login nothing used. -3. **baserow cannot.** It has a username setting, and its caches, task queue, results, scheduler and - channel layer each take a prefix a settings module could set. But its own code also writes - through raw connections under fixed names: rate-limit keys, singleton locks, a queue-size health - check, an async client. Those cannot be moved without changing baserow. Its container, run - without a cache address, starts a cache of its own inside the container, under a password it - makes itself; that is its standard single-container deployment. +3. **baserow cannot.** It has a username setting, and its caches, task queue, results, scheduler, + channel layer and task locks each take a prefix a settings module could set — none of them from + the environment. But its own code also writes under names fixed in the code: presence keys, a + recently-used-workspaces key that is updated by default, rate-limit keys, and a health check that + reads the task queue by its unprefixed name. Those cannot be moved without changing baserow. Its + container, run without a cache address, starts a cache of its own inside the container, under a + password it generates on first start and keeps on its data volume; that is its standard + single-container deployment. 4. **So neither module takes the shared cache.** baserow uses its own; n8n needs none until someone runs it in queue mode, which would need a cache of its own for the same pub/sub reason. The shared cache remains for software that keeps its keys under the login it is given, which the @@ -32,3 +34,10 @@ passed only because it never looked at the cache. **Located in:** the two modules. Not a decision: a module choosing its own cache over a shared one it cannot share is the module's. + +*On review.* An independent review re-read both images and confirmed the chain above step by step, +corrected which baserow keys are movable (the task-lock prefix is, through a settings module), and +tightened the bed: it now asserts baserow said it chose its own cache and that the cache challenges +for its password, rather than looking for the absence of errors, which a startup race could have +produced. Upgrading a running node drops any task still queued in the shared cache; the provider +removes the consumer's login when the grant goes, a path this change does not exercise. diff --git a/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md b/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md new file mode 100644 index 0000000..b794135 --- /dev/null +++ b/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md @@ -0,0 +1,33 @@ +--- +status: resolved +opened: 2026-09-22 +located-in: [mesh-controller internal/link (serve, enrolment), mesh-controller internal/inventory] +fixed-by: mesh-controller multiple-fixes (a report the store cannot take right now is handed back to the broker after a short pause; any other failure is still acknowledged); unit tests on both halves; the two-node bed +--- + +# 082 — A report that arrives while the store restarts is lost, and the node is never heard from + +## Symptom + +A push that adopts the foundation's store on the control-node recreates the store. The node +applies the declaration and publishes its report once. If the report reaches the control plane +while the store is still starting, recording it fails — "the database system is starting up" — +and the control plane logs the failure and acknowledges the report anyway. The node never +reports that apply again: its periodic re-apply does not report. The mesh shows the node as sent +and never answered, while every container on it is up and doing what it was told. + +Found as an intermittent timeout in the two-node lab bed — twice in three runs — read off the +control plane's own log on machines kept for the purpose. + +## Why it matters beyond the instance + +Adopting the store and broker is the first thing a control-node does at genesis and the first +thing a migration does. Anything that restarts the store the control plane writes into — an +upgrade of it, a host reboot, a failover — opens the same window. A report is the only way the +mesh learns a node did what it was told, and it is sent exactly once. + +## What would close it + +A failure that means "not now" — the store unreachable or starting — leaves the report with the +broker to be asked again, while a failure that is an answer is still acknowledged so it cannot +come back for ever. A test on each half, and the bed that found it passing. diff --git a/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md b/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md new file mode 100644 index 0000000..2f20b8d --- /dev/null +++ b/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md @@ -0,0 +1,21 @@ +# Diagnosis — 2026-09-22 + +1. The host publishes a report once per apply and acknowledges the declaration after publishing; + that half is sound, and the host logged no failure. The control plane's consumer takes one + message at a time and acknowledged every report, recorded or not. +2. The store it records into is the foundation store, which the adoption push had just recreated. + The control plane's log carried the proof: the report was received and could not be written + because the store was starting up. +3. Fixed in the control plane. The inventory says whether an error means the store could not be + asked right now — a connection that failed, or the server's "connection exception" and + "operator intervention" classes, which include starting and shutting down. The recorder marks + such a failure as worth another attempt; the consumer waits a moment and hands the report back + to the broker, which delivers it again. Any other failure is acknowledged as before, so a + report for a node the mesh does not know cannot spin. + +**How it is checked:** a unit test classifies a starting store and a lost connection as "not now" +and a constraint the store enforced as an answer; a unit test holds the consumer to handing back a +report marked "not now" and acknowledging one refused or recorded. The two-node bed, which found +it, runs the adoption push. + +**Located in:** the control plane's report consumer and its recorder. Not a decision.