From 5d1a0372cb48479cde51b72d502129cd741b37b1 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 22 Sep 2026 12:27:12 +0200 Subject: [PATCH 1/4] Issue 081 resolved: neither cache consumer can keep its keys under its login, so neither takes the shared cache --- .../00-report.md | 3 +- .../01-diagnosis.md | 34 +++++++++++++++++++ 2 files changed, 36 insertions(+), 1 deletion(-) create mode 100644 04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md diff --git a/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/00-report.md b/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/00-report.md index 6e1d91d..c26f051 100644 --- a/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/00-report.md +++ b/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/00-report.md @@ -1,7 +1,8 @@ --- -status: open +status: resolved opened: 2026-09-22 located-in: [mesh-catalog modules/n8n, mesh-catalog modules/baserow] +fixed-by: mesh-catalog multiple-fixes (neither module takes the shared cache — baserow runs its own, n8n in its shipped mode needs none); mesh-lab multiple-fixes (a check that every catalogue module taking the cache presents the login; the two-node bed asserts baserow's own cache answers) --- # 081 — A cache consumer does not use the login it was granted diff --git a/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md b/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md new file mode 100644 index 0000000..b7c8c97 --- /dev/null +++ b/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md @@ -0,0 +1,34 @@ +# Diagnosis — 2026-09-22 + +Read from the two pinned images rather than from their documentation. + +1. **Presenting the login is not enough.** The cache confines each consumer to keys under its own + login (issue 080). A consumer that sends the login but writes its keys under names of its own + choosing is refused on every write, which is worse than the default user it replaced. So the + question for each module was whether its software can put every key, and every pub/sub channel, + under the login. +2. **n8n cannot, and in its shipped mode does not need to.** It has a username setting, and its + queue and cache keys take configurable prefixes, so the keys could move. Its pub/sub channels + are fixed names in its code, and a per-consumer grant allows no channels. Those channels are + used only in queue mode. The catalogue runs n8n in its regular mode, where it opens no connection + to the cache at all: the requirement bought a login nothing used. +3. **baserow cannot.** It has a username setting, and its caches, task queue, results, scheduler and + channel layer each take a prefix a settings module could set. But its own code also writes + through raw connections under fixed names: rate-limit keys, singleton locks, a queue-size health + check, an async client. Those cannot be moved without changing baserow. Its container, run + without a cache address, starts a cache of its own inside the container, under a password it + makes itself; that is its standard single-container deployment. +4. **So neither module takes the shared cache.** baserow uses its own; n8n needs none until + someone runs it in queue mode, which would need a cache of its own for the same pub/sub reason. + The shared cache remains for software that keeps its keys under the login it is given, which the + grant bed proves. + +**How it is checked:** a mesh-lab unit test reads every catalogue module that requires the shared +cache and fails unless it hands its software the login it was granted; it fails on the catalogue as +it was, naming both modules, and passes now. The two-node bed runs baserow as the catalogue shapes +it and asserts its own cache answers inside its container and that baserow logged no refusal. The +same bed had carried a copy of baserow that took the shared cache with the password alone, and +passed only because it never looked at the cache. + +**Located in:** the two modules. Not a decision: a module choosing its own cache over a shared one it +cannot share is the module's. -- 2.54.0 From 9a75f2b6a907bc4087b26053a7ddaff4ee8aca42 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 22 Sep 2026 13:25:33 +0200 Subject: [PATCH 2/4] Issue 082: a report that arrives while the store restarts was lost; 081's diagnosis corrected on review --- .../01-diagnosis.md | 21 ++++++++---- .../00-report.md | 33 +++++++++++++++++++ .../01-diagnosis.md | 21 ++++++++++++ 3 files changed, 69 insertions(+), 6 deletions(-) create mode 100644 04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md create mode 100644 04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md diff --git a/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md b/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md index b7c8c97..114aa7c 100644 --- a/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md +++ b/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md @@ -12,12 +12,14 @@ Read from the two pinned images rather than from their documentation. are fixed names in its code, and a per-consumer grant allows no channels. Those channels are used only in queue mode. The catalogue runs n8n in its regular mode, where it opens no connection to the cache at all: the requirement bought a login nothing used. -3. **baserow cannot.** It has a username setting, and its caches, task queue, results, scheduler and - channel layer each take a prefix a settings module could set. But its own code also writes - through raw connections under fixed names: rate-limit keys, singleton locks, a queue-size health - check, an async client. Those cannot be moved without changing baserow. Its container, run - without a cache address, starts a cache of its own inside the container, under a password it - makes itself; that is its standard single-container deployment. +3. **baserow cannot.** It has a username setting, and its caches, task queue, results, scheduler, + channel layer and task locks each take a prefix a settings module could set — none of them from + the environment. But its own code also writes under names fixed in the code: presence keys, a + recently-used-workspaces key that is updated by default, rate-limit keys, and a health check that + reads the task queue by its unprefixed name. Those cannot be moved without changing baserow. Its + container, run without a cache address, starts a cache of its own inside the container, under a + password it generates on first start and keeps on its data volume; that is its standard + single-container deployment. 4. **So neither module takes the shared cache.** baserow uses its own; n8n needs none until someone runs it in queue mode, which would need a cache of its own for the same pub/sub reason. The shared cache remains for software that keeps its keys under the login it is given, which the @@ -32,3 +34,10 @@ passed only because it never looked at the cache. **Located in:** the two modules. Not a decision: a module choosing its own cache over a shared one it cannot share is the module's. + +*On review.* An independent review re-read both images and confirmed the chain above step by step, +corrected which baserow keys are movable (the task-lock prefix is, through a settings module), and +tightened the bed: it now asserts baserow said it chose its own cache and that the cache challenges +for its password, rather than looking for the absence of errors, which a startup race could have +produced. Upgrading a running node drops any task still queued in the shared cache; the provider +removes the consumer's login when the grant goes, a path this change does not exercise. diff --git a/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md b/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md new file mode 100644 index 0000000..b794135 --- /dev/null +++ b/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md @@ -0,0 +1,33 @@ +--- +status: resolved +opened: 2026-09-22 +located-in: [mesh-controller internal/link (serve, enrolment), mesh-controller internal/inventory] +fixed-by: mesh-controller multiple-fixes (a report the store cannot take right now is handed back to the broker after a short pause; any other failure is still acknowledged); unit tests on both halves; the two-node bed +--- + +# 082 — A report that arrives while the store restarts is lost, and the node is never heard from + +## Symptom + +A push that adopts the foundation's store on the control-node recreates the store. The node +applies the declaration and publishes its report once. If the report reaches the control plane +while the store is still starting, recording it fails — "the database system is starting up" — +and the control plane logs the failure and acknowledges the report anyway. The node never +reports that apply again: its periodic re-apply does not report. The mesh shows the node as sent +and never answered, while every container on it is up and doing what it was told. + +Found as an intermittent timeout in the two-node lab bed — twice in three runs — read off the +control plane's own log on machines kept for the purpose. + +## Why it matters beyond the instance + +Adopting the store and broker is the first thing a control-node does at genesis and the first +thing a migration does. Anything that restarts the store the control plane writes into — an +upgrade of it, a host reboot, a failover — opens the same window. A report is the only way the +mesh learns a node did what it was told, and it is sent exactly once. + +## What would close it + +A failure that means "not now" — the store unreachable or starting — leaves the report with the +broker to be asked again, while a failure that is an answer is still acknowledged so it cannot +come back for ever. A test on each half, and the bed that found it passing. diff --git a/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md b/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md new file mode 100644 index 0000000..2f20b8d --- /dev/null +++ b/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md @@ -0,0 +1,21 @@ +# Diagnosis — 2026-09-22 + +1. The host publishes a report once per apply and acknowledges the declaration after publishing; + that half is sound, and the host logged no failure. The control plane's consumer takes one + message at a time and acknowledged every report, recorded or not. +2. The store it records into is the foundation store, which the adoption push had just recreated. + The control plane's log carried the proof: the report was received and could not be written + because the store was starting up. +3. Fixed in the control plane. The inventory says whether an error means the store could not be + asked right now — a connection that failed, or the server's "connection exception" and + "operator intervention" classes, which include starting and shutting down. The recorder marks + such a failure as worth another attempt; the consumer waits a moment and hands the report back + to the broker, which delivers it again. Any other failure is acknowledged as before, so a + report for a node the mesh does not know cannot spin. + +**How it is checked:** a unit test classifies a starting store and a lost connection as "not now" +and a constraint the store enforced as an answer; a unit test holds the consumer to handing back a +report marked "not now" and acknowledging one refused or recorded. The two-node bed, which found +it, runs the adoption push. + +**Located in:** the control plane's report consumer and its recorder. Not a decision. -- 2.54.0 From f6ded3102dab31bd54f4cc9f3bef067d00f9e2b3 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 22 Sep 2026 13:34:05 +0200 Subject: [PATCH 3/4] Issue 083 opened (other control messages lost while the store restarts); 082's diagnosis carries its review --- .../01-diagnosis.md | 13 +++++++ .../00-report.md | 39 +++++++++++++++++++ 2 files changed, 52 insertions(+) create mode 100644 04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/00-report.md diff --git a/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md b/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md index 2f20b8d..64ec06a 100644 --- a/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md +++ b/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md @@ -19,3 +19,16 @@ report marked "not now" and acknowledging one refused or recorded. The two-node it, runs the adoption push. **Located in:** the control plane's report consumer and its recorder. Not a decision. + +*On review.* An independent review stopped, killed and restarted a real store under the recorder +and found two gaps. A store killed rather than stopped surfaces as a bare "unexpected EOF", which +was not recognised as an outage, so the report was still lost; and a failed connection was always +taken for an outage, so a wrong password would have been asked again every two seconds for ever, +holding every enrolment and report behind it. The classification now checks the server's answer +first — a wrong password, a cancelled statement, a dropped database are answers — and counts a +network error, a connection that ended mid-conversation, a timeout, or anything the driver marks +safe to retry as an outage. One report holds the queue at most two minutes, then is let go with a +line saying it was lost; shutting down does not wait out the pause. What remains: a failed report +retried after a partial write counts its failure twice (the recorder's writes are not one +transaction), and the other control messages have the same loss +([issue 083](../083-other-control-messages-are-lost-while-the-store-restarts/00-report.md)). diff --git a/04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/00-report.md b/04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/00-report.md new file mode 100644 index 0000000..f8fbe68 --- /dev/null +++ b/04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/00-report.md @@ -0,0 +1,39 @@ +--- +status: open +opened: 2026-09-22 +located-in: [mesh-controller internal/link (serve: enrolment, build results, upgrade and catch-up announcements)] +--- + +# 083 — Other control messages are lost while the store restarts + +## Symptom + +[Issue 082](../082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md) kept a +node's report for another attempt while the store is restarting. The same control queue carries +four other kinds of message, and each is still settled on a failed write as if it had been +handled: + +- **An enrolment** is answered with a refusal when the store cannot be asked, so a valid token is + refused. Worse, if the store goes after the token is spent and before the node's key is kept, + the node holds no key and needs a new token — and trying the enrolment again is not safe, + because spending a token cannot be done twice. +- **A build result** is rejected on any failure to record it, and nothing announces it again. +- **An upgrade announcement** is acknowledged whatever happens, so a module that moved while the + store was down is never sent to the machines that run it. +- **A catch-up announcement** is acknowledged before it is acted on, so it is lost until the + catalogue restarts. + +Found by review of 082, by reading the consumer. Not reproduced. + +## Why it matters beyond the instance + +A migration adopts the store while nodes enrol, builds run and modules move — the window 082 +found for reports, for every other message the control plane cannot afford to drop. + +## What would close it + +Each kind decided on its own: which can be handed back and asked again (build results and +announcements, once recording them is shown to be safe to repeat), and which needs a different +shape (enrolment, whose token must be spent once — spent and keyed in one write, or answered "not +now" so the node asks again with the same token). A test for each, and a bed that restarts the +store while a node enrols and a build lands. -- 2.54.0 From 0d40918e92947339c184f02d8a36be8cda7ac1a1 Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 22 Sep 2026 13:53:00 +0200 Subject: [PATCH 4/4] Issue 081: proven by the two-node bed, and what it found about baserow's data directory --- .../01-diagnosis.md | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md b/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md index 114aa7c..6beeda3 100644 --- a/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md +++ b/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md @@ -41,3 +41,9 @@ tightened the bed: it now asserts baserow said it chose its own cache and that t for its password, rather than looking for the absence of errors, which a startup race could have produced. Upgrading a running node drops any task still queued in the shared cache; the provider removes the consumer's login when the grant goes, a path this change does not exercise. + +*Proven, 2026-09-22.* The two-node bed passes with baserow as the catalogue shapes it. On the way it +found the module's data directory narrower than the image ships it: the cache baserow runs for +itself runs as another user, who could not reach its own directory under a directory closed to +everyone but baserow's user. The module declares the image's mode now. The shared cache had hidden +this: baserow never ran its own until this change. -- 2.54.0