diff --git a/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/00-report.md b/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/00-report.md index 6e1d91d..c26f051 100644 --- a/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/00-report.md +++ b/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/00-report.md @@ -1,7 +1,8 @@ --- -status: open +status: resolved opened: 2026-09-22 located-in: [mesh-catalog modules/n8n, mesh-catalog modules/baserow] +fixed-by: mesh-catalog multiple-fixes (neither module takes the shared cache — baserow runs its own, n8n in its shipped mode needs none); mesh-lab multiple-fixes (a check that every catalogue module taking the cache presents the login; the two-node bed asserts baserow's own cache answers) --- # 081 — A cache consumer does not use the login it was granted diff --git a/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md b/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md new file mode 100644 index 0000000..6beeda3 --- /dev/null +++ b/04-ISSUES/081-a-cache-consumer-does-not-use-the-login-it-was-granted/01-diagnosis.md @@ -0,0 +1,49 @@ +# Diagnosis — 2026-09-22 + +Read from the two pinned images rather than from their documentation. + +1. **Presenting the login is not enough.** The cache confines each consumer to keys under its own + login (issue 080). A consumer that sends the login but writes its keys under names of its own + choosing is refused on every write, which is worse than the default user it replaced. So the + question for each module was whether its software can put every key, and every pub/sub channel, + under the login. +2. **n8n cannot, and in its shipped mode does not need to.** It has a username setting, and its + queue and cache keys take configurable prefixes, so the keys could move. Its pub/sub channels + are fixed names in its code, and a per-consumer grant allows no channels. Those channels are + used only in queue mode. The catalogue runs n8n in its regular mode, where it opens no connection + to the cache at all: the requirement bought a login nothing used. +3. **baserow cannot.** It has a username setting, and its caches, task queue, results, scheduler, + channel layer and task locks each take a prefix a settings module could set — none of them from + the environment. But its own code also writes under names fixed in the code: presence keys, a + recently-used-workspaces key that is updated by default, rate-limit keys, and a health check that + reads the task queue by its unprefixed name. Those cannot be moved without changing baserow. Its + container, run without a cache address, starts a cache of its own inside the container, under a + password it generates on first start and keeps on its data volume; that is its standard + single-container deployment. +4. **So neither module takes the shared cache.** baserow uses its own; n8n needs none until + someone runs it in queue mode, which would need a cache of its own for the same pub/sub reason. + The shared cache remains for software that keeps its keys under the login it is given, which the + grant bed proves. + +**How it is checked:** a mesh-lab unit test reads every catalogue module that requires the shared +cache and fails unless it hands its software the login it was granted; it fails on the catalogue as +it was, naming both modules, and passes now. The two-node bed runs baserow as the catalogue shapes +it and asserts its own cache answers inside its container and that baserow logged no refusal. The +same bed had carried a copy of baserow that took the shared cache with the password alone, and +passed only because it never looked at the cache. + +**Located in:** the two modules. Not a decision: a module choosing its own cache over a shared one it +cannot share is the module's. + +*On review.* An independent review re-read both images and confirmed the chain above step by step, +corrected which baserow keys are movable (the task-lock prefix is, through a settings module), and +tightened the bed: it now asserts baserow said it chose its own cache and that the cache challenges +for its password, rather than looking for the absence of errors, which a startup race could have +produced. Upgrading a running node drops any task still queued in the shared cache; the provider +removes the consumer's login when the grant goes, a path this change does not exercise. + +*Proven, 2026-09-22.* The two-node bed passes with baserow as the catalogue shapes it. On the way it +found the module's data directory narrower than the image ships it: the cache baserow runs for +itself runs as another user, who could not reach its own directory under a directory closed to +everyone but baserow's user. The module declares the image's mode now. The shared cache had hidden +this: baserow never ran its own until this change. diff --git a/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md b/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md new file mode 100644 index 0000000..b794135 --- /dev/null +++ b/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md @@ -0,0 +1,33 @@ +--- +status: resolved +opened: 2026-09-22 +located-in: [mesh-controller internal/link (serve, enrolment), mesh-controller internal/inventory] +fixed-by: mesh-controller multiple-fixes (a report the store cannot take right now is handed back to the broker after a short pause; any other failure is still acknowledged); unit tests on both halves; the two-node bed +--- + +# 082 — A report that arrives while the store restarts is lost, and the node is never heard from + +## Symptom + +A push that adopts the foundation's store on the control-node recreates the store. The node +applies the declaration and publishes its report once. If the report reaches the control plane +while the store is still starting, recording it fails — "the database system is starting up" — +and the control plane logs the failure and acknowledges the report anyway. The node never +reports that apply again: its periodic re-apply does not report. The mesh shows the node as sent +and never answered, while every container on it is up and doing what it was told. + +Found as an intermittent timeout in the two-node lab bed — twice in three runs — read off the +control plane's own log on machines kept for the purpose. + +## Why it matters beyond the instance + +Adopting the store and broker is the first thing a control-node does at genesis and the first +thing a migration does. Anything that restarts the store the control plane writes into — an +upgrade of it, a host reboot, a failover — opens the same window. A report is the only way the +mesh learns a node did what it was told, and it is sent exactly once. + +## What would close it + +A failure that means "not now" — the store unreachable or starting — leaves the report with the +broker to be asked again, while a failure that is an answer is still acknowledged so it cannot +come back for ever. A test on each half, and the bed that found it passing. diff --git a/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md b/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md new file mode 100644 index 0000000..64ec06a --- /dev/null +++ b/04-ISSUES/082-a-report-that-arrives-while-the-store-restarts-is-lost/01-diagnosis.md @@ -0,0 +1,34 @@ +# Diagnosis — 2026-09-22 + +1. The host publishes a report once per apply and acknowledges the declaration after publishing; + that half is sound, and the host logged no failure. The control plane's consumer takes one + message at a time and acknowledged every report, recorded or not. +2. The store it records into is the foundation store, which the adoption push had just recreated. + The control plane's log carried the proof: the report was received and could not be written + because the store was starting up. +3. Fixed in the control plane. The inventory says whether an error means the store could not be + asked right now — a connection that failed, or the server's "connection exception" and + "operator intervention" classes, which include starting and shutting down. The recorder marks + such a failure as worth another attempt; the consumer waits a moment and hands the report back + to the broker, which delivers it again. Any other failure is acknowledged as before, so a + report for a node the mesh does not know cannot spin. + +**How it is checked:** a unit test classifies a starting store and a lost connection as "not now" +and a constraint the store enforced as an answer; a unit test holds the consumer to handing back a +report marked "not now" and acknowledging one refused or recorded. The two-node bed, which found +it, runs the adoption push. + +**Located in:** the control plane's report consumer and its recorder. Not a decision. + +*On review.* An independent review stopped, killed and restarted a real store under the recorder +and found two gaps. A store killed rather than stopped surfaces as a bare "unexpected EOF", which +was not recognised as an outage, so the report was still lost; and a failed connection was always +taken for an outage, so a wrong password would have been asked again every two seconds for ever, +holding every enrolment and report behind it. The classification now checks the server's answer +first — a wrong password, a cancelled statement, a dropped database are answers — and counts a +network error, a connection that ended mid-conversation, a timeout, or anything the driver marks +safe to retry as an outage. One report holds the queue at most two minutes, then is let go with a +line saying it was lost; shutting down does not wait out the pause. What remains: a failed report +retried after a partial write counts its failure twice (the recorder's writes are not one +transaction), and the other control messages have the same loss +([issue 083](../083-other-control-messages-are-lost-while-the-store-restarts/00-report.md)). diff --git a/04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/00-report.md b/04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/00-report.md new file mode 100644 index 0000000..f8fbe68 --- /dev/null +++ b/04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/00-report.md @@ -0,0 +1,39 @@ +--- +status: open +opened: 2026-09-22 +located-in: [mesh-controller internal/link (serve: enrolment, build results, upgrade and catch-up announcements)] +--- + +# 083 — Other control messages are lost while the store restarts + +## Symptom + +[Issue 082](../082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md) kept a +node's report for another attempt while the store is restarting. The same control queue carries +four other kinds of message, and each is still settled on a failed write as if it had been +handled: + +- **An enrolment** is answered with a refusal when the store cannot be asked, so a valid token is + refused. Worse, if the store goes after the token is spent and before the node's key is kept, + the node holds no key and needs a new token — and trying the enrolment again is not safe, + because spending a token cannot be done twice. +- **A build result** is rejected on any failure to record it, and nothing announces it again. +- **An upgrade announcement** is acknowledged whatever happens, so a module that moved while the + store was down is never sent to the machines that run it. +- **A catch-up announcement** is acknowledged before it is acted on, so it is lost until the + catalogue restarts. + +Found by review of 082, by reading the consumer. Not reproduced. + +## Why it matters beyond the instance + +A migration adopts the store while nodes enrol, builds run and modules move — the window 082 +found for reports, for every other message the control plane cannot afford to drop. + +## What would close it + +Each kind decided on its own: which can be handed back and asked again (build results and +announcements, once recording them is shown to be safe to repeat), and which needs a different +shape (enrolment, whose token must be spent once — spent and keyed in one write, or answered "not +now" so the node asks again with the same token). A test for each, and a bed that restarts the +store while a node enrols and a build lands.