Merge pull request 'Issues 081 and 082 resolved; 083 opened' (#72) from multiple-fixes into main

This commit was merged in pull request #72.
This commit is contained in:
2026-09-22 13:53:16 +02:00
5 changed files with 157 additions and 1 deletions
@@ -1,7 +1,8 @@
---
status: open
status: resolved
opened: 2026-09-22
located-in: [mesh-catalog modules/n8n, mesh-catalog modules/baserow]
fixed-by: mesh-catalog multiple-fixes (neither module takes the shared cache — baserow runs its own, n8n in its shipped mode needs none); mesh-lab multiple-fixes (a check that every catalogue module taking the cache presents the login; the two-node bed asserts baserow's own cache answers)
---
# 081 — A cache consumer does not use the login it was granted
@@ -0,0 +1,49 @@
# Diagnosis — 2026-09-22
Read from the two pinned images rather than from their documentation.
1. **Presenting the login is not enough.** The cache confines each consumer to keys under its own
login (issue 080). A consumer that sends the login but writes its keys under names of its own
choosing is refused on every write, which is worse than the default user it replaced. So the
question for each module was whether its software can put every key, and every pub/sub channel,
under the login.
2. **n8n cannot, and in its shipped mode does not need to.** It has a username setting, and its
queue and cache keys take configurable prefixes, so the keys could move. Its pub/sub channels
are fixed names in its code, and a per-consumer grant allows no channels. Those channels are
used only in queue mode. The catalogue runs n8n in its regular mode, where it opens no connection
to the cache at all: the requirement bought a login nothing used.
3. **baserow cannot.** It has a username setting, and its caches, task queue, results, scheduler,
channel layer and task locks each take a prefix a settings module could set — none of them from
the environment. But its own code also writes under names fixed in the code: presence keys, a
recently-used-workspaces key that is updated by default, rate-limit keys, and a health check that
reads the task queue by its unprefixed name. Those cannot be moved without changing baserow. Its
container, run without a cache address, starts a cache of its own inside the container, under a
password it generates on first start and keeps on its data volume; that is its standard
single-container deployment.
4. **So neither module takes the shared cache.** baserow uses its own; n8n needs none until
someone runs it in queue mode, which would need a cache of its own for the same pub/sub reason.
The shared cache remains for software that keeps its keys under the login it is given, which the
grant bed proves.
**How it is checked:** a mesh-lab unit test reads every catalogue module that requires the shared
cache and fails unless it hands its software the login it was granted; it fails on the catalogue as
it was, naming both modules, and passes now. The two-node bed runs baserow as the catalogue shapes
it and asserts its own cache answers inside its container and that baserow logged no refusal. The
same bed had carried a copy of baserow that took the shared cache with the password alone, and
passed only because it never looked at the cache.
**Located in:** the two modules. Not a decision: a module choosing its own cache over a shared one it
cannot share is the module's.
*On review.* An independent review re-read both images and confirmed the chain above step by step,
corrected which baserow keys are movable (the task-lock prefix is, through a settings module), and
tightened the bed: it now asserts baserow said it chose its own cache and that the cache challenges
for its password, rather than looking for the absence of errors, which a startup race could have
produced. Upgrading a running node drops any task still queued in the shared cache; the provider
removes the consumer's login when the grant goes, a path this change does not exercise.
*Proven, 2026-09-22.* The two-node bed passes with baserow as the catalogue shapes it. On the way it
found the module's data directory narrower than the image ships it: the cache baserow runs for
itself runs as another user, who could not reach its own directory under a directory closed to
everyone but baserow's user. The module declares the image's mode now. The shared cache had hidden
this: baserow never ran its own until this change.
@@ -0,0 +1,33 @@
---
status: resolved
opened: 2026-09-22
located-in: [mesh-controller internal/link (serve, enrolment), mesh-controller internal/inventory]
fixed-by: mesh-controller multiple-fixes (a report the store cannot take right now is handed back to the broker after a short pause; any other failure is still acknowledged); unit tests on both halves; the two-node bed
---
# 082 — A report that arrives while the store restarts is lost, and the node is never heard from
## Symptom
A push that adopts the foundation's store on the control-node recreates the store. The node
applies the declaration and publishes its report once. If the report reaches the control plane
while the store is still starting, recording it fails — "the database system is starting up" —
and the control plane logs the failure and acknowledges the report anyway. The node never
reports that apply again: its periodic re-apply does not report. The mesh shows the node as sent
and never answered, while every container on it is up and doing what it was told.
Found as an intermittent timeout in the two-node lab bed — twice in three runs — read off the
control plane's own log on machines kept for the purpose.
## Why it matters beyond the instance
Adopting the store and broker is the first thing a control-node does at genesis and the first
thing a migration does. Anything that restarts the store the control plane writes into — an
upgrade of it, a host reboot, a failover — opens the same window. A report is the only way the
mesh learns a node did what it was told, and it is sent exactly once.
## What would close it
A failure that means "not now" — the store unreachable or starting — leaves the report with the
broker to be asked again, while a failure that is an answer is still acknowledged so it cannot
come back for ever. A test on each half, and the bed that found it passing.
@@ -0,0 +1,34 @@
# Diagnosis — 2026-09-22
1. The host publishes a report once per apply and acknowledges the declaration after publishing;
that half is sound, and the host logged no failure. The control plane's consumer takes one
message at a time and acknowledged every report, recorded or not.
2. The store it records into is the foundation store, which the adoption push had just recreated.
The control plane's log carried the proof: the report was received and could not be written
because the store was starting up.
3. Fixed in the control plane. The inventory says whether an error means the store could not be
asked right now — a connection that failed, or the server's "connection exception" and
"operator intervention" classes, which include starting and shutting down. The recorder marks
such a failure as worth another attempt; the consumer waits a moment and hands the report back
to the broker, which delivers it again. Any other failure is acknowledged as before, so a
report for a node the mesh does not know cannot spin.
**How it is checked:** a unit test classifies a starting store and a lost connection as "not now"
and a constraint the store enforced as an answer; a unit test holds the consumer to handing back a
report marked "not now" and acknowledging one refused or recorded. The two-node bed, which found
it, runs the adoption push.
**Located in:** the control plane's report consumer and its recorder. Not a decision.
*On review.* An independent review stopped, killed and restarted a real store under the recorder
and found two gaps. A store killed rather than stopped surfaces as a bare "unexpected EOF", which
was not recognised as an outage, so the report was still lost; and a failed connection was always
taken for an outage, so a wrong password would have been asked again every two seconds for ever,
holding every enrolment and report behind it. The classification now checks the server's answer
first — a wrong password, a cancelled statement, a dropped database are answers — and counts a
network error, a connection that ended mid-conversation, a timeout, or anything the driver marks
safe to retry as an outage. One report holds the queue at most two minutes, then is let go with a
line saying it was lost; shutting down does not wait out the pause. What remains: a failed report
retried after a partial write counts its failure twice (the recorder's writes are not one
transaction), and the other control messages have the same loss
([issue 083](../083-other-control-messages-are-lost-while-the-store-restarts/00-report.md)).
@@ -0,0 +1,39 @@
---
status: open
opened: 2026-09-22
located-in: [mesh-controller internal/link (serve: enrolment, build results, upgrade and catch-up announcements)]
---
# 083 — Other control messages are lost while the store restarts
## Symptom
[Issue 082](../082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md) kept a
node's report for another attempt while the store is restarting. The same control queue carries
four other kinds of message, and each is still settled on a failed write as if it had been
handled:
- **An enrolment** is answered with a refusal when the store cannot be asked, so a valid token is
refused. Worse, if the store goes after the token is spent and before the node's key is kept,
the node holds no key and needs a new token — and trying the enrolment again is not safe,
because spending a token cannot be done twice.
- **A build result** is rejected on any failure to record it, and nothing announces it again.
- **An upgrade announcement** is acknowledged whatever happens, so a module that moved while the
store was down is never sent to the machines that run it.
- **A catch-up announcement** is acknowledged before it is acted on, so it is lost until the
catalogue restarts.
Found by review of 082, by reading the consumer. Not reproduced.
## Why it matters beyond the instance
A migration adopts the store while nodes enrol, builds run and modules move — the window 082
found for reports, for every other message the control plane cannot afford to drop.
## What would close it
Each kind decided on its own: which can be handed back and asked again (build results and
announcements, once recording them is shown to be safe to repeat), and which needs a different
shape (enrolment, whose token must be spent once — spent and keyed in one write, or answered "not
now" so the node asks again with the same token). A test for each, and a bed that restarts the
store while a node enrols and a build lands.