Issues 081 and 082 resolved; 083 opened #72
@@ -1,7 +1,8 @@
|
|||||||
---
|
---
|
||||||
status: open
|
status: resolved
|
||||||
opened: 2026-09-22
|
opened: 2026-09-22
|
||||||
located-in: [mesh-catalog modules/n8n, mesh-catalog modules/baserow]
|
located-in: [mesh-catalog modules/n8n, mesh-catalog modules/baserow]
|
||||||
|
fixed-by: mesh-catalog multiple-fixes (neither module takes the shared cache — baserow runs its own, n8n in its shipped mode needs none); mesh-lab multiple-fixes (a check that every catalogue module taking the cache presents the login; the two-node bed asserts baserow's own cache answers)
|
||||||
---
|
---
|
||||||
|
|
||||||
# 081 — A cache consumer does not use the login it was granted
|
# 081 — A cache consumer does not use the login it was granted
|
||||||
|
|||||||
@@ -0,0 +1,49 @@
|
|||||||
|
# Diagnosis — 2026-09-22
|
||||||
|
|
||||||
|
Read from the two pinned images rather than from their documentation.
|
||||||
|
|
||||||
|
1. **Presenting the login is not enough.** The cache confines each consumer to keys under its own
|
||||||
|
login (issue 080). A consumer that sends the login but writes its keys under names of its own
|
||||||
|
choosing is refused on every write, which is worse than the default user it replaced. So the
|
||||||
|
question for each module was whether its software can put every key, and every pub/sub channel,
|
||||||
|
under the login.
|
||||||
|
2. **n8n cannot, and in its shipped mode does not need to.** It has a username setting, and its
|
||||||
|
queue and cache keys take configurable prefixes, so the keys could move. Its pub/sub channels
|
||||||
|
are fixed names in its code, and a per-consumer grant allows no channels. Those channels are
|
||||||
|
used only in queue mode. The catalogue runs n8n in its regular mode, where it opens no connection
|
||||||
|
to the cache at all: the requirement bought a login nothing used.
|
||||||
|
3. **baserow cannot.** It has a username setting, and its caches, task queue, results, scheduler,
|
||||||
|
channel layer and task locks each take a prefix a settings module could set — none of them from
|
||||||
|
the environment. But its own code also writes under names fixed in the code: presence keys, a
|
||||||
|
recently-used-workspaces key that is updated by default, rate-limit keys, and a health check that
|
||||||
|
reads the task queue by its unprefixed name. Those cannot be moved without changing baserow. Its
|
||||||
|
container, run without a cache address, starts a cache of its own inside the container, under a
|
||||||
|
password it generates on first start and keeps on its data volume; that is its standard
|
||||||
|
single-container deployment.
|
||||||
|
4. **So neither module takes the shared cache.** baserow uses its own; n8n needs none until
|
||||||
|
someone runs it in queue mode, which would need a cache of its own for the same pub/sub reason.
|
||||||
|
The shared cache remains for software that keeps its keys under the login it is given, which the
|
||||||
|
grant bed proves.
|
||||||
|
|
||||||
|
**How it is checked:** a mesh-lab unit test reads every catalogue module that requires the shared
|
||||||
|
cache and fails unless it hands its software the login it was granted; it fails on the catalogue as
|
||||||
|
it was, naming both modules, and passes now. The two-node bed runs baserow as the catalogue shapes
|
||||||
|
it and asserts its own cache answers inside its container and that baserow logged no refusal. The
|
||||||
|
same bed had carried a copy of baserow that took the shared cache with the password alone, and
|
||||||
|
passed only because it never looked at the cache.
|
||||||
|
|
||||||
|
**Located in:** the two modules. Not a decision: a module choosing its own cache over a shared one it
|
||||||
|
cannot share is the module's.
|
||||||
|
|
||||||
|
*On review.* An independent review re-read both images and confirmed the chain above step by step,
|
||||||
|
corrected which baserow keys are movable (the task-lock prefix is, through a settings module), and
|
||||||
|
tightened the bed: it now asserts baserow said it chose its own cache and that the cache challenges
|
||||||
|
for its password, rather than looking for the absence of errors, which a startup race could have
|
||||||
|
produced. Upgrading a running node drops any task still queued in the shared cache; the provider
|
||||||
|
removes the consumer's login when the grant goes, a path this change does not exercise.
|
||||||
|
|
||||||
|
*Proven, 2026-09-22.* The two-node bed passes with baserow as the catalogue shapes it. On the way it
|
||||||
|
found the module's data directory narrower than the image ships it: the cache baserow runs for
|
||||||
|
itself runs as another user, who could not reach its own directory under a directory closed to
|
||||||
|
everyone but baserow's user. The module declares the image's mode now. The shared cache had hidden
|
||||||
|
this: baserow never ran its own until this change.
|
||||||
@@ -0,0 +1,33 @@
|
|||||||
|
---
|
||||||
|
status: resolved
|
||||||
|
opened: 2026-09-22
|
||||||
|
located-in: [mesh-controller internal/link (serve, enrolment), mesh-controller internal/inventory]
|
||||||
|
fixed-by: mesh-controller multiple-fixes (a report the store cannot take right now is handed back to the broker after a short pause; any other failure is still acknowledged); unit tests on both halves; the two-node bed
|
||||||
|
---
|
||||||
|
|
||||||
|
# 082 — A report that arrives while the store restarts is lost, and the node is never heard from
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
A push that adopts the foundation's store on the control-node recreates the store. The node
|
||||||
|
applies the declaration and publishes its report once. If the report reaches the control plane
|
||||||
|
while the store is still starting, recording it fails — "the database system is starting up" —
|
||||||
|
and the control plane logs the failure and acknowledges the report anyway. The node never
|
||||||
|
reports that apply again: its periodic re-apply does not report. The mesh shows the node as sent
|
||||||
|
and never answered, while every container on it is up and doing what it was told.
|
||||||
|
|
||||||
|
Found as an intermittent timeout in the two-node lab bed — twice in three runs — read off the
|
||||||
|
control plane's own log on machines kept for the purpose.
|
||||||
|
|
||||||
|
## Why it matters beyond the instance
|
||||||
|
|
||||||
|
Adopting the store and broker is the first thing a control-node does at genesis and the first
|
||||||
|
thing a migration does. Anything that restarts the store the control plane writes into — an
|
||||||
|
upgrade of it, a host reboot, a failover — opens the same window. A report is the only way the
|
||||||
|
mesh learns a node did what it was told, and it is sent exactly once.
|
||||||
|
|
||||||
|
## What would close it
|
||||||
|
|
||||||
|
A failure that means "not now" — the store unreachable or starting — leaves the report with the
|
||||||
|
broker to be asked again, while a failure that is an answer is still acknowledged so it cannot
|
||||||
|
come back for ever. A test on each half, and the bed that found it passing.
|
||||||
@@ -0,0 +1,34 @@
|
|||||||
|
# Diagnosis — 2026-09-22
|
||||||
|
|
||||||
|
1. The host publishes a report once per apply and acknowledges the declaration after publishing;
|
||||||
|
that half is sound, and the host logged no failure. The control plane's consumer takes one
|
||||||
|
message at a time and acknowledged every report, recorded or not.
|
||||||
|
2. The store it records into is the foundation store, which the adoption push had just recreated.
|
||||||
|
The control plane's log carried the proof: the report was received and could not be written
|
||||||
|
because the store was starting up.
|
||||||
|
3. Fixed in the control plane. The inventory says whether an error means the store could not be
|
||||||
|
asked right now — a connection that failed, or the server's "connection exception" and
|
||||||
|
"operator intervention" classes, which include starting and shutting down. The recorder marks
|
||||||
|
such a failure as worth another attempt; the consumer waits a moment and hands the report back
|
||||||
|
to the broker, which delivers it again. Any other failure is acknowledged as before, so a
|
||||||
|
report for a node the mesh does not know cannot spin.
|
||||||
|
|
||||||
|
**How it is checked:** a unit test classifies a starting store and a lost connection as "not now"
|
||||||
|
and a constraint the store enforced as an answer; a unit test holds the consumer to handing back a
|
||||||
|
report marked "not now" and acknowledging one refused or recorded. The two-node bed, which found
|
||||||
|
it, runs the adoption push.
|
||||||
|
|
||||||
|
**Located in:** the control plane's report consumer and its recorder. Not a decision.
|
||||||
|
|
||||||
|
*On review.* An independent review stopped, killed and restarted a real store under the recorder
|
||||||
|
and found two gaps. A store killed rather than stopped surfaces as a bare "unexpected EOF", which
|
||||||
|
was not recognised as an outage, so the report was still lost; and a failed connection was always
|
||||||
|
taken for an outage, so a wrong password would have been asked again every two seconds for ever,
|
||||||
|
holding every enrolment and report behind it. The classification now checks the server's answer
|
||||||
|
first — a wrong password, a cancelled statement, a dropped database are answers — and counts a
|
||||||
|
network error, a connection that ended mid-conversation, a timeout, or anything the driver marks
|
||||||
|
safe to retry as an outage. One report holds the queue at most two minutes, then is let go with a
|
||||||
|
line saying it was lost; shutting down does not wait out the pause. What remains: a failed report
|
||||||
|
retried after a partial write counts its failure twice (the recorder's writes are not one
|
||||||
|
transaction), and the other control messages have the same loss
|
||||||
|
([issue 083](../083-other-control-messages-are-lost-while-the-store-restarts/00-report.md)).
|
||||||
@@ -0,0 +1,39 @@
|
|||||||
|
---
|
||||||
|
status: open
|
||||||
|
opened: 2026-09-22
|
||||||
|
located-in: [mesh-controller internal/link (serve: enrolment, build results, upgrade and catch-up announcements)]
|
||||||
|
---
|
||||||
|
|
||||||
|
# 083 — Other control messages are lost while the store restarts
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
[Issue 082](../082-a-report-that-arrives-while-the-store-restarts-is-lost/00-report.md) kept a
|
||||||
|
node's report for another attempt while the store is restarting. The same control queue carries
|
||||||
|
four other kinds of message, and each is still settled on a failed write as if it had been
|
||||||
|
handled:
|
||||||
|
|
||||||
|
- **An enrolment** is answered with a refusal when the store cannot be asked, so a valid token is
|
||||||
|
refused. Worse, if the store goes after the token is spent and before the node's key is kept,
|
||||||
|
the node holds no key and needs a new token — and trying the enrolment again is not safe,
|
||||||
|
because spending a token cannot be done twice.
|
||||||
|
- **A build result** is rejected on any failure to record it, and nothing announces it again.
|
||||||
|
- **An upgrade announcement** is acknowledged whatever happens, so a module that moved while the
|
||||||
|
store was down is never sent to the machines that run it.
|
||||||
|
- **A catch-up announcement** is acknowledged before it is acted on, so it is lost until the
|
||||||
|
catalogue restarts.
|
||||||
|
|
||||||
|
Found by review of 082, by reading the consumer. Not reproduced.
|
||||||
|
|
||||||
|
## Why it matters beyond the instance
|
||||||
|
|
||||||
|
A migration adopts the store while nodes enrol, builds run and modules move — the window 082
|
||||||
|
found for reports, for every other message the control plane cannot afford to drop.
|
||||||
|
|
||||||
|
## What would close it
|
||||||
|
|
||||||
|
Each kind decided on its own: which can be handed back and asked again (build results and
|
||||||
|
announcements, once recording them is shown to be safe to repeat), and which needs a different
|
||||||
|
shape (enrolment, whose token must be spent once — spent and keyed in one write, or answered "not
|
||||||
|
now" so the node asks again with the same token). A test for each, and a bed that restarts the
|
||||||
|
store while a node enrols and a build lands.
|
||||||
Reference in New Issue
Block a user