From fcf3b6f4d559ce177fef3025da52842ec3cb5294 Mon Sep 17 00:00:00 2001 From: jochen Date: Wed, 16 Sep 2026 22:04:28 +0200 Subject: [PATCH] =?UTF-8?q?Issues=20054,=20055=20=E2=80=94=20the=20debt=20?= =?UTF-8?q?adopting=20the=20store=20and=20broker=20leaves?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 054: the adopted servers bind 0.0.0.0 from genesis but the packet filter is installed later, so there is a window where they are open with only bootstrap credentials. 055: the servers bind on the control-node and the one-node bed cannot prove a consumer on another machine can reach them over the overlay. Both are follow-ups to issue 051's adoption (WBS Phase 3), tracked rather than rushed. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx --- .../00-report.md | 48 +++++++++++++++++++ .../00-report.md | 40 ++++++++++++++++ 2 files changed, 88 insertions(+) create mode 100644 04-ISSUES/054-the-adopted-store-and-broker-are-open-before-the-filter/00-report.md create mode 100644 04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md diff --git a/04-ISSUES/054-the-adopted-store-and-broker-are-open-before-the-filter/00-report.md b/04-ISSUES/054-the-adopted-store-and-broker-are-open-before-the-filter/00-report.md new file mode 100644 index 0000000..d2bc8d1 --- /dev/null +++ b/04-ISSUES/054-the-adopted-store-and-broker-are-open-before-the-filter/00-report.md @@ -0,0 +1,48 @@ +--- +status: open +opened: 2026-09-16 +located-in: [] +fixed-by: +amended-design: +--- + +# 054 — The adopted store and broker are open before the packet filter exists + +## Symptom + +Adopting the store and broker as ordinary modules (issue 051) needs them reachable by +consumers across the mesh, so genesis now raises `mesh-store` and `mesh-broker` bound to +`0.0.0.0` rather than to loopback as the foundation used to. The packet filter is a module, +installed several steps after the store, the broker and the catalogue are already running. + +Between each server coming up and the filter being applied, both listen on every interface +of the machine with only the genesis bootstrap credentials, and nothing drops traffic to +them. On a control-node that faces the network while it is being adopted, that is postgres +(bootstrap superuser) and a message broker open to anyone who can reach the machine, for the +length of the install. + +## Why it matters + +The design's rule is that what a port is reachable from is decided by the firewall, computed +from each module's `listens.from` — `mesh` for both of these. A rule enforced by nothing is +indistinguishable from a wrong one, and for the duration of this window that rule is enforced +by nothing: the thing that would apply it does not exist yet. It is the same window the ssh +rule already reasons about ("reached over the network, before the private network exists"), +but ssh is one guarded port and this is the mesh's whole store. + +The foundation used to sidestep this by binding the store to loopback — only the co-located +control plane reached it — and adoption trades that away, because a module that adopts the +container in place must declare the same bind, and the module has to serve consumers. + +## Open questions + +- Can a default-deny base filter (drop everything but loopback, established, and ssh) be + applied at genesis, before the store and broker come up, and the mesh-scoped rules refined + once the private network has addresses to name? The consumers that must reach the store + (the catalogue) come up before the network step, so the `from: mesh` rule would have to be + in place by then. +- Or should the servers bind narrowly at genesis (loopback plus the container bridge) and + widen only once the filter that protects them exists — accepting that a bind change is a + recreate, so this would mean the store is recreated once during install? +- Is the window acceptable as-is, given the machine is mid-bootstrap and the exposure matches + what the pre-adoption `postgres`/`lavinmq` modules already had in steady state? diff --git a/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md b/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md new file mode 100644 index 0000000..d856dc0 --- /dev/null +++ b/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md @@ -0,0 +1,40 @@ +--- +status: open +opened: 2026-09-16 +located-in: [] +fixed-by: +amended-design: +--- + +# 055 — The adopted store and broker may be reachable on the control-node only + +## Symptom + +The adopted `mesh-store` and `mesh-broker` bind `0.0.0.0` on the control-node. The +co-located provisioner and control plane reach them over loopback, and a co-located consumer +reaches them through the container bridge — which is what the one-node bed proves. A consumer +on ANOTHER machine reaches a provider by its `.internal` name over the private network, and +whether that path resolves to the control-node's bind is unproven: the one-node bed cannot +exercise it, and no multi-node bed installs the adopted store or broker. + +## Why it matters + +Phase 3's stated goal is a mesh — of any size — that runs one postgres and one lavinmq. The +whole point of a shared store and broker is that a module on any machine that is granted a +database or a vhost can open it. If the adopted servers are reachable only on the machine +they run on, a grant to a module placed elsewhere names an endpoint that machine cannot dial, +and the failure surfaces far from here as a consumer that cannot connect. + +Before adoption this was a non-question: the store served the control plane alone and the +`postgres` module raised a second server that published mesh-wide. Collapsing to one server +means the one server has to be the mesh-wide one, reachable across the overlay. + +## Open questions + +- Does the mesh deliver the store and broker endpoint to a remote consumer as an address that + consumer can dial — the provider's overlay address — rather than a loopback or bridge + address meaningful only on the control-node? +- Is a `0.0.0.0` bind on the control-node reachable over the WireGuard overlay from a joined + machine, and is the packet filter's `from: mesh` rule enough to let it through? +- What is the smallest multi-node bed that would prove a database granted to a module on a + joined machine can be opened from there?