Issues 054, 055 — the debt adopting the store and broker leaves

054: the adopted servers bind 0.0.0.0 from genesis but the packet filter is
installed later, so there is a window where they are open with only bootstrap
credentials. 055: the servers bind on the control-node and the one-node bed
cannot prove a consumer on another machine can reach them over the overlay.

Both are follow-ups to issue 051's adoption (WBS Phase 3), tracked rather than
rushed.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
This commit is contained in:
2026-09-16 22:04:28 +02:00
parent 43f3565617
commit fcf3b6f4d5
2 changed files with 88 additions and 0 deletions
@@ -0,0 +1,48 @@
---
status: open
opened: 2026-09-16
located-in: []
fixed-by:
amended-design:
---
# 054 — The adopted store and broker are open before the packet filter exists
## Symptom
Adopting the store and broker as ordinary modules (issue 051) needs them reachable by
consumers across the mesh, so genesis now raises `mesh-store` and `mesh-broker` bound to
`0.0.0.0` rather than to loopback as the foundation used to. The packet filter is a module,
installed several steps after the store, the broker and the catalogue are already running.
Between each server coming up and the filter being applied, both listen on every interface
of the machine with only the genesis bootstrap credentials, and nothing drops traffic to
them. On a control-node that faces the network while it is being adopted, that is postgres
(bootstrap superuser) and a message broker open to anyone who can reach the machine, for the
length of the install.
## Why it matters
The design's rule is that what a port is reachable from is decided by the firewall, computed
from each module's `listens.from` — `mesh` for both of these. A rule enforced by nothing is
indistinguishable from a wrong one, and for the duration of this window that rule is enforced
by nothing: the thing that would apply it does not exist yet. It is the same window the ssh
rule already reasons about ("reached over the network, before the private network exists"),
but ssh is one guarded port and this is the mesh's whole store.
The foundation used to sidestep this by binding the store to loopback — only the co-located
control plane reached it — and adoption trades that away, because a module that adopts the
container in place must declare the same bind, and the module has to serve consumers.
## Open questions
- Can a default-deny base filter (drop everything but loopback, established, and ssh) be
applied at genesis, before the store and broker come up, and the mesh-scoped rules refined
once the private network has addresses to name? The consumers that must reach the store
(the catalogue) come up before the network step, so the `from: mesh` rule would have to be
in place by then.
- Or should the servers bind narrowly at genesis (loopback plus the container bridge) and
widen only once the filter that protects them exists — accepting that a bind change is a
recreate, so this would mean the store is recreated once during install?
- Is the window acceptable as-is, given the machine is mid-bootstrap and the exposure matches
what the pre-adoption `postgres`/`lavinmq` modules already had in steady state?
@@ -0,0 +1,40 @@
---
status: open
opened: 2026-09-16
located-in: []
fixed-by:
amended-design:
---
# 055 — The adopted store and broker may be reachable on the control-node only
## Symptom
The adopted `mesh-store` and `mesh-broker` bind `0.0.0.0` on the control-node. The
co-located provisioner and control plane reach them over loopback, and a co-located consumer
reaches them through the container bridge — which is what the one-node bed proves. A consumer
on ANOTHER machine reaches a provider by its `.internal` name over the private network, and
whether that path resolves to the control-node's bind is unproven: the one-node bed cannot
exercise it, and no multi-node bed installs the adopted store or broker.
## Why it matters
Phase 3's stated goal is a mesh — of any size — that runs one postgres and one lavinmq. The
whole point of a shared store and broker is that a module on any machine that is granted a
database or a vhost can open it. If the adopted servers are reachable only on the machine
they run on, a grant to a module placed elsewhere names an endpoint that machine cannot dial,
and the failure surfaces far from here as a consumer that cannot connect.
Before adoption this was a non-question: the store served the control plane alone and the
`postgres` module raised a second server that published mesh-wide. Collapsing to one server
means the one server has to be the mesh-wide one, reachable across the overlay.
## Open questions
- Does the mesh deliver the store and broker endpoint to a remote consumer as an address that
consumer can dial — the provider's overlay address — rather than a loopback or bridge
address meaningful only on the control-node?
- Is a `0.0.0.0` bind on the control-node reachable over the WireGuard overlay from a joined
machine, and is the packet filter's `from: mesh` rule enough to let it through?
- What is the smallest multi-node bed that would prove a database granted to a module on a
joined machine can be opened from there?