Files
hq/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md
T
jschoubben fcf3b6f4d5 Issues 054, 055 — the debt adopting the store and broker leaves
054: the adopted servers bind 0.0.0.0 from genesis but the packet filter is
installed later, so there is a window where they are open with only bootstrap
credentials. 055: the servers bind on the control-node and the one-node bed
cannot prove a consumer on another machine can reach them over the overlay.

Both are follow-ups to issue 051's adoption (WBS Phase 3), tracked rather than
rushed.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-16 22:04:28 +02:00

1.9 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
open 2026-09-16

055 — The adopted store and broker may be reachable on the control-node only

Symptom

The adopted mesh-store and mesh-broker bind 0.0.0.0 on the control-node. The co-located provisioner and control plane reach them over loopback, and a co-located consumer reaches them through the container bridge — which is what the one-node bed proves. A consumer on ANOTHER machine reaches a provider by its .internal name over the private network, and whether that path resolves to the control-node's bind is unproven: the one-node bed cannot exercise it, and no multi-node bed installs the adopted store or broker.

Why it matters

Phase 3's stated goal is a mesh — of any size — that runs one postgres and one lavinmq. The whole point of a shared store and broker is that a module on any machine that is granted a database or a vhost can open it. If the adopted servers are reachable only on the machine they run on, a grant to a module placed elsewhere names an endpoint that machine cannot dial, and the failure surfaces far from here as a consumer that cannot connect.

Before adoption this was a non-question: the store served the control plane alone and the postgres module raised a second server that published mesh-wide. Collapsing to one server means the one server has to be the mesh-wide one, reachable across the overlay.

Open questions

  • Does the mesh deliver the store and broker endpoint to a remote consumer as an address that consumer can dial — the provider's overlay address — rather than a loopback or bridge address meaningful only on the control-node?
  • Is a 0.0.0.0 bind on the control-node reachable over the WireGuard overlay from a joined machine, and is the packet filter's from: mesh rule enough to let it through?
  • What is the smallest multi-node bed that would prove a database granted to a module on a joined machine can be opened from there?