Files
hq/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md
T
jschoubben 94fee5d849 Fix what the records review found
The glossary's authority page still named the controller's seat the-controller in two
entries, contradicting its own seat section after ADR 0079; issue 058's heading kept the
pre-renumber 059; 055's fixed-by named branches that stop existing after merge (now merge
commits/PRs) and its located-in listed file paths where the convention wants repos; 056's
located-in named mesh-host, which received no fix, instead of mesh-catalog; and the design
layer never said the one-store/one-broker property is enforced — 07-the-foundation and the
installation table now state the seats.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:24:57 +02:00

6.2 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
resolved 2026-09-16
mesh-controller
mesh-catalog
mesh-controller PR 27 (5718add); mesh-catalog PR 24 (a24362b)

055 — The adopted store and broker may be reachable on the control-node only

Symptom

The adopted mesh-store and mesh-broker bind 0.0.0.0 on the control-node. The co-located provisioner and control plane reach them over loopback, and a co-located consumer reaches them through the container bridge — which is what the one-node bed proves. A consumer on ANOTHER machine reaches a provider by its .internal name over the private network, and whether that path resolves to the control-node's bind is unproven: the one-node bed cannot exercise it, and no multi-node bed installs the adopted store or broker.

Why it matters

Phase 3's stated goal is a mesh — of any size — that runs one postgres and one lavinmq. The whole point of a shared store and broker is that a module on any machine that is granted a database or a vhost can open it. If the adopted servers are reachable only on the machine they run on, a grant to a module placed elsewhere names an endpoint that machine cannot dial, and the failure surfaces far from here as a consumer that cannot connect.

Before adoption this was a non-question: the store served the control plane alone and the postgres module raised a second server that published mesh-wide. Collapsing to one server means the one server has to be the mesh-wide one, reachable across the overlay.

Open questions

  • Does the mesh deliver the store and broker endpoint to a remote consumer as an address that consumer can dial — the provider's overlay address — rather than a loopback or bridge address meaningful only on the control-node?
  • Is a 0.0.0.0 bind on the control-node reachable over the WireGuard overlay from a joined machine, and is the packet filter's from: mesh rule enough to let it through?
  • What is the smallest multi-node bed that would prove a database granted to a module on a joined machine can be opened from there?

Diagnosis (2026-09-17)

A focused two-node bed (mesh-lab test/integration/adopted-store-cross-node.test.ts, scenario adopted-store-cross-node.yml) puts the adopted broker on anchor and runs amqp-ping on the joined node2. It fails, and the cause is two-layered — the first fixed, the second the real one:

  1. Firewall (fixed). The adopted lavinmq module declared only 5672 in listens, not the amqps bus port 5671. mesh-broker's ports are published (DNAT'd → the nftables forward chain), so with no 5671 forward rule the bus was dropped cross-node. Fix: the module declares 5671 from mesh too (mesh-catalog fix/broker-declares-amqps-port). Confirmed: after it, node2 -> 10.42.0.1:5671 (overlay) is OPEN and the forward chain carries proto-dst 5671.

  2. The broker address is the genesis public one, not the overlay (the real bug). The broker secret handed to a module is built at cmd/mesh-controller/modules.go:294 (and the builder's at build.go:231) as amqps://<acct>:<pw>@<known.Address>/, where known.Address is MESH_BROKER_ADDRESS — the genesis PUBLIC endpoint (e.g. 192.0.2.10:5671). amqp-ping on node2 therefore dials 192.0.2.10:5671, which the firewall does NOT admit cross-node (it admits the overlay 10.42.0.1:5671), and times out "fetching the broker's certificate." Its grant never completes, so the provisioner's receives file shows given: [] and no vhost is minted — the downstream symptom.

The consumer already resolves the provider's .internal overlay name for its provision binding (bound.at == anchor.internal) — that half works. What does not is the bus account URL, which uses the public address for every module regardless of node.

Open questions

  • The broker URL for a module on a joined node should name the control-node's overlay address (<control-node>.internal:5671), which resolves (via --add-host) and the firewall admits.
  • But genesis-time accounts (the builder, issued on the control-node before the overlay exists) must keep a locally-reachable address, or genesis breaks. So the fix is conditional on the target node being on the overlay, not a blanket switch — where is that known at module issue?
  • Should the control-node's own modules also move to the overlay name for uniformity, or keep the local address they already reach?

Resolution (2026-09-17)

A two-node bed (mesh-lab test/integration/adopted-store-cross-node.test.ts) proves it: a consumer (amqp-ping) on a joined node reaches the adopted broker (mesh-broker) on the control-node over the overlay, its binding names anchor.internal, its vhost is minted, and it holds the connection. Three things were wrong, now fixed:

  1. The broker's amqps port was not in the firewall. The lavinmq module declared only 5672 in listens; the bus is 5671. Published container ports are matched in the nftables forward chain, so 5671 was dropped cross-node. The module now declares 5671 from mesh (mesh-catalog fix/broker-declares-amqps-port).

  2. The broker credential named the public address. module issue and builder issue built the URL with MESH_BROKER_ADDRESS — the broker's genesis public endpoint, which a joined node's firewall does not admit. A new brokerReachableAt returns the broker as the target node can reach it: a node on the overlay gets the hub's .internal name (fingerprint pinning makes the host swap safe for TLS); a node not yet on the overlay keeps the genesis address, so bring-up is unchanged (mesh-controller multi-node/broker-reaches-over-overlay).

  3. The provider was not recomposed after the remote consumer arrived (operational). A provision secret is minted as a side-effect of composing the CONSUMER's plan; the provider's grant list is a pure read of secrets already issued from it. So the provider's provisioner learns of a cross-node consumer only when the provider node is composed again. Pushing the provider node after adding the remote consumer mints the vhost. This is not a code fix — it is an ordering the operator must follow, and its silent-failure edge is opened as issue 057.