Files
hq/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md
T
jschoubben 812dc3b303 Issue 055 resolved, and 057 opened for its silent operational edge
055 is proven end-to-end by the two-node bed: a consumer on a joined node reaches the
adopted broker over the overlay, its binding names the control-node, and its vhost is
minted. Three fixes — the broker's amqps port in the firewall, the broker credential
naming the overlay not the public address, and pushing the provider node after the remote
consumer arrives. The last is an operator ordering, not a code fix, and its silent-failure
edge (a cross-node consumer that never provisions and only crash-loops) is opened as 057.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 01:35:47 +02:00

6.3 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
resolved 2026-09-16
mesh-controller/cmd/mesh-controller/modules.go
mesh-controller/cmd/mesh-controller/build.go
mesh-catalog/modules/lavinmq/module.json
mesh-controller multi-node/broker-reaches-over-overlay; mesh-catalog fix/broker-declares-amqps-port

055 — The adopted store and broker may be reachable on the control-node only

Symptom

The adopted mesh-store and mesh-broker bind 0.0.0.0 on the control-node. The co-located provisioner and control plane reach them over loopback, and a co-located consumer reaches them through the container bridge — which is what the one-node bed proves. A consumer on ANOTHER machine reaches a provider by its .internal name over the private network, and whether that path resolves to the control-node's bind is unproven: the one-node bed cannot exercise it, and no multi-node bed installs the adopted store or broker.

Why it matters

Phase 3's stated goal is a mesh — of any size — that runs one postgres and one lavinmq. The whole point of a shared store and broker is that a module on any machine that is granted a database or a vhost can open it. If the adopted servers are reachable only on the machine they run on, a grant to a module placed elsewhere names an endpoint that machine cannot dial, and the failure surfaces far from here as a consumer that cannot connect.

Before adoption this was a non-question: the store served the control plane alone and the postgres module raised a second server that published mesh-wide. Collapsing to one server means the one server has to be the mesh-wide one, reachable across the overlay.

Open questions

  • Does the mesh deliver the store and broker endpoint to a remote consumer as an address that consumer can dial — the provider's overlay address — rather than a loopback or bridge address meaningful only on the control-node?
  • Is a 0.0.0.0 bind on the control-node reachable over the WireGuard overlay from a joined machine, and is the packet filter's from: mesh rule enough to let it through?
  • What is the smallest multi-node bed that would prove a database granted to a module on a joined machine can be opened from there?

Diagnosis (2026-09-17)

A focused two-node bed (mesh-lab test/integration/adopted-store-cross-node.test.ts, scenario adopted-store-cross-node.yml) puts the adopted broker on anchor and runs amqp-ping on the joined node2. It fails, and the cause is two-layered — the first fixed, the second the real one:

  1. Firewall (fixed). The adopted lavinmq module declared only 5672 in listens, not the amqps bus port 5671. mesh-broker's ports are published (DNAT'd → the nftables forward chain), so with no 5671 forward rule the bus was dropped cross-node. Fix: the module declares 5671 from mesh too (mesh-catalog fix/broker-declares-amqps-port). Confirmed: after it, node2 -> 10.42.0.1:5671 (overlay) is OPEN and the forward chain carries proto-dst 5671.

  2. The broker address is the genesis public one, not the overlay (the real bug). The broker secret handed to a module is built at cmd/mesh-controller/modules.go:294 (and the builder's at build.go:231) as amqps://<acct>:<pw>@<known.Address>/, where known.Address is MESH_BROKER_ADDRESS — the genesis PUBLIC endpoint (e.g. 192.0.2.10:5671). amqp-ping on node2 therefore dials 192.0.2.10:5671, which the firewall does NOT admit cross-node (it admits the overlay 10.42.0.1:5671), and times out "fetching the broker's certificate." Its grant never completes, so the provisioner's receives file shows given: [] and no vhost is minted — the downstream symptom.

The consumer already resolves the provider's .internal overlay name for its provision binding (bound.at == anchor.internal) — that half works. What does not is the bus account URL, which uses the public address for every module regardless of node.

Open questions

  • The broker URL for a module on a joined node should name the control-node's overlay address (<control-node>.internal:5671), which resolves (via --add-host) and the firewall admits.
  • But genesis-time accounts (the builder, issued on the control-node before the overlay exists) must keep a locally-reachable address, or genesis breaks. So the fix is conditional on the target node being on the overlay, not a blanket switch — where is that known at module issue?
  • Should the control-node's own modules also move to the overlay name for uniformity, or keep the local address they already reach?

Resolution (2026-09-17)

A two-node bed (mesh-lab test/integration/adopted-store-cross-node.test.ts) proves it: a consumer (amqp-ping) on a joined node reaches the adopted broker (mesh-broker) on the control-node over the overlay, its binding names anchor.internal, its vhost is minted, and it holds the connection. Three things were wrong, now fixed:

  1. The broker's amqps port was not in the firewall. The lavinmq module declared only 5672 in listens; the bus is 5671. Published container ports are matched in the nftables forward chain, so 5671 was dropped cross-node. The module now declares 5671 from mesh (mesh-catalog fix/broker-declares-amqps-port).

  2. The broker credential named the public address. module issue and builder issue built the URL with MESH_BROKER_ADDRESS — the broker's genesis public endpoint, which a joined node's firewall does not admit. A new brokerReachableAt returns the broker as the target node can reach it: a node on the overlay gets the hub's .internal name (fingerprint pinning makes the host swap safe for TLS); a node not yet on the overlay keeps the genesis address, so bring-up is unchanged (mesh-controller multi-node/broker-reaches-over-overlay).

  3. The provider was not recomposed after the remote consumer arrived (operational). A provision secret is minted as a side-effect of composing the CONSUMER's plan; the provider's grant list is a pure read of secrets already issued from it. So the provider's provisioner learns of a cross-node consumer only when the provider node is composed again. Pushing the provider node after adding the remote consumer mints the vhost. This is not a code fix — it is an ordering the operator must follow, and its silent-failure edge is opened as issue 057.