diff --git a/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md b/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md index d856dc0..c6a33ab 100644 --- a/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md +++ b/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md @@ -1,5 +1,5 @@ --- -status: open +status: diagnosing opened: 2026-09-16 located-in: [] fixed-by: @@ -38,3 +38,38 @@ means the one server has to be the mesh-wide one, reachable across the overlay. machine, and is the packet filter's `from: mesh` rule enough to let it through? - What is the smallest multi-node bed that would prove a database granted to a module on a joined machine can be opened from there? + +## Diagnosis (2026-09-17) + +A focused two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`, scenario +`adopted-store-cross-node.yml`) puts the adopted broker on `anchor` and runs `amqp-ping` on the +joined `node2`. It fails, and the cause is two-layered — the first fixed, the second the real one: + +1. **Firewall (fixed).** The adopted `lavinmq` module declared only `5672` in `listens`, not the + amqps bus port `5671`. mesh-broker's ports are published (DNAT'd → the nftables *forward* chain), + so with no `5671` forward rule the bus was dropped cross-node. Fix: the module declares `5671` + from `mesh` too (mesh-catalog `fix/broker-declares-amqps-port`). Confirmed: after it, `node2 -> + 10.42.0.1:5671` (overlay) is OPEN and the forward chain carries `proto-dst 5671`. + +2. **The broker address is the genesis public one, not the overlay (the real bug).** The broker + secret handed to a module is built at `cmd/mesh-controller/modules.go:294` (and the builder's at + `build.go:231`) as `amqps://:@/`, where `known.Address` is + `MESH_BROKER_ADDRESS` — the genesis PUBLIC endpoint (e.g. `192.0.2.10:5671`). `amqp-ping` on + node2 therefore dials `192.0.2.10:5671`, which the firewall does NOT admit cross-node (it admits + the overlay `10.42.0.1:5671`), and times out "fetching the broker's certificate." Its grant + never completes, so the provisioner's receives file shows `given: []` and no vhost is minted — + the downstream symptom. + +The consumer already resolves the provider's `.internal` overlay name for its *provision* binding +(`bound.at == anchor.internal`) — that half works. What does not is the *bus account* URL, which +uses the public address for every module regardless of node. + +## Open questions + +- The broker URL for a module on a joined node should name the control-node's overlay address + (`.internal:5671`), which resolves (via `--add-host`) and the firewall admits. +- But genesis-time accounts (the builder, issued on the control-node before the overlay exists) + must keep a locally-reachable address, or genesis breaks. So the fix is conditional on the + target node being on the overlay, not a blanket switch — where is that known at `module issue`? +- Should the control-node's own modules also move to the overlay name for uniformity, or keep the + local address they already reach?