Issue 055 diagnosed — the module broker URL uses the public address, not the overlay

A two-node bed pins it: the firewall gap (broker's 5671 not in listens) is fixed
in mesh-catalog; the real bug is modules.go:294/build.go:231 building the broker
URL with the genesis public MESH_BROKER_ADDRESS, which a joined node cannot reach.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
This commit is contained in:
2026-09-17 01:02:07 +02:00
parent bec1dd0082
commit 33fc322f4d
@@ -1,5 +1,5 @@
--- ---
status: open status: diagnosing
opened: 2026-09-16 opened: 2026-09-16
located-in: [] located-in: []
fixed-by: fixed-by:
@@ -38,3 +38,38 @@ means the one server has to be the mesh-wide one, reachable across the overlay.
machine, and is the packet filter's `from: mesh` rule enough to let it through? machine, and is the packet filter's `from: mesh` rule enough to let it through?
- What is the smallest multi-node bed that would prove a database granted to a module on a - What is the smallest multi-node bed that would prove a database granted to a module on a
joined machine can be opened from there? joined machine can be opened from there?
## Diagnosis (2026-09-17)
A focused two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`, scenario
`adopted-store-cross-node.yml`) puts the adopted broker on `anchor` and runs `amqp-ping` on the
joined `node2`. It fails, and the cause is two-layered — the first fixed, the second the real one:
1. **Firewall (fixed).** The adopted `lavinmq` module declared only `5672` in `listens`, not the
amqps bus port `5671`. mesh-broker's ports are published (DNAT'd → the nftables *forward* chain),
so with no `5671` forward rule the bus was dropped cross-node. Fix: the module declares `5671`
from `mesh` too (mesh-catalog `fix/broker-declares-amqps-port`). Confirmed: after it, `node2 ->
10.42.0.1:5671` (overlay) is OPEN and the forward chain carries `proto-dst 5671`.
2. **The broker address is the genesis public one, not the overlay (the real bug).** The broker
secret handed to a module is built at `cmd/mesh-controller/modules.go:294` (and the builder's at
`build.go:231`) as `amqps://<acct>:<pw>@<known.Address>/`, where `known.Address` is
`MESH_BROKER_ADDRESS` — the genesis PUBLIC endpoint (e.g. `192.0.2.10:5671`). `amqp-ping` on
node2 therefore dials `192.0.2.10:5671`, which the firewall does NOT admit cross-node (it admits
the overlay `10.42.0.1:5671`), and times out "fetching the broker's certificate." Its grant
never completes, so the provisioner's receives file shows `given: []` and no vhost is minted —
the downstream symptom.
The consumer already resolves the provider's `.internal` overlay name for its *provision* binding
(`bound.at == anchor.internal`) — that half works. What does not is the *bus account* URL, which
uses the public address for every module regardless of node.
## Open questions
- The broker URL for a module on a joined node should name the control-node's overlay address
(`<control-node>.internal:5671`), which resolves (via `--add-host`) and the firewall admits.
- But genesis-time accounts (the builder, issued on the control-node before the overlay exists)
must keep a locally-reachable address, or genesis breaks. So the fix is conditional on the
target node being on the overlay, not a blanket switch — where is that known at `module issue`?
- Should the control-node's own modules also move to the overlay name for uniformity, or keep the
local address they already reach?