Issue 055 diagnosed — the module broker URL uses the public address, not the overlay
A two-node bed pins it: the firewall gap (broker's 5671 not in listens) is fixed in mesh-catalog; the real bug is modules.go:294/build.go:231 building the broker URL with the genesis public MESH_BROKER_ADDRESS, which a joined node cannot reach. https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
This commit is contained in:
+36
-1
@@ -1,5 +1,5 @@
|
||||
---
|
||||
status: open
|
||||
status: diagnosing
|
||||
opened: 2026-09-16
|
||||
located-in: []
|
||||
fixed-by:
|
||||
@@ -38,3 +38,38 @@ means the one server has to be the mesh-wide one, reachable across the overlay.
|
||||
machine, and is the packet filter's `from: mesh` rule enough to let it through?
|
||||
- What is the smallest multi-node bed that would prove a database granted to a module on a
|
||||
joined machine can be opened from there?
|
||||
|
||||
## Diagnosis (2026-09-17)
|
||||
|
||||
A focused two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`, scenario
|
||||
`adopted-store-cross-node.yml`) puts the adopted broker on `anchor` and runs `amqp-ping` on the
|
||||
joined `node2`. It fails, and the cause is two-layered — the first fixed, the second the real one:
|
||||
|
||||
1. **Firewall (fixed).** The adopted `lavinmq` module declared only `5672` in `listens`, not the
|
||||
amqps bus port `5671`. mesh-broker's ports are published (DNAT'd → the nftables *forward* chain),
|
||||
so with no `5671` forward rule the bus was dropped cross-node. Fix: the module declares `5671`
|
||||
from `mesh` too (mesh-catalog `fix/broker-declares-amqps-port`). Confirmed: after it, `node2 ->
|
||||
10.42.0.1:5671` (overlay) is OPEN and the forward chain carries `proto-dst 5671`.
|
||||
|
||||
2. **The broker address is the genesis public one, not the overlay (the real bug).** The broker
|
||||
secret handed to a module is built at `cmd/mesh-controller/modules.go:294` (and the builder's at
|
||||
`build.go:231`) as `amqps://<acct>:<pw>@<known.Address>/`, where `known.Address` is
|
||||
`MESH_BROKER_ADDRESS` — the genesis PUBLIC endpoint (e.g. `192.0.2.10:5671`). `amqp-ping` on
|
||||
node2 therefore dials `192.0.2.10:5671`, which the firewall does NOT admit cross-node (it admits
|
||||
the overlay `10.42.0.1:5671`), and times out "fetching the broker's certificate." Its grant
|
||||
never completes, so the provisioner's receives file shows `given: []` and no vhost is minted —
|
||||
the downstream symptom.
|
||||
|
||||
The consumer already resolves the provider's `.internal` overlay name for its *provision* binding
|
||||
(`bound.at == anchor.internal`) — that half works. What does not is the *bus account* URL, which
|
||||
uses the public address for every module regardless of node.
|
||||
|
||||
## Open questions
|
||||
|
||||
- The broker URL for a module on a joined node should name the control-node's overlay address
|
||||
(`<control-node>.internal:5671`), which resolves (via `--add-host`) and the firewall admits.
|
||||
- But genesis-time accounts (the builder, issued on the control-node before the overlay exists)
|
||||
must keep a locally-reachable address, or genesis breaks. So the fix is conditional on the
|
||||
target node being on the overlay, not a blanket switch — where is that known at `module issue`?
|
||||
- Should the control-node's own modules also move to the overlay name for uniformity, or keep the
|
||||
local address they already reach?
|
||||
|
||||
Reference in New Issue
Block a user