Files
hq/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md
T
jschoubben 812dc3b303 Issue 055 resolved, and 057 opened for its silent operational edge
055 is proven end-to-end by the two-node bed: a consumer on a joined node reaches the
adopted broker over the overlay, its binding names the control-node, and its vhost is
minted. Three fixes — the broker's amqps port in the firewall, the broker credential
naming the overlay not the public address, and pushing the provider node after the remote
consumer arrives. The last is an operator ordering, not a code fix, and its silent-failure
edge (a cross-node consumer that never provisions and only crash-loops) is opened as 057.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 01:35:47 +02:00

105 lines
6.3 KiB
Markdown

---
status: resolved
opened: 2026-09-16
located-in:
- mesh-controller/cmd/mesh-controller/modules.go
- mesh-controller/cmd/mesh-controller/build.go
- mesh-catalog/modules/lavinmq/module.json
fixed-by: mesh-controller multi-node/broker-reaches-over-overlay; mesh-catalog fix/broker-declares-amqps-port
amended-design:
---
# 055 — The adopted store and broker may be reachable on the control-node only
## Symptom
The adopted `mesh-store` and `mesh-broker` bind `0.0.0.0` on the control-node. The
co-located provisioner and control plane reach them over loopback, and a co-located consumer
reaches them through the container bridge — which is what the one-node bed proves. A consumer
on ANOTHER machine reaches a provider by its `.internal` name over the private network, and
whether that path resolves to the control-node's bind is unproven: the one-node bed cannot
exercise it, and no multi-node bed installs the adopted store or broker.
## Why it matters
Phase 3's stated goal is a mesh — of any size — that runs one postgres and one lavinmq. The
whole point of a shared store and broker is that a module on any machine that is granted a
database or a vhost can open it. If the adopted servers are reachable only on the machine
they run on, a grant to a module placed elsewhere names an endpoint that machine cannot dial,
and the failure surfaces far from here as a consumer that cannot connect.
Before adoption this was a non-question: the store served the control plane alone and the
`postgres` module raised a second server that published mesh-wide. Collapsing to one server
means the one server has to be the mesh-wide one, reachable across the overlay.
## Open questions
- Does the mesh deliver the store and broker endpoint to a remote consumer as an address that
consumer can dial — the provider's overlay address — rather than a loopback or bridge
address meaningful only on the control-node?
- Is a `0.0.0.0` bind on the control-node reachable over the WireGuard overlay from a joined
machine, and is the packet filter's `from: mesh` rule enough to let it through?
- What is the smallest multi-node bed that would prove a database granted to a module on a
joined machine can be opened from there?
## Diagnosis (2026-09-17)
A focused two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`, scenario
`adopted-store-cross-node.yml`) puts the adopted broker on `anchor` and runs `amqp-ping` on the
joined `node2`. It fails, and the cause is two-layered — the first fixed, the second the real one:
1. **Firewall (fixed).** The adopted `lavinmq` module declared only `5672` in `listens`, not the
amqps bus port `5671`. mesh-broker's ports are published (DNAT'd → the nftables *forward* chain),
so with no `5671` forward rule the bus was dropped cross-node. Fix: the module declares `5671`
from `mesh` too (mesh-catalog `fix/broker-declares-amqps-port`). Confirmed: after it, `node2 ->
10.42.0.1:5671` (overlay) is OPEN and the forward chain carries `proto-dst 5671`.
2. **The broker address is the genesis public one, not the overlay (the real bug).** The broker
secret handed to a module is built at `cmd/mesh-controller/modules.go:294` (and the builder's at
`build.go:231`) as `amqps://<acct>:<pw>@<known.Address>/`, where `known.Address` is
`MESH_BROKER_ADDRESS` — the genesis PUBLIC endpoint (e.g. `192.0.2.10:5671`). `amqp-ping` on
node2 therefore dials `192.0.2.10:5671`, which the firewall does NOT admit cross-node (it admits
the overlay `10.42.0.1:5671`), and times out "fetching the broker's certificate." Its grant
never completes, so the provisioner's receives file shows `given: []` and no vhost is minted —
the downstream symptom.
The consumer already resolves the provider's `.internal` overlay name for its *provision* binding
(`bound.at == anchor.internal`) — that half works. What does not is the *bus account* URL, which
uses the public address for every module regardless of node.
## Open questions
- The broker URL for a module on a joined node should name the control-node's overlay address
(`<control-node>.internal:5671`), which resolves (via `--add-host`) and the firewall admits.
- But genesis-time accounts (the builder, issued on the control-node before the overlay exists)
must keep a locally-reachable address, or genesis breaks. So the fix is conditional on the
target node being on the overlay, not a blanket switch — where is that known at `module issue`?
- Should the control-node's own modules also move to the overlay name for uniformity, or keep the
local address they already reach?
## Resolution (2026-09-17)
A two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`) proves it: a
consumer (`amqp-ping`) on a joined node reaches the adopted broker (`mesh-broker`) on the
control-node over the overlay, its binding names `anchor.internal`, its vhost is minted, and it
holds the connection. Three things were wrong, now fixed:
1. **The broker's amqps port was not in the firewall.** The `lavinmq` module declared only `5672`
in `listens`; the bus is `5671`. Published container ports are matched in the nftables *forward*
chain, so `5671` was dropped cross-node. The module now declares `5671` from `mesh`
(mesh-catalog `fix/broker-declares-amqps-port`).
2. **The broker credential named the public address.** `module issue` and `builder issue` built
the URL with `MESH_BROKER_ADDRESS` — the broker's genesis public endpoint, which a joined
node's firewall does not admit. A new `brokerReachableAt` returns the broker as the target node
can reach it: a node on the overlay gets the hub's `.internal` name (fingerprint pinning makes
the host swap safe for TLS); a node not yet on the overlay keeps the genesis address, so
bring-up is unchanged (mesh-controller `multi-node/broker-reaches-over-overlay`).
3. **The provider was not recomposed after the remote consumer arrived (operational).** A
provision secret is minted as a side-effect of composing the CONSUMER's plan; the provider's
grant list is a pure read of secrets already issued from it. So the provider's provisioner
learns of a cross-node consumer only when the provider node is composed again. Pushing the
provider node after adding the remote consumer mints the vhost. This is not a code fix — it is
an ordering the operator must follow, and its silent-failure edge is opened as issue 057.