The glossary's authority page still named the controller's seat the-controller in two entries, contradicting its own seat section after ADR 0079; issue 058's heading kept the pre-renumber 059; 055's fixed-by named branches that stop existing after merge (now merge commits/PRs) and its located-in listed file paths where the convention wants repos; 056's located-in named mesh-host, which received no fix, instead of mesh-catalog; and the design layer never said the one-store/one-broker property is enforced — 07-the-foundation and the installation table now state the seats. https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
6.2 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design | ||
|---|---|---|---|---|---|---|
| resolved | 2026-09-16 |
|
mesh-controller PR 27 (5718add); mesh-catalog PR 24 (a24362b) |
055 — The adopted store and broker may be reachable on the control-node only
Symptom
The adopted mesh-store and mesh-broker bind 0.0.0.0 on the control-node. The
co-located provisioner and control plane reach them over loopback, and a co-located consumer
reaches them through the container bridge — which is what the one-node bed proves. A consumer
on ANOTHER machine reaches a provider by its .internal name over the private network, and
whether that path resolves to the control-node's bind is unproven: the one-node bed cannot
exercise it, and no multi-node bed installs the adopted store or broker.
Why it matters
Phase 3's stated goal is a mesh — of any size — that runs one postgres and one lavinmq. The whole point of a shared store and broker is that a module on any machine that is granted a database or a vhost can open it. If the adopted servers are reachable only on the machine they run on, a grant to a module placed elsewhere names an endpoint that machine cannot dial, and the failure surfaces far from here as a consumer that cannot connect.
Before adoption this was a non-question: the store served the control plane alone and the
postgres module raised a second server that published mesh-wide. Collapsing to one server
means the one server has to be the mesh-wide one, reachable across the overlay.
Open questions
- Does the mesh deliver the store and broker endpoint to a remote consumer as an address that consumer can dial — the provider's overlay address — rather than a loopback or bridge address meaningful only on the control-node?
- Is a
0.0.0.0bind on the control-node reachable over the WireGuard overlay from a joined machine, and is the packet filter'sfrom: meshrule enough to let it through? - What is the smallest multi-node bed that would prove a database granted to a module on a joined machine can be opened from there?
Diagnosis (2026-09-17)
A focused two-node bed (mesh-lab test/integration/adopted-store-cross-node.test.ts, scenario
adopted-store-cross-node.yml) puts the adopted broker on anchor and runs amqp-ping on the
joined node2. It fails, and the cause is two-layered — the first fixed, the second the real one:
-
Firewall (fixed). The adopted
lavinmqmodule declared only5672inlistens, not the amqps bus port5671. mesh-broker's ports are published (DNAT'd → the nftables forward chain), so with no5671forward rule the bus was dropped cross-node. Fix: the module declares5671frommeshtoo (mesh-catalogfix/broker-declares-amqps-port). Confirmed: after it,node2 -> 10.42.0.1:5671(overlay) is OPEN and the forward chain carriesproto-dst 5671. -
The broker address is the genesis public one, not the overlay (the real bug). The broker secret handed to a module is built at
cmd/mesh-controller/modules.go:294(and the builder's atbuild.go:231) asamqps://<acct>:<pw>@<known.Address>/, whereknown.AddressisMESH_BROKER_ADDRESS— the genesis PUBLIC endpoint (e.g.192.0.2.10:5671).amqp-pingon node2 therefore dials192.0.2.10:5671, which the firewall does NOT admit cross-node (it admits the overlay10.42.0.1:5671), and times out "fetching the broker's certificate." Its grant never completes, so the provisioner's receives file showsgiven: []and no vhost is minted — the downstream symptom.
The consumer already resolves the provider's .internal overlay name for its provision binding
(bound.at == anchor.internal) — that half works. What does not is the bus account URL, which
uses the public address for every module regardless of node.
Open questions
- The broker URL for a module on a joined node should name the control-node's overlay address
(
<control-node>.internal:5671), which resolves (via--add-host) and the firewall admits. - But genesis-time accounts (the builder, issued on the control-node before the overlay exists)
must keep a locally-reachable address, or genesis breaks. So the fix is conditional on the
target node being on the overlay, not a blanket switch — where is that known at
module issue? - Should the control-node's own modules also move to the overlay name for uniformity, or keep the local address they already reach?
Resolution (2026-09-17)
A two-node bed (mesh-lab test/integration/adopted-store-cross-node.test.ts) proves it: a
consumer (amqp-ping) on a joined node reaches the adopted broker (mesh-broker) on the
control-node over the overlay, its binding names anchor.internal, its vhost is minted, and it
holds the connection. Three things were wrong, now fixed:
-
The broker's amqps port was not in the firewall. The
lavinmqmodule declared only5672inlistens; the bus is5671. Published container ports are matched in the nftables forward chain, so5671was dropped cross-node. The module now declares5671frommesh(mesh-catalogfix/broker-declares-amqps-port). -
The broker credential named the public address.
module issueandbuilder issuebuilt the URL withMESH_BROKER_ADDRESS— the broker's genesis public endpoint, which a joined node's firewall does not admit. A newbrokerReachableAtreturns the broker as the target node can reach it: a node on the overlay gets the hub's.internalname (fingerprint pinning makes the host swap safe for TLS); a node not yet on the overlay keeps the genesis address, so bring-up is unchanged (mesh-controllermulti-node/broker-reaches-over-overlay). -
The provider was not recomposed after the remote consumer arrived (operational). A provision secret is minted as a side-effect of composing the CONSUMER's plan; the provider's grant list is a pure read of secrets already issued from it. So the provider's provisioner learns of a cross-node consumer only when the provider node is composed again. Pushing the provider node after adding the remote consumer mints the vhost. This is not a code fix — it is an ordering the operator must follow, and its silent-failure edge is opened as issue 057.