The glossary's authority page still named the controller's seat the-controller in two entries, contradicting its own seat section after ADR 0079; issue 058's heading kept the pre-renumber 059; 055's fixed-by named branches that stop existing after merge (now merge commits/PRs) and its located-in listed file paths where the convention wants repos; 056's located-in named mesh-host, which received no fix, instead of mesh-catalog; and the design layer never said the one-store/one-broker property is enforced — 07-the-foundation and the installation table now state the seats. https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
104 lines
6.2 KiB
Markdown
104 lines
6.2 KiB
Markdown
---
|
|
status: resolved
|
|
opened: 2026-09-16
|
|
located-in:
|
|
- mesh-controller
|
|
- mesh-catalog
|
|
fixed-by: mesh-controller PR 27 (5718add); mesh-catalog PR 24 (a24362b)
|
|
amended-design:
|
|
---
|
|
|
|
# 055 — The adopted store and broker may be reachable on the control-node only
|
|
|
|
## Symptom
|
|
|
|
The adopted `mesh-store` and `mesh-broker` bind `0.0.0.0` on the control-node. The
|
|
co-located provisioner and control plane reach them over loopback, and a co-located consumer
|
|
reaches them through the container bridge — which is what the one-node bed proves. A consumer
|
|
on ANOTHER machine reaches a provider by its `.internal` name over the private network, and
|
|
whether that path resolves to the control-node's bind is unproven: the one-node bed cannot
|
|
exercise it, and no multi-node bed installs the adopted store or broker.
|
|
|
|
## Why it matters
|
|
|
|
Phase 3's stated goal is a mesh — of any size — that runs one postgres and one lavinmq. The
|
|
whole point of a shared store and broker is that a module on any machine that is granted a
|
|
database or a vhost can open it. If the adopted servers are reachable only on the machine
|
|
they run on, a grant to a module placed elsewhere names an endpoint that machine cannot dial,
|
|
and the failure surfaces far from here as a consumer that cannot connect.
|
|
|
|
Before adoption this was a non-question: the store served the control plane alone and the
|
|
`postgres` module raised a second server that published mesh-wide. Collapsing to one server
|
|
means the one server has to be the mesh-wide one, reachable across the overlay.
|
|
|
|
## Open questions
|
|
|
|
- Does the mesh deliver the store and broker endpoint to a remote consumer as an address that
|
|
consumer can dial — the provider's overlay address — rather than a loopback or bridge
|
|
address meaningful only on the control-node?
|
|
- Is a `0.0.0.0` bind on the control-node reachable over the WireGuard overlay from a joined
|
|
machine, and is the packet filter's `from: mesh` rule enough to let it through?
|
|
- What is the smallest multi-node bed that would prove a database granted to a module on a
|
|
joined machine can be opened from there?
|
|
|
|
## Diagnosis (2026-09-17)
|
|
|
|
A focused two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`, scenario
|
|
`adopted-store-cross-node.yml`) puts the adopted broker on `anchor` and runs `amqp-ping` on the
|
|
joined `node2`. It fails, and the cause is two-layered — the first fixed, the second the real one:
|
|
|
|
1. **Firewall (fixed).** The adopted `lavinmq` module declared only `5672` in `listens`, not the
|
|
amqps bus port `5671`. mesh-broker's ports are published (DNAT'd → the nftables *forward* chain),
|
|
so with no `5671` forward rule the bus was dropped cross-node. Fix: the module declares `5671`
|
|
from `mesh` too (mesh-catalog `fix/broker-declares-amqps-port`). Confirmed: after it, `node2 ->
|
|
10.42.0.1:5671` (overlay) is OPEN and the forward chain carries `proto-dst 5671`.
|
|
|
|
2. **The broker address is the genesis public one, not the overlay (the real bug).** The broker
|
|
secret handed to a module is built at `cmd/mesh-controller/modules.go:294` (and the builder's at
|
|
`build.go:231`) as `amqps://<acct>:<pw>@<known.Address>/`, where `known.Address` is
|
|
`MESH_BROKER_ADDRESS` — the genesis PUBLIC endpoint (e.g. `192.0.2.10:5671`). `amqp-ping` on
|
|
node2 therefore dials `192.0.2.10:5671`, which the firewall does NOT admit cross-node (it admits
|
|
the overlay `10.42.0.1:5671`), and times out "fetching the broker's certificate." Its grant
|
|
never completes, so the provisioner's receives file shows `given: []` and no vhost is minted —
|
|
the downstream symptom.
|
|
|
|
The consumer already resolves the provider's `.internal` overlay name for its *provision* binding
|
|
(`bound.at == anchor.internal`) — that half works. What does not is the *bus account* URL, which
|
|
uses the public address for every module regardless of node.
|
|
|
|
## Open questions
|
|
|
|
- The broker URL for a module on a joined node should name the control-node's overlay address
|
|
(`<control-node>.internal:5671`), which resolves (via `--add-host`) and the firewall admits.
|
|
- But genesis-time accounts (the builder, issued on the control-node before the overlay exists)
|
|
must keep a locally-reachable address, or genesis breaks. So the fix is conditional on the
|
|
target node being on the overlay, not a blanket switch — where is that known at `module issue`?
|
|
- Should the control-node's own modules also move to the overlay name for uniformity, or keep the
|
|
local address they already reach?
|
|
|
|
## Resolution (2026-09-17)
|
|
|
|
A two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`) proves it: a
|
|
consumer (`amqp-ping`) on a joined node reaches the adopted broker (`mesh-broker`) on the
|
|
control-node over the overlay, its binding names `anchor.internal`, its vhost is minted, and it
|
|
holds the connection. Three things were wrong, now fixed:
|
|
|
|
1. **The broker's amqps port was not in the firewall.** The `lavinmq` module declared only `5672`
|
|
in `listens`; the bus is `5671`. Published container ports are matched in the nftables *forward*
|
|
chain, so `5671` was dropped cross-node. The module now declares `5671` from `mesh`
|
|
(mesh-catalog `fix/broker-declares-amqps-port`).
|
|
|
|
2. **The broker credential named the public address.** `module issue` and `builder issue` built
|
|
the URL with `MESH_BROKER_ADDRESS` — the broker's genesis public endpoint, which a joined
|
|
node's firewall does not admit. A new `brokerReachableAt` returns the broker as the target node
|
|
can reach it: a node on the overlay gets the hub's `.internal` name (fingerprint pinning makes
|
|
the host swap safe for TLS); a node not yet on the overlay keeps the genesis address, so
|
|
bring-up is unchanged (mesh-controller `multi-node/broker-reaches-over-overlay`).
|
|
|
|
3. **The provider was not recomposed after the remote consumer arrived (operational).** A
|
|
provision secret is minted as a side-effect of composing the CONSUMER's plan; the provider's
|
|
grant list is a pure read of secrets already issued from it. So the provider's provisioner
|
|
learns of a cross-node consumer only when the provider node is composed again. Pushing the
|
|
provider node after adding the remote consumer mints the vhost. This is not a code fix — it is
|
|
an ordering the operator must follow, and its silent-failure edge is opened as issue 057.
|