Files
jschoubben 94fee5d849 Fix what the records review found
The glossary's authority page still named the controller's seat the-controller in two
entries, contradicting its own seat section after ADR 0079; issue 058's heading kept the
pre-renumber 059; 055's fixed-by named branches that stop existing after merge (now merge
commits/PRs) and its located-in listed file paths where the convention wants repos; 056's
located-in named mesh-host, which received no fix, instead of mesh-catalog; and the design
layer never said the one-store/one-broker property is enforced — 07-the-foundation and the
installation table now state the seats.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:24:57 +02:00

104 lines
6.2 KiB
Markdown

---
status: resolved
opened: 2026-09-16
located-in:
- mesh-controller
- mesh-catalog
fixed-by: mesh-controller PR 27 (5718add); mesh-catalog PR 24 (a24362b)
amended-design:
---
# 055 — The adopted store and broker may be reachable on the control-node only
## Symptom
The adopted `mesh-store` and `mesh-broker` bind `0.0.0.0` on the control-node. The
co-located provisioner and control plane reach them over loopback, and a co-located consumer
reaches them through the container bridge — which is what the one-node bed proves. A consumer
on ANOTHER machine reaches a provider by its `.internal` name over the private network, and
whether that path resolves to the control-node's bind is unproven: the one-node bed cannot
exercise it, and no multi-node bed installs the adopted store or broker.
## Why it matters
Phase 3's stated goal is a mesh — of any size — that runs one postgres and one lavinmq. The
whole point of a shared store and broker is that a module on any machine that is granted a
database or a vhost can open it. If the adopted servers are reachable only on the machine
they run on, a grant to a module placed elsewhere names an endpoint that machine cannot dial,
and the failure surfaces far from here as a consumer that cannot connect.
Before adoption this was a non-question: the store served the control plane alone and the
`postgres` module raised a second server that published mesh-wide. Collapsing to one server
means the one server has to be the mesh-wide one, reachable across the overlay.
## Open questions
- Does the mesh deliver the store and broker endpoint to a remote consumer as an address that
consumer can dial — the provider's overlay address — rather than a loopback or bridge
address meaningful only on the control-node?
- Is a `0.0.0.0` bind on the control-node reachable over the WireGuard overlay from a joined
machine, and is the packet filter's `from: mesh` rule enough to let it through?
- What is the smallest multi-node bed that would prove a database granted to a module on a
joined machine can be opened from there?
## Diagnosis (2026-09-17)
A focused two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`, scenario
`adopted-store-cross-node.yml`) puts the adopted broker on `anchor` and runs `amqp-ping` on the
joined `node2`. It fails, and the cause is two-layered — the first fixed, the second the real one:
1. **Firewall (fixed).** The adopted `lavinmq` module declared only `5672` in `listens`, not the
amqps bus port `5671`. mesh-broker's ports are published (DNAT'd → the nftables *forward* chain),
so with no `5671` forward rule the bus was dropped cross-node. Fix: the module declares `5671`
from `mesh` too (mesh-catalog `fix/broker-declares-amqps-port`). Confirmed: after it, `node2 ->
10.42.0.1:5671` (overlay) is OPEN and the forward chain carries `proto-dst 5671`.
2. **The broker address is the genesis public one, not the overlay (the real bug).** The broker
secret handed to a module is built at `cmd/mesh-controller/modules.go:294` (and the builder's at
`build.go:231`) as `amqps://<acct>:<pw>@<known.Address>/`, where `known.Address` is
`MESH_BROKER_ADDRESS` — the genesis PUBLIC endpoint (e.g. `192.0.2.10:5671`). `amqp-ping` on
node2 therefore dials `192.0.2.10:5671`, which the firewall does NOT admit cross-node (it admits
the overlay `10.42.0.1:5671`), and times out "fetching the broker's certificate." Its grant
never completes, so the provisioner's receives file shows `given: []` and no vhost is minted —
the downstream symptom.
The consumer already resolves the provider's `.internal` overlay name for its *provision* binding
(`bound.at == anchor.internal`) — that half works. What does not is the *bus account* URL, which
uses the public address for every module regardless of node.
## Open questions
- The broker URL for a module on a joined node should name the control-node's overlay address
(`<control-node>.internal:5671`), which resolves (via `--add-host`) and the firewall admits.
- But genesis-time accounts (the builder, issued on the control-node before the overlay exists)
must keep a locally-reachable address, or genesis breaks. So the fix is conditional on the
target node being on the overlay, not a blanket switch — where is that known at `module issue`?
- Should the control-node's own modules also move to the overlay name for uniformity, or keep the
local address they already reach?
## Resolution (2026-09-17)
A two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`) proves it: a
consumer (`amqp-ping`) on a joined node reaches the adopted broker (`mesh-broker`) on the
control-node over the overlay, its binding names `anchor.internal`, its vhost is minted, and it
holds the connection. Three things were wrong, now fixed:
1. **The broker's amqps port was not in the firewall.** The `lavinmq` module declared only `5672`
in `listens`; the bus is `5671`. Published container ports are matched in the nftables *forward*
chain, so `5671` was dropped cross-node. The module now declares `5671` from `mesh`
(mesh-catalog `fix/broker-declares-amqps-port`).
2. **The broker credential named the public address.** `module issue` and `builder issue` built
the URL with `MESH_BROKER_ADDRESS` — the broker's genesis public endpoint, which a joined
node's firewall does not admit. A new `brokerReachableAt` returns the broker as the target node
can reach it: a node on the overlay gets the hub's `.internal` name (fingerprint pinning makes
the host swap safe for TLS); a node not yet on the overlay keeps the genesis address, so
bring-up is unchanged (mesh-controller `multi-node/broker-reaches-over-overlay`).
3. **The provider was not recomposed after the remote consumer arrived (operational).** A
provision secret is minted as a side-effect of composing the CONSUMER's plan; the provider's
grant list is a pure read of secrets already issued from it. So the provider's provisioner
learns of a cross-node consumer only when the provider node is composed again. Pushing the
provider node after adding the remote consumer mints the vhost. This is not a code fix — it is
an ordering the operator must follow, and its silent-failure edge is opened as issue 057.