Merge pull request 'Issue 055 resolved and proven; 056 and 057 opened' (#45) from multi-node/harden-and-prove into main

This commit was merged in pull request #45.
This commit is contained in:
2026-09-17 01:44:59 +02:00
3 changed files with 155 additions and 3 deletions
@@ -1,8 +1,11 @@
---
status: open
status: resolved
opened: 2026-09-16
located-in: []
fixed-by:
located-in:
- mesh-controller/cmd/mesh-controller/modules.go
- mesh-controller/cmd/mesh-controller/build.go
- mesh-catalog/modules/lavinmq/module.json
fixed-by: mesh-controller multi-node/broker-reaches-over-overlay; mesh-catalog fix/broker-declares-amqps-port
amended-design:
---
@@ -38,3 +41,64 @@ means the one server has to be the mesh-wide one, reachable across the overlay.
machine, and is the packet filter's `from: mesh` rule enough to let it through?
- What is the smallest multi-node bed that would prove a database granted to a module on a
joined machine can be opened from there?
## Diagnosis (2026-09-17)
A focused two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`, scenario
`adopted-store-cross-node.yml`) puts the adopted broker on `anchor` and runs `amqp-ping` on the
joined `node2`. It fails, and the cause is two-layered — the first fixed, the second the real one:
1. **Firewall (fixed).** The adopted `lavinmq` module declared only `5672` in `listens`, not the
amqps bus port `5671`. mesh-broker's ports are published (DNAT'd → the nftables *forward* chain),
so with no `5671` forward rule the bus was dropped cross-node. Fix: the module declares `5671`
from `mesh` too (mesh-catalog `fix/broker-declares-amqps-port`). Confirmed: after it, `node2 ->
10.42.0.1:5671` (overlay) is OPEN and the forward chain carries `proto-dst 5671`.
2. **The broker address is the genesis public one, not the overlay (the real bug).** The broker
secret handed to a module is built at `cmd/mesh-controller/modules.go:294` (and the builder's at
`build.go:231`) as `amqps://<acct>:<pw>@<known.Address>/`, where `known.Address` is
`MESH_BROKER_ADDRESS` — the genesis PUBLIC endpoint (e.g. `192.0.2.10:5671`). `amqp-ping` on
node2 therefore dials `192.0.2.10:5671`, which the firewall does NOT admit cross-node (it admits
the overlay `10.42.0.1:5671`), and times out "fetching the broker's certificate." Its grant
never completes, so the provisioner's receives file shows `given: []` and no vhost is minted —
the downstream symptom.
The consumer already resolves the provider's `.internal` overlay name for its *provision* binding
(`bound.at == anchor.internal`) — that half works. What does not is the *bus account* URL, which
uses the public address for every module regardless of node.
## Open questions
- The broker URL for a module on a joined node should name the control-node's overlay address
(`<control-node>.internal:5671`), which resolves (via `--add-host`) and the firewall admits.
- But genesis-time accounts (the builder, issued on the control-node before the overlay exists)
must keep a locally-reachable address, or genesis breaks. So the fix is conditional on the
target node being on the overlay, not a blanket switch — where is that known at `module issue`?
- Should the control-node's own modules also move to the overlay name for uniformity, or keep the
local address they already reach?
## Resolution (2026-09-17)
A two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`) proves it: a
consumer (`amqp-ping`) on a joined node reaches the adopted broker (`mesh-broker`) on the
control-node over the overlay, its binding names `anchor.internal`, its vhost is minted, and it
holds the connection. Three things were wrong, now fixed:
1. **The broker's amqps port was not in the firewall.** The `lavinmq` module declared only `5672`
in `listens`; the bus is `5671`. Published container ports are matched in the nftables *forward*
chain, so `5671` was dropped cross-node. The module now declares `5671` from `mesh`
(mesh-catalog `fix/broker-declares-amqps-port`).
2. **The broker credential named the public address.** `module issue` and `builder issue` built
the URL with `MESH_BROKER_ADDRESS` — the broker's genesis public endpoint, which a joined
node's firewall does not admit. A new `brokerReachableAt` returns the broker as the target node
can reach it: a node on the overlay gets the hub's `.internal` name (fingerprint pinning makes
the host swap safe for TLS); a node not yet on the overlay keeps the genesis address, so
bring-up is unchanged (mesh-controller `multi-node/broker-reaches-over-overlay`).
3. **The provider was not recomposed after the remote consumer arrived (operational).** A
provision secret is minted as a side-effect of composing the CONSUMER's plan; the provider's
grant list is a pure read of secrets already issued from it. So the provider's provisioner
learns of a cross-node consumer only when the provider node is composed again. Pushing the
provider node after adding the remote consumer mints the vhost. This is not a code fix — it is
an ordering the operator must follow, and its silent-failure edge is opened as issue 057.
@@ -0,0 +1,42 @@
---
status: open
opened: 2026-09-17
located-in: [mesh-controller, mesh-host]
fixed-by:
amended-design:
---
# 056 — An adopted module assigned to a second node raises a second server
## Symptom
The store and broker are adopted in place ([ADR 0078](../../02-DECISIONS/0078-the-store-and-broker-are-modules.md),
issue 051): genesis raises `mesh-store`/`mesh-broker` on the control-node, and assigning the
`postgres`/`lavinmq` module there makes the module reconcile the already-running container instead
of raising a new one. Adoption works *because the container is already running on that node*.
Nothing stops the same module being assigned to a **second** node. On a node where no
`mesh-store` is running, the applier finds no container by that name and raises one — a second
postgres, a second broker. The whole point of Phase 3 — "one postgres, one lavinmq" — holds only
by the convention that genesis assigns these modules to the control-node alone; it is not enforced.
## Why it matters
"There is one store and one broker" is stated as a property of the mesh, and a property enforced by
nothing is indistinguishable from a wrong one. A single extra `assign` — the ordinary verb an
operator types — silently produces a divergent second server holding none of the first's data, and
consumers resolved to it get an empty store. The failure is silent and far from its cause.
The exclusive foundation modules are the mesh's clearest case of "there is one of me," yet unlike a
mesh-scoped seat, an adopted module carries no claim that the resolver would refuse a second holder
for.
## Open questions
- Should `postgres`/`lavinmq` claim a mesh-scoped exclusive seat (the way `mesh-controller` claims
`the-controller`), so the resolver refuses a second assignment by the same mechanism that keeps
one controller?
- Or should adoption be explicit — a module that adopts a foundation container declares it, and the
controller refuses to place it on a node whose foundation did not raise that container?
- How is "one postgres, one lavinmq" checked, rather than assumed — in the resolver, in `status`, or
in a bed that tries the second assignment and asserts the refusal?
@@ -0,0 +1,46 @@
---
status: open
opened: 2026-09-17
located-in: []
fixed-by:
amended-design:
---
# 057 — A cross-node consumer is provisioned only when the provider is pushed again
## Symptom
Adding a module that requires a mesh-scoped provision (a broker vhost, a database) on one node,
and pushing that node, is not enough for the provision to be made. The consumer starts, cannot
connect, and retries forever, while nothing says why. The provision appears only after the
*provider's* node is pushed a second time — an action the operator has no signal to take.
Observed while proving [issue 055](../055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md).
A consumer was assigned to a joined node and that node was pushed; the consumer came up but
logged `amqp connection closed before the round-trip completed` on a loop. The provider's grant
file listed no consumers (`given: []`), so its provisioner minted nothing. Pushing the provider's
node again populated the grant, the vhost was minted, and the consumer connected on its next
retry. No error was raised at any point — the only symptom was a consumer that never became ready.
## Why it matters beyond the instance
A provision secret is minted as a side-effect of composing the **consumer's** plan; the provider's
grant list is a pure read of secrets already issued from it. So composing the provider *before* a
new cross-node consumer exists — or never recomposing it — leaves the provider blind to that
consumer. Co-located consumer and provider hide this: one push composes both. The failure is
therefore invisible until the mesh is actually multi-node, which is exactly when it is hardest to
diagnose, and it fails silently — the one shape [AGENTS.md](../../AGENTS.md) says belongs here.
The rule "push the provider node too" is real and undocumented, and a rule enforced by nothing is
indistinguishable from a wrong one.
## Open questions
- Should composing any node that adds or removes a cross-node consumer also recompose the
providers it now binds to, so a single push is sufficient? What is the blast radius of that?
- Failing that, should composing a consumer whose provider is on another node *report* that the
provider must be pushed, rather than issue a grant no one will read?
- A consumer that cannot reach its provision retries silently. Should a consumer that has waited
past some bound surface as un-ready in the mesh's own view, not only in its container log?
- Does the same gap withdraw provisions late — is a provider told a cross-node consumer left only
when the provider is next pushed?