Issue 055 resolved, and 057 opened for its silent operational edge

055 is proven end-to-end by the two-node bed: a consumer on a joined node reaches the
adopted broker over the overlay, its binding names the control-node, and its vhost is
minted. Three fixes — the broker's amqps port in the firewall, the broker credential
naming the overlay not the public address, and pushing the provider node after the remote
consumer arrives. The last is an operator ordering, not a code fix, and its silent-failure
edge (a cross-node consumer that never provisions and only crash-loops) is opened as 057.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
This commit is contained in:
2026-09-17 01:35:47 +02:00
parent 33fc322f4d
commit 812dc3b303
2 changed files with 78 additions and 3 deletions
@@ -1,8 +1,11 @@
---
status: diagnosing
status: resolved
opened: 2026-09-16
located-in: []
fixed-by:
located-in:
- mesh-controller/cmd/mesh-controller/modules.go
- mesh-controller/cmd/mesh-controller/build.go
- mesh-catalog/modules/lavinmq/module.json
fixed-by: mesh-controller multi-node/broker-reaches-over-overlay; mesh-catalog fix/broker-declares-amqps-port
amended-design:
---
@@ -73,3 +76,29 @@ uses the public address for every module regardless of node.
target node being on the overlay, not a blanket switch — where is that known at `module issue`?
- Should the control-node's own modules also move to the overlay name for uniformity, or keep the
local address they already reach?
## Resolution (2026-09-17)
A two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`) proves it: a
consumer (`amqp-ping`) on a joined node reaches the adopted broker (`mesh-broker`) on the
control-node over the overlay, its binding names `anchor.internal`, its vhost is minted, and it
holds the connection. Three things were wrong, now fixed:
1. **The broker's amqps port was not in the firewall.** The `lavinmq` module declared only `5672`
in `listens`; the bus is `5671`. Published container ports are matched in the nftables *forward*
chain, so `5671` was dropped cross-node. The module now declares `5671` from `mesh`
(mesh-catalog `fix/broker-declares-amqps-port`).
2. **The broker credential named the public address.** `module issue` and `builder issue` built
the URL with `MESH_BROKER_ADDRESS` — the broker's genesis public endpoint, which a joined
node's firewall does not admit. A new `brokerReachableAt` returns the broker as the target node
can reach it: a node on the overlay gets the hub's `.internal` name (fingerprint pinning makes
the host swap safe for TLS); a node not yet on the overlay keeps the genesis address, so
bring-up is unchanged (mesh-controller `multi-node/broker-reaches-over-overlay`).
3. **The provider was not recomposed after the remote consumer arrived (operational).** A
provision secret is minted as a side-effect of composing the CONSUMER's plan; the provider's
grant list is a pure read of secrets already issued from it. So the provider's provisioner
learns of a cross-node consumer only when the provider node is composed again. Pushing the
provider node after adding the remote consumer mints the vhost. This is not a code fix — it is
an ordering the operator must follow, and its silent-failure edge is opened as issue 057.
@@ -0,0 +1,46 @@
---
status: open
opened: 2026-09-17
located-in: []
fixed-by:
amended-design:
---
# 057 — A cross-node consumer is provisioned only when the provider is pushed again
## Symptom
Adding a module that requires a mesh-scoped provision (a broker vhost, a database) on one node,
and pushing that node, is not enough for the provision to be made. The consumer starts, cannot
connect, and retries forever, while nothing says why. The provision appears only after the
*provider's* node is pushed a second time — an action the operator has no signal to take.
Observed while proving [issue 055](../055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md).
A consumer was assigned to a joined node and that node was pushed; the consumer came up but
logged `amqp connection closed before the round-trip completed` on a loop. The provider's grant
file listed no consumers (`given: []`), so its provisioner minted nothing. Pushing the provider's
node again populated the grant, the vhost was minted, and the consumer connected on its next
retry. No error was raised at any point — the only symptom was a consumer that never became ready.
## Why it matters beyond the instance
A provision secret is minted as a side-effect of composing the **consumer's** plan; the provider's
grant list is a pure read of secrets already issued from it. So composing the provider *before* a
new cross-node consumer exists — or never recomposing it — leaves the provider blind to that
consumer. Co-located consumer and provider hide this: one push composes both. The failure is
therefore invisible until the mesh is actually multi-node, which is exactly when it is hardest to
diagnose, and it fails silently — the one shape [AGENTS.md](../../AGENTS.md) says belongs here.
The rule "push the provider node too" is real and undocumented, and a rule enforced by nothing is
indistinguishable from a wrong one.
## Open questions
- Should composing any node that adds or removes a cross-node consumer also recompose the
providers it now binds to, so a single push is sufficient? What is the blast radius of that?
- Failing that, should composing a consumer whose provider is on another node *report* that the
provider must be pushed, rather than issue a grant no one will read?
- A consumer that cannot reach its provision retries silently. Should a consumer that has waited
past some bound surface as un-ready in the mesh's own view, not only in its container log?
- Does the same gap withdraw provisions late — is a provider told a cross-node consumer left only
when the provider is next pushed?