diff --git a/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md b/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md index c6a33ab..8cb1ea0 100644 --- a/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md +++ b/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md @@ -1,8 +1,11 @@ --- -status: diagnosing +status: resolved opened: 2026-09-16 -located-in: [] -fixed-by: +located-in: + - mesh-controller/cmd/mesh-controller/modules.go + - mesh-controller/cmd/mesh-controller/build.go + - mesh-catalog/modules/lavinmq/module.json +fixed-by: mesh-controller multi-node/broker-reaches-over-overlay; mesh-catalog fix/broker-declares-amqps-port amended-design: --- @@ -73,3 +76,29 @@ uses the public address for every module regardless of node. target node being on the overlay, not a blanket switch — where is that known at `module issue`? - Should the control-node's own modules also move to the overlay name for uniformity, or keep the local address they already reach? + +## Resolution (2026-09-17) + +A two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`) proves it: a +consumer (`amqp-ping`) on a joined node reaches the adopted broker (`mesh-broker`) on the +control-node over the overlay, its binding names `anchor.internal`, its vhost is minted, and it +holds the connection. Three things were wrong, now fixed: + +1. **The broker's amqps port was not in the firewall.** The `lavinmq` module declared only `5672` + in `listens`; the bus is `5671`. Published container ports are matched in the nftables *forward* + chain, so `5671` was dropped cross-node. The module now declares `5671` from `mesh` + (mesh-catalog `fix/broker-declares-amqps-port`). + +2. **The broker credential named the public address.** `module issue` and `builder issue` built + the URL with `MESH_BROKER_ADDRESS` — the broker's genesis public endpoint, which a joined + node's firewall does not admit. A new `brokerReachableAt` returns the broker as the target node + can reach it: a node on the overlay gets the hub's `.internal` name (fingerprint pinning makes + the host swap safe for TLS); a node not yet on the overlay keeps the genesis address, so + bring-up is unchanged (mesh-controller `multi-node/broker-reaches-over-overlay`). + +3. **The provider was not recomposed after the remote consumer arrived (operational).** A + provision secret is minted as a side-effect of composing the CONSUMER's plan; the provider's + grant list is a pure read of secrets already issued from it. So the provider's provisioner + learns of a cross-node consumer only when the provider node is composed again. Pushing the + provider node after adding the remote consumer mints the vhost. This is not a code fix — it is + an ordering the operator must follow, and its silent-failure edge is opened as issue 057. diff --git a/04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md b/04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md new file mode 100644 index 0000000..a818050 --- /dev/null +++ b/04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md @@ -0,0 +1,46 @@ +--- +status: open +opened: 2026-09-17 +located-in: [] +fixed-by: +amended-design: +--- + +# 057 — A cross-node consumer is provisioned only when the provider is pushed again + +## Symptom + +Adding a module that requires a mesh-scoped provision (a broker vhost, a database) on one node, +and pushing that node, is not enough for the provision to be made. The consumer starts, cannot +connect, and retries forever, while nothing says why. The provision appears only after the +*provider's* node is pushed a second time — an action the operator has no signal to take. + +Observed while proving [issue 055](../055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md). +A consumer was assigned to a joined node and that node was pushed; the consumer came up but +logged `amqp connection closed before the round-trip completed` on a loop. The provider's grant +file listed no consumers (`given: []`), so its provisioner minted nothing. Pushing the provider's +node again populated the grant, the vhost was minted, and the consumer connected on its next +retry. No error was raised at any point — the only symptom was a consumer that never became ready. + +## Why it matters beyond the instance + +A provision secret is minted as a side-effect of composing the **consumer's** plan; the provider's +grant list is a pure read of secrets already issued from it. So composing the provider *before* a +new cross-node consumer exists — or never recomposing it — leaves the provider blind to that +consumer. Co-located consumer and provider hide this: one push composes both. The failure is +therefore invisible until the mesh is actually multi-node, which is exactly when it is hardest to +diagnose, and it fails silently — the one shape [AGENTS.md](../../AGENTS.md) says belongs here. + +The rule "push the provider node too" is real and undocumented, and a rule enforced by nothing is +indistinguishable from a wrong one. + +## Open questions + +- Should composing any node that adds or removes a cross-node consumer also recompose the + providers it now binds to, so a single push is sufficient? What is the blast radius of that? +- Failing that, should composing a consumer whose provider is on another node *report* that the + provider must be pushed, rather than issue a grant no one will read? +- A consumer that cannot reach its provision retries silently. Should a consumer that has waited + past some bound surface as un-ready in the mesh's own view, not only in its container log? +- Does the same gap withdraw provisions late — is a provider told a cross-node consumer left only + when the provider is next pushed?