Issue 055 resolved and proven; 056 and 057 opened #45
+32
-3
@@ -1,8 +1,11 @@
|
||||
---
|
||||
status: diagnosing
|
||||
status: resolved
|
||||
opened: 2026-09-16
|
||||
located-in: []
|
||||
fixed-by:
|
||||
located-in:
|
||||
- mesh-controller/cmd/mesh-controller/modules.go
|
||||
- mesh-controller/cmd/mesh-controller/build.go
|
||||
- mesh-catalog/modules/lavinmq/module.json
|
||||
fixed-by: mesh-controller multi-node/broker-reaches-over-overlay; mesh-catalog fix/broker-declares-amqps-port
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -73,3 +76,29 @@ uses the public address for every module regardless of node.
|
||||
target node being on the overlay, not a blanket switch — where is that known at `module issue`?
|
||||
- Should the control-node's own modules also move to the overlay name for uniformity, or keep the
|
||||
local address they already reach?
|
||||
|
||||
## Resolution (2026-09-17)
|
||||
|
||||
A two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`) proves it: a
|
||||
consumer (`amqp-ping`) on a joined node reaches the adopted broker (`mesh-broker`) on the
|
||||
control-node over the overlay, its binding names `anchor.internal`, its vhost is minted, and it
|
||||
holds the connection. Three things were wrong, now fixed:
|
||||
|
||||
1. **The broker's amqps port was not in the firewall.** The `lavinmq` module declared only `5672`
|
||||
in `listens`; the bus is `5671`. Published container ports are matched in the nftables *forward*
|
||||
chain, so `5671` was dropped cross-node. The module now declares `5671` from `mesh`
|
||||
(mesh-catalog `fix/broker-declares-amqps-port`).
|
||||
|
||||
2. **The broker credential named the public address.** `module issue` and `builder issue` built
|
||||
the URL with `MESH_BROKER_ADDRESS` — the broker's genesis public endpoint, which a joined
|
||||
node's firewall does not admit. A new `brokerReachableAt` returns the broker as the target node
|
||||
can reach it: a node on the overlay gets the hub's `.internal` name (fingerprint pinning makes
|
||||
the host swap safe for TLS); a node not yet on the overlay keeps the genesis address, so
|
||||
bring-up is unchanged (mesh-controller `multi-node/broker-reaches-over-overlay`).
|
||||
|
||||
3. **The provider was not recomposed after the remote consumer arrived (operational).** A
|
||||
provision secret is minted as a side-effect of composing the CONSUMER's plan; the provider's
|
||||
grant list is a pure read of secrets already issued from it. So the provider's provisioner
|
||||
learns of a cross-node consumer only when the provider node is composed again. Pushing the
|
||||
provider node after adding the remote consumer mints the vhost. This is not a code fix — it is
|
||||
an ordering the operator must follow, and its silent-failure edge is opened as issue 057.
|
||||
|
||||
+46
@@ -0,0 +1,46 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-17
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 057 — A cross-node consumer is provisioned only when the provider is pushed again
|
||||
|
||||
## Symptom
|
||||
|
||||
Adding a module that requires a mesh-scoped provision (a broker vhost, a database) on one node,
|
||||
and pushing that node, is not enough for the provision to be made. The consumer starts, cannot
|
||||
connect, and retries forever, while nothing says why. The provision appears only after the
|
||||
*provider's* node is pushed a second time — an action the operator has no signal to take.
|
||||
|
||||
Observed while proving [issue 055](../055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md).
|
||||
A consumer was assigned to a joined node and that node was pushed; the consumer came up but
|
||||
logged `amqp connection closed before the round-trip completed` on a loop. The provider's grant
|
||||
file listed no consumers (`given: []`), so its provisioner minted nothing. Pushing the provider's
|
||||
node again populated the grant, the vhost was minted, and the consumer connected on its next
|
||||
retry. No error was raised at any point — the only symptom was a consumer that never became ready.
|
||||
|
||||
## Why it matters beyond the instance
|
||||
|
||||
A provision secret is minted as a side-effect of composing the **consumer's** plan; the provider's
|
||||
grant list is a pure read of secrets already issued from it. So composing the provider *before* a
|
||||
new cross-node consumer exists — or never recomposing it — leaves the provider blind to that
|
||||
consumer. Co-located consumer and provider hide this: one push composes both. The failure is
|
||||
therefore invisible until the mesh is actually multi-node, which is exactly when it is hardest to
|
||||
diagnose, and it fails silently — the one shape [AGENTS.md](../../AGENTS.md) says belongs here.
|
||||
|
||||
The rule "push the provider node too" is real and undocumented, and a rule enforced by nothing is
|
||||
indistinguishable from a wrong one.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should composing any node that adds or removes a cross-node consumer also recompose the
|
||||
providers it now binds to, so a single push is sufficient? What is the blast radius of that?
|
||||
- Failing that, should composing a consumer whose provider is on another node *report* that the
|
||||
provider must be pushed, rather than issue a grant no one will read?
|
||||
- A consumer that cannot reach its provision retries silently. Should a consumer that has waited
|
||||
past some bound surface as un-ready in the mesh's own view, not only in its container log?
|
||||
- Does the same gap withdraw provisions late — is a provider told a cross-node consumer left only
|
||||
when the provider is next pushed?
|
||||
Reference in New Issue
Block a user