From bec1dd008219f90beaa6f5175a65f9c4b9d804aa Mon Sep 17 00:00:00 2001 From: jochen Date: Thu, 17 Sep 2026 00:22:18 +0200 Subject: [PATCH 1/3] =?UTF-8?q?Issue=20056=20=E2=80=94=20an=20adopted=20mo?= =?UTF-8?q?dule=20assigned=20to=20a=20second=20node=20raises=20a=20second?= =?UTF-8?q?=20server?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit "One postgres, one lavinmq" (ADR 0078) holds only by genesis assigning the adopted modules to the control-node alone; nothing enforces it. A second assign raises a divergent second server, silently. Open questions: a mesh-scoped exclusive seat, or explicit adoption the controller refuses to place elsewhere. https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx --- .../00-report.md | 42 +++++++++++++++++++ 1 file changed, 42 insertions(+) create mode 100644 04-ISSUES/056-an-adopted-module-assigned-to-a-second-node-raises-a-second-server/00-report.md diff --git a/04-ISSUES/056-an-adopted-module-assigned-to-a-second-node-raises-a-second-server/00-report.md b/04-ISSUES/056-an-adopted-module-assigned-to-a-second-node-raises-a-second-server/00-report.md new file mode 100644 index 0000000..5088ec3 --- /dev/null +++ b/04-ISSUES/056-an-adopted-module-assigned-to-a-second-node-raises-a-second-server/00-report.md @@ -0,0 +1,42 @@ +--- +status: open +opened: 2026-09-17 +located-in: [mesh-controller, mesh-host] +fixed-by: +amended-design: +--- + +# 056 — An adopted module assigned to a second node raises a second server + +## Symptom + +The store and broker are adopted in place ([ADR 0078](../../02-DECISIONS/0078-the-store-and-broker-are-modules.md), +issue 051): genesis raises `mesh-store`/`mesh-broker` on the control-node, and assigning the +`postgres`/`lavinmq` module there makes the module reconcile the already-running container instead +of raising a new one. Adoption works *because the container is already running on that node*. + +Nothing stops the same module being assigned to a **second** node. On a node where no +`mesh-store` is running, the applier finds no container by that name and raises one — a second +postgres, a second broker. The whole point of Phase 3 — "one postgres, one lavinmq" — holds only +by the convention that genesis assigns these modules to the control-node alone; it is not enforced. + +## Why it matters + +"There is one store and one broker" is stated as a property of the mesh, and a property enforced by +nothing is indistinguishable from a wrong one. A single extra `assign` — the ordinary verb an +operator types — silently produces a divergent second server holding none of the first's data, and +consumers resolved to it get an empty store. The failure is silent and far from its cause. + +The exclusive foundation modules are the mesh's clearest case of "there is one of me," yet unlike a +mesh-scoped seat, an adopted module carries no claim that the resolver would refuse a second holder +for. + +## Open questions + +- Should `postgres`/`lavinmq` claim a mesh-scoped exclusive seat (the way `mesh-controller` claims + `the-controller`), so the resolver refuses a second assignment by the same mechanism that keeps + one controller? +- Or should adoption be explicit — a module that adopts a foundation container declares it, and the + controller refuses to place it on a node whose foundation did not raise that container? +- How is "one postgres, one lavinmq" checked, rather than assumed — in the resolver, in `status`, or + in a bed that tries the second assignment and asserts the refusal? From 33fc322f4da6d2e9dd8434b8f364d655d2871743 Mon Sep 17 00:00:00 2001 From: jochen Date: Thu, 17 Sep 2026 01:02:07 +0200 Subject: [PATCH 2/3] =?UTF-8?q?Issue=20055=20diagnosed=20=E2=80=94=20the?= =?UTF-8?q?=20module=20broker=20URL=20uses=20the=20public=20address,=20not?= =?UTF-8?q?=20the=20overlay?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A two-node bed pins it: the firewall gap (broker's 5671 not in listens) is fixed in mesh-catalog; the real bug is modules.go:294/build.go:231 building the broker URL with the genesis public MESH_BROKER_ADDRESS, which a joined node cannot reach. https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx --- .../00-report.md | 37 ++++++++++++++++++- 1 file changed, 36 insertions(+), 1 deletion(-) diff --git a/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md b/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md index d856dc0..c6a33ab 100644 --- a/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md +++ b/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md @@ -1,5 +1,5 @@ --- -status: open +status: diagnosing opened: 2026-09-16 located-in: [] fixed-by: @@ -38,3 +38,38 @@ means the one server has to be the mesh-wide one, reachable across the overlay. machine, and is the packet filter's `from: mesh` rule enough to let it through? - What is the smallest multi-node bed that would prove a database granted to a module on a joined machine can be opened from there? + +## Diagnosis (2026-09-17) + +A focused two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`, scenario +`adopted-store-cross-node.yml`) puts the adopted broker on `anchor` and runs `amqp-ping` on the +joined `node2`. It fails, and the cause is two-layered — the first fixed, the second the real one: + +1. **Firewall (fixed).** The adopted `lavinmq` module declared only `5672` in `listens`, not the + amqps bus port `5671`. mesh-broker's ports are published (DNAT'd → the nftables *forward* chain), + so with no `5671` forward rule the bus was dropped cross-node. Fix: the module declares `5671` + from `mesh` too (mesh-catalog `fix/broker-declares-amqps-port`). Confirmed: after it, `node2 -> + 10.42.0.1:5671` (overlay) is OPEN and the forward chain carries `proto-dst 5671`. + +2. **The broker address is the genesis public one, not the overlay (the real bug).** The broker + secret handed to a module is built at `cmd/mesh-controller/modules.go:294` (and the builder's at + `build.go:231`) as `amqps://:@/`, where `known.Address` is + `MESH_BROKER_ADDRESS` — the genesis PUBLIC endpoint (e.g. `192.0.2.10:5671`). `amqp-ping` on + node2 therefore dials `192.0.2.10:5671`, which the firewall does NOT admit cross-node (it admits + the overlay `10.42.0.1:5671`), and times out "fetching the broker's certificate." Its grant + never completes, so the provisioner's receives file shows `given: []` and no vhost is minted — + the downstream symptom. + +The consumer already resolves the provider's `.internal` overlay name for its *provision* binding +(`bound.at == anchor.internal`) — that half works. What does not is the *bus account* URL, which +uses the public address for every module regardless of node. + +## Open questions + +- The broker URL for a module on a joined node should name the control-node's overlay address + (`.internal:5671`), which resolves (via `--add-host`) and the firewall admits. +- But genesis-time accounts (the builder, issued on the control-node before the overlay exists) + must keep a locally-reachable address, or genesis breaks. So the fix is conditional on the + target node being on the overlay, not a blanket switch — where is that known at `module issue`? +- Should the control-node's own modules also move to the overlay name for uniformity, or keep the + local address they already reach? From 812dc3b30356abedbf9ffb738d51513167224c96 Mon Sep 17 00:00:00 2001 From: jochen Date: Thu, 17 Sep 2026 01:35:47 +0200 Subject: [PATCH 3/3] Issue 055 resolved, and 057 opened for its silent operational edge MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 055 is proven end-to-end by the two-node bed: a consumer on a joined node reaches the adopted broker over the overlay, its binding names the control-node, and its vhost is minted. Three fixes — the broker's amqps port in the firewall, the broker credential naming the overlay not the public address, and pushing the provider node after the remote consumer arrives. The last is an operator ordering, not a code fix, and its silent-failure edge (a cross-node consumer that never provisions and only crash-loops) is opened as 057. https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx --- .../00-report.md | 35 ++++++++++++-- .../00-report.md | 46 +++++++++++++++++++ 2 files changed, 78 insertions(+), 3 deletions(-) create mode 100644 04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md diff --git a/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md b/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md index c6a33ab..8cb1ea0 100644 --- a/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md +++ b/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md @@ -1,8 +1,11 @@ --- -status: diagnosing +status: resolved opened: 2026-09-16 -located-in: [] -fixed-by: +located-in: + - mesh-controller/cmd/mesh-controller/modules.go + - mesh-controller/cmd/mesh-controller/build.go + - mesh-catalog/modules/lavinmq/module.json +fixed-by: mesh-controller multi-node/broker-reaches-over-overlay; mesh-catalog fix/broker-declares-amqps-port amended-design: --- @@ -73,3 +76,29 @@ uses the public address for every module regardless of node. target node being on the overlay, not a blanket switch — where is that known at `module issue`? - Should the control-node's own modules also move to the overlay name for uniformity, or keep the local address they already reach? + +## Resolution (2026-09-17) + +A two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`) proves it: a +consumer (`amqp-ping`) on a joined node reaches the adopted broker (`mesh-broker`) on the +control-node over the overlay, its binding names `anchor.internal`, its vhost is minted, and it +holds the connection. Three things were wrong, now fixed: + +1. **The broker's amqps port was not in the firewall.** The `lavinmq` module declared only `5672` + in `listens`; the bus is `5671`. Published container ports are matched in the nftables *forward* + chain, so `5671` was dropped cross-node. The module now declares `5671` from `mesh` + (mesh-catalog `fix/broker-declares-amqps-port`). + +2. **The broker credential named the public address.** `module issue` and `builder issue` built + the URL with `MESH_BROKER_ADDRESS` — the broker's genesis public endpoint, which a joined + node's firewall does not admit. A new `brokerReachableAt` returns the broker as the target node + can reach it: a node on the overlay gets the hub's `.internal` name (fingerprint pinning makes + the host swap safe for TLS); a node not yet on the overlay keeps the genesis address, so + bring-up is unchanged (mesh-controller `multi-node/broker-reaches-over-overlay`). + +3. **The provider was not recomposed after the remote consumer arrived (operational).** A + provision secret is minted as a side-effect of composing the CONSUMER's plan; the provider's + grant list is a pure read of secrets already issued from it. So the provider's provisioner + learns of a cross-node consumer only when the provider node is composed again. Pushing the + provider node after adding the remote consumer mints the vhost. This is not a code fix — it is + an ordering the operator must follow, and its silent-failure edge is opened as issue 057. diff --git a/04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md b/04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md new file mode 100644 index 0000000..a818050 --- /dev/null +++ b/04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md @@ -0,0 +1,46 @@ +--- +status: open +opened: 2026-09-17 +located-in: [] +fixed-by: +amended-design: +--- + +# 057 — A cross-node consumer is provisioned only when the provider is pushed again + +## Symptom + +Adding a module that requires a mesh-scoped provision (a broker vhost, a database) on one node, +and pushing that node, is not enough for the provision to be made. The consumer starts, cannot +connect, and retries forever, while nothing says why. The provision appears only after the +*provider's* node is pushed a second time — an action the operator has no signal to take. + +Observed while proving [issue 055](../055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md). +A consumer was assigned to a joined node and that node was pushed; the consumer came up but +logged `amqp connection closed before the round-trip completed` on a loop. The provider's grant +file listed no consumers (`given: []`), so its provisioner minted nothing. Pushing the provider's +node again populated the grant, the vhost was minted, and the consumer connected on its next +retry. No error was raised at any point — the only symptom was a consumer that never became ready. + +## Why it matters beyond the instance + +A provision secret is minted as a side-effect of composing the **consumer's** plan; the provider's +grant list is a pure read of secrets already issued from it. So composing the provider *before* a +new cross-node consumer exists — or never recomposing it — leaves the provider blind to that +consumer. Co-located consumer and provider hide this: one push composes both. The failure is +therefore invisible until the mesh is actually multi-node, which is exactly when it is hardest to +diagnose, and it fails silently — the one shape [AGENTS.md](../../AGENTS.md) says belongs here. + +The rule "push the provider node too" is real and undocumented, and a rule enforced by nothing is +indistinguishable from a wrong one. + +## Open questions + +- Should composing any node that adds or removes a cross-node consumer also recompose the + providers it now binds to, so a single push is sufficient? What is the blast radius of that? +- Failing that, should composing a consumer whose provider is on another node *report* that the + provider must be pushed, rather than issue a grant no one will read? +- A consumer that cannot reach its provision retries silently. Should a consumer that has waited + past some bound surface as un-ready in the mesh's own view, not only in its container log? +- Does the same gap withdraw provisions late — is a provider told a cross-node consumer left only + when the provider is next pushed?