diff --git a/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md b/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md index d856dc0..8cb1ea0 100644 --- a/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md +++ b/04-ISSUES/055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md @@ -1,8 +1,11 @@ --- -status: open +status: resolved opened: 2026-09-16 -located-in: [] -fixed-by: +located-in: + - mesh-controller/cmd/mesh-controller/modules.go + - mesh-controller/cmd/mesh-controller/build.go + - mesh-catalog/modules/lavinmq/module.json +fixed-by: mesh-controller multi-node/broker-reaches-over-overlay; mesh-catalog fix/broker-declares-amqps-port amended-design: --- @@ -38,3 +41,64 @@ means the one server has to be the mesh-wide one, reachable across the overlay. machine, and is the packet filter's `from: mesh` rule enough to let it through? - What is the smallest multi-node bed that would prove a database granted to a module on a joined machine can be opened from there? + +## Diagnosis (2026-09-17) + +A focused two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`, scenario +`adopted-store-cross-node.yml`) puts the adopted broker on `anchor` and runs `amqp-ping` on the +joined `node2`. It fails, and the cause is two-layered — the first fixed, the second the real one: + +1. **Firewall (fixed).** The adopted `lavinmq` module declared only `5672` in `listens`, not the + amqps bus port `5671`. mesh-broker's ports are published (DNAT'd → the nftables *forward* chain), + so with no `5671` forward rule the bus was dropped cross-node. Fix: the module declares `5671` + from `mesh` too (mesh-catalog `fix/broker-declares-amqps-port`). Confirmed: after it, `node2 -> + 10.42.0.1:5671` (overlay) is OPEN and the forward chain carries `proto-dst 5671`. + +2. **The broker address is the genesis public one, not the overlay (the real bug).** The broker + secret handed to a module is built at `cmd/mesh-controller/modules.go:294` (and the builder's at + `build.go:231`) as `amqps://:@/`, where `known.Address` is + `MESH_BROKER_ADDRESS` — the genesis PUBLIC endpoint (e.g. `192.0.2.10:5671`). `amqp-ping` on + node2 therefore dials `192.0.2.10:5671`, which the firewall does NOT admit cross-node (it admits + the overlay `10.42.0.1:5671`), and times out "fetching the broker's certificate." Its grant + never completes, so the provisioner's receives file shows `given: []` and no vhost is minted — + the downstream symptom. + +The consumer already resolves the provider's `.internal` overlay name for its *provision* binding +(`bound.at == anchor.internal`) — that half works. What does not is the *bus account* URL, which +uses the public address for every module regardless of node. + +## Open questions + +- The broker URL for a module on a joined node should name the control-node's overlay address + (`.internal:5671`), which resolves (via `--add-host`) and the firewall admits. +- But genesis-time accounts (the builder, issued on the control-node before the overlay exists) + must keep a locally-reachable address, or genesis breaks. So the fix is conditional on the + target node being on the overlay, not a blanket switch — where is that known at `module issue`? +- Should the control-node's own modules also move to the overlay name for uniformity, or keep the + local address they already reach? + +## Resolution (2026-09-17) + +A two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`) proves it: a +consumer (`amqp-ping`) on a joined node reaches the adopted broker (`mesh-broker`) on the +control-node over the overlay, its binding names `anchor.internal`, its vhost is minted, and it +holds the connection. Three things were wrong, now fixed: + +1. **The broker's amqps port was not in the firewall.** The `lavinmq` module declared only `5672` + in `listens`; the bus is `5671`. Published container ports are matched in the nftables *forward* + chain, so `5671` was dropped cross-node. The module now declares `5671` from `mesh` + (mesh-catalog `fix/broker-declares-amqps-port`). + +2. **The broker credential named the public address.** `module issue` and `builder issue` built + the URL with `MESH_BROKER_ADDRESS` — the broker's genesis public endpoint, which a joined + node's firewall does not admit. A new `brokerReachableAt` returns the broker as the target node + can reach it: a node on the overlay gets the hub's `.internal` name (fingerprint pinning makes + the host swap safe for TLS); a node not yet on the overlay keeps the genesis address, so + bring-up is unchanged (mesh-controller `multi-node/broker-reaches-over-overlay`). + +3. **The provider was not recomposed after the remote consumer arrived (operational).** A + provision secret is minted as a side-effect of composing the CONSUMER's plan; the provider's + grant list is a pure read of secrets already issued from it. So the provider's provisioner + learns of a cross-node consumer only when the provider node is composed again. Pushing the + provider node after adding the remote consumer mints the vhost. This is not a code fix — it is + an ordering the operator must follow, and its silent-failure edge is opened as issue 057. diff --git a/04-ISSUES/056-an-adopted-module-assigned-to-a-second-node-raises-a-second-server/00-report.md b/04-ISSUES/056-an-adopted-module-assigned-to-a-second-node-raises-a-second-server/00-report.md new file mode 100644 index 0000000..5088ec3 --- /dev/null +++ b/04-ISSUES/056-an-adopted-module-assigned-to-a-second-node-raises-a-second-server/00-report.md @@ -0,0 +1,42 @@ +--- +status: open +opened: 2026-09-17 +located-in: [mesh-controller, mesh-host] +fixed-by: +amended-design: +--- + +# 056 — An adopted module assigned to a second node raises a second server + +## Symptom + +The store and broker are adopted in place ([ADR 0078](../../02-DECISIONS/0078-the-store-and-broker-are-modules.md), +issue 051): genesis raises `mesh-store`/`mesh-broker` on the control-node, and assigning the +`postgres`/`lavinmq` module there makes the module reconcile the already-running container instead +of raising a new one. Adoption works *because the container is already running on that node*. + +Nothing stops the same module being assigned to a **second** node. On a node where no +`mesh-store` is running, the applier finds no container by that name and raises one — a second +postgres, a second broker. The whole point of Phase 3 — "one postgres, one lavinmq" — holds only +by the convention that genesis assigns these modules to the control-node alone; it is not enforced. + +## Why it matters + +"There is one store and one broker" is stated as a property of the mesh, and a property enforced by +nothing is indistinguishable from a wrong one. A single extra `assign` — the ordinary verb an +operator types — silently produces a divergent second server holding none of the first's data, and +consumers resolved to it get an empty store. The failure is silent and far from its cause. + +The exclusive foundation modules are the mesh's clearest case of "there is one of me," yet unlike a +mesh-scoped seat, an adopted module carries no claim that the resolver would refuse a second holder +for. + +## Open questions + +- Should `postgres`/`lavinmq` claim a mesh-scoped exclusive seat (the way `mesh-controller` claims + `the-controller`), so the resolver refuses a second assignment by the same mechanism that keeps + one controller? +- Or should adoption be explicit — a module that adopts a foundation container declares it, and the + controller refuses to place it on a node whose foundation did not raise that container? +- How is "one postgres, one lavinmq" checked, rather than assumed — in the resolver, in `status`, or + in a bed that tries the second assignment and asserts the refusal? diff --git a/04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md b/04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md new file mode 100644 index 0000000..a818050 --- /dev/null +++ b/04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md @@ -0,0 +1,46 @@ +--- +status: open +opened: 2026-09-17 +located-in: [] +fixed-by: +amended-design: +--- + +# 057 — A cross-node consumer is provisioned only when the provider is pushed again + +## Symptom + +Adding a module that requires a mesh-scoped provision (a broker vhost, a database) on one node, +and pushing that node, is not enough for the provision to be made. The consumer starts, cannot +connect, and retries forever, while nothing says why. The provision appears only after the +*provider's* node is pushed a second time — an action the operator has no signal to take. + +Observed while proving [issue 055](../055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md). +A consumer was assigned to a joined node and that node was pushed; the consumer came up but +logged `amqp connection closed before the round-trip completed` on a loop. The provider's grant +file listed no consumers (`given: []`), so its provisioner minted nothing. Pushing the provider's +node again populated the grant, the vhost was minted, and the consumer connected on its next +retry. No error was raised at any point — the only symptom was a consumer that never became ready. + +## Why it matters beyond the instance + +A provision secret is minted as a side-effect of composing the **consumer's** plan; the provider's +grant list is a pure read of secrets already issued from it. So composing the provider *before* a +new cross-node consumer exists — or never recomposing it — leaves the provider blind to that +consumer. Co-located consumer and provider hide this: one push composes both. The failure is +therefore invisible until the mesh is actually multi-node, which is exactly when it is hardest to +diagnose, and it fails silently — the one shape [AGENTS.md](../../AGENTS.md) says belongs here. + +The rule "push the provider node too" is real and undocumented, and a rule enforced by nothing is +indistinguishable from a wrong one. + +## Open questions + +- Should composing any node that adds or removes a cross-node consumer also recompose the + providers it now binds to, so a single push is sufficient? What is the blast radius of that? +- Failing that, should composing a consumer whose provider is on another node *report* that the + provider must be pushed, rather than issue a grant no one will read? +- A consumer that cannot reach its provision retries silently. Should a consumer that has waited + past some bound surface as un-ready in the mesh's own view, not only in its container log? +- Does the same gap withdraw provisions late — is a provider told a cross-node consumer left only + when the provider is next pushed?