Issue 055 resolved and proven; 056 and 057 opened #45
+67
-3
@@ -1,8 +1,11 @@
|
||||
---
|
||||
status: open
|
||||
status: resolved
|
||||
opened: 2026-09-16
|
||||
located-in: []
|
||||
fixed-by:
|
||||
located-in:
|
||||
- mesh-controller/cmd/mesh-controller/modules.go
|
||||
- mesh-controller/cmd/mesh-controller/build.go
|
||||
- mesh-catalog/modules/lavinmq/module.json
|
||||
fixed-by: mesh-controller multi-node/broker-reaches-over-overlay; mesh-catalog fix/broker-declares-amqps-port
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -38,3 +41,64 @@ means the one server has to be the mesh-wide one, reachable across the overlay.
|
||||
machine, and is the packet filter's `from: mesh` rule enough to let it through?
|
||||
- What is the smallest multi-node bed that would prove a database granted to a module on a
|
||||
joined machine can be opened from there?
|
||||
|
||||
## Diagnosis (2026-09-17)
|
||||
|
||||
A focused two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`, scenario
|
||||
`adopted-store-cross-node.yml`) puts the adopted broker on `anchor` and runs `amqp-ping` on the
|
||||
joined `node2`. It fails, and the cause is two-layered — the first fixed, the second the real one:
|
||||
|
||||
1. **Firewall (fixed).** The adopted `lavinmq` module declared only `5672` in `listens`, not the
|
||||
amqps bus port `5671`. mesh-broker's ports are published (DNAT'd → the nftables *forward* chain),
|
||||
so with no `5671` forward rule the bus was dropped cross-node. Fix: the module declares `5671`
|
||||
from `mesh` too (mesh-catalog `fix/broker-declares-amqps-port`). Confirmed: after it, `node2 ->
|
||||
10.42.0.1:5671` (overlay) is OPEN and the forward chain carries `proto-dst 5671`.
|
||||
|
||||
2. **The broker address is the genesis public one, not the overlay (the real bug).** The broker
|
||||
secret handed to a module is built at `cmd/mesh-controller/modules.go:294` (and the builder's at
|
||||
`build.go:231`) as `amqps://<acct>:<pw>@<known.Address>/`, where `known.Address` is
|
||||
`MESH_BROKER_ADDRESS` — the genesis PUBLIC endpoint (e.g. `192.0.2.10:5671`). `amqp-ping` on
|
||||
node2 therefore dials `192.0.2.10:5671`, which the firewall does NOT admit cross-node (it admits
|
||||
the overlay `10.42.0.1:5671`), and times out "fetching the broker's certificate." Its grant
|
||||
never completes, so the provisioner's receives file shows `given: []` and no vhost is minted —
|
||||
the downstream symptom.
|
||||
|
||||
The consumer already resolves the provider's `.internal` overlay name for its *provision* binding
|
||||
(`bound.at == anchor.internal`) — that half works. What does not is the *bus account* URL, which
|
||||
uses the public address for every module regardless of node.
|
||||
|
||||
## Open questions
|
||||
|
||||
- The broker URL for a module on a joined node should name the control-node's overlay address
|
||||
(`<control-node>.internal:5671`), which resolves (via `--add-host`) and the firewall admits.
|
||||
- But genesis-time accounts (the builder, issued on the control-node before the overlay exists)
|
||||
must keep a locally-reachable address, or genesis breaks. So the fix is conditional on the
|
||||
target node being on the overlay, not a blanket switch — where is that known at `module issue`?
|
||||
- Should the control-node's own modules also move to the overlay name for uniformity, or keep the
|
||||
local address they already reach?
|
||||
|
||||
## Resolution (2026-09-17)
|
||||
|
||||
A two-node bed (`mesh-lab test/integration/adopted-store-cross-node.test.ts`) proves it: a
|
||||
consumer (`amqp-ping`) on a joined node reaches the adopted broker (`mesh-broker`) on the
|
||||
control-node over the overlay, its binding names `anchor.internal`, its vhost is minted, and it
|
||||
holds the connection. Three things were wrong, now fixed:
|
||||
|
||||
1. **The broker's amqps port was not in the firewall.** The `lavinmq` module declared only `5672`
|
||||
in `listens`; the bus is `5671`. Published container ports are matched in the nftables *forward*
|
||||
chain, so `5671` was dropped cross-node. The module now declares `5671` from `mesh`
|
||||
(mesh-catalog `fix/broker-declares-amqps-port`).
|
||||
|
||||
2. **The broker credential named the public address.** `module issue` and `builder issue` built
|
||||
the URL with `MESH_BROKER_ADDRESS` — the broker's genesis public endpoint, which a joined
|
||||
node's firewall does not admit. A new `brokerReachableAt` returns the broker as the target node
|
||||
can reach it: a node on the overlay gets the hub's `.internal` name (fingerprint pinning makes
|
||||
the host swap safe for TLS); a node not yet on the overlay keeps the genesis address, so
|
||||
bring-up is unchanged (mesh-controller `multi-node/broker-reaches-over-overlay`).
|
||||
|
||||
3. **The provider was not recomposed after the remote consumer arrived (operational).** A
|
||||
provision secret is minted as a side-effect of composing the CONSUMER's plan; the provider's
|
||||
grant list is a pure read of secrets already issued from it. So the provider's provisioner
|
||||
learns of a cross-node consumer only when the provider node is composed again. Pushing the
|
||||
provider node after adding the remote consumer mints the vhost. This is not a code fix — it is
|
||||
an ordering the operator must follow, and its silent-failure edge is opened as issue 057.
|
||||
|
||||
+42
@@ -0,0 +1,42 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-17
|
||||
located-in: [mesh-controller, mesh-host]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 056 — An adopted module assigned to a second node raises a second server
|
||||
|
||||
## Symptom
|
||||
|
||||
The store and broker are adopted in place ([ADR 0078](../../02-DECISIONS/0078-the-store-and-broker-are-modules.md),
|
||||
issue 051): genesis raises `mesh-store`/`mesh-broker` on the control-node, and assigning the
|
||||
`postgres`/`lavinmq` module there makes the module reconcile the already-running container instead
|
||||
of raising a new one. Adoption works *because the container is already running on that node*.
|
||||
|
||||
Nothing stops the same module being assigned to a **second** node. On a node where no
|
||||
`mesh-store` is running, the applier finds no container by that name and raises one — a second
|
||||
postgres, a second broker. The whole point of Phase 3 — "one postgres, one lavinmq" — holds only
|
||||
by the convention that genesis assigns these modules to the control-node alone; it is not enforced.
|
||||
|
||||
## Why it matters
|
||||
|
||||
"There is one store and one broker" is stated as a property of the mesh, and a property enforced by
|
||||
nothing is indistinguishable from a wrong one. A single extra `assign` — the ordinary verb an
|
||||
operator types — silently produces a divergent second server holding none of the first's data, and
|
||||
consumers resolved to it get an empty store. The failure is silent and far from its cause.
|
||||
|
||||
The exclusive foundation modules are the mesh's clearest case of "there is one of me," yet unlike a
|
||||
mesh-scoped seat, an adopted module carries no claim that the resolver would refuse a second holder
|
||||
for.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should `postgres`/`lavinmq` claim a mesh-scoped exclusive seat (the way `mesh-controller` claims
|
||||
`the-controller`), so the resolver refuses a second assignment by the same mechanism that keeps
|
||||
one controller?
|
||||
- Or should adoption be explicit — a module that adopts a foundation container declares it, and the
|
||||
controller refuses to place it on a node whose foundation did not raise that container?
|
||||
- How is "one postgres, one lavinmq" checked, rather than assumed — in the resolver, in `status`, or
|
||||
in a bed that tries the second assignment and asserts the refusal?
|
||||
+46
@@ -0,0 +1,46 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-17
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 057 — A cross-node consumer is provisioned only when the provider is pushed again
|
||||
|
||||
## Symptom
|
||||
|
||||
Adding a module that requires a mesh-scoped provision (a broker vhost, a database) on one node,
|
||||
and pushing that node, is not enough for the provision to be made. The consumer starts, cannot
|
||||
connect, and retries forever, while nothing says why. The provision appears only after the
|
||||
*provider's* node is pushed a second time — an action the operator has no signal to take.
|
||||
|
||||
Observed while proving [issue 055](../055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md).
|
||||
A consumer was assigned to a joined node and that node was pushed; the consumer came up but
|
||||
logged `amqp connection closed before the round-trip completed` on a loop. The provider's grant
|
||||
file listed no consumers (`given: []`), so its provisioner minted nothing. Pushing the provider's
|
||||
node again populated the grant, the vhost was minted, and the consumer connected on its next
|
||||
retry. No error was raised at any point — the only symptom was a consumer that never became ready.
|
||||
|
||||
## Why it matters beyond the instance
|
||||
|
||||
A provision secret is minted as a side-effect of composing the **consumer's** plan; the provider's
|
||||
grant list is a pure read of secrets already issued from it. So composing the provider *before* a
|
||||
new cross-node consumer exists — or never recomposing it — leaves the provider blind to that
|
||||
consumer. Co-located consumer and provider hide this: one push composes both. The failure is
|
||||
therefore invisible until the mesh is actually multi-node, which is exactly when it is hardest to
|
||||
diagnose, and it fails silently — the one shape [AGENTS.md](../../AGENTS.md) says belongs here.
|
||||
|
||||
The rule "push the provider node too" is real and undocumented, and a rule enforced by nothing is
|
||||
indistinguishable from a wrong one.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should composing any node that adds or removes a cross-node consumer also recompose the
|
||||
providers it now binds to, so a single push is sufficient? What is the blast radius of that?
|
||||
- Failing that, should composing a consumer whose provider is on another node *report* that the
|
||||
provider must be pushed, rather than issue a grant no one will read?
|
||||
- A consumer that cannot reach its provision retries silently. Should a consumer that has waited
|
||||
past some bound surface as un-ready in the mesh's own view, not only in its container log?
|
||||
- Does the same gap withdraw provisions late — is a provider told a cross-node consumer left only
|
||||
when the provider is next pushed?
|
||||
Reference in New Issue
Block a user