The P1-sweep PR left both at 'located'; the fixes have since merged (mesh-controller #31, mesh-tools #10) and the bed proves them. Flip to resolved with their fixed-by.
48 lines
2.7 KiB
Markdown
48 lines
2.7 KiB
Markdown
---
|
|
status: resolved
|
|
opened: 2026-09-17
|
|
located-in:
|
|
- mesh-controller
|
|
fixed-by: mesh-controller PR 31 (5dd2432); ADR 0083; proven by the built-store-cross-node bed
|
|
amended-design: 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
|
|
---
|
|
|
|
# 057 — A cross-node consumer is provisioned only when the provider is pushed again
|
|
|
|
## Symptom
|
|
|
|
Adding a module that requires a mesh-scoped provision (a broker vhost, a database) on one node,
|
|
and pushing that node, is not enough for the provision to be made. The consumer starts, cannot
|
|
connect, and retries forever, while nothing says why. The provision appears only after the
|
|
*provider's* node is pushed a second time — an action the operator has no signal to take.
|
|
|
|
Observed while proving [issue 055](../055-the-adopted-store-and-broker-may-be-reachable-on-one-node-only/00-report.md).
|
|
A consumer was assigned to a joined node and that node was pushed; the consumer came up but
|
|
logged `amqp connection closed before the round-trip completed` on a loop. The provider's grant
|
|
file listed no consumers (`given: []`), so its provisioner minted nothing. Pushing the provider's
|
|
node again populated the grant, the vhost was minted, and the consumer connected on its next
|
|
retry. No error was raised at any point — the only symptom was a consumer that never became ready.
|
|
|
|
## Why it matters beyond the instance
|
|
|
|
A provision secret is minted as a side-effect of composing the **consumer's** plan; the provider's
|
|
grant list is a pure read of secrets already issued from it. So composing the provider *before* a
|
|
new cross-node consumer exists — or never recomposing it — leaves the provider blind to that
|
|
consumer. Co-located consumer and provider hide this: one push composes both. The failure is
|
|
therefore invisible until the mesh is actually multi-node, which is exactly when it is hardest to
|
|
diagnose, and it fails silently — the one shape [AGENTS.md](../../AGENTS.md) says belongs here.
|
|
|
|
The rule "push the provider node too" is real and undocumented, and a rule enforced by nothing is
|
|
indistinguishable from a wrong one.
|
|
|
|
## Open questions
|
|
|
|
- Should composing any node that adds or removes a cross-node consumer also recompose the
|
|
providers it now binds to, so a single push is sufficient? What is the blast radius of that?
|
|
- Failing that, should composing a consumer whose provider is on another node *report* that the
|
|
provider must be pushed, rather than issue a grant no one will read?
|
|
- A consumer that cannot reach its provision retries silently. Should a consumer that has waited
|
|
past some bound surface as un-ready in the mesh's own view, not only in its container log?
|
|
- Does the same gap withdraw provisions late — is a provider told a cross-node consumer left only
|
|
when the provider is next pushed?
|