ADR 0083 proposed; 057/058 diagnosed and located
One push leaves the mesh consistent (the 057 decision, proposed for acceptance); the shared runtime waits for its broker (058). Fixes on mesh-control fix/one-push-is-enough and mesh-tools fix/the-runtime-waits-for-its-broker; the built-store-cross-node bed enforces both.
This commit is contained in:
@@ -0,0 +1,52 @@
|
||||
---
|
||||
topic: the mesh
|
||||
status: proposed
|
||||
date: 2026-09-18
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 0010-delivery.md
|
||||
---
|
||||
|
||||
# 83. One push leaves the mesh consistent
|
||||
|
||||
## Context
|
||||
|
||||
A provision is minted while composing the *consumer's* node; the provider's grant list is a pure
|
||||
read of secrets already issued from it. So assigning a cross-node consumer and pushing its node
|
||||
produced a consumer that retried forever against a provider that had never heard of it, until the
|
||||
provider's node was pushed a second time — an action with no signal to take, documented nowhere,
|
||||
and invisible whenever consumer and provider share a machine
|
||||
([issue 057](../04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md)).
|
||||
|
||||
Two remedies were on the table: **cascade** — a push also delivers to the machines its compose
|
||||
changed — or **report** — a push says "now push the provider" and leaves the act to the operator.
|
||||
|
||||
## Decision
|
||||
|
||||
A push finishes what it starts: after composing and sending the named node, the controller
|
||||
recomputes what every machine should be, and any machine whose declaration changed *because of
|
||||
this push* is sent its declaration too — by name, in the push's own output, converging over a
|
||||
bounded number of rounds (a cascaded send may itself mint).
|
||||
|
||||
"Changed because of this push" is a comparison, not a guess: the digest of what each machine
|
||||
should be is captured before the named compose and recomputed after. Machines that were already
|
||||
behind for unrelated reasons are not swept in — that remains `push --behind`, the explicit
|
||||
whole-mesh reconcile.
|
||||
|
||||
Reporting alone was rejected because it converts a derived fact the controller already holds into
|
||||
an operator obligation, and an obligation enforced by nothing is issue 057 restated. The
|
||||
declaration is computed from the whole mesh; delivering a mesh that is knowingly inconsistent and
|
||||
merely saying so would make "push succeeded" mean less than it says.
|
||||
|
||||
## Consequences
|
||||
|
||||
- One push is sufficient for a cross-node consumer: the provider's grants arrive from the same
|
||||
act that minted the provision. The undocumented rule "push the provider node too" ceases to
|
||||
exist rather than becoming documentation.
|
||||
- A named push may deliver to machines the operator did not name. This is bounded to machines
|
||||
whose declarations this push changed, and every one is named in the output — never silent.
|
||||
- The blast radius question from the issue is answered by the comparison: nothing is recomposed
|
||||
into delivery except what the named compose provably changed.
|
||||
- How this is checked: the built-store-cross-node bed registers a cross-node consumer, pushes
|
||||
only the consumer's node, and asserts the provider minted its vhost — the workaround push is
|
||||
removed, so a regression fails the bed.
|
||||
@@ -85,6 +85,7 @@ python3 00-META/checks/index.py fail if stale
|
||||
- **0002** — [Nodes communicate over a message broker, not over HTTP](0002-nodes-communicate-over-a-broker.md)
|
||||
- **0003** — [An agent is a persistent employee, not an instance of a pool](0003-agents-are-persistent-employees.md)
|
||||
- **0077** — [The parts are named controller, foundation, node — not control plane, substrate, master](0077-the-controller-and-the-foundation.md)
|
||||
- **0083** — [One push leaves the mesh consistent](0083-one-push-leaves-the-mesh-consistent.md) *(proposed)*
|
||||
|
||||
### Its tiers, from the bottom up
|
||||
|
||||
|
||||
+4
-3
@@ -1,9 +1,10 @@
|
||||
---
|
||||
status: open
|
||||
status: located
|
||||
opened: 2026-09-17
|
||||
located-in: []
|
||||
located-in:
|
||||
- mesh-controller
|
||||
fixed-by:
|
||||
amended-design:
|
||||
amended-design: 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
|
||||
---
|
||||
|
||||
# 057 — A cross-node consumer is provisioned only when the provider is pushed again
|
||||
|
||||
+16
@@ -0,0 +1,16 @@
|
||||
# Diagnosis — 2026-09-18
|
||||
|
||||
The mechanism was already understood when the issue was opened: a provision secret is minted as a
|
||||
side-effect of composing the consumer's node, and the provider's grant list is a pure read of
|
||||
secrets already issued — so a provider composed before the consumer existed, and never again, is
|
||||
blind to it. The open question was the remedy: cascade the push, or report the obligation.
|
||||
|
||||
Decided as [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md): a push
|
||||
finishes what it starts. The controller captures the digest of what every machine should be
|
||||
before composing the named node, recomputes after, and delivers to exactly the machines whose
|
||||
declaration changed because of this push — named in the output, bounded rounds, converging.
|
||||
Machines behind for unrelated reasons stay the business of `push --behind`.
|
||||
|
||||
**Located in:** mesh-controller (the push command). Checked by the built-store-cross-node bed,
|
||||
which now pushes only the consumer's node and asserts the provider minted the vhost — the
|
||||
workaround push is removed, so a regression fails the bed.
|
||||
+3
-2
@@ -1,7 +1,8 @@
|
||||
---
|
||||
status: open
|
||||
status: located
|
||||
opened: 2026-09-17
|
||||
located-in: []
|
||||
located-in:
|
||||
- mesh-tools
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
+17
@@ -0,0 +1,17 @@
|
||||
# Diagnosis — 2026-09-18
|
||||
|
||||
1. The exit is the shared runtime's: serve mode connected the broker once and treated any failure
|
||||
as fatal, so "the overlay tunnel came up a moment after the container" — the normal case on a
|
||||
joined node — became an exit, a container-runtime restart, and a visible crash-loop.
|
||||
2. The failing step is the certificate pre-fetch that pins the broker (a 15s TLS probe), but the
|
||||
shape is general: any reachability failure at startup had the same consequence.
|
||||
3. Answering the report's open questions: the shared runtime now retries the broker connection
|
||||
in-process with capped backoff (2s doubling to 30s), aloud, indefinitely — "how long before it
|
||||
is fatal" is answered *never*, because the failure modes that do not heal are not reachability:
|
||||
a pinned-certificate mismatch still refuses immediately (an impostor does not become the
|
||||
broker by being asked again), and configuration errors still exit at once. One-shot commands
|
||||
(emit, invoke, run) still fail fast — a person is waiting.
|
||||
4. Checked by the built-store-cross-node bed: the joined consumer must report a container-runtime
|
||||
restart count of zero.
|
||||
|
||||
**Located in:** mesh-tools (the shared runtime's serve mode).
|
||||
Reference in New Issue
Block a user