ADR 0083 and issues 057/058/060/063/064 — the P1 sweep #54

Merged
jschoubben merged 5 commits from issue/057-058-one-push-patient-runtime into main 2026-09-20 10:53:25 +00:00
6 changed files with 93 additions and 5 deletions
Showing only changes of commit 9fd7e6c458 - Show all commits
@@ -0,0 +1,52 @@
---
topic: the mesh
status: proposed
date: 2026-09-18
deciders: jochen
reconstructed: false
extends: 0010-delivery.md
---
# 83. One push leaves the mesh consistent
## Context
A provision is minted while composing the *consumer's* node; the provider's grant list is a pure
read of secrets already issued from it. So assigning a cross-node consumer and pushing its node
produced a consumer that retried forever against a provider that had never heard of it, until the
provider's node was pushed a second time — an action with no signal to take, documented nowhere,
and invisible whenever consumer and provider share a machine
([issue 057](../04-ISSUES/057-a-cross-node-consumer-is-provisioned-only-when-the-provider-is-pushed-again/00-report.md)).
Two remedies were on the table: **cascade** — a push also delivers to the machines its compose
changed — or **report** — a push says "now push the provider" and leaves the act to the operator.
## Decision
A push finishes what it starts: after composing and sending the named node, the controller
recomputes what every machine should be, and any machine whose declaration changed *because of
this push* is sent its declaration too — by name, in the push's own output, converging over a
bounded number of rounds (a cascaded send may itself mint).
"Changed because of this push" is a comparison, not a guess: the digest of what each machine
should be is captured before the named compose and recomputed after. Machines that were already
behind for unrelated reasons are not swept in — that remains `push --behind`, the explicit
whole-mesh reconcile.
Reporting alone was rejected because it converts a derived fact the controller already holds into
an operator obligation, and an obligation enforced by nothing is issue 057 restated. The
declaration is computed from the whole mesh; delivering a mesh that is knowingly inconsistent and
merely saying so would make "push succeeded" mean less than it says.
## Consequences
- One push is sufficient for a cross-node consumer: the provider's grants arrive from the same
act that minted the provision. The undocumented rule "push the provider node too" ceases to
exist rather than becoming documentation.
- A named push may deliver to machines the operator did not name. This is bounded to machines
whose declarations this push changed, and every one is named in the output — never silent.
- The blast radius question from the issue is answered by the comparison: nothing is recomposed
into delivery except what the named compose provably changed.
- How this is checked: the built-store-cross-node bed registers a cross-node consumer, pushes
only the consumer's node, and asserts the provider minted its vhost — the workaround push is
removed, so a regression fails the bed.
+1
View File
@@ -85,6 +85,7 @@ python3 00-META/checks/index.py fail if stale
- **0002** — [Nodes communicate over a message broker, not over HTTP](0002-nodes-communicate-over-a-broker.md)
- **0003** — [An agent is a persistent employee, not an instance of a pool](0003-agents-are-persistent-employees.md)
- **0077** — [The parts are named controller, foundation, node — not control plane, substrate, master](0077-the-controller-and-the-foundation.md)
- **0083** — [One push leaves the mesh consistent](0083-one-push-leaves-the-mesh-consistent.md) *(proposed)*
### Its tiers, from the bottom up
@@ -1,9 +1,10 @@
---
status: open
status: located
opened: 2026-09-17
located-in: []
located-in:
- mesh-controller
fixed-by:
amended-design:
amended-design: 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
---
# 057 — A cross-node consumer is provisioned only when the provider is pushed again
@@ -0,0 +1,16 @@
# Diagnosis — 2026-09-18
The mechanism was already understood when the issue was opened: a provision secret is minted as a
side-effect of composing the consumer's node, and the provider's grant list is a pure read of
secrets already issued — so a provider composed before the consumer existed, and never again, is
blind to it. The open question was the remedy: cascade the push, or report the obligation.
Decided as [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md): a push
finishes what it starts. The controller captures the digest of what every machine should be
before composing the named node, recomputes after, and delivers to exactly the machines whose
declaration changed because of this push — named in the output, bounded rounds, converging.
Machines behind for unrelated reasons stay the business of `push --behind`.
**Located in:** mesh-controller (the push command). Checked by the built-store-cross-node bed,
which now pushes only the consumer's node and asserts the provider minted the vhost — the
workaround push is removed, so a regression fails the bed.
@@ -1,7 +1,8 @@
---
status: open
status: located
opened: 2026-09-17
located-in: []
located-in:
- mesh-tools
fixed-by:
amended-design:
---
@@ -0,0 +1,17 @@
# Diagnosis — 2026-09-18
1. The exit is the shared runtime's: serve mode connected the broker once and treated any failure
as fatal, so "the overlay tunnel came up a moment after the container" — the normal case on a
joined node — became an exit, a container-runtime restart, and a visible crash-loop.
2. The failing step is the certificate pre-fetch that pins the broker (a 15s TLS probe), but the
shape is general: any reachability failure at startup had the same consequence.
3. Answering the report's open questions: the shared runtime now retries the broker connection
in-process with capped backoff (2s doubling to 30s), aloud, indefinitely — "how long before it
is fatal" is answered *never*, because the failure modes that do not heal are not reachability:
a pinned-certificate mismatch still refuses immediately (an impostor does not become the
broker by being asked again), and configuration errors still exit at once. One-shot commands
(emit, invoke, run) still fail fast — a person is waiting.
4. Checked by the built-store-cross-node bed: the joined consumer must report a container-runtime
restart count of zero.
**Located in:** mesh-tools (the shared runtime's serve mode).