ADR 0083 proposed; 057/058 diagnosed and located

One push leaves the mesh consistent (the 057 decision, proposed for
acceptance); the shared runtime waits for its broker (058). Fixes on
mesh-control fix/one-push-is-enough and mesh-tools
fix/the-runtime-waits-for-its-broker; the built-store-cross-node bed
enforces both.
This commit is contained in:
2026-09-18 02:10:01 +02:00
parent 5703c86cf6
commit 9fd7e6c458
6 changed files with 93 additions and 5 deletions
@@ -1,9 +1,10 @@
---
status: open
status: located
opened: 2026-09-17
located-in: []
located-in:
- mesh-controller
fixed-by:
amended-design:
amended-design: 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
---
# 057 — A cross-node consumer is provisioned only when the provider is pushed again
@@ -0,0 +1,16 @@
# Diagnosis — 2026-09-18
The mechanism was already understood when the issue was opened: a provision secret is minted as a
side-effect of composing the consumer's node, and the provider's grant list is a pure read of
secrets already issued — so a provider composed before the consumer existed, and never again, is
blind to it. The open question was the remedy: cascade the push, or report the obligation.
Decided as [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md): a push
finishes what it starts. The controller captures the digest of what every machine should be
before composing the named node, recomputes after, and delivers to exactly the machines whose
declaration changed because of this push — named in the output, bounded rounds, converging.
Machines behind for unrelated reasons stay the business of `push --behind`.
**Located in:** mesh-controller (the push command). Checked by the built-store-cross-node bed,
which now pushes only the consumer's node and asserts the provider minted the vhost — the
workaround push is removed, so a regression fails the bed.
@@ -1,7 +1,8 @@
---
status: open
status: located
opened: 2026-09-17
located-in: []
located-in:
- mesh-tools
fixed-by:
amended-design:
---
@@ -0,0 +1,17 @@
# Diagnosis — 2026-09-18
1. The exit is the shared runtime's: serve mode connected the broker once and treated any failure
as fatal, so "the overlay tunnel came up a moment after the container" — the normal case on a
joined node — became an exit, a container-runtime restart, and a visible crash-loop.
2. The failing step is the certificate pre-fetch that pins the broker (a 15s TLS probe), but the
shape is general: any reachability failure at startup had the same consequence.
3. Answering the report's open questions: the shared runtime now retries the broker connection
in-process with capped backoff (2s doubling to 30s), aloud, indefinitely — "how long before it
is fatal" is answered *never*, because the failure modes that do not heal are not reachability:
a pinned-certificate mismatch still refuses immediately (an impostor does not become the
broker by being asked again), and configuration errors still exit at once. One-shot commands
(emit, invoke, run) still fail fast — a person is waiting.
4. Checked by the built-store-cross-node bed: the joined consumer must report a container-runtime
restart count of zero.
**Located in:** mesh-tools (the shared runtime's serve mode).