ADR 0083 proposed; 057/058 diagnosed and located
One push leaves the mesh consistent (the 057 decision, proposed for acceptance); the shared runtime waits for its broker (058). Fixes on mesh-control fix/one-push-is-enough and mesh-tools fix/the-runtime-waits-for-its-broker; the built-store-cross-node bed enforces both.
This commit is contained in:
+4
-3
@@ -1,9 +1,10 @@
|
||||
---
|
||||
status: open
|
||||
status: located
|
||||
opened: 2026-09-17
|
||||
located-in: []
|
||||
located-in:
|
||||
- mesh-controller
|
||||
fixed-by:
|
||||
amended-design:
|
||||
amended-design: 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
|
||||
---
|
||||
|
||||
# 057 — A cross-node consumer is provisioned only when the provider is pushed again
|
||||
|
||||
+16
@@ -0,0 +1,16 @@
|
||||
# Diagnosis — 2026-09-18
|
||||
|
||||
The mechanism was already understood when the issue was opened: a provision secret is minted as a
|
||||
side-effect of composing the consumer's node, and the provider's grant list is a pure read of
|
||||
secrets already issued — so a provider composed before the consumer existed, and never again, is
|
||||
blind to it. The open question was the remedy: cascade the push, or report the obligation.
|
||||
|
||||
Decided as [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md): a push
|
||||
finishes what it starts. The controller captures the digest of what every machine should be
|
||||
before composing the named node, recomputes after, and delivers to exactly the machines whose
|
||||
declaration changed because of this push — named in the output, bounded rounds, converging.
|
||||
Machines behind for unrelated reasons stay the business of `push --behind`.
|
||||
|
||||
**Located in:** mesh-controller (the push command). Checked by the built-store-cross-node bed,
|
||||
which now pushes only the consumer's node and asserts the provider minted the vhost — the
|
||||
workaround push is removed, so a regression fails the bed.
|
||||
+3
-2
@@ -1,7 +1,8 @@
|
||||
---
|
||||
status: open
|
||||
status: located
|
||||
opened: 2026-09-17
|
||||
located-in: []
|
||||
located-in:
|
||||
- mesh-tools
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
+17
@@ -0,0 +1,17 @@
|
||||
# Diagnosis — 2026-09-18
|
||||
|
||||
1. The exit is the shared runtime's: serve mode connected the broker once and treated any failure
|
||||
as fatal, so "the overlay tunnel came up a moment after the container" — the normal case on a
|
||||
joined node — became an exit, a container-runtime restart, and a visible crash-loop.
|
||||
2. The failing step is the certificate pre-fetch that pins the broker (a 15s TLS probe), but the
|
||||
shape is general: any reachability failure at startup had the same consequence.
|
||||
3. Answering the report's open questions: the shared runtime now retries the broker connection
|
||||
in-process with capped backoff (2s doubling to 30s), aloud, indefinitely — "how long before it
|
||||
is fatal" is answered *never*, because the failure modes that do not heal are not reachability:
|
||||
a pinned-certificate mismatch still refuses immediately (an impostor does not become the
|
||||
broker by being asked again), and configuration errors still exit at once. One-shot commands
|
||||
(emit, invoke, run) still fail fast — a person is waiting.
|
||||
4. Checked by the built-store-cross-node bed: the joined consumer must report a container-runtime
|
||||
restart count of zero.
|
||||
|
||||
**Located in:** mesh-tools (the shared runtime's serve mode).
|
||||
Reference in New Issue
Block a user