One push leaves the mesh consistent (the 057 decision, proposed for acceptance); the shared runtime waits for its broker (058). Fixes on mesh-control fix/one-push-is-enough and mesh-tools fix/the-runtime-waits-for-its-broker; the built-store-cross-node bed enforces both.
1.2 KiB
1.2 KiB
Diagnosis — 2026-09-18
- The exit is the shared runtime's: serve mode connected the broker once and treated any failure as fatal, so "the overlay tunnel came up a moment after the container" — the normal case on a joined node — became an exit, a container-runtime restart, and a visible crash-loop.
- The failing step is the certificate pre-fetch that pins the broker (a 15s TLS probe), but the shape is general: any reachability failure at startup had the same consequence.
- Answering the report's open questions: the shared runtime now retries the broker connection in-process with capped backoff (2s doubling to 30s), aloud, indefinitely — "how long before it is fatal" is answered never, because the failure modes that do not heal are not reachability: a pinned-certificate mismatch still refuses immediately (an impostor does not become the broker by being asked again), and configuration errors still exit at once. One-shot commands (emit, invoke, run) still fail fast — a person is waiting.
- Checked by the built-store-cross-node bed: the joined consumer must report a container-runtime restart count of zero.
Located in: mesh-tools (the shared runtime's serve mode).