--- status: resolved opened: 2026-09-17 located-in: - mesh-tools fixed-by: mesh-tools PR 10 (0ea3db3); proven by the built-store-cross-node bed (restart count 0) amended-design: --- # 058 — A provisioner runtime crash-loops until the overlay is up ## Symptom A module runtime whose broker is on another node exits at startup — `timed out fetching the broker's certificate` — and is restarted by the container runtime until the overlay tunnel is established, after which it connects and works. Observed on a joined node during the two-node database bed: the runtime's container restarted several times in its first minute, then settled and provisioned correctly. ## Why it matters beyond the instance The runtime treats "the broker is not reachable yet" as fatal, but at startup it is the *normal* case: a container comes up in seconds, the overlay a moment later. The consequence is not just noise — every restart-counting health check reads the churn as a crash-loop, and a person watching `docker ps` sees a failing module that is actually fine. The lab's beds had to learn to tell "churn that stops" from "a real crash-loop" (a settle-based wait); every operator will face the same reading. A process whose dependency may appear seconds later should wait for it, with backoff, in-process — exiting delegates the retry to the container runtime and turns an ordinary ordering into a visible failure. ## Open questions - Should the shared runtime bootstrap (the code every `mesh-runtime-` starts with) retry the broker connection with backoff instead of exiting, and for how long before it *is* fatal? - Is the certificate fetch the only startup step with this shape, or do provision waits and grant reads exit the same way? - How is this checked once fixed — a bed asserting restart count stays 0 on a joined node?