diff --git a/04-ISSUES/059-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md b/04-ISSUES/059-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md new file mode 100644 index 0000000..1adf39d --- /dev/null +++ b/04-ISSUES/059-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md @@ -0,0 +1,38 @@ +--- +status: open +opened: 2026-09-17 +located-in: [] +fixed-by: +amended-design: +--- + +# 059 — A provisioner runtime crash-loops until the overlay is up + +## Symptom + +A module runtime whose broker is on another node exits at startup — `timed out fetching the +broker's certificate` — and is restarted by the container runtime until the overlay tunnel is +established, after which it connects and works. Observed on a joined node during the two-node +database bed: the runtime's container restarted several times in its first minute, then settled and +provisioned correctly. + +## Why it matters beyond the instance + +The runtime treats "the broker is not reachable yet" as fatal, but at startup it is the *normal* +case: a container comes up in seconds, the overlay a moment later. The consequence is not just +noise — every restart-counting health check reads the churn as a crash-loop, and a person watching +`docker ps` sees a failing module that is actually fine. The lab's beds had to learn to tell +"churn that stops" from "a real crash-loop" (a settle-based wait); every operator will face the +same reading. + +A process whose dependency may appear seconds later should wait for it, with backoff, in-process — +exiting delegates the retry to the container runtime and turns an ordinary ordering into a visible +failure. + +## Open questions + +- Should the shared runtime bootstrap (the code every `mesh-runtime-` starts with) retry + the broker connection with backoff instead of exiting, and for how long before it *is* fatal? +- Is the certificate fetch the only startup step with this shape, or do provision waits and grant + reads exit the same way? +- How is this checked once fixed — a bed asserting restart count stays 0 on a joined node?