Files
hq/04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md
T
jschoubben 82c03a8d10 057 and 058 resolved — their fixes merged and proven
The P1-sweep PR left both at 'located'; the fixes have since merged
(mesh-controller #31, mesh-tools #10) and the bed proves them. Flip to
resolved with their fixed-by.
2026-09-20 12:55:00 +02:00

1.8 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
resolved 2026-09-17
mesh-tools
mesh-tools PR 10 (0ea3db3); proven by the built-store-cross-node bed (restart count 0)

058 — A provisioner runtime crash-loops until the overlay is up

Symptom

A module runtime whose broker is on another node exits at startup — timed out fetching the broker's certificate — and is restarted by the container runtime until the overlay tunnel is established, after which it connects and works. Observed on a joined node during the two-node database bed: the runtime's container restarted several times in its first minute, then settled and provisioned correctly.

Why it matters beyond the instance

The runtime treats "the broker is not reachable yet" as fatal, but at startup it is the normal case: a container comes up in seconds, the overlay a moment later. The consequence is not just noise — every restart-counting health check reads the churn as a crash-loop, and a person watching docker ps sees a failing module that is actually fine. The lab's beds had to learn to tell "churn that stops" from "a real crash-loop" (a settle-based wait); every operator will face the same reading.

A process whose dependency may appear seconds later should wait for it, with backoff, in-process — exiting delegates the retry to the container runtime and turns an ordinary ordering into a visible failure.

Open questions

  • Should the shared runtime bootstrap (the code every mesh-runtime-<module> starts with) retry the broker connection with backoff instead of exiting, and for how long before it is fatal?
  • Is the certificate fetch the only startup step with this shape, or do provision waits and grant reads exit the same way?
  • How is this checked once fixed — a bed asserting restart count stays 0 on a joined node?