Observed in the two-node bed: a cross-node runtime exits on 'timed out fetching the broker's certificate' until the tunnel forms, then settles. The retry belongs in-process. https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
1.7 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design |
|---|---|---|---|---|
| open | 2026-09-17 |
059 — A provisioner runtime crash-loops until the overlay is up
Symptom
A module runtime whose broker is on another node exits at startup — timed out fetching the broker's certificate — and is restarted by the container runtime until the overlay tunnel is
established, after which it connects and works. Observed on a joined node during the two-node
database bed: the runtime's container restarted several times in its first minute, then settled and
provisioned correctly.
Why it matters beyond the instance
The runtime treats "the broker is not reachable yet" as fatal, but at startup it is the normal
case: a container comes up in seconds, the overlay a moment later. The consequence is not just
noise — every restart-counting health check reads the churn as a crash-loop, and a person watching
docker ps sees a failing module that is actually fine. The lab's beds had to learn to tell
"churn that stops" from "a real crash-loop" (a settle-based wait); every operator will face the
same reading.
A process whose dependency may appear seconds later should wait for it, with backoff, in-process — exiting delegates the retry to the container runtime and turns an ordinary ordering into a visible failure.
Open questions
- Should the shared runtime bootstrap (the code every
mesh-runtime-<module>starts with) retry the broker connection with backoff instead of exiting, and for how long before it is fatal? - Is the certificate fetch the only startup step with this shape, or do provision waits and grant reads exit the same way?
- How is this checked once fixed — a bed asserting restart count stays 0 on a joined node?