Issue 059 — a provisioner runtime crash-loops until the overlay is up

Observed in the two-node bed: a cross-node runtime exits on 'timed out fetching the
broker's certificate' until the tunnel forms, then settles. The retry belongs in-process.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
This commit is contained in:
2026-09-17 21:56:40 +02:00
parent 0e9d4ab86a
commit dd5bf9cd8c
@@ -0,0 +1,38 @@
---
status: open
opened: 2026-09-17
located-in: []
fixed-by:
amended-design:
---
# 059 — A provisioner runtime crash-loops until the overlay is up
## Symptom
A module runtime whose broker is on another node exits at startup — `timed out fetching the
broker's certificate` — and is restarted by the container runtime until the overlay tunnel is
established, after which it connects and works. Observed on a joined node during the two-node
database bed: the runtime's container restarted several times in its first minute, then settled and
provisioned correctly.
## Why it matters beyond the instance
The runtime treats "the broker is not reachable yet" as fatal, but at startup it is the *normal*
case: a container comes up in seconds, the overlay a moment later. The consequence is not just
noise — every restart-counting health check reads the churn as a crash-loop, and a person watching
`docker ps` sees a failing module that is actually fine. The lab's beds had to learn to tell
"churn that stops" from "a real crash-loop" (a settle-based wait); every operator will face the
same reading.
A process whose dependency may appear seconds later should wait for it, with backoff, in-process —
exiting delegates the retry to the container runtime and turns an ordinary ordering into a visible
failure.
## Open questions
- Should the shared runtime bootstrap (the code every `mesh-runtime-<module>` starts with) retry
the broker connection with backoff instead of exiting, and for how long before it *is* fatal?
- Is the certificate fetch the only startup step with this shape, or do provision waits and grant
reads exit the same way?
- How is this checked once fixed — a bed asserting restart count stays 0 on a joined node?