Issue 059 — a provisioner runtime crash-loops until the overlay is up
Observed in the two-node bed: a cross-node runtime exits on 'timed out fetching the broker's certificate' until the tunnel forms, then settles. The retry belongs in-process. https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
This commit is contained in:
@@ -0,0 +1,38 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-17
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 059 — A provisioner runtime crash-loops until the overlay is up
|
||||
|
||||
## Symptom
|
||||
|
||||
A module runtime whose broker is on another node exits at startup — `timed out fetching the
|
||||
broker's certificate` — and is restarted by the container runtime until the overlay tunnel is
|
||||
established, after which it connects and works. Observed on a joined node during the two-node
|
||||
database bed: the runtime's container restarted several times in its first minute, then settled and
|
||||
provisioned correctly.
|
||||
|
||||
## Why it matters beyond the instance
|
||||
|
||||
The runtime treats "the broker is not reachable yet" as fatal, but at startup it is the *normal*
|
||||
case: a container comes up in seconds, the overlay a moment later. The consequence is not just
|
||||
noise — every restart-counting health check reads the churn as a crash-loop, and a person watching
|
||||
`docker ps` sees a failing module that is actually fine. The lab's beds had to learn to tell
|
||||
"churn that stops" from "a real crash-loop" (a settle-based wait); every operator will face the
|
||||
same reading.
|
||||
|
||||
A process whose dependency may appear seconds later should wait for it, with backoff, in-process —
|
||||
exiting delegates the retry to the container runtime and turns an ordinary ordering into a visible
|
||||
failure.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should the shared runtime bootstrap (the code every `mesh-runtime-<module>` starts with) retry
|
||||
the broker connection with backoff instead of exiting, and for how long before it *is* fatal?
|
||||
- Is the certificate fetch the only startup step with this shape, or do provision waits and grant
|
||||
reads exit the same way?
|
||||
- How is this checked once fixed — a bed asserting restart count stays 0 on a joined node?
|
||||
Reference in New Issue
Block a user