The glossary's authority page still named the controller's seat the-controller in two entries, contradicting its own seat section after ADR 0079; issue 058's heading kept the pre-renumber 059; 055's fixed-by named branches that stop existing after merge (now merge commits/PRs) and its located-in listed file paths where the convention wants repos; 056's located-in named mesh-host, which received no fix, instead of mesh-catalog; and the design layer never said the one-store/one-broker property is enforced — 07-the-foundation and the installation table now state the seats. https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
1.7 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design |
|---|---|---|---|---|
| open | 2026-09-17 |
058 — A provisioner runtime crash-loops until the overlay is up
Symptom
A module runtime whose broker is on another node exits at startup — timed out fetching the broker's certificate — and is restarted by the container runtime until the overlay tunnel is
established, after which it connects and works. Observed on a joined node during the two-node
database bed: the runtime's container restarted several times in its first minute, then settled and
provisioned correctly.
Why it matters beyond the instance
The runtime treats "the broker is not reachable yet" as fatal, but at startup it is the normal
case: a container comes up in seconds, the overlay a moment later. The consequence is not just
noise — every restart-counting health check reads the churn as a crash-loop, and a person watching
docker ps sees a failing module that is actually fine. The lab's beds had to learn to tell
"churn that stops" from "a real crash-loop" (a settle-based wait); every operator will face the
same reading.
A process whose dependency may appear seconds later should wait for it, with backoff, in-process — exiting delegates the retry to the container runtime and turns an ordinary ordering into a visible failure.
Open questions
- Should the shared runtime bootstrap (the code every
mesh-runtime-<module>starts with) retry the broker connection with backoff instead of exiting, and for how long before it is fatal? - Is the certificate fetch the only startup step with this shape, or do provision waits and grant reads exit the same way?
- How is this checked once fixed — a bed asserting restart count stays 0 on a joined node?