From dd5bf9cd8c7cd9249098174b2dcf080f6447ed26 Mon Sep 17 00:00:00 2001 From: jochen Date: Thu, 17 Sep 2026 21:56:40 +0200 Subject: [PATCH 1/2] =?UTF-8?q?Issue=20059=20=E2=80=94=20a=20provisioner?= =?UTF-8?q?=20runtime=20crash-loops=20until=20the=20overlay=20is=20up?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Observed in the two-node bed: a cross-node runtime exits on 'timed out fetching the broker's certificate' until the tunnel forms, then settles. The retry belongs in-process. https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx --- .../00-report.md | 38 +++++++++++++++++++ 1 file changed, 38 insertions(+) create mode 100644 04-ISSUES/059-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md diff --git a/04-ISSUES/059-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md b/04-ISSUES/059-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md new file mode 100644 index 0000000..1adf39d --- /dev/null +++ b/04-ISSUES/059-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md @@ -0,0 +1,38 @@ +--- +status: open +opened: 2026-09-17 +located-in: [] +fixed-by: +amended-design: +--- + +# 059 — A provisioner runtime crash-loops until the overlay is up + +## Symptom + +A module runtime whose broker is on another node exits at startup — `timed out fetching the +broker's certificate` — and is restarted by the container runtime until the overlay tunnel is +established, after which it connects and works. Observed on a joined node during the two-node +database bed: the runtime's container restarted several times in its first minute, then settled and +provisioned correctly. + +## Why it matters beyond the instance + +The runtime treats "the broker is not reachable yet" as fatal, but at startup it is the *normal* +case: a container comes up in seconds, the overlay a moment later. The consequence is not just +noise — every restart-counting health check reads the churn as a crash-loop, and a person watching +`docker ps` sees a failing module that is actually fine. The lab's beds had to learn to tell +"churn that stops" from "a real crash-loop" (a settle-based wait); every operator will face the +same reading. + +A process whose dependency may appear seconds later should wait for it, with backoff, in-process — +exiting delegates the retry to the container runtime and turns an ordinary ordering into a visible +failure. + +## Open questions + +- Should the shared runtime bootstrap (the code every `mesh-runtime-` starts with) retry + the broker connection with backoff instead of exiting, and for how long before it *is* fatal? +- Is the certificate fetch the only startup step with this shape, or do provision waits and grant + reads exit the same way? +- How is this checked once fixed — a bed asserting restart count stays 0 on a joined node? -- 2.54.0 From 0ffdf4e5811f7eaf7f388e4b624770cb5d3d4e78 Mon Sep 17 00:00:00 2001 From: jochen Date: Thu, 17 Sep 2026 21:56:49 +0200 Subject: [PATCH 2/2] The issue takes the next free number, 058 https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx --- .../00-report.md | 0 1 file changed, 0 insertions(+), 0 deletions(-) rename 04-ISSUES/{059-a-provisioner-runtime-crash-loops-until-the-overlay-is-up => 058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up}/00-report.md (100%) diff --git a/04-ISSUES/059-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md b/04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md similarity index 100% rename from 04-ISSUES/059-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md rename to 04-ISSUES/058-a-provisioner-runtime-crash-loops-until-the-overlay-is-up/00-report.md -- 2.54.0