From 6165a7ae020492cfd49ef19c8c1c61ea3d94504e Mon Sep 17 00:00:00 2001 From: jochen Date: Tue, 22 Sep 2026 17:09:21 +0200 Subject: [PATCH] Issue 084: taking networking on an adopted node restarts every container, and the held runtime file blocks pulling --- .../00-report.md | 57 +++++++++++++++++++ 1 file changed, 57 insertions(+) create mode 100644 04-ISSUES/084-taking-networking-on-an-adopted-node-restarts-every-container/00-report.md diff --git a/04-ISSUES/084-taking-networking-on-an-adopted-node-restarts-every-container/00-report.md b/04-ISSUES/084-taking-networking-on-an-adopted-node-restarts-every-container/00-report.md new file mode 100644 index 0000000..95d3da0 --- /dev/null +++ b/04-ISSUES/084-taking-networking-on-an-adopted-node-restarts-every-container/00-report.md @@ -0,0 +1,57 @@ +--- +status: open +opened: 2026-09-22 +located-in: [] +fixed-by: +amended-design: +--- + +# 084 — Taking the networking module on an adopted node restarts every container on the machine + +## What was observed + +Found while planning the build of adoption mode +([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)), by +reading the controller's networking module rather than by running it. + +The networking module declares two machine-wide files, and each one **whole**: + +- the container runtime's configuration file. It carries the trust for the mesh's plain-HTTP + registry. The runtime's service is declared to restart whenever this file changes. +- the machine's hosts file, as the fact that gives nodes their names. + +On a node the predecessor still serves, both files already exist, written by the predecessor. +Under ADR 0100 they are *found*, and they are held until the networking module is taken. Holding +them is safe. The trouble starts on either side of the hold: + +1. **Taking networking is not "one service replaced".** The first write of the runtime's + configuration restarts the container runtime. That restarts every container on the machine, + and any predecessor container without a restart policy stays down. The record promises that + taking a module replaces one service and that each step says what it changes first. This step + would replace one file and stop everything the machine serves. +2. **While the file is held, the node cannot pull the mesh's images.** The registry's trust lives + in the held file. An adopted machine that joins can therefore not pull from the mesh's + registry over the private network until networking is taken, and nothing in the modules it + is assigned says why. + +The lab does not show either problem. Its base image writes the runtime's configuration before +any bed starts, so the file the module declares is never found. + +## Why it matters beyond this instance + +A module that owns a whole machine-wide file shared with software it did not install makes +taking that module a cutover for everything on the machine. Adoption mode assumes the opposite: +that modules can be taken one at a time, each when its own data has moved. Every such file, +including the hosts file and the resolver's configuration, breaks that assumption in the same way, +and nothing checks for it today. + +## Open questions + +- Should the mesh write its part of a shared machine-wide file, beside what is already there, + instead of the whole file? And does the container runtime offer a way to add registry trust + without its main configuration file? +- If networking must replace the file whole, should taking it be a cutover of its own? Its + preview would name the restart and every container it will stop. +- Which other modules declare a whole file that other software on the machine also writes? +- Should an adopted node that cannot trust the registry be refused a module that needs to pull? + Or should the refusal come earlier, when the node joins?