From 61e4e971a4077430910d5e1f18657d2b9d1a0428 Mon Sep 17 00:00:00 2001 From: jochen Date: Wed, 23 Sep 2026 20:13:53 +0200 Subject: [PATCH] Issue 101: taking a service reached by container name cuts its neighbours off MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Found checking the third cutover rather than running it. The first two were safe by accident — both are reached through a host port, which survives a change of owner. This is the first constraint found that decides the order of the migration. --- .../00-report.md | 63 +++++++++++++++++++ 1 file changed, 63 insertions(+) create mode 100644 04-ISSUES/101-taking-a-service-reached-by-container-name-cuts-its-neighbours-off/00-report.md diff --git a/04-ISSUES/101-taking-a-service-reached-by-container-name-cuts-its-neighbours-off/00-report.md b/04-ISSUES/101-taking-a-service-reached-by-container-name-cuts-its-neighbours-off/00-report.md new file mode 100644 index 0000000..6b8da96 --- /dev/null +++ b/04-ISSUES/101-taking-a-service-reached-by-container-name-cuts-its-neighbours-off/00-report.md @@ -0,0 +1,63 @@ +--- +status: open +opened: 2026-09-23 +located-in: [] +fixed-by: +amended-design: +--- + +# 101 — Taking a service its neighbours reach by container name cuts them off, and nothing says so + +## What was observed + +Checking the third cutover on the control-node, 2026-09-23, before running it. + +The service looked like the cleanest candidate yet: the same image the machine runs, the same +published port, the same seven data directories, no configuration file to overwrite. Every check +the previous two cutovers had taught agreed. + +Then the network. The predecessor's container sits on a shared network — fifteen of its services +do — and the one that uses this service reaches it **by container name**: + +``` +DocumentServerInternalUrl = http:/// +``` + +The module declares a network of its own, as a module should. Taking it would move the container +onto that network, the name would stop resolving for the neighbour, and the neighbour would fail — +not at the cutover, but the next time a person tried to open a document. + +The two services already migrated were safe by accident. Both are reached **through a host port**, +which survives the change of owner because the port is the machine's, not the network's. Nothing +distinguished the two cases; the difference only appeared because a fourth thing was checked. + +## Why it matters beyond this instance + +This decides the **order** of a migration, and it is the first constraint found that does. + +Ports, versions, data directories and configuration are all properties of the service being taken. +This one is a property of everything *else*: a service can be taken when the things that call it +call it by an address that survives, and not when they call it by a name only the predecessor's +network resolves. On the machine measured, fifteen services share one network, so the question is +not rare — it is most of what is left. + +It also cannot be answered by looking at the module. The module is correct: a module declares its +own network, because a catalogue shared by every mesh cannot name a network one installation +happens to have. The answer is in the **consumer's configuration**, which belongs to a service that +has not been migrated and may not even be a module yet. + +And the failure is quiet. Nothing fails at the cutover. The service answers on its port, the mesh +reports it healthy, and the breakage is in a neighbour, later, on a code path a person has to +exercise. + +## Open questions + +- What tells the operator, before a cutover, that something reaches this service by a name that + will stop resolving? Nothing on the machine currently answers "who talks to this, and how". +- Should a taken container be able to join a network the mesh did not create — named per node, as + a setting — so a half-migrated machine still resolves? That trades the clean boundary for a + bridge with an end date. +- Or should the rule simply be that such services are taken **together**, as a group, and the mesh + learn to take a group atomically? Nothing takes more than one module at a time today. +- Does the same hole exist for anything else the predecessor's runtime resolves and the mesh's does + not — a network alias, a `depends_on`, a name in a shared `/etc/hosts`? -- 2.54.0