Merge pull request 'Issue 101: taking a service its neighbours reach by container name cuts them off' (#87) from issues/101-a-service-reached-by-name-loses-its-network into main

This commit was merged in pull request #87.
This commit is contained in:
2026-09-23 18:14:08 +00:00
@@ -0,0 +1,63 @@
---
status: open
opened: 2026-09-23
located-in: []
fixed-by:
amended-design:
---
# 101 — Taking a service its neighbours reach by container name cuts them off, and nothing says so
## What was observed
Checking the third cutover on the control-node, 2026-09-23, before running it.
The service looked like the cleanest candidate yet: the same image the machine runs, the same
published port, the same seven data directories, no configuration file to overwrite. Every check
the previous two cutovers had taught agreed.
Then the network. The predecessor's container sits on a shared network — fifteen of its services
do — and the one that uses this service reaches it **by container name**:
```
DocumentServerInternalUrl = http://<service-name>/
```
The module declares a network of its own, as a module should. Taking it would move the container
onto that network, the name would stop resolving for the neighbour, and the neighbour would fail —
not at the cutover, but the next time a person tried to open a document.
The two services already migrated were safe by accident. Both are reached **through a host port**,
which survives the change of owner because the port is the machine's, not the network's. Nothing
distinguished the two cases; the difference only appeared because a fourth thing was checked.
## Why it matters beyond this instance
This decides the **order** of a migration, and it is the first constraint found that does.
Ports, versions, data directories and configuration are all properties of the service being taken.
This one is a property of everything *else*: a service can be taken when the things that call it
call it by an address that survives, and not when they call it by a name only the predecessor's
network resolves. On the machine measured, fifteen services share one network, so the question is
not rare — it is most of what is left.
It also cannot be answered by looking at the module. The module is correct: a module declares its
own network, because a catalogue shared by every mesh cannot name a network one installation
happens to have. The answer is in the **consumer's configuration**, which belongs to a service that
has not been migrated and may not even be a module yet.
And the failure is quiet. Nothing fails at the cutover. The service answers on its port, the mesh
reports it healthy, and the breakage is in a neighbour, later, on a code path a person has to
exercise.
## Open questions
- What tells the operator, before a cutover, that something reaches this service by a name that
will stop resolving? Nothing on the machine currently answers "who talks to this, and how".
- Should a taken container be able to join a network the mesh did not create — named per node, as
a setting — so a half-migrated machine still resolves? That trades the clean boundary for a
bridge with an end date.
- Or should the rule simply be that such services are taken **together**, as a group, and the mesh
learn to take a group atomically? Nothing takes more than one module at a time today.
- Does the same hole exist for anything else the predecessor's runtime resolves and the mesh's does
not — a network alias, a `depends_on`, a name in a shared `/etc/hosts`?