Issue 101: taking a service reached by container name cuts its neighbours off
Found checking the third cutover rather than running it. The first two were safe by accident — both are reached through a host port, which survives a change of owner. This is the first constraint found that decides the order of the migration.
This commit is contained in:
+63
@@ -0,0 +1,63 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-23
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 101 — Taking a service its neighbours reach by container name cuts them off, and nothing says so
|
||||
|
||||
## What was observed
|
||||
|
||||
Checking the third cutover on the control-node, 2026-09-23, before running it.
|
||||
|
||||
The service looked like the cleanest candidate yet: the same image the machine runs, the same
|
||||
published port, the same seven data directories, no configuration file to overwrite. Every check
|
||||
the previous two cutovers had taught agreed.
|
||||
|
||||
Then the network. The predecessor's container sits on a shared network — fifteen of its services
|
||||
do — and the one that uses this service reaches it **by container name**:
|
||||
|
||||
```
|
||||
DocumentServerInternalUrl = http://<service-name>/
|
||||
```
|
||||
|
||||
The module declares a network of its own, as a module should. Taking it would move the container
|
||||
onto that network, the name would stop resolving for the neighbour, and the neighbour would fail —
|
||||
not at the cutover, but the next time a person tried to open a document.
|
||||
|
||||
The two services already migrated were safe by accident. Both are reached **through a host port**,
|
||||
which survives the change of owner because the port is the machine's, not the network's. Nothing
|
||||
distinguished the two cases; the difference only appeared because a fourth thing was checked.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
This decides the **order** of a migration, and it is the first constraint found that does.
|
||||
|
||||
Ports, versions, data directories and configuration are all properties of the service being taken.
|
||||
This one is a property of everything *else*: a service can be taken when the things that call it
|
||||
call it by an address that survives, and not when they call it by a name only the predecessor's
|
||||
network resolves. On the machine measured, fifteen services share one network, so the question is
|
||||
not rare — it is most of what is left.
|
||||
|
||||
It also cannot be answered by looking at the module. The module is correct: a module declares its
|
||||
own network, because a catalogue shared by every mesh cannot name a network one installation
|
||||
happens to have. The answer is in the **consumer's configuration**, which belongs to a service that
|
||||
has not been migrated and may not even be a module yet.
|
||||
|
||||
And the failure is quiet. Nothing fails at the cutover. The service answers on its port, the mesh
|
||||
reports it healthy, and the breakage is in a neighbour, later, on a code path a person has to
|
||||
exercise.
|
||||
|
||||
## Open questions
|
||||
|
||||
- What tells the operator, before a cutover, that something reaches this service by a name that
|
||||
will stop resolving? Nothing on the machine currently answers "who talks to this, and how".
|
||||
- Should a taken container be able to join a network the mesh did not create — named per node, as
|
||||
a setting — so a half-migrated machine still resolves? That trades the clean boundary for a
|
||||
bridge with an end date.
|
||||
- Or should the rule simply be that such services are taken **together**, as a group, and the mesh
|
||||
learn to take a group atomically? Nothing takes more than one module at a time today.
|
||||
- Does the same hole exist for anything else the predecessor's runtime resolves and the mesh's does
|
||||
not — a network alias, a `depends_on`, a name in a shared `/etc/hosts`?
|
||||
Reference in New Issue
Block a user