Issue 101: taking a service its neighbours reach by container name cuts them off #87
+63
@@ -0,0 +1,63 @@
|
|||||||
|
---
|
||||||
|
status: open
|
||||||
|
opened: 2026-09-23
|
||||||
|
located-in: []
|
||||||
|
fixed-by:
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 101 — Taking a service its neighbours reach by container name cuts them off, and nothing says so
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
Checking the third cutover on the control-node, 2026-09-23, before running it.
|
||||||
|
|
||||||
|
The service looked like the cleanest candidate yet: the same image the machine runs, the same
|
||||||
|
published port, the same seven data directories, no configuration file to overwrite. Every check
|
||||||
|
the previous two cutovers had taught agreed.
|
||||||
|
|
||||||
|
Then the network. The predecessor's container sits on a shared network — fifteen of its services
|
||||||
|
do — and the one that uses this service reaches it **by container name**:
|
||||||
|
|
||||||
|
```
|
||||||
|
DocumentServerInternalUrl = http://<service-name>/
|
||||||
|
```
|
||||||
|
|
||||||
|
The module declares a network of its own, as a module should. Taking it would move the container
|
||||||
|
onto that network, the name would stop resolving for the neighbour, and the neighbour would fail —
|
||||||
|
not at the cutover, but the next time a person tried to open a document.
|
||||||
|
|
||||||
|
The two services already migrated were safe by accident. Both are reached **through a host port**,
|
||||||
|
which survives the change of owner because the port is the machine's, not the network's. Nothing
|
||||||
|
distinguished the two cases; the difference only appeared because a fourth thing was checked.
|
||||||
|
|
||||||
|
## Why it matters beyond this instance
|
||||||
|
|
||||||
|
This decides the **order** of a migration, and it is the first constraint found that does.
|
||||||
|
|
||||||
|
Ports, versions, data directories and configuration are all properties of the service being taken.
|
||||||
|
This one is a property of everything *else*: a service can be taken when the things that call it
|
||||||
|
call it by an address that survives, and not when they call it by a name only the predecessor's
|
||||||
|
network resolves. On the machine measured, fifteen services share one network, so the question is
|
||||||
|
not rare — it is most of what is left.
|
||||||
|
|
||||||
|
It also cannot be answered by looking at the module. The module is correct: a module declares its
|
||||||
|
own network, because a catalogue shared by every mesh cannot name a network one installation
|
||||||
|
happens to have. The answer is in the **consumer's configuration**, which belongs to a service that
|
||||||
|
has not been migrated and may not even be a module yet.
|
||||||
|
|
||||||
|
And the failure is quiet. Nothing fails at the cutover. The service answers on its port, the mesh
|
||||||
|
reports it healthy, and the breakage is in a neighbour, later, on a code path a person has to
|
||||||
|
exercise.
|
||||||
|
|
||||||
|
## Open questions
|
||||||
|
|
||||||
|
- What tells the operator, before a cutover, that something reaches this service by a name that
|
||||||
|
will stop resolving? Nothing on the machine currently answers "who talks to this, and how".
|
||||||
|
- Should a taken container be able to join a network the mesh did not create — named per node, as
|
||||||
|
a setting — so a half-migrated machine still resolves? That trades the clean boundary for a
|
||||||
|
bridge with an end date.
|
||||||
|
- Or should the rule simply be that such services are taken **together**, as a group, and the mesh
|
||||||
|
learn to take a group atomically? Nothing takes more than one module at a time today.
|
||||||
|
- Does the same hole exist for anything else the predecessor's runtime resolves and the mesh's does
|
||||||
|
not — a network alias, a `depends_on`, a name in a shared `/etc/hosts`?
|
||||||
Reference in New Issue
Block a user