Files
hq/04-ISSUES/101-taking-a-service-reached-by-container-name-cuts-its-neighbours-off/00-report.md
T
jschoubben 61e4e971a4 Issue 101: taking a service reached by container name cuts its neighbours off
Found checking the third cutover rather than running it. The first two were safe by
accident — both are reached through a host port, which survives a change of owner.
This is the first constraint found that decides the order of the migration.
2026-09-23 20:13:53 +02:00

3.2 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
open 2026-09-23

101 — Taking a service its neighbours reach by container name cuts them off, and nothing says so

What was observed

Checking the third cutover on the control-node, 2026-09-23, before running it.

The service looked like the cleanest candidate yet: the same image the machine runs, the same published port, the same seven data directories, no configuration file to overwrite. Every check the previous two cutovers had taught agreed.

Then the network. The predecessor's container sits on a shared network — fifteen of its services do — and the one that uses this service reaches it by container name:

DocumentServerInternalUrl = http://<service-name>/

The module declares a network of its own, as a module should. Taking it would move the container onto that network, the name would stop resolving for the neighbour, and the neighbour would fail — not at the cutover, but the next time a person tried to open a document.

The two services already migrated were safe by accident. Both are reached through a host port, which survives the change of owner because the port is the machine's, not the network's. Nothing distinguished the two cases; the difference only appeared because a fourth thing was checked.

Why it matters beyond this instance

This decides the order of a migration, and it is the first constraint found that does.

Ports, versions, data directories and configuration are all properties of the service being taken. This one is a property of everything else: a service can be taken when the things that call it call it by an address that survives, and not when they call it by a name only the predecessor's network resolves. On the machine measured, fifteen services share one network, so the question is not rare — it is most of what is left.

It also cannot be answered by looking at the module. The module is correct: a module declares its own network, because a catalogue shared by every mesh cannot name a network one installation happens to have. The answer is in the consumer's configuration, which belongs to a service that has not been migrated and may not even be a module yet.

And the failure is quiet. Nothing fails at the cutover. The service answers on its port, the mesh reports it healthy, and the breakage is in a neighbour, later, on a code path a person has to exercise.

Open questions

  • What tells the operator, before a cutover, that something reaches this service by a name that will stop resolving? Nothing on the machine currently answers "who talks to this, and how".
  • Should a taken container be able to join a network the mesh did not create — named per node, as a setting — so a half-migrated machine still resolves? That trades the clean boundary for a bridge with an end date.
  • Or should the rule simply be that such services are taken together, as a group, and the mesh learn to take a group atomically? Nothing takes more than one module at a time today.
  • Does the same hole exist for anything else the predecessor's runtime resolves and the mesh's does not — a network alias, a depends_on, a name in a shared /etc/hosts?