Files
hq/04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md
T
jschoubben 295bf2f3ab Issue 118 is resolved: it was issue 135, and umami is healthy
Verified on the machine: umami has zero restarts, applies all 26
migrations, reports the store up to date, and umami.novox.be answers 200
where it had answered 502 since 2026-09-25.

It was the same fault as 135, filed three days earlier and diagnosed
without either record noticing the other. 135's container IS umami — it
is named in 135's own evidence table, holding novox.internal:10.42.0.1
after the overlay range moved. The dial appeared to succeed and the first
real query timed out because the name pointed at an address that no
longer existed; the store, the path and the credential were all fine.

118's own reasoning is kept as a warning, because it is careful and
wrong: it argued path-MTU and conntrack, and it named the move that would
have found it — compare umami against a container created that week —
and did not make it. A stale name presents as a network fault, which is
why ADR 0148 stops copying names into containers rather than detecting
when the copies go bad.

Also: 118's located-in named mesh-catalog modules/umami, which was never
at fault; it names mesh-host's comparison now. And 135 gains the pointer
to 118, which is the back-reference I have now missed three times.
2026-09-30 01:54:19 +02:00

99 lines
5.7 KiB
Markdown

---
status: resolved
opened: 2026-09-28
located-in: [mesh-host internal/apply]
fixed-by: mesh-host — a container's mesh names are part of the spec digest the host compares, sorted so the digest does not move for a reordering. A container whose names moved is now recreated like a container whose image moved, and the test fails against the previous behaviour.
amended-design:
---
# 135 — A container's mesh names are not compared, so a moved address is never noticed
## What was observed
One container on this mesh had been restarting every thirty seconds for five days — 2286 times — and
the mesh reported the machine as doing what it was told.
Its logs said its database connected and then a query timed out. The database was reachable: the same
query from the same network, with the same credential, answered in milliseconds. What differed was the
name. Inside that container, `novox.internal` resolved to `10.42.0.1`; in every other container on the
machine it resolved to `10.10.0.1`. The mesh's overlay range had moved, and this container still held
the old one:
```
umami created 2026-09-23 novox.internal:10.42.0.1
mesh-catalog created today novox.internal:10.10.0.1
```
A container resolves other machines and public names through the entries the mesh gives it when it is
created, and nothing re-reads them afterwards. The host compares a container against what was declared
by a digest of its spec — image, name, environment, ports, volumes, arguments, resolver, address, and
what it reads — and **the mesh's names were not in it**. So this container matched what was declared,
was left alone, and kept an address that had not existed for five days.
Forty-eight other containers had current names. Not because anything corrected them: each had been
recreated for some other reason — a new image, a changed file — and picked up the current roster on the
way. This one's image is an upstream release that had not moved, and nothing else about it changed, so
nothing ever recreated it.
## Why it matters beyond this instance
**It is the exact fault [issue 045](../045-a-container-keeps-the-values-it-started-with/00-report.md)
named, in the one field that was left out.** That issue is why the digest carries what a container
reads: "a container whose configuration had since been rewritten compared equal and was left alone —
running values the machine no longer holds, while every check reported success." The same sentence
describes this, with *names* in place of *files*.
**The failure is invisible in exactly the way that matters.** The container runs, so the machine
reports it applied. It restarts, but a restarting container is a normal sight during an upgrade. The
only account of the fault is inside the container's own log, in the words of the application rather
than of the mesh — and what it says is that a query timed out, which points at the database.
**And it is most likely to bite what changes least.** Every container that is rebuilt often repairs
itself by accident. The victim is the module whose image is stable — which is to say, the module that
was working fine.
## What was done
The mesh's names are part of the digest, sorted so the digest does not move for a reordering nobody
made. A container whose names moved is now recreated exactly as one whose image moved.
The first apply after this recreates every container that carries mesh names — one restart each,
already the price the mesh pays for any image update — because their recorded digests predate the
field.
## What is still true
The mesh gives a container its names at creation and has no way to change them in place. That is the
container runtime's shape, not a choice; the answer is to recreate, which is what this does. A module
that would rather re-read a roster from a file can already ask for one as a fact
([ADR 0120](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md)) and restart on
it.
## The container was umami, and it had its own record (2026-09-30)
The container in the table above is umami, and its symptom had already been filed three days earlier as
[issue 118](../118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md) — the store
answering the dial and timing out the query, a restart loop, and a public name answering `502`. Neither
record noticed the other: 118 was reasoning about path-MTU and conntrack on the network between the
container and the store, which is what a stale name looks like from inside the container.
118 is resolved by this record's fix and says so. Noted here because a symptom filed twice, diagnosed
once, and closed in one place is how a repository comes to disagree with itself.
## What replaced this fix (2026-09-30)
The fix here — putting the mesh's names into the digest the host compares, so a container whose names
moved is recreated like one whose image moved — worked, and cost more than it was worth. It made the
roster part of every container's identity, so one name moving replaced every container in the mesh: a
module assigned on one machine restarted the store, the registry, the edge and mail on another
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)).
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) goes at
the cause this record only described: **a copy taken at creation is stale the moment the roster moves,
and detecting that is not as good as not copying.** A container resolves the mesh's names through its
machine's resolver, at the moment it asks, so the fault this record reports cannot occur rather than
being noticed a restart later.
Recorded here because this is where somebody arrives to find out why the digest carries names, and the
answer is that it did, for two days short of a month, and stopped.