Verified on the machine: umami has zero restarts, applies all 26 migrations, reports the store up to date, and umami.novox.be answers 200 where it had answered 502 since 2026-09-25. It was the same fault as 135, filed three days earlier and diagnosed without either record noticing the other. 135's container IS umami — it is named in 135's own evidence table, holding novox.internal:10.42.0.1 after the overlay range moved. The dial appeared to succeed and the first real query timed out because the name pointed at an address that no longer existed; the store, the path and the credential were all fine. 118's own reasoning is kept as a warning, because it is careful and wrong: it argued path-MTU and conntrack, and it named the move that would have found it — compare umami against a container created that week — and did not make it. A stale name presents as a network fault, which is why ADR 0148 stops copying names into containers rather than detecting when the copies go bad. Also: 118's located-in named mesh-catalog modules/umami, which was never at fault; it names mesh-host's comparison now. And 135 gains the pointer to 118, which is the back-reference I have now missed three times.
99 lines
5.7 KiB
Markdown
99 lines
5.7 KiB
Markdown
---
|
|
status: resolved
|
|
opened: 2026-09-28
|
|
located-in: [mesh-host internal/apply]
|
|
fixed-by: mesh-host — a container's mesh names are part of the spec digest the host compares, sorted so the digest does not move for a reordering. A container whose names moved is now recreated like a container whose image moved, and the test fails against the previous behaviour.
|
|
amended-design:
|
|
---
|
|
|
|
# 135 — A container's mesh names are not compared, so a moved address is never noticed
|
|
|
|
## What was observed
|
|
|
|
One container on this mesh had been restarting every thirty seconds for five days — 2286 times — and
|
|
the mesh reported the machine as doing what it was told.
|
|
|
|
Its logs said its database connected and then a query timed out. The database was reachable: the same
|
|
query from the same network, with the same credential, answered in milliseconds. What differed was the
|
|
name. Inside that container, `novox.internal` resolved to `10.42.0.1`; in every other container on the
|
|
machine it resolved to `10.10.0.1`. The mesh's overlay range had moved, and this container still held
|
|
the old one:
|
|
|
|
```
|
|
umami created 2026-09-23 novox.internal:10.42.0.1
|
|
mesh-catalog created today novox.internal:10.10.0.1
|
|
```
|
|
|
|
A container resolves other machines and public names through the entries the mesh gives it when it is
|
|
created, and nothing re-reads them afterwards. The host compares a container against what was declared
|
|
by a digest of its spec — image, name, environment, ports, volumes, arguments, resolver, address, and
|
|
what it reads — and **the mesh's names were not in it**. So this container matched what was declared,
|
|
was left alone, and kept an address that had not existed for five days.
|
|
|
|
Forty-eight other containers had current names. Not because anything corrected them: each had been
|
|
recreated for some other reason — a new image, a changed file — and picked up the current roster on the
|
|
way. This one's image is an upstream release that had not moved, and nothing else about it changed, so
|
|
nothing ever recreated it.
|
|
|
|
## Why it matters beyond this instance
|
|
|
|
**It is the exact fault [issue 045](../045-a-container-keeps-the-values-it-started-with/00-report.md)
|
|
named, in the one field that was left out.** That issue is why the digest carries what a container
|
|
reads: "a container whose configuration had since been rewritten compared equal and was left alone —
|
|
running values the machine no longer holds, while every check reported success." The same sentence
|
|
describes this, with *names* in place of *files*.
|
|
|
|
**The failure is invisible in exactly the way that matters.** The container runs, so the machine
|
|
reports it applied. It restarts, but a restarting container is a normal sight during an upgrade. The
|
|
only account of the fault is inside the container's own log, in the words of the application rather
|
|
than of the mesh — and what it says is that a query timed out, which points at the database.
|
|
|
|
**And it is most likely to bite what changes least.** Every container that is rebuilt often repairs
|
|
itself by accident. The victim is the module whose image is stable — which is to say, the module that
|
|
was working fine.
|
|
|
|
## What was done
|
|
|
|
The mesh's names are part of the digest, sorted so the digest does not move for a reordering nobody
|
|
made. A container whose names moved is now recreated exactly as one whose image moved.
|
|
|
|
The first apply after this recreates every container that carries mesh names — one restart each,
|
|
already the price the mesh pays for any image update — because their recorded digests predate the
|
|
field.
|
|
|
|
## What is still true
|
|
|
|
The mesh gives a container its names at creation and has no way to change them in place. That is the
|
|
container runtime's shape, not a choice; the answer is to recreate, which is what this does. A module
|
|
that would rather re-read a roster from a file can already ask for one as a fact
|
|
([ADR 0120](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md)) and restart on
|
|
it.
|
|
|
|
## The container was umami, and it had its own record (2026-09-30)
|
|
|
|
The container in the table above is umami, and its symptom had already been filed three days earlier as
|
|
[issue 118](../118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md) — the store
|
|
answering the dial and timing out the query, a restart loop, and a public name answering `502`. Neither
|
|
record noticed the other: 118 was reasoning about path-MTU and conntrack on the network between the
|
|
container and the store, which is what a stale name looks like from inside the container.
|
|
|
|
118 is resolved by this record's fix and says so. Noted here because a symptom filed twice, diagnosed
|
|
once, and closed in one place is how a repository comes to disagree with itself.
|
|
|
|
## What replaced this fix (2026-09-30)
|
|
|
|
The fix here — putting the mesh's names into the digest the host compares, so a container whose names
|
|
moved is recreated like one whose image moved — worked, and cost more than it was worth. It made the
|
|
roster part of every container's identity, so one name moving replaced every container in the mesh: a
|
|
module assigned on one machine restarted the store, the registry, the edge and mail on another
|
|
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)).
|
|
|
|
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) goes at
|
|
the cause this record only described: **a copy taken at creation is stale the moment the roster moves,
|
|
and detecting that is not as good as not copying.** A container resolves the mesh's names through its
|
|
machine's resolver, at the moment it asks, so the fault this record reports cannot occur rather than
|
|
being noticed a restart later.
|
|
|
|
Recorded here because this is where somebody arrives to find out why the digest carries names, and the
|
|
answer is that it did, for two days short of a month, and stopped.
|