Three records were left describing a mechanism a new record had moved, and a reader arrives at them by following a citation: 0066 still said a routed name is written into every container after 0148 replaced that with resolution; 0016 still read as though the lab were the test bed after 0149; and issues 109 and 135 said nothing about 0148 ending the copying that 135's own fix made comparable. Each was a citation leading to the wrong answer in a record that was not wrong about anything it decided. This is the second time in one session. The playbook rule I added last round did not stop it, so the convention is now written where the record conventions live, with the shape to use and three worked examples — and with the honest note that it is NOT machine-checked and cannot be from `extends:` alone: 102 records extend another, 87 have no back-reference, and that is correct, because extending usually means building on a context. Making it mechanical means a record declaring the relationship in frontmatter, which is a schema change and is not mine to decide. Also: designs 18 and 20 claimed `updated:` dates from before I edited them, and 117's `fixed-by` gained the commit beside the record.
88 lines
5.0 KiB
Markdown
88 lines
5.0 KiB
Markdown
---
|
|
status: resolved
|
|
opened: 2026-09-28
|
|
located-in: [mesh-host internal/apply]
|
|
fixed-by: mesh-host — a container's mesh names are part of the spec digest the host compares, sorted so the digest does not move for a reordering. A container whose names moved is now recreated like a container whose image moved, and the test fails against the previous behaviour.
|
|
amended-design:
|
|
---
|
|
|
|
# 135 — A container's mesh names are not compared, so a moved address is never noticed
|
|
|
|
## What was observed
|
|
|
|
One container on this mesh had been restarting every thirty seconds for five days — 2286 times — and
|
|
the mesh reported the machine as doing what it was told.
|
|
|
|
Its logs said its database connected and then a query timed out. The database was reachable: the same
|
|
query from the same network, with the same credential, answered in milliseconds. What differed was the
|
|
name. Inside that container, `novox.internal` resolved to `10.42.0.1`; in every other container on the
|
|
machine it resolved to `10.10.0.1`. The mesh's overlay range had moved, and this container still held
|
|
the old one:
|
|
|
|
```
|
|
umami created 2026-09-23 novox.internal:10.42.0.1
|
|
mesh-catalog created today novox.internal:10.10.0.1
|
|
```
|
|
|
|
A container resolves other machines and public names through the entries the mesh gives it when it is
|
|
created, and nothing re-reads them afterwards. The host compares a container against what was declared
|
|
by a digest of its spec — image, name, environment, ports, volumes, arguments, resolver, address, and
|
|
what it reads — and **the mesh's names were not in it**. So this container matched what was declared,
|
|
was left alone, and kept an address that had not existed for five days.
|
|
|
|
Forty-eight other containers had current names. Not because anything corrected them: each had been
|
|
recreated for some other reason — a new image, a changed file — and picked up the current roster on the
|
|
way. This one's image is an upstream release that had not moved, and nothing else about it changed, so
|
|
nothing ever recreated it.
|
|
|
|
## Why it matters beyond this instance
|
|
|
|
**It is the exact fault [issue 045](../045-a-container-keeps-the-values-it-started-with/00-report.md)
|
|
named, in the one field that was left out.** That issue is why the digest carries what a container
|
|
reads: "a container whose configuration had since been rewritten compared equal and was left alone —
|
|
running values the machine no longer holds, while every check reported success." The same sentence
|
|
describes this, with *names* in place of *files*.
|
|
|
|
**The failure is invisible in exactly the way that matters.** The container runs, so the machine
|
|
reports it applied. It restarts, but a restarting container is a normal sight during an upgrade. The
|
|
only account of the fault is inside the container's own log, in the words of the application rather
|
|
than of the mesh — and what it says is that a query timed out, which points at the database.
|
|
|
|
**And it is most likely to bite what changes least.** Every container that is rebuilt often repairs
|
|
itself by accident. The victim is the module whose image is stable — which is to say, the module that
|
|
was working fine.
|
|
|
|
## What was done
|
|
|
|
The mesh's names are part of the digest, sorted so the digest does not move for a reordering nobody
|
|
made. A container whose names moved is now recreated exactly as one whose image moved.
|
|
|
|
The first apply after this recreates every container that carries mesh names — one restart each,
|
|
already the price the mesh pays for any image update — because their recorded digests predate the
|
|
field.
|
|
|
|
## What is still true
|
|
|
|
The mesh gives a container its names at creation and has no way to change them in place. That is the
|
|
container runtime's shape, not a choice; the answer is to recreate, which is what this does. A module
|
|
that would rather re-read a roster from a file can already ask for one as a fact
|
|
([ADR 0120](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md)) and restart on
|
|
it.
|
|
|
|
## What replaced this fix (2026-09-30)
|
|
|
|
The fix here — putting the mesh's names into the digest the host compares, so a container whose names
|
|
moved is recreated like one whose image moved — worked, and cost more than it was worth. It made the
|
|
roster part of every container's identity, so one name moving replaced every container in the mesh: a
|
|
module assigned on one machine restarted the store, the registry, the edge and mail on another
|
|
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)).
|
|
|
|
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) goes at
|
|
the cause this record only described: **a copy taken at creation is stale the moment the roster moves,
|
|
and detecting that is not as good as not copying.** A container resolves the mesh's names through its
|
|
machine's resolver, at the moment it asks, so the fault this record reports cannot occur rather than
|
|
being noticed a restart later.
|
|
|
|
Recorded here because this is where somebody arrives to find out why the digest carries names, and the
|
|
answer is that it did, for two days short of a month, and stopped.
|