Issue 118 is resolved: it was issue 135, and umami is healthy #199

Merged
jschoubben merged 1 commits from issue/118-is-135-and-is-resolved into main 2026-09-29 23:54:39 +00:00
3 changed files with 93 additions and 3 deletions
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-25
located-in: [mesh-catalog modules/umami]
fixed-by:
located-in: [mesh-host internal/apply (the spec comparison, via issue 135) — not mesh-catalog modules/umami, which this record first named and which was never at fault]
fixed-by: mesh-host e82789a (issue 135's fix) — recreating the container with a current roster ended it; verified live 2026-09-30
amended-design:
---
@@ -0,0 +1,79 @@
# 118 — resolved: it was issue 135, and it is over
*Verified on the machine, 2026-09-30.*
## It no longer happens
```
$ docker inspect -f '{{.RestartCount}}' umami
0
$ docker logs umami --tail 12
26 migrations found in prisma/migrations
No pending migrations to apply.
✓ Database is up to date.
▲ Next.js 16.3.4
✓ Ready in 0ms
$ curl -o /dev/null -w '%{http_code}' https://umami.novox.be/
200
```
No restart loop, no `prisma.$queryRaw()` timeout, and the public name that had answered `502` since
2026-09-25 answers `200`. The raw query that could not complete now runs twenty-six migrations and
reports the store up to date.
## What it was
**The same fault as [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md), which
was diagnosed three days later without either record noticing the other.** 135's container is this
one, named in its own evidence table:
```
umami created 2026-09-23 novox.internal:10.42.0.1
mesh-catalog created today novox.internal:10.10.0.1
```
The mesh's overlay range had moved. Umami had been created before the move and held the store's name
at an address that no longer existed, while every container made after the move held the current one.
That is why the dial appeared to succeed and the first real query timed out, and why the same query
from the same network with the same credential answered in milliseconds — **what differed was the
name, not the path, not the credential and not the store.**
Today it holds `novox.internal:10.10.0.1`.
## Why this record did not find it
The report's own reasoning is worth keeping as a warning, because it is careful and it is wrong:
> A dial that succeeds and a query that times out, from a container on one network to a store on
> another, has the shape of a path-MTU or conntrack fault (large response packets dropped after the
> small handshake ones pass), or of the store accepting the TCP connection while the backend it
> proxies for is wedged.
Both are good hypotheses about a network path. Neither is the answer, and the record also names the
move that would have found it — *"what is known to differ for umami against every working consumer of
the same store tonight is nothing yet — that comparison is the first move"* — and then did not make it.
Comparing umami's hosts entries against any container created that week would have shown a five-day-old
address in one field.
**A stale name presents as a network fault.** That is the lesson, and it is the reason
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) stops
copying names into containers at all: not because detecting staleness is hard, but because it disguises
itself as something else for five days while every check reports success.
## What actually ended it
Issue 135's fix — `mesh-host` e82789a, *a container's mesh names are part of what it is* — put the
roster into the spec digest the host compares, so a container whose names moved is recreated like one
whose image moved. That recreated umami with a current roster and ended this.
That fix has since been superseded in turn, by 0148, because making the roster part of every
container's identity meant one name moving replaced every container in the mesh
([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). So this record
closes on a fix that is itself on the way out — which does not make it less closed, and is worth saying
plainly rather than leaving a reader to find out.
## Not carried forward
The `502` had one other contributor worth recording as ruled out: a stale duplicate Traefik router for
this name, hand-authored in the predecessor's era beside the mesh-written one, removed on 2026-09-25
and — as the report says — changing nothing. It was not the cause and it is gone.
@@ -69,6 +69,17 @@ that would rather re-read a roster from a file can already ask for one as a fact
([ADR 0120](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md)) and restart on
it.
## The container was umami, and it had its own record (2026-09-30)
The container in the table above is umami, and its symptom had already been filed three days earlier as
[issue 118](../118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md) — the store
answering the dial and timing out the query, a restart loop, and a public name answering `502`. Neither
record noticed the other: 118 was reasoning about path-MTU and conntrack on the network between the
container and the store, which is what a stale name looks like from inside the container.
118 is resolved by this record's fix and says so. Noted here because a symptom filed twice, diagnosed
once, and closed in one place is how a repository comes to disagree with itself.
## What replaced this fix (2026-09-30)
The fix here — putting the mesh's names into the digest the host compares, so a container whose names