Files
hq/04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/01-resolution.md
T
jschoubben 295bf2f3ab Issue 118 is resolved: it was issue 135, and umami is healthy
Verified on the machine: umami has zero restarts, applies all 26
migrations, reports the store up to date, and umami.novox.be answers 200
where it had answered 502 since 2026-09-25.

It was the same fault as 135, filed three days earlier and diagnosed
without either record noticing the other. 135's container IS umami — it
is named in 135's own evidence table, holding novox.internal:10.42.0.1
after the overlay range moved. The dial appeared to succeed and the first
real query timed out because the name pointed at an address that no
longer existed; the store, the path and the credential were all fine.

118's own reasoning is kept as a warning, because it is careful and
wrong: it argued path-MTU and conntrack, and it named the move that would
have found it — compare umami against a container created that week —
and did not make it. A stale name presents as a network fault, which is
why ADR 0148 stops copying names into containers rather than detecting
when the copies go bad.

Also: 118's located-in named mesh-catalog modules/umami, which was never
at fault; it names mesh-host's comparison now. And 135 gains the pointer
to 118, which is the back-reference I have now missed three times.
2026-09-30 01:54:19 +02:00

3.8 KiB

118 — resolved: it was issue 135, and it is over

Verified on the machine, 2026-09-30.

It no longer happens

$ docker inspect -f '{{.RestartCount}}' umami
0
$ docker logs umami --tail 12
26 migrations found in prisma/migrations
No pending migrations to apply.
✓ Database is up to date.
▲ Next.js 16.3.4
✓ Ready in 0ms
$ curl -o /dev/null -w '%{http_code}' https://umami.novox.be/
200

No restart loop, no prisma.$queryRaw() timeout, and the public name that had answered 502 since 2026-09-25 answers 200. The raw query that could not complete now runs twenty-six migrations and reports the store up to date.

What it was

The same fault as issue 135, which was diagnosed three days later without either record noticing the other. 135's container is this one, named in its own evidence table:

umami          created 2026-09-23   novox.internal:10.42.0.1
mesh-catalog   created today        novox.internal:10.10.0.1

The mesh's overlay range had moved. Umami had been created before the move and held the store's name at an address that no longer existed, while every container made after the move held the current one. That is why the dial appeared to succeed and the first real query timed out, and why the same query from the same network with the same credential answered in milliseconds — what differed was the name, not the path, not the credential and not the store.

Today it holds novox.internal:10.10.0.1.

Why this record did not find it

The report's own reasoning is worth keeping as a warning, because it is careful and it is wrong:

A dial that succeeds and a query that times out, from a container on one network to a store on another, has the shape of a path-MTU or conntrack fault (large response packets dropped after the small handshake ones pass), or of the store accepting the TCP connection while the backend it proxies for is wedged.

Both are good hypotheses about a network path. Neither is the answer, and the record also names the move that would have found it — "what is known to differ for umami against every working consumer of the same store tonight is nothing yet — that comparison is the first move" — and then did not make it. Comparing umami's hosts entries against any container created that week would have shown a five-day-old address in one field.

A stale name presents as a network fault. That is the lesson, and it is the reason ADR 0148 stops copying names into containers at all: not because detecting staleness is hard, but because it disguises itself as something else for five days while every check reports success.

What actually ended it

Issue 135's fix — mesh-host e82789a, a container's mesh names are part of what it is — put the roster into the spec digest the host compares, so a container whose names moved is recreated like one whose image moved. That recreated umami with a current roster and ended this.

That fix has since been superseded in turn, by 0148, because making the roster part of every container's identity meant one name moving replaced every container in the mesh (issue 151). So this record closes on a fix that is itself on the way out — which does not make it less closed, and is worth saying plainly rather than leaving a reader to find out.

Not carried forward

The 502 had one other contributor worth recording as ruled out: a stale duplicate Traefik router for this name, hand-authored in the predecessor's era beside the mesh-written one, removed on 2026-09-25 and — as the report says — changing nothing. It was not the cause and it is gone.