diff --git a/04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md b/04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md index 27b4801..fa36401 100644 --- a/04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md +++ b/04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md @@ -1,8 +1,8 @@ --- -status: located +status: resolved opened: 2026-09-25 -located-in: [mesh-catalog modules/umami] -fixed-by: +located-in: [mesh-host internal/apply (the spec comparison, via issue 135) — not mesh-catalog modules/umami, which this record first named and which was never at fault] +fixed-by: mesh-host e82789a (issue 135's fix) — recreating the container with a current roster ended it; verified live 2026-09-30 amended-design: --- diff --git a/04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/01-resolution.md b/04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/01-resolution.md new file mode 100644 index 0000000..8eeac7c --- /dev/null +++ b/04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/01-resolution.md @@ -0,0 +1,79 @@ +# 118 — resolved: it was issue 135, and it is over + +*Verified on the machine, 2026-09-30.* + +## It no longer happens + +``` +$ docker inspect -f '{{.RestartCount}}' umami +0 +$ docker logs umami --tail 12 +26 migrations found in prisma/migrations +No pending migrations to apply. +✓ Database is up to date. +▲ Next.js 16.3.4 +✓ Ready in 0ms +$ curl -o /dev/null -w '%{http_code}' https://umami.novox.be/ +200 +``` + +No restart loop, no `prisma.$queryRaw()` timeout, and the public name that had answered `502` since +2026-09-25 answers `200`. The raw query that could not complete now runs twenty-six migrations and +reports the store up to date. + +## What it was + +**The same fault as [issue 135](../135-a-containers-mesh-names-are-not-compared/00-report.md), which +was diagnosed three days later without either record noticing the other.** 135's container is this +one, named in its own evidence table: + +``` +umami created 2026-09-23 novox.internal:10.42.0.1 +mesh-catalog created today novox.internal:10.10.0.1 +``` + +The mesh's overlay range had moved. Umami had been created before the move and held the store's name +at an address that no longer existed, while every container made after the move held the current one. +That is why the dial appeared to succeed and the first real query timed out, and why the same query +from the same network with the same credential answered in milliseconds — **what differed was the +name, not the path, not the credential and not the store.** + +Today it holds `novox.internal:10.10.0.1`. + +## Why this record did not find it + +The report's own reasoning is worth keeping as a warning, because it is careful and it is wrong: + +> A dial that succeeds and a query that times out, from a container on one network to a store on +> another, has the shape of a path-MTU or conntrack fault (large response packets dropped after the +> small handshake ones pass), or of the store accepting the TCP connection while the backend it +> proxies for is wedged. + +Both are good hypotheses about a network path. Neither is the answer, and the record also names the +move that would have found it — *"what is known to differ for umami against every working consumer of +the same store tonight is nothing yet — that comparison is the first move"* — and then did not make it. +Comparing umami's hosts entries against any container created that week would have shown a five-day-old +address in one field. + +**A stale name presents as a network fault.** That is the lesson, and it is the reason +[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) stops +copying names into containers at all: not because detecting staleness is hard, but because it disguises +itself as something else for five days while every check reports success. + +## What actually ended it + +Issue 135's fix — `mesh-host` e82789a, *a container's mesh names are part of what it is* — put the +roster into the spec digest the host compares, so a container whose names moved is recreated like one +whose image moved. That recreated umami with a current roster and ended this. + +That fix has since been superseded in turn, by 0148, because making the roster part of every +container's identity meant one name moving replaced every container in the mesh +([issue 151](../151-a-new-name-recreates-every-container-in-the-mesh/00-report.md)). So this record +closes on a fix that is itself on the way out — which does not make it less closed, and is worth saying +plainly rather than leaving a reader to find out. + +## Not carried forward + +The `502` had one other contributor worth recording as ruled out: a stale duplicate Traefik router for +this name, hand-authored in the predecessor's era beside the mesh-written one, removed on 2026-09-25 +and — as the report says — changing nothing. It was not the cause and it is gone. diff --git a/04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md b/04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md index 30cc47b..19cc402 100644 --- a/04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md +++ b/04-ISSUES/135-a-containers-mesh-names-are-not-compared/00-report.md @@ -69,6 +69,17 @@ that would rather re-read a roster from a file can already ask for one as a fact ([ADR 0120](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md)) and restart on it. +## The container was umami, and it had its own record (2026-09-30) + +The container in the table above is umami, and its symptom had already been filed three days earlier as +[issue 118](../118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md) — the store +answering the dial and timing out the query, a restart loop, and a public name answering `502`. Neither +record noticed the other: 118 was reasoning about path-MTU and conntrack on the network between the +container and the store, which is what a stale name looks like from inside the container. + +118 is resolved by this record's fix and says so. Noted here because a symptom filed twice, diagnosed +once, and closed in one place is how a repository comes to disagree with itself. + ## What replaced this fix (2026-09-30) The fix here — putting the mesh's names into the digest the host compares, so a container whose names