Verified on the machine: umami has zero restarts, applies all 26 migrations, reports the store up to date, and umami.novox.be answers 200 where it had answered 502 since 2026-09-25. It was the same fault as 135, filed three days earlier and diagnosed without either record noticing the other. 135's container IS umami — it is named in 135's own evidence table, holding novox.internal:10.42.0.1 after the overlay range moved. The dial appeared to succeed and the first real query timed out because the name pointed at an address that no longer existed; the store, the path and the credential were all fine. 118's own reasoning is kept as a warning, because it is careful and wrong: it argued path-MTU and conntrack, and it named the move that would have found it — compare umami against a container created that week — and did not make it. A stale name presents as a network fault, which is why ADR 0148 stops copying names into containers rather than detecting when the copies go bad. Also: 118's located-in named mesh-catalog modules/umami, which was never at fault; it names mesh-host's comparison now. And 135 gains the pointer to 118, which is the back-reference I have now missed three times.
2.2 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design | |||
|---|---|---|---|---|---|---|---|
| resolved | 2026-09-25 |
|
mesh-host e82789a (issue 135's fix) — recreating the container with a current roster ended it; verified live 2026-09-30 |
118 — umami's store answers the dial and times out the query
What was observed
The analytics service's public name has answered 502 through the whole of 2026-09-25's migration session
(first noted mid-afternoon, still true at night). The container restart-loops on a
timescale of about a minute. Its own log, every cycle:
✓ DATABASE_URL is defined.
✓ Database connection successful.
Invalid `prisma.$queryRaw()` invocation:
Raw query failed. Code: `N/A`. Message: `Operation has timed out`
The connection is established — the dial succeeds — and the first raw query then times out. This is not a credentials fault and not an unreachable store.
What it is not
- Not the routing layer: the
502is Traefik faithfully reporting a backend that is restart-looping. The stale duplicate Traefik router for this name (a HAL-era hand-authored file beside the mesh-written one) was removed the same night and changed nothing, as expected. - Not the mesh's grant machinery: the binding and sealed secret compose, and the store accepts the login — a wrong credential refuses the dial, and this dial succeeds.
Where to look
A dial that succeeds and a query that times out, from a container on one network to a store on another, has the shape of a path-MTU or conntrack fault (large response packets dropped after the small handshake ones pass), or of the store accepting the TCP connection while the backend it proxies for is wedged. Neither is proven. What is known to differ for umami against every working consumer of the same store tonight is nothing yet — that comparison is the first move.
Why it is filed rather than chased
The 2026-09-25 session's scope was routing and the build chain; this fault predates the night's changes, survived them unchanged, and needs its own sitting with the store's own logs beside the consumer's.