Files
hq/04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md
T
jschoubben 295bf2f3ab Issue 118 is resolved: it was issue 135, and umami is healthy
Verified on the machine: umami has zero restarts, applies all 26
migrations, reports the store up to date, and umami.novox.be answers 200
where it had answered 502 since 2026-09-25.

It was the same fault as 135, filed three days earlier and diagnosed
without either record noticing the other. 135's container IS umami — it
is named in 135's own evidence table, holding novox.internal:10.42.0.1
after the overlay range moved. The dial appeared to succeed and the first
real query timed out because the name pointed at an address that no
longer existed; the store, the path and the credential were all fine.

118's own reasoning is kept as a warning, because it is careful and
wrong: it argued path-MTU and conntrack, and it named the move that would
have found it — compare umami against a container created that week —
and did not make it. A stale name presents as a network fault, which is
why ADR 0148 stops copying names into containers rather than detecting
when the copies go bad.

Also: 118's located-in named mesh-catalog modules/umami, which was never
at fault; it names mesh-host's comparison now. And 135 gains the pointer
to 118, which is the back-reference I have now missed three times.
2026-09-30 01:54:19 +02:00

2.2 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
resolved 2026-09-25
mesh-host internal/apply (the spec comparison
via issue 135) — not mesh-catalog modules/umami
which this record first named and which was never at fault
mesh-host e82789a (issue 135's fix) — recreating the container with a current roster ended it; verified live 2026-09-30

118 — umami's store answers the dial and times out the query

What was observed

The analytics service's public name has answered 502 through the whole of 2026-09-25's migration session (first noted mid-afternoon, still true at night). The container restart-loops on a timescale of about a minute. Its own log, every cycle:

✓ DATABASE_URL is defined.
✓ Database connection successful.
Invalid `prisma.$queryRaw()` invocation:
Raw query failed. Code: `N/A`. Message: `Operation has timed out`

The connection is established — the dial succeeds — and the first raw query then times out. This is not a credentials fault and not an unreachable store.

What it is not

  • Not the routing layer: the 502 is Traefik faithfully reporting a backend that is restart-looping. The stale duplicate Traefik router for this name (a HAL-era hand-authored file beside the mesh-written one) was removed the same night and changed nothing, as expected.
  • Not the mesh's grant machinery: the binding and sealed secret compose, and the store accepts the login — a wrong credential refuses the dial, and this dial succeeds.

Where to look

A dial that succeeds and a query that times out, from a container on one network to a store on another, has the shape of a path-MTU or conntrack fault (large response packets dropped after the small handshake ones pass), or of the store accepting the TCP connection while the backend it proxies for is wedged. Neither is proven. What is known to differ for umami against every working consumer of the same store tonight is nothing yet — that comparison is the first move.

Why it is filed rather than chased

The 2026-09-25 session's scope was routing and the build chain; this fault predates the night's changes, survived them unchanged, and needs its own sitting with the store's own logs beside the consumer's.