Files
hq/04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md
T
jschoubben 295bf2f3ab Issue 118 is resolved: it was issue 135, and umami is healthy
Verified on the machine: umami has zero restarts, applies all 26
migrations, reports the store up to date, and umami.novox.be answers 200
where it had answered 502 since 2026-09-25.

It was the same fault as 135, filed three days earlier and diagnosed
without either record noticing the other. 135's container IS umami — it
is named in 135's own evidence table, holding novox.internal:10.42.0.1
after the overlay range moved. The dial appeared to succeed and the first
real query timed out because the name pointed at an address that no
longer existed; the store, the path and the credential were all fine.

118's own reasoning is kept as a warning, because it is careful and
wrong: it argued path-MTU and conntrack, and it named the move that would
have found it — compare umami against a container created that week —
and did not make it. A stale name presents as a network fault, which is
why ADR 0148 stops copying names into containers rather than detecting
when the copies go bad.

Also: 118's located-in named mesh-catalog modules/umami, which was never
at fault; it names mesh-host's comparison now. And 135 gains the pointer
to 118, which is the back-reference I have now missed three times.
2026-09-30 01:54:19 +02:00

50 lines
2.2 KiB
Markdown

---
status: resolved
opened: 2026-09-25
located-in: [mesh-host internal/apply (the spec comparison, via issue 135) — not mesh-catalog modules/umami, which this record first named and which was never at fault]
fixed-by: mesh-host e82789a (issue 135's fix) — recreating the container with a current roster ended it; verified live 2026-09-30
amended-design:
---
# 118 — umami's store answers the dial and times out the query
## What was observed
The analytics service's public name has answered `502` through the whole of 2026-09-25's migration session
(first noted mid-afternoon, still true at night). The container restart-loops on a
timescale of about a minute. Its own log, every cycle:
```
✓ DATABASE_URL is defined.
✓ Database connection successful.
Invalid `prisma.$queryRaw()` invocation:
Raw query failed. Code: `N/A`. Message: `Operation has timed out`
```
The connection is established — the dial succeeds — and the first raw query then times
out. This is not a credentials fault and not an unreachable store.
## What it is not
- Not the routing layer: the `502` is Traefik faithfully reporting a backend that is
restart-looping. The stale duplicate Traefik router for this name (a HAL-era
hand-authored file beside the mesh-written one) was removed the same night and changed
nothing, as expected.
- Not the mesh's grant machinery: the binding and sealed secret compose, and the store
accepts the login — a wrong credential refuses the dial, and this dial succeeds.
## Where to look
A dial that succeeds and a query that times out, from a container on one network to a
store on another, has the shape of a path-MTU or conntrack fault (large response packets
dropped after the small handshake ones pass), or of the store accepting the TCP
connection while the backend it proxies for is wedged. Neither is proven. What is known
to differ for umami against every working consumer of the same store tonight is nothing
yet — that comparison is the first move.
## Why it is filed rather than chased
The 2026-09-25 session's scope was routing and the build chain; this fault predates the
night's changes, survived them unchanged, and needs its own sitting with the store's own
logs beside the consumer's.