Verified on the machine: umami has zero restarts, applies all 26 migrations, reports the store up to date, and umami.novox.be answers 200 where it had answered 502 since 2026-09-25. It was the same fault as 135, filed three days earlier and diagnosed without either record noticing the other. 135's container IS umami — it is named in 135's own evidence table, holding novox.internal:10.42.0.1 after the overlay range moved. The dial appeared to succeed and the first real query timed out because the name pointed at an address that no longer existed; the store, the path and the credential were all fine. 118's own reasoning is kept as a warning, because it is careful and wrong: it argued path-MTU and conntrack, and it named the move that would have found it — compare umami against a container created that week — and did not make it. A stale name presents as a network fault, which is why ADR 0148 stops copying names into containers rather than detecting when the copies go bad. Also: 118's located-in named mesh-catalog modules/umami, which was never at fault; it names mesh-host's comparison now. And 135 gains the pointer to 118, which is the back-reference I have now missed three times.
50 lines
2.2 KiB
Markdown
50 lines
2.2 KiB
Markdown
---
|
|
status: resolved
|
|
opened: 2026-09-25
|
|
located-in: [mesh-host internal/apply (the spec comparison, via issue 135) — not mesh-catalog modules/umami, which this record first named and which was never at fault]
|
|
fixed-by: mesh-host e82789a (issue 135's fix) — recreating the container with a current roster ended it; verified live 2026-09-30
|
|
amended-design:
|
|
---
|
|
|
|
# 118 — umami's store answers the dial and times out the query
|
|
|
|
## What was observed
|
|
|
|
The analytics service's public name has answered `502` through the whole of 2026-09-25's migration session
|
|
(first noted mid-afternoon, still true at night). The container restart-loops on a
|
|
timescale of about a minute. Its own log, every cycle:
|
|
|
|
```
|
|
✓ DATABASE_URL is defined.
|
|
✓ Database connection successful.
|
|
Invalid `prisma.$queryRaw()` invocation:
|
|
Raw query failed. Code: `N/A`. Message: `Operation has timed out`
|
|
```
|
|
|
|
The connection is established — the dial succeeds — and the first raw query then times
|
|
out. This is not a credentials fault and not an unreachable store.
|
|
|
|
## What it is not
|
|
|
|
- Not the routing layer: the `502` is Traefik faithfully reporting a backend that is
|
|
restart-looping. The stale duplicate Traefik router for this name (a HAL-era
|
|
hand-authored file beside the mesh-written one) was removed the same night and changed
|
|
nothing, as expected.
|
|
- Not the mesh's grant machinery: the binding and sealed secret compose, and the store
|
|
accepts the login — a wrong credential refuses the dial, and this dial succeeds.
|
|
|
|
## Where to look
|
|
|
|
A dial that succeeds and a query that times out, from a container on one network to a
|
|
store on another, has the shape of a path-MTU or conntrack fault (large response packets
|
|
dropped after the small handshake ones pass), or of the store accepting the TCP
|
|
connection while the backend it proxies for is wedged. Neither is proven. What is known
|
|
to differ for umami against every working consumer of the same store tonight is nothing
|
|
yet — that comparison is the first move.
|
|
|
|
## Why it is filed rather than chased
|
|
|
|
The 2026-09-25 session's scope was routing and the build chain; this fault predates the
|
|
night's changes, survived them unchanged, and needs its own sitting with the store's own
|
|
logs beside the consumer's.
|