From 74ae0609cb784344bc508a044824a14b8758c52f Mon Sep 17 00:00:00 2001 From: jochen Date: Fri, 25 Sep 2026 22:01:28 +0200 Subject: [PATCH 1/2] issue 118: umami's store answers the dial and times out the query --- .../00-report.md | 47 +++++++++++++++++++ 1 file changed, 47 insertions(+) create mode 100644 04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md diff --git a/04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md b/04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md new file mode 100644 index 0000000..e14e43c --- /dev/null +++ b/04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md @@ -0,0 +1,47 @@ +--- +status: located +opened: 2026-09-25 +located-in: [mesh-catalog modules/umami] +--- + +# 118 — umami's store answers the dial and times out the query + +## What was observed + +`umami.novox.be` has answered `502` through the whole of 2026-09-25's migration session +(first noted mid-afternoon, still true at night). The container restart-loops on a +timescale of about a minute. Its own log, every cycle: + +``` +✓ DATABASE_URL is defined. +✓ Database connection successful. +Invalid `prisma.$queryRaw()` invocation: +Raw query failed. Code: `N/A`. Message: `Operation has timed out` +``` + +The connection is established — the dial succeeds — and the first raw query then times +out. This is not a credentials fault and not an unreachable store. + +## What it is not + +- Not the routing layer: the `502` is Traefik faithfully reporting a backend that is + restart-looping. The stale duplicate Traefik router for this name (a HAL-era + hand-authored file beside the mesh-written one) was removed the same night and changed + nothing, as expected. +- Not the mesh's grant machinery: the binding and sealed secret compose, and the store + accepts the login — a wrong credential refuses the dial, and this dial succeeds. + +## Where to look + +A dial that succeeds and a query that times out, from a container on one network to a +store on another, has the shape of a path-MTU or conntrack fault (large response packets +dropped after the small handshake ones pass), or of the store accepting the TCP +connection while the backend it proxies for is wedged. Neither is proven. What is known +to differ for umami against every working consumer of the same store tonight is nothing +yet — that comparison is the first move. + +## Why it is filed rather than chased + +The 2026-09-25 session's scope was routing and the build chain; this fault predates the +night's changes, survived them unchanged, and needs its own sitting with the store's own +logs beside the consumer's. From b083790b21afb713220d4c88f9923dfbfd01f95a Mon Sep 17 00:00:00 2001 From: jochen Date: Sat, 26 Sep 2026 14:19:40 +0200 Subject: [PATCH 2/2] Issue 118: the analytics store answers the dial and times out the query MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Keeps 118: the other claimant to this number is on main as issue 119, where ADR 0112 points. Scrubbed the service's public name — this repository is public — and completed the report's frontmatter with the fixed-by and amended-design keys every other report carries. --- .../00-report.md | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md b/04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md index e14e43c..27b4801 100644 --- a/04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md +++ b/04-ISSUES/118-umamis-store-answers-the-dial-and-times-out-the-query/00-report.md @@ -2,13 +2,15 @@ status: located opened: 2026-09-25 located-in: [mesh-catalog modules/umami] +fixed-by: +amended-design: --- # 118 — umami's store answers the dial and times out the query ## What was observed -`umami.novox.be` has answered `502` through the whole of 2026-09-25's migration session +The analytics service's public name has answered `502` through the whole of 2026-09-25's migration session (first noted mid-afternoon, still true at night). The container restart-loops on a timescale of about a minute. Its own log, every cycle: