From 315bd7118a3d121870d9d0919fda444dfbc7f4ad Mon Sep 17 00:00:00 2001 From: jochen Date: Wed, 9 Sep 2026 22:29:04 +0200 Subject: [PATCH] =?UTF-8?q?Issue=20021=20=E2=80=94=20a=20provider=20is=20a?= =?UTF-8?q?nnounced=20at=20a=20name=20its=20port=20is=20not=20bound=20to?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A from:mesh provider is announced (per 018's fix) at the node's private-network name, but its port is published bound to loopback, so consumers dialing the announced .internal:port reach nothing. Diagnosed from the lab: DB consumers that require the database at startup crash-loop; the provider is healthy on loopback only. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF --- .../00-report.md | 62 +++++++++++++++++++ 1 file changed, 62 insertions(+) create mode 100644 04-ISSUES/021-a-provider-is-announced-at-a-name-its-port-is-not-bound-to/00-report.md diff --git a/04-ISSUES/021-a-provider-is-announced-at-a-name-its-port-is-not-bound-to/00-report.md b/04-ISSUES/021-a-provider-is-announced-at-a-name-its-port-is-not-bound-to/00-report.md new file mode 100644 index 0000000..f14bdce --- /dev/null +++ b/04-ISSUES/021-a-provider-is-announced-at-a-name-its-port-is-not-bound-to/00-report.md @@ -0,0 +1,62 @@ +--- +status: open +opened: 2026-09-09 +located-in: [] +fixed-by: +amended-design: +--- + +# 021 — A provider is announced at a name its port is not bound to + +## Symptom + +A module that provides a `from: mesh` provision (observed with the database provider) is +announced to its consumers, by [issue 018](../018-a-provider-on-the-same-machine-was-never-announced/00-report.md)'s +fix, at the node's private-network name — the binding a consumer reads carries +`at: .internal` and `serves.port: 5432`. + +But the provider's container port is **published bound to loopback** (`127.0.0.1:`), +not to the address `.internal` resolves to. So every consumer dials the announced +`.internal:5432`, which resolves to the node's private-network address, where **nothing is +listening** — the port is open only on `127.0.0.1`. + +Observed on a node hosting the provider and several consumers: + +- The consumer's binding file says `"at": ".internal"`, `"serves": { "port": 5432 }`. +- Inside a consumer container, that name resolves to the node's private-network address. +- A connection test from the node: the private-network address on port 5432 is **CLOSED**; only + `127.0.0.1` (on the assigned host port) is OPEN. +- Consumers that touch the database only lazily serve a landing page and *look* healthy; consumers + that require the database at startup crash-loop — one with "acquisition timeout while waiting for + a new connection", another connecting and then timing out on its first query. +- The provider itself is healthy: a direct client on loopback answers instantly, few connections, + no locks. + +## Why this matters + +The announced address and the actual listener disagree, so the binding is a promise the mesh does +not keep. It is not one module's misconfiguration: it is the port-publishing step choosing a bind +address that does not match the `at` the resolver hands consumers, so it fails the same way for +**every** `from: mesh` provider with an off-node-reachable consumer — and, on a single node, for +same-node consumers too. + +It hides well. The provider is up, the credential is correct, the database exists, a manual client +works — every part a person checks in isolation passes. Only a consumer that must use the provision +before it can serve anything reveals it, and it reveals it as a timeout, which reads as slowness or +load rather than "the address was never listening". A mesh that co-locates a provider with its +consumers (the ordinary small-mesh case) is exactly where it bites. + +It also blocks anything that must *reach* a routed/served name from inside the mesh, not just +application traffic — see the internal-CA validation dependency noted in the connectivity design. + +## Open questions + +- Should a `from: mesh` provision publish on the node's private-network address specifically, on + all interfaces, or on whatever address the resolver will hand consumers as `at` — i.e. should the + publish bind be derived from the *same* decision that fills `at`, so the two cannot drift? +- What is correct for a `from: machine` provision by contrast — is loopback right there, and is the + bug simply that both are being treated the same? +- How is this checked so it cannot regress: a test that resolves a co-located provider/consumer and + asserts the announced `at:port` is actually connectable (not merely that a binding file exists, which + is what 018 asserts)? +- Does the same address/publish mismatch affect the substrate's own ports, or only module provisions?