Issue 021 — a provider is announced at a name its port is not bound to
A from:mesh provider is announced (per 018's fix) at the node's private-network name, but its port is published bound to loopback, so consumers dialing the announced <node>.internal:port reach nothing. Diagnosed from the lab: DB consumers that require the database at startup crash-loop; the provider is healthy on loopback only. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
This commit is contained in:
@@ -0,0 +1,62 @@
|
|||||||
|
---
|
||||||
|
status: open
|
||||||
|
opened: 2026-09-09
|
||||||
|
located-in: []
|
||||||
|
fixed-by:
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 021 — A provider is announced at a name its port is not bound to
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
A module that provides a `from: mesh` provision (observed with the database provider) is
|
||||||
|
announced to its consumers, by [issue 018](../018-a-provider-on-the-same-machine-was-never-announced/00-report.md)'s
|
||||||
|
fix, at the node's private-network name — the binding a consumer reads carries
|
||||||
|
`at: <node>.internal` and `serves.port: 5432`.
|
||||||
|
|
||||||
|
But the provider's container port is **published bound to loopback** (`127.0.0.1:<assigned>`),
|
||||||
|
not to the address `<node>.internal` resolves to. So every consumer dials the announced
|
||||||
|
`<node>.internal:5432`, which resolves to the node's private-network address, where **nothing is
|
||||||
|
listening** — the port is open only on `127.0.0.1`.
|
||||||
|
|
||||||
|
Observed on a node hosting the provider and several consumers:
|
||||||
|
|
||||||
|
- The consumer's binding file says `"at": "<node>.internal"`, `"serves": { "port": 5432 }`.
|
||||||
|
- Inside a consumer container, that name resolves to the node's private-network address.
|
||||||
|
- A connection test from the node: the private-network address on port 5432 is **CLOSED**; only
|
||||||
|
`127.0.0.1` (on the assigned host port) is OPEN.
|
||||||
|
- Consumers that touch the database only lazily serve a landing page and *look* healthy; consumers
|
||||||
|
that require the database at startup crash-loop — one with "acquisition timeout while waiting for
|
||||||
|
a new connection", another connecting and then timing out on its first query.
|
||||||
|
- The provider itself is healthy: a direct client on loopback answers instantly, few connections,
|
||||||
|
no locks.
|
||||||
|
|
||||||
|
## Why this matters
|
||||||
|
|
||||||
|
The announced address and the actual listener disagree, so the binding is a promise the mesh does
|
||||||
|
not keep. It is not one module's misconfiguration: it is the port-publishing step choosing a bind
|
||||||
|
address that does not match the `at` the resolver hands consumers, so it fails the same way for
|
||||||
|
**every** `from: mesh` provider with an off-node-reachable consumer — and, on a single node, for
|
||||||
|
same-node consumers too.
|
||||||
|
|
||||||
|
It hides well. The provider is up, the credential is correct, the database exists, a manual client
|
||||||
|
works — every part a person checks in isolation passes. Only a consumer that must use the provision
|
||||||
|
before it can serve anything reveals it, and it reveals it as a timeout, which reads as slowness or
|
||||||
|
load rather than "the address was never listening". A mesh that co-locates a provider with its
|
||||||
|
consumers (the ordinary small-mesh case) is exactly where it bites.
|
||||||
|
|
||||||
|
It also blocks anything that must *reach* a routed/served name from inside the mesh, not just
|
||||||
|
application traffic — see the internal-CA validation dependency noted in the connectivity design.
|
||||||
|
|
||||||
|
## Open questions
|
||||||
|
|
||||||
|
- Should a `from: mesh` provision publish on the node's private-network address specifically, on
|
||||||
|
all interfaces, or on whatever address the resolver will hand consumers as `at` — i.e. should the
|
||||||
|
publish bind be derived from the *same* decision that fills `at`, so the two cannot drift?
|
||||||
|
- What is correct for a `from: machine` provision by contrast — is loopback right there, and is the
|
||||||
|
bug simply that both are being treated the same?
|
||||||
|
- How is this checked so it cannot regress: a test that resolves a co-located provider/consumer and
|
||||||
|
asserts the announced `at:port` is actually connectable (not merely that a binding file exists, which
|
||||||
|
is what 018 asserts)?
|
||||||
|
- Does the same address/publish mismatch affect the substrate's own ports, or only module provisions?
|
||||||
Reference in New Issue
Block a user