4.0 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design |
|---|---|---|---|---|
| open | 2026-09-09 |
021 — A provider is announced at a name its port is not bound to
Symptom
A module that provides a from: mesh provision (observed with the database provider) is
announced to its consumers, by issue 018's
fix, at the node's private-network name — the binding a consumer reads carries
at: <node>.internal and serves.port: 5432.
But the provider's container port is published bound to loopback (127.0.0.1:<assigned>),
not to the address <node>.internal resolves to. So every consumer dials the announced
<node>.internal:5432, which resolves to the node's private-network address, where nothing is
listening — the port is open only on 127.0.0.1.
Observed on a node hosting the provider and several consumers:
- The consumer's binding file says
"at": "<node>.internal","serves": { "port": 5432 }. - Inside a consumer container, that name resolves to the node's private-network address.
- A connection test from the node: the private-network address on port 5432 is CLOSED; only
127.0.0.1(on the assigned host port) is OPEN. - Consumers that touch the database only lazily serve a landing page and look healthy; consumers that require the database at startup crash-loop — one with "acquisition timeout while waiting for a new connection", another connecting and then timing out on its first query.
- The provider itself is healthy: a direct client on loopback answers instantly, few connections, no locks.
Why this matters
The announced address and the actual listener disagree, so the binding is a promise the mesh does
not keep. It is not one module's misconfiguration: it is the port-publishing step choosing a bind
address that does not match the at the resolver hands consumers, so it fails the same way for
every from: mesh provider with an off-node-reachable consumer — and, on a single node, for
same-node consumers too.
It hides well. The provider is up, the credential is correct, the database exists, a manual client works — every part a person checks in isolation passes. Only a consumer that must use the provision before it can serve anything reveals it, and it reveals it as a timeout, which reads as slowness or load rather than "the address was never listening". A mesh that co-locates a provider with its consumers (the ordinary small-mesh case) is exactly where it bites.
It also blocks anything that must reach a routed/served name from inside the mesh, not just application traffic — see the internal-CA validation dependency noted in the connectivity design.
Narrowing
The loopback bind appears tied to mesh-assigned host ports, not to publishing as such. A
container that declares a bare port ("5432") — which the mesh assigns a host port for — is bound
to 127.0.0.1:<assigned>. A container that declares an explicit host mapping ("17672:15672") is
bound to 0.0.0.0:17672 and is reachable at the node's private-network address. So the defect is
in the path that assigns and binds a host port for a bare declaration, not in the general publish
step — which is also why a route to an explicitly-mapped port works while a database on an assigned
port does not.
Open questions
- Should a
from: meshprovision publish on the node's private-network address specifically, on all interfaces, or on whatever address the resolver will hand consumers asat— i.e. should the publish bind be derived from the same decision that fillsat, so the two cannot drift? - What is correct for a
from: machineprovision by contrast — is loopback right there, and is the bug simply that both are being treated the same? - How is this checked so it cannot regress: a test that resolves a co-located provider/consumer and
asserts the announced
at:portis actually connectable (not merely that a binding file exists, which is what 018 asserts)? - Does the same address/publish mismatch affect the substrate's own ports, or only module provisions?