The real defect is same-node consumers announced the declared port instead of the assigned/published one; the loopback observation was a stale pre-0038 build. Fixed in mesh-control fix/same-node-provider-announced-port (c147a26) with a regression test. Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
92 lines
5.3 KiB
Markdown
92 lines
5.3 KiB
Markdown
---
|
|
status: resolved
|
|
opened: 2026-09-09
|
|
located-in: [mesh-control]
|
|
fixed-by: mesh-control — a same-node provider is announced at the port it is published on
|
|
amended-design:
|
|
---
|
|
|
|
# 038 — A provider is announced at a name its port is not bound to
|
|
|
|
## Symptom
|
|
|
|
A module that provides a `from: mesh` provision (observed with the database provider) is
|
|
announced to its consumers, by [issue 018](../018-a-provider-on-the-same-machine-was-never-announced/00-report.md)'s
|
|
fix, at the node's private-network name — the binding a consumer reads carries
|
|
`at: <node>.internal` and `serves.port: 5432`.
|
|
|
|
But the provider's container port is **published bound to loopback** (`127.0.0.1:<assigned>`),
|
|
not to the address `<node>.internal` resolves to. So every consumer dials the announced
|
|
`<node>.internal:5432`, which resolves to the node's private-network address, where **nothing is
|
|
listening** — the port is open only on `127.0.0.1`.
|
|
|
|
Observed on a node hosting the provider and several consumers:
|
|
|
|
- The consumer's binding file says `"at": "<node>.internal"`, `"serves": { "port": 5432 }`.
|
|
- Inside a consumer container, that name resolves to the node's private-network address.
|
|
- A connection test from the node: the private-network address on port 5432 is **CLOSED**; only
|
|
`127.0.0.1` (on the assigned host port) is OPEN.
|
|
- Consumers that touch the database only lazily serve a landing page and *look* healthy; consumers
|
|
that require the database at startup crash-loop — one with "acquisition timeout while waiting for
|
|
a new connection", another connecting and then timing out on its first query.
|
|
- The provider itself is healthy: a direct client on loopback answers instantly, few connections,
|
|
no locks.
|
|
|
|
## Why this matters
|
|
|
|
The announced address and the actual listener disagree, so the binding is a promise the mesh does
|
|
not keep. It is not one module's misconfiguration: it is the port-publishing step choosing a bind
|
|
address that does not match the `at` the resolver hands consumers, so it fails the same way for
|
|
**every** `from: mesh` provider with an off-node-reachable consumer — and, on a single node, for
|
|
same-node consumers too.
|
|
|
|
It hides well. The provider is up, the credential is correct, the database exists, a manual client
|
|
works — every part a person checks in isolation passes. Only a consumer that must use the provision
|
|
before it can serve anything reveals it, and it reveals it as a timeout, which reads as slowness or
|
|
load rather than "the address was never listening". A mesh that co-locates a provider with its
|
|
consumers (the ordinary small-mesh case) is exactly where it bites.
|
|
|
|
It also blocks anything that must *reach* a routed/served name from inside the mesh, not just
|
|
application traffic — see the internal-CA validation dependency noted in the connectivity design.
|
|
|
|
## Diagnosis
|
|
|
|
The symptom's first reading — "published on loopback" — was **partly a red herring**. Two things
|
|
were tangled:
|
|
|
|
1. **The real, current-code defect is a served-*port* mismatch, not a bind address.** A bare
|
|
`ports: ["5432"]` is assigned a host port and published as `"15432:5432"` — no bind IP, so on
|
|
**all interfaces**, reachable at the node's private-network address. But the port a consumer is
|
|
*told* is only re-derived from the assignment on the **cross-node** path. The **same-node** paths
|
|
(the resolver's `servedHere`, and the `here()` fallback) settle their served facts *while
|
|
resolving* — before the host port is assigned — so they carry the **declared** port (5432), not
|
|
the **assigned** one (15432). A co-located consumer is therefore announced
|
|
`<node>.internal:5432` while the provider is published on `<node>.internal:15432`, and dials a
|
|
port nothing listens on. Cross-node consumers were always fine, which is why it read as "the
|
|
small-mesh case."
|
|
|
|
2. **The `127.0.0.1:15432` seen in the running lab was a stale build.** Current code's publish step
|
|
binds all interfaces; the running instance was raised from a mesh-control predating the ADR 0038
|
|
publish rewrite. The substrate's own store *is* deliberately `127.0.0.1:5432` (a private store
|
|
must not be exposed) — correct, and not this bug.
|
|
|
|
## Fixed by
|
|
|
|
`mesh-control` branch `fix/same-node-provider-announced-port` (`c147a26`): after the host port is
|
|
assigned, same-node needs (and the `here()` fallback) are redirected through the same
|
|
provision→module→assigned-port lookup the cross-node path already uses, so a co-located consumer is
|
|
announced the port that is actually published. Idempotent (keyed by the declared port). Regression
|
|
test `TestASameNodeProviderIsAnnouncedAtThePortItIsPublishedOn` asserts the announced port equals
|
|
the published host port for a co-located provider/consumer — the next assertion after 018's (which
|
|
only checked a binding file exists); verified failing without the change.
|
|
|
|
*Not yet merged, and the running lab is additionally stale — proving it end-to-end there needs
|
|
mesh-control rebuilt and the affected consumer containers recreated.*
|
|
|
|
## Noted, not taken
|
|
|
|
Binding the assigned port to the node's private-network address specifically (rather than all
|
|
interfaces) would be defence-in-depth and would make the publish address match `at` by construction
|
|
— but it is a larger behavioural change entangled with the unenforced firewall scope
|
|
([003](../003-firewall-scope-is-read-by-no-code/00-report.md)), so it is left as an option.
|