Files
hq/04-ISSUES/038-a-provider-is-announced-at-a-name-its-port-is-not-bound-to/00-report.md
T
jschoubben a4440c9acf Issue 038 resolved — corrected root cause (same-node served-port mismatch, not loopback) + fix reference
The real defect is same-node consumers announced the declared port instead of
the assigned/published one; the loopback observation was a stale pre-0038 build.
Fixed in mesh-control fix/same-node-provider-announced-port (c147a26) with a
regression test.

Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
2026-09-09 23:06:58 +02:00

5.3 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
resolved 2026-09-09
mesh-control
mesh-control — a same-node provider is announced at the port it is published on

038 — A provider is announced at a name its port is not bound to

Symptom

A module that provides a from: mesh provision (observed with the database provider) is announced to its consumers, by issue 018's fix, at the node's private-network name — the binding a consumer reads carries at: <node>.internal and serves.port: 5432.

But the provider's container port is published bound to loopback (127.0.0.1:<assigned>), not to the address <node>.internal resolves to. So every consumer dials the announced <node>.internal:5432, which resolves to the node's private-network address, where nothing is listening — the port is open only on 127.0.0.1.

Observed on a node hosting the provider and several consumers:

  • The consumer's binding file says "at": "<node>.internal", "serves": { "port": 5432 }.
  • Inside a consumer container, that name resolves to the node's private-network address.
  • A connection test from the node: the private-network address on port 5432 is CLOSED; only 127.0.0.1 (on the assigned host port) is OPEN.
  • Consumers that touch the database only lazily serve a landing page and look healthy; consumers that require the database at startup crash-loop — one with "acquisition timeout while waiting for a new connection", another connecting and then timing out on its first query.
  • The provider itself is healthy: a direct client on loopback answers instantly, few connections, no locks.

Why this matters

The announced address and the actual listener disagree, so the binding is a promise the mesh does not keep. It is not one module's misconfiguration: it is the port-publishing step choosing a bind address that does not match the at the resolver hands consumers, so it fails the same way for every from: mesh provider with an off-node-reachable consumer — and, on a single node, for same-node consumers too.

It hides well. The provider is up, the credential is correct, the database exists, a manual client works — every part a person checks in isolation passes. Only a consumer that must use the provision before it can serve anything reveals it, and it reveals it as a timeout, which reads as slowness or load rather than "the address was never listening". A mesh that co-locates a provider with its consumers (the ordinary small-mesh case) is exactly where it bites.

It also blocks anything that must reach a routed/served name from inside the mesh, not just application traffic — see the internal-CA validation dependency noted in the connectivity design.

Diagnosis

The symptom's first reading — "published on loopback" — was partly a red herring. Two things were tangled:

  1. The real, current-code defect is a served-port mismatch, not a bind address. A bare ports: ["5432"] is assigned a host port and published as "15432:5432" — no bind IP, so on all interfaces, reachable at the node's private-network address. But the port a consumer is told is only re-derived from the assignment on the cross-node path. The same-node paths (the resolver's servedHere, and the here() fallback) settle their served facts while resolving — before the host port is assigned — so they carry the declared port (5432), not the assigned one (15432). A co-located consumer is therefore announced <node>.internal:5432 while the provider is published on <node>.internal:15432, and dials a port nothing listens on. Cross-node consumers were always fine, which is why it read as "the small-mesh case."

  2. The 127.0.0.1:15432 seen in the running lab was a stale build. Current code's publish step binds all interfaces; the running instance was raised from a mesh-control predating the ADR 0038 publish rewrite. The substrate's own store is deliberately 127.0.0.1:5432 (a private store must not be exposed) — correct, and not this bug.

Fixed by

mesh-control branch fix/same-node-provider-announced-port (c147a26): after the host port is assigned, same-node needs (and the here() fallback) are redirected through the same provision→module→assigned-port lookup the cross-node path already uses, so a co-located consumer is announced the port that is actually published. Idempotent (keyed by the declared port). Regression test TestASameNodeProviderIsAnnouncedAtThePortItIsPublishedOn asserts the announced port equals the published host port for a co-located provider/consumer — the next assertion after 018's (which only checked a binding file exists); verified failing without the change.

Not yet merged, and the running lab is additionally stale — proving it end-to-end there needs mesh-control rebuilt and the affected consumer containers recreated.

Noted, not taken

Binding the assigned port to the node's private-network address specifically (rather than all interfaces) would be defence-in-depth and would make the publish address match at by construction — but it is a larger behavioural change entangled with the unenforced firewall scope (003), so it is left as an option.