Issue 110's cause was not the filter: the runtime had never been told, and the resolver dropped a query arriving on a bridge. ADR 0148 step 3 landed once it did (109, 151 resolved). ADR 0151 composes a route's internal name under the serving node and drops the suffixed alias (139, 157 resolved). Design 08 amended; a fact in 0148 corrected.
3.5 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design | |
|---|---|---|---|---|---|
| resolved | 2026-09-28 |
|
mesh-controller PR 163 — the internal name composes under the node whose proxy serves the route (ADR 0151, 2026-09-30) | 03-DESIGN/01-to-be/08-connectivity.md |
139 — An internal route name resolves to the consumer's node, not the one that serves it
What was observed
A module that requires a route is given two names: a public one composed under the serving node's
domain, and an internal one composed under the consumer's own machine — <label>.<node>.internal.
The two are published differently:
- The public name is written into every machine's hosts file at the address of the node whose proxy answers it. The mesh computes that deliberately, so any container resolving a routed name reaches the proxy.
- The internal name is resolved by the machine's own resolver, which answers every name under
<node>.internalwith that node's address — the consumer's, because the name was composed from it.
Where the proxy runs beside the consumer these are the same machine, which is every case on this mesh today, and both names work. Measured on 2026-09-28: the internal name of a service on the control node answers with a certificate from the mesh's internal authority, and the public name with one from the public authority.
Where the proxy is on another machine they disagree. The internal name sends the client to a machine that runs no proxy and has nothing listening on the port, while the public name sends it to the one that does.
Why it matters beyond this instance
It is latent exactly where the mesh is heading. route is provided mesh-wide precisely so a
module can be routed by a proxy on another machine. The first module assigned that way gets an
internal name that does not work, and the public one that does — with no error anywhere, because both
names resolve.
A per-machine name is what an operator will reach for. <service>.<machine>.internal reads like a
promise that the service on that machine is reachable there, and the wildcard makes every such name
resolve whether or not anything answers.
Open questions
- Should the internal name be composed under the serving node, like the public one, or should it stay the consumer's and be published at the serving node's address like the public name is?
- Is a per-node route holder the real answer — a proxy on every machine that serves its own names —
and if so, is
routestill one mesh-wide provision or a node-scoped seat with a mesh-wide fallback? - What certifies the name in either case? The certificate is obtained by whoever terminates TLS, and that is the question above in another form.
Answered (2026-09-30)
ADR 0151:
the internal name is composed under the node that serves the route — the machine the request arrives
at — because <x>.<node>.internal means goes to that node and nothing else. The public name stays
the consumer node's, which is where the operator put it. The first question is answered that way; the
second, a per-node route holder, is a decision about seats and is left where it is; the third is
unchanged, since the proxy that terminates the name is given it and certifies it.
mesh-controller PR 163 carries it. On this mesh every route is served beside its module, so no name changed; the controller's tests hold the case where it would.