Group 2 is resolved: containers resolve, nothing is copied, a route's name says where it arrives

Issue 110's cause was not the filter: the runtime had never been told,
and the resolver dropped a query arriving on a bridge. ADR 0148 step 3
landed once it did (109, 151 resolved). ADR 0151 composes a route's
internal name under the serving node and drops the suffixed alias
(139, 157 resolved). Design 08 amended; a fact in 0148 corrected.
This commit is contained in:
2026-09-30 14:56:43 +02:00
parent 846c1f85f2
commit 04c9500b5b
10 changed files with 247 additions and 19 deletions
@@ -1,8 +1,8 @@
---
status: located
status: resolved
opened: 2026-09-24
located-in: [mesh-host internal/apply]
fixed-by:
located-in: [mesh-controller internal/catalogue/declaration.go (every container was given the roster at creation)]
fixed-by: mesh-controller PR 161 — no container is given a mesh name; it resolves through its machine's resolver (ADR 0148, landed 2026-09-30 once issue 110 did)
amended-design:
---
@@ -73,3 +73,13 @@ copying: a container resolves through its machine's resolver at the moment it as
record reports then has nowhere to occur. It is gated on
[issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md), so
until that lands the mesh still copies and still compares.
## Resolved (2026-09-30)
110 landed the same day ([its resolution](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/01-resolution.md)),
and mesh-controller PR 161 then removed the copy: no container is given a mesh name or a mesh address,
and a module's own declared entries are the only `host` lines it carries. Verified on the control-node
after its containers were recreated once — the last time a name will do that: the forge's container
carries no extra hosts and resolves another machine and a routed name through the machine's resolver,
so the shape this record describes has nowhere to occur. Checked in the controller's tests: a
container's declaration is byte-for-byte the same under a roster of one machine and a roster of three.
@@ -1,8 +1,9 @@
---
status: open
status: resolved
opened: 2026-09-24
located-in: []
fixed-by:
located-in:
- mesh-catalog modules/dnsmasq (the runtime was never told; the resolver answered by interface)
fixed-by: mesh-catalog PR 175 (the runtime is reloaded and keeps its containers over a restart) and PR 176 (the resolver answers by address, so a query from a bridge is admitted) — measured 2026-09-30, 01-resolution.md
amended-design:
---
@@ -67,3 +68,6 @@ and [135](../135-a-containers-mesh-names-are-not-compared/00-report.md)).
network is the case with no DNS at all, and it is the case the mesh's own forge runs in. Two of four
machines also bind the resolver to loopback only, so the runtime hands their containers a public
resolver. Both halves are this issue.
*Later the same day: the second half was wrong, and the first had a different cause than the one above.
[01-resolution.md](01-resolution.md) has what was actually found.*
@@ -0,0 +1,64 @@
# 110 — resolved: a container on any network reaches the resolver, and is answered
*2026-09-30. Measured on the three converged machines; the adopted one holds its resolver module until it
is taken and is not covered.*
## What was actually wrong
Not what the report predicted. The report named the filter: a container on the runtime's default
network asks from a bridge address, and the converged filter admitted queries by source address only.
That was true when it was written and was fixed before this issue was ever tested — the filter admits
by the link a packet arrives on ([ADR 0144](../../02-DECISIONS/0144-anything-on-a-machine-may-call-anything-on-it.md)),
and a container's bridge is admitted whole. Tested on every machine: the query arrives, the filter
passes it.
Three other things were wrong, each hiding the next.
**The runtime had never been told.** The resolver module writes the runtime's `dns` key into the
runtime's own configuration file. The runtime reads that key when it starts and not on a reload, and on
two machines the runtime predated the file — so every container they started got a public resolver, and
`novox.internal` came back as not existing. Nothing reported this: the file was present and current,
the resolver ran, and a name not existing is a valid answer. Fixed in mesh-catalog PR 175: the module
also sets `live-restore` and reloads the runtime when its file changes, so the one restart the `dns` key
needs no longer stops every container. The restart is then the operator's, once per machine; done on
both today, with every running container kept.
**The resolver dropped the query.** With the runtime corrected, a container's query reached the resolver
— and got no answer, on every machine, including the one whose runtime had been right all along. The
socket was bound to the private address; the filter admitted the packet; dnsmasq received it and
discarded it without a line of log. Its configuration said `interface=mesh0`, and dnsmasq admits a
query by the interface it arrives on when told an interface: a container's query is addressed to the
private address but arrives on the runtime's bridge, and the bridge is not `mesh0`. Fixed in mesh-catalog
PR 176: the resolver is told the address to answer on, not the interface that carries it, and a query to
that address is admitted whatever bridge brings it. The bridges are the runtime's to name.
**The report's second half was wrong.** "Two of four machines bind the resolver to loopback only" was
an inference from the containers' behaviour, and the behaviour had the cause above. The resolver bound
the private address on all four; nothing had asked it there.
## What is verified
From a container on the runtime's default network, started by hand and given nothing, on each of the
three converged machines: `novox.internal` answers with the hub's private address, through the machine's
own resolver. That is the fourth check of
[ADR 0148](../../02-DECISIONS/0148-the-meshs-names-are-resolved-not-copied-into-containers.md) — "on every
network the runtime offers" — and its first step; the record's step 2 (the runtime told per machine, as
a file) was already how the module works. Step 3 may now begin.
## What checks it
By hand, today. Nothing in the mesh asserts that a container can resolve a mesh name: the resolver's
own tests cover what it answers, not who can ask. The check that would have caught all three faults is
the one the report asked for and 0148 lists — a container on the default network resolving a mesh name
— and it is not built. It belongs with the reachability check of
[issue 145](../145-a-machine-reads-healthy-while-its-modules-cannot-reach-each-other/00-report.md),
which is parked; until then this is a thing a person verifies after touching the resolver, the filter,
or the runtime's configuration.
## What this cost to find
The three faults produced one symptom — a container that cannot resolve — and each fix revealed the
next. The first was found by reading the runtime's own view of its configuration rather than the file;
the second by capturing the query on the bridge and finding it arrive and go unanswered; the third only
by admitting the first belief was wrong. A machine that had been believed to work all day had never
worked either.
@@ -1,9 +1,9 @@
---
status: open
status: resolved
opened: 2026-09-28
located-in: [mesh-controller internal/catalogue]
fixed-by:
amended-design:
located-in: [mesh-controller internal/catalogue/declaration.go (composeName took the consumer's own name as the internal domain)]
fixed-by: mesh-controller PR 163 — the internal name composes under the node whose proxy serves the route (ADR 0151, 2026-09-30)
amended-design: 03-DESIGN/01-to-be/08-connectivity.md
---
# 139 — An internal route name resolves to the consumer's node, not the one that serves it
@@ -49,3 +49,15 @@ resolve whether or not anything answers.
and if so, is `route` still one mesh-wide provision or a node-scoped seat with a mesh-wide fallback?
- What certifies the name in either case? The certificate is obtained by whoever terminates TLS, and
that is the question above in another form.
## Answered (2026-09-30)
[ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md):
the internal name is composed under the node that serves the route — the machine the request arrives
at — because `<x>.<node>.internal` means *goes to that node* and nothing else. The public name stays
the consumer node's, which is where the operator put it. The first question is answered that way; the
second, a per-node route holder, is a decision about seats and is left where it is; the third is
unchanged, since the proxy that terminates the name is given it and certifies it.
mesh-controller PR 163 carries it. On this mesh every route is served beside its module, so no name
changed; the controller's tests hold the case where it would.
@@ -1,10 +1,10 @@
---
status: open
status: resolved
opened: 2026-09-29
located-in:
- mesh-host internal/apply/apply.go (containerSpecReading hashes every `host` entry)
- mesh-controller internal/catalogue/declaration.go (withMeshNames gives every container the mesh's names)
fixed-by:
fixed-by: mesh-controller PR 161 — the roster left every container's declaration and so its digest (ADR 0148 step 3, 2026-09-30)
---
# 151 — A new name recreates every container in the mesh
@@ -92,3 +92,15 @@ copying names until a container can reach the resolver from any of the runtime's
([issue 110](../110-a-container-on-the-runtimes-own-network-cannot-reach-the-resolver/00-report.md)),
which it cannot on two of four machines today. Removing the copy first reintroduces 109 and 135
silently, on a live mesh, which is how both were found. The order is in the record.
## Resolved (2026-09-30)
Step 3 landed the day 110 did. mesh-controller PR 161 stops writing the roster into any container, so
a container's digest no longer carries a name that is not its own. The controller's tests hold the
record's check — a container's declaration does not move when the mesh's roster does, and does move
when the module's own declared entries do.
The first push after the change recreated every container once, because every digest lost its host
entries at the same moment. That was the last such event: from here a name added or moved on one
machine changes no container anywhere, and the record's second check — add a routed name, watch every
other machine's apply report say nothing changed — is what the next module assignment will show.
@@ -1,9 +1,9 @@
---
status: located
status: resolved
opened: 2026-09-30
located-in:
- mesh-controller internal/catalogue/roster.go (the roster's entries for a routed name)
fixed-by:
fixed-by: mesh-controller PR 163 — a routed name is published as itself, once, with no suffixed alias (ADR 0151, 2026-09-30)
amended-design:
---
@@ -61,3 +61,10 @@ composition produces `keycloak.novox.be.internal`, which is not a name anything
The fix is a judgement about what a routed name's internal form is, and 139 is the record that asks it;
this one is the evidence that the current answer publishes a third thing that is neither.
## Resolved (2026-09-30)
The judgement 139 asked for is [ADR 0151](../../02-DECISIONS/0151-a-routes-internal-name-is-composed-under-the-node-that-serves-it.md):
a routed name has no mesh form. The roster now publishes it as itself, once, at the serving node's
address; the `<domain>.internal` line is gone from every machine's hosts file, and a controller test
refuses it coming back.