Answers issue 151. Copying the roster into each container made the roster part of each container's identity, so one name moving replaced every container in the mesh — and it never stopped the staleness it was for, since a copy taken at creation is stale the moment the roster moves (109, 135). A container resolves through its machine's resolver instead, and nothing is copied. Staleness stops being possible rather than detected, and a name's blast radius becomes nothing. Scoping each container to the names it binds was the close call and is rejected: it contradicts anything-calls-anything, and leaves the roster in the digest so the churn returns for a widely-bound name. Gated on issue 110 — a container on the runtime's default network has no DNS at all today. Removing the copy first reintroduces 109 and 135 silently on a live mesh. 151 stays open until the code lands; design 08's file-not-resolver passage is narrowed to the machine's own roster.
3.6 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design |
|---|---|---|---|---|
| open | 2026-09-24 |
110 — On a converged node, a container on the runtime's own network cannot reach the resolver
What was observed
Reviewing the resolver's conversion from the predecessor's, 2026-09-24, before it is assigned anywhere.
The resolver answers on the private network's interface and on loopback, and it declares that it listens from the mesh — so the filter a converged node loads admits queries whose source is a private-network address. A container asks in one of two ways:
- on a network the module declared, the runtime answers from the container's own namespace and forwards to the resolver from the machine itself, which the filter admits;
- on the runtime's default network, the container is handed the resolver's address directly and asks from its own address on that network — which is not a private-network address, and the filter drops it.
So on a converged node a container on the default network has no DNS. It is not hypothetical: the forge's container on this machine is on the default network, and the service beside it is not — which is exactly why one survived the hub's address change and the other did not (issue 109).
Nothing fails today, because the node is adopted and the predecessor's firewall is still in force. It fails at the flip.
Why it matters beyond this instance
Two of the mesh's answers depend on this working. The resolver exists so that a machine and its containers resolve the mesh's names; and issue 109 asks whether a container should be given no address at all and always ask the resolver. That answer is only available if every container can reach it.
It also means the flip is not as previewed. Converging lists the ports it will close; it does not say "and the containers on the runtime's default network will stop resolving names", because nothing knows that is what the rule means.
Open questions
- Should the resolver's
listenssay it is reachable from the container runtime's own networks — the same exception the mesh's guard already makes for them, by the interface a packet arrives on rather than by its source address? - Or should every module be required to declare a network, so no container of the mesh's is ever on the runtime's default one? That is a stronger rule and would have prevented 109 as well.
- What checks it? A converged bed with a container on the default network resolving a mesh name is the missing assertion; nothing in the resolver's own beds covers the filter.
What now depends on this (2026-09-30)
This stopped being a container-DNS inconvenience. ADR 0148 decides that a container resolves the mesh's names rather than being given a copy of them, which is what stops one name moving from replacing every container in the mesh (issue 151) and what makes a stale address impossible rather than merely noticed (issues 109 and 135).
That decision cannot land until this one does, and not partly: a container on the runtime's default network is the case with no DNS at all, and it is the case the mesh's own forge runs in. Two of four machines also bind the resolver to loopback only, so the runtime hands their containers a public resolver. Both halves are this issue.