Answers issue 151. Copying the roster into each container made the roster part of each container's identity, so one name moving replaced every container in the mesh — and it never stopped the staleness it was for, since a copy taken at creation is stale the moment the roster moves (109, 135). A container resolves through its machine's resolver instead, and nothing is copied. Staleness stops being possible rather than detected, and a name's blast radius becomes nothing. Scoping each container to the names it binds was the close call and is rejected: it contradicts anything-calls-anything, and leaves the roster in the digest so the churn returns for a widely-bound name. Gated on issue 110 — a container on the runtime's default network has no DNS at all today. Removing the copy first reintroduces 109 and 135 silently on a live mesh. 151 stays open until the code lands; design 08's file-not-resolver passage is narrowed to the machine's own roster.
10 KiB
topic, status, date, deciders, reconstructed, extends
| topic | status | date | deciders | reconstructed | extends |
|---|---|---|---|---|---|
| the tiers | accepted | 2026-09-30 | jochen | false | 02-DECISIONS/0066-public-routing-is-name-agnostic.md |
148. The mesh's names are resolved, not copied into every container
Context
The mesh gives every container it declares the whole roster of mesh names as entries written into the container's own hosts file at creation (design 08 §2, ADR 0066). A container takes those entries once and never looks again.
Three issues are the same fact arriving three times.
A container keeps the address it was made with. Adopting the predecessor's tunnel moved the hub's private address; the declaration followed it within one push and nothing on the machine did. The forge's container held the old address, lost its database, reported healthy while its existing connections lasted, and then the public name went down (issue 109).
The same fault, four days later, undetected for five days. One container had restarted 2286 times
against a database it could no longer find, while the mesh reported the machine as doing what it was
told. Inside it, novox.internal was an address that had not existed for five days
(issue 135). Forty-eight
other containers were current, none of them corrected — each had been recreated for some other
reason and picked up the roster on the way.
135 was fixed by putting the roster into the digest the host compares a container against, so a container whose names moved is recreated like one whose image moved. That made the roster part of every container's identity, which is the third arrival:
One name moving replaces every container in the mesh. Migrating one small module on one machine took four routine actions; each changed the roster, and each replaced every container on the control node — its own store, the registry, the edge proxy, the forge, the directory, mail. The control plane was unreachable twice while its own store came back through crash recovery (issue 151). None of the replaced containers had anything to do with the module being migrated, or with its machine.
The blast radius of a name is now every container that carries the list, which is all of them. The node-by-node migration ahead adds names one module at a time — on one machine alone that is around twenty-five — and each would be a full restart of every service on the hub.
Considered Options
1. Keep the roster in every container and accept the churn. Rejected. It is not a cost that can be paid down: the mesh gets more names as it grows, and every name costs a restart of everything. A rollback costs another.
2. Scope each container's entries to the names it actually binds. A container is given the names of the things it declared a requirement on, so a name's blast radius is its consumers. Tidy, needs no new mechanism, and keeps 135's guarantee exactly.
Rejected, and this is the close one. It contradicts the standing intent that anything on the mesh can call anything on it — three cases, same machine, the private network, the public network, and no fourth. Scoping resolution to declared couplings makes a name reachable only where the mesh was told in advance that it would be wanted, and a person debugging inside a container would find names missing that exist everywhere else on the machine. It also leaves the roster in the digest, so the churn returns the moment a widely-bound name moves — smaller, not gone.
3. Resolve at lookup time through the machine's resolver, and copy nothing. Chosen.
Decision
A container resolves the mesh's names through its machine's resolver, at the moment it asks. No mesh name and no mesh address is written into a container, and none is part of a container's identity.
The three consequences that make this worth doing:
- Staleness stops being possible, rather than being detected. 109 and 135 are not bugs that were fixed; they are a shape that no longer exists. A name that moves is answered differently by the next lookup, in every container, with nothing recreated and nothing restarted.
- A name's blast radius becomes nothing. Assigning a module on one machine does not touch a container on another.
- Anything can still call anything, which option 2 gave up. The resolver answers every mesh name to every asker on the machine, exactly as it answers the machine itself.
The resolver is a machine-level process, not a container — one of the modules that is not a container at all — so a container depending on it is not the circularity it would be if the mesh's own store had to resolve a name through something the store's own runtime had to start first.
What a module declares for itself is untouched. Entries a manifest asks for are the module's own, stay in the container, and stay in its identity: they are part of what the module is, they do not move when the mesh's roster does, and the mesh does not know what they mean.
The machine's own roster file is untouched. It is a file, rewritten in place, read by processes and people; nothing restarts when it changes. It is only the copy into each container that this ends.
The order this lands in, which is not a preference
Nothing may stop copying names until resolution works from a container. Removing the copy first reintroduces 109 and 135 — silently, and on a live mesh, which is exactly how both were found.
- A container on any network can reach the resolver. Today a container on the runtime's default network asks from an address the converged filter drops, so it has no DNS at all (issue 110); and on two of four machines the resolver binds loopback only, so the runtime hands containers a public resolver instead. Both are prerequisites, not related work.
- The runtime is told which resolver to use, per machine, as a file — not per container as a creation-time argument, or the resolver's address is back in every container's identity and the problem has only got smaller.
- Then, and only then, the roster leaves the declaration and the digest.
Until step 3 the mesh keeps copying, and keeps comparing. 151 stays open until step 3 lands; it is not closed by this record, only answered by it.
How this is checked
- A container resolves a name that moved, without being recreated. Move a name the mesh serves; from a container that was running before the move and has not been touched since, the name answers with the new address. This is the one 109 and 135 would both have failed.
- A name's blast radius is nothing. Add a routed name on one machine; no container on any other machine is recreated. The apply report on each machine says nothing changed. This is 151.
- Anything calls anything. From a container on any machine, every
<node>.internalname and every routed name the mesh serves resolves — including names the module never declared a requirement on, which is the guarantee option 2 would have given up. - On every network the runtime offers. The first three hold for a container on the runtime's default network as well as one on a declared network, because the default network is the case that has no DNS today.
- No mesh name is in a container's spec. A test asserts the digest a host computes for a container does not move when the mesh's roster does, and does move when the module's own declared entries do.
Consequences
- The resolver becomes load-bearing for every container, where before it was load-bearing for the machine. This is a real cost and is accepted: a resolver that is down is a machine that cannot resolve, which is already true of the machine itself, and is a smaller event than a roster change destroying and recreating every container on the machine.
- Design 08's "a file rather than a resolver" no longer describes containers. It was written when
the mesh had no resolver and it gave the right answer then. The reasoning it rested on — every Linux
has a hosts file, no package needed — was already overtaken by names a hosts file cannot express:
service names and wildcards under
<node>.internal, which is why the resolver was built. - ADR 0066's mesh-wide propagation is kept and its mechanism changes. A routed name still reaches every asker in the mesh; it reaches them through the resolver rather than by being written into each container. The consequence 0066 records — that an internal issuer's challenge needs the routed name resolvable inside the mesh — holds unchanged and by the same means the machine already uses.
- Issue 110 stops being a container-DNS inconvenience and becomes a prerequisite for the mesh not restarting itself whenever it learns a name.
- A container started by hand gets the mesh's names too, where before only declared containers did. Design 08 drew that boundary deliberately, on the grounds that reaching into every container is what a nameserver would be for. This record accepts that consequence rather than working around it: a person debugging in a hand-started container resolving the same names as everything else is the behaviour worth having, and it is what "anything can call anything" means.