ADR 0194 rejected sending every query to the mesh's resolver because a node with its tunnel down would resolve nothing; a public resolver listed second answers exactly then. That drops the systemd-resolved stub and the runtime's dns: containers copy the machine's resolvers. Narrows 0194; amends connectivity §2.
5.6 KiB
topic, status, date, deciders, reconstructed, supersedes-in-part
| topic | status | date | deciders | reconstructed | supersedes-in-part | |
|---|---|---|---|---|---|---|
| the tiers | accepted | 2026-10-03 | jochen | false |
|
196. A node asks the mesh's resolver first, and a public one only when it is silent
Context
ADR 0194 gave the
mesh one resolver and had each node ask it for the mesh's names only. Because resolv.conf cannot
route by domain, that needed a stub on every node — a systemd-resolved module — and a separate
dns for the container runtime, which cannot use a loopback stub. It rejected the simpler shape,
every node sending every query to the mesh's resolver, on the grounds that "a laptop whose tunnel is
down could resolve nothing at all."
That is true only of a resolv.conf naming the mesh's resolver alone. The C library asks the
servers it lists in order and moves to the next when one does not answer within its timeout. A public
resolver listed second is asked exactly when the mesh's is unreachable — the anchor down, the tunnel
down, a laptop behind a captive portal that has not let the tunnel up — and never otherwise. An answer
from the first, including "no such name", is final, so .internal is never asked of a public resolver
while the mesh's answers.
And the container runtime copies a machine's resolvers into its containers when they are not loopback addresses. With the mesh's resolver and a public one listed, every container gets both, as they are, with nothing configured for the runtime.
Considered Options
1. Keep ADR 0194's stub. Public names never touch the mesh, and a node with the anchor down resolves public names at full speed. It costs a module and a running service on every node, a second configuration for containers, and the one asymmetry ADR 0194 had to state — containers' public names through the mesh, nodes' not.
2. Every node asks the mesh's resolver for everything, with a public resolver as the silent fallback. One server answers every node and every container; nothing on a node routes, holds names, or runs. Chosen.
Decision
A node's /etc/resolv.conf names the mesh's resolver first and a public resolver second, with a
short timeout and a single attempt. It is written by the module holding node-resolver-config — the
existing resolv-conf — which now names mesh-resolver's address instead of the machine's own. The
mesh's resolver answers the mesh's names from what it holds and forwards every other name, giving the
public answer (ADR 0191 is unchanged: no
public name gets a private answer).
Containers take the same two resolvers from their machine. The container runtime's own dns
setting is not written; the runtime copies the machine's non-loopback resolvers into every container.
This replaces, from ADR 0194: the asking side as a stub ("So the asking side is a
systemd-resolved module"), the container runtime's dns naming mesh-resolver, and step 2 of the
migration as written. There is no systemd-resolved module. Everything else in ADR 0194 stands — one
mesh-resolver, on the node every tunnel converges on, holding each node's internal domain, the
retirement of node-dns-resolver and every per-node copy, and a LAN's resolver not being the mesh's.
The migration, as it now reads:
mesh-resolveris assigned and answers on the private network.- Each node's
resolv-confnamesmesh-resolverfirst and a public resolver second. - A LAN whose router points at a node's resolver is pointed at its router.
node-dns-resolveris unassigned from every node, and the hosts region is withdrawn.
Consequences
- Every name a node or container asks goes through the anchor while it is up. A public lookup takes a few milliseconds longer than asking a public resolver directly, and the mesh's resolver sees every name its nodes look up. It is the operator's own server.
- With the anchor unreachable, each lookup waits out one timeout, then resolves publicly.
.internalnames fail then — as.internaltraffic does, every tunnel going through the anchor. - A LAN is unaffected by this choice. Devices that are not members never read a node's
resolv.conf; they get their resolver from their router, which step 3 points at itself. - Nothing new runs on a node. No stub, no module, no per-node configuration for containers.
- The runtime's
dnskey goes withnode-dns-resolver. The dnsmasq module wrote it into the runtime's configuration; unassigning that module in step 4 withdraws it, and the runtime reads the change only when it next starts — withlive-restoreon, that restart keeps every container running.
How each is checked:
- Order: each node's
/etc/resolv.conflistsmesh-resolver's private address first and a public resolver second, and nothing else. - Fallback: with
mesh-resolverunreachable from a node, a public name still resolves there, after the timeout. - Containers: a container started on a node lists the same two resolvers.
- A LAN: the router's DHCP DNS option names the router, not a node.
References
- ADR 0194 — the one resolver; this record replaces how nodes and containers ask it.
- ADR 0191 — what the resolver holds.
- Connectivity §2, amended alongside this record.