Files
hq/02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md
T
jschoubben f5d518d256 ADR 0196: a node asks the mesh's resolver first, and a public one only when it is silent
ADR 0194 rejected sending every query to the mesh's resolver because a node with its tunnel down
would resolve nothing; a public resolver listed second answers exactly then. That drops the
systemd-resolved stub and the runtime's dns: containers copy the machine's resolvers. Narrows 0194;
amends connectivity §2.
2026-10-03 21:56:33 +02:00

10 KiB

topic, status, date, deciders, reconstructed, supersedes-in-part, extends
topic status date deciders reconstructed supersedes-in-part extends
the tiers accepted 2026-10-03 jochen false
0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md
0191-the-meshs-resolver-holds-only-the-meshs-own-names.md

194. The mesh has one resolver, and every node asks it for the mesh's names

Narrowed, not replaced — 2026-10-03. How a node asks is decided again by ADR 0196: every node and container asks mesh-resolver first and a public resolver only when it is silent. There is no systemd-resolved module and no runtime dns naming mesh-resolver, and step 2 of the migration reads as 0196 states it. Option 2 below was rejected for a laptop with its tunnel down resolving nothing; a public resolver listed second answers exactly then. The one resolver, its placement and the retirement of every per-node copy stand.

Context

Every node runs its own resolver and holds its own copy of the mesh's names. On the production mesh on 2026-10-03, each of the four nodes held node-dns-resolver with dnsmasq, fed on every push with a zones file (one wildcard per node) and a region of /etc/hosts (the machines), and pointed its own /etc/resolv.conf at itself. The controller computes the names once; four daemons then hold four copies, each read in its own way.

Every resolution fault found that day was a copy disagreeing with the truth, not the truth being wrong:

  • A copy read once. dnsmasq reads /etc/hosts at start. After the controller stopped publishing public names (ADR 0191), every node's hosts file was right and every resolver still answered the mail server's public name with a tunnel address, until each was restarted.
  • A copy beside other copies. On the workstation, a name resolved to two addresses in rotation: the mesh's region gave the tunnel address, and two lines the operator had written before the mesh existed — one in /etc/hosts, one in a file the resolver also reads — gave the LAN address. A comment beside one of them said to delete it once the mesh took over; nothing made that happen.
  • A copy that became somebody else's resolver. The home server's resolver also answers its LAN (a listen address added 2026-10-02), and the LAN's router hands that address out as the only DNS server. Every phone and television on the LAN resolved through a mesh node's private copy, which is how ADR 0191's outage reached them.

And the overlay already has one centre. Every node has exactly one tunnel peer — the anchor — and routes the whole private range through it. Two nodes on the same LAN reach each other through the anchor. So a name under .internal is only ever useful while the anchor is reachable: a resolver anywhere else adds a copy without adding an answer anybody can use.

What a node asks is already a separate role. The connectivity design split serving (answers the names) from asking (decides what the machine asks), because systemd-resolved cannot answer a wildcard and can only route the mesh's suffix to something that can (connectivity §2). ADR 0121 kept them as two seats, node-dns-resolver and node-resolver-config, both at node scope. No node runs systemd-resolved today; each writes /etc/resolv.conf as a plain file pointing at its own dnsmasq.

Considered Options

1. Keep a resolver on every node, and make the copies more careful. Restart on every file it reads, own every file it reads, refuse to listen on a LAN. Each is a fix for one way a copy goes stale, and the next way is not on the list yet. It keeps four answers to one question.

2. One resolver for the mesh, and every node sends it every query. The simplest asking side — resolv.conf names the mesh's resolver and nothing else. Rejected: public resolution then depends on the tunnel. A laptop whose tunnel is down could resolve nothing at all, and a public name would take a detour through the anchor for no reason ADR 0191 left standing.

3. One resolver for the mesh's names; each node asks it for those only. The mesh's resolver holds every node's internal domain. Each node's asking role routes the mesh's suffix to it and every other name to public resolvers. Chosen.

Decision

The mesh has one resolver. It is a module holding a new mesh-scoped seat, mesh-resolver (capacity one). It holds each node's internal domain — <node>.internal and everything under it, at that node's private address (ADR 0191) — and listens on the private network only. It is placed on the node every tunnel converges on, so that it shares the overlay's single point rather than adding one. Which daemon fills the seat stays the module's business, as the connectivity design says.

Every node asks it for the mesh's names and nothing else. node-resolver-config routes the mesh's suffix to mesh-resolver and leaves every other name with public resolvers.

Why a stub on every node. resolv.conf cannot route by domain: the C library asks the servers it lists in order, for every name, and moves to the next only when one does not answer — an NXDOMAIN from the first is final. Listing mesh-resolver first sends every public name through the tunnel (option 2); listing a public resolver first means .internal is never asked of the mesh. Something on the node has to look at the name before choosing a server, and that is a stub resolver. Keeping dnsmasq for it would keep a daemon that reads hosts files and can be told to answer a LAN — the two ways copies went wrong. systemd-resolved holds no names of its own, routes by domain natively (a routing domain ~<suffix> on the server that answers it), and is part of systemd, already installed on every node and enabled on none.

So the asking side is a systemd-resolved module, claiming node-resolver-config — the same claim as the resolv-conf module it replaces, so the mesh refuses both on one node. It enables the service, writes its configuration (the mesh resolver for the suffix, public resolvers for everything else), and writes /etc/resolv.conf as a file naming the stub — a file the module owns, not a link to one.

A container asks the mesh's resolver directly. The container runtime cannot use a loopback stub and drops its routing domains, so the runtime's dns names mesh-resolver, which forwards public names for the containers that ask it. This is the one place a public name passes through the mesh, and it is stated rather than hidden.

node-dns-resolver is retired, and with it every per-node copy: the zones file, the mesh's region of /etc/hosts (the floor connectivity §2 already planned to remove), and the daemon on every node but the one holding mesh-resolver. This narrows ADR 0121's "the-dns-port → node-dns-resolver": the serving role keeps its distinction from the asking role and moves to mesh scope, as ADR 0121 did for the private network.

A LAN's resolver is not the mesh's. No device that is not a member can reach a private address, so no member's resolver answers a LAN on the mesh's behalf. A router that hands out a node's address as a LAN's DNS server is pointed elsewhere before that node stops answering.

The order is fixed, because every step before the last leaves a working resolver:

  1. mesh-resolver is assigned and answers on the private network.
  2. Each node's node-resolver-config moves from resolv-conf to systemd-resolved, and the container runtime's dns to mesh-resolver.
  3. A LAN whose router points at a node's resolver is pointed at its router or a public resolver.
  4. node-dns-resolver is unassigned from every node, and the hosts region is withdrawn.

Consequences

  • One answer per name. A name is wrong in one place or right everywhere; no node can hold a copy that disagrees, and no operator file on a node is read by the mesh's resolver.
  • The anchor down means no .internal names — which it already meant for .internal traffic, since every tunnel goes through it. Public resolution on every node is unaffected.
  • A container's public resolution depends on the mesh's resolver. Accepted, and named in the decision; a container that must resolve public names with the tunnel down is the case it costs.
  • Every node runs systemd-resolved, through the systemd-resolved module. It is installed everywhere already and enabled nowhere; the mesh still ships no resolver of its own.
  • The runtime's dns changes once per node, which the runtime reads only at start. With live-restore already on, that restart keeps every container running.
  • A LAN loses a resolver it had borrowed. The router change is an explicit step, done through the module that manages the router, before the node's resolver goes.

How each is checked:

  • One holder: the seat has capacity one, so a second assignment is refused by the controller.
  • Asking: on each node, resolvectl shows the tunnel's link with mesh-resolver and the suffix as its routing domain; a name under .internal is answered by it, and a public name is answered without it (its query log shows no public name from a node).
  • No copies: no node but the holder answers DNS on a private or LAN address — every other node's port 53 is systemd-resolved's loopback stub and nothing else — and no node's /etc/hosts carries a mesh region.
  • A LAN: the router's DHCP DNS option names no node's address.

References

  • ADR 0191 — what the mesh's resolver holds; this record decides where it runs and how nodes reach it.
  • ADR 0121 — the two resolver seats, and the private network's move to mesh scope this mirrors.
  • Connectivity §2 — serving and asking as two roles; amended alongside this record.
  • The seats — the seat table, amended alongside.