ADR 0194: why every node needs a stub, and the systemd-resolved module that provides it

This commit is contained in:
2026-10-03 21:50:33 +02:00
parent 1de4a5f25e
commit 6b6ff76a19
2 changed files with 23 additions and 10 deletions
@@ -72,10 +72,22 @@ that it shares the overlay's single point rather than adding one. Which daemon f
the module's business, as the connectivity design says.
**Every node asks it for the mesh's names and nothing else.** `node-resolver-config` routes the
mesh's suffix to `mesh-resolver` and leaves every other name with public resolvers. Plain
`resolv.conf` cannot route by domain, so the asking side is a stub that can — systemd-resolved, with
the tunnel's link carrying the mesh resolver and the suffix as its routing domain. The stub holds no
names; it decides only where a question goes.
mesh's suffix to `mesh-resolver` and leaves every other name with public resolvers.
**Why a stub on every node.** `resolv.conf` cannot route by domain: the C library asks the servers it
lists in order, for every name, and moves to the next only when one does not answer — an NXDOMAIN
from the first is final. Listing `mesh-resolver` first sends every public name through the tunnel
(option 2); listing a public resolver first means `.internal` is never asked of the mesh. Something on
the node has to look at the name before choosing a server, and that is a stub resolver. Keeping
dnsmasq for it would keep a daemon that reads hosts files and can be told to answer a LAN — the two
ways copies went wrong. systemd-resolved holds no names of its own, routes by domain natively (a
routing domain `~<suffix>` on the server that answers it), and is part of systemd, already installed
on every node and enabled on none.
**So the asking side is a `systemd-resolved` module**, claiming `node-resolver-config` — the same claim
as the `resolv-conf` module it replaces, so the mesh refuses both on one node. It enables the service,
writes its configuration (the mesh resolver for the suffix, public resolvers for everything else), and
writes `/etc/resolv.conf` as a file naming the stub — a file the module owns, not a link to one.
**A container asks the mesh's resolver directly.** The container runtime cannot use a loopback stub
and drops its routing domains, so the runtime's `dns` names `mesh-resolver`, which forwards public
@@ -95,8 +107,8 @@ as a LAN's DNS server is pointed elsewhere before that node stops answering.
**The order is fixed, because every step before the last leaves a working resolver:**
1. `mesh-resolver` is assigned and answers on the private network.
2. Each node's `node-resolver-config` switches to the routing stub, and the container runtime's `dns`
to `mesh-resolver`.
2. Each node's `node-resolver-config` moves from `resolv-conf` to `systemd-resolved`, and the container
runtime's `dns` to `mesh-resolver`.
3. A LAN whose router points at a node's resolver is pointed at its router or a public resolver.
4. `node-dns-resolver` is unassigned from every node, and the hosts region is withdrawn.
@@ -108,8 +120,8 @@ as a LAN's DNS server is pointed elsewhere before that node stops answering.
since every tunnel goes through it. Public resolution on every node is unaffected.
- **A container's public resolution depends on the mesh's resolver.** Accepted, and named in the
decision; a container that must resolve public names with the tunnel down is the case it costs.
- **Every node gains systemd-resolved as its asking stub**, which today none runs. It is configured
by the module holding `node-resolver-config`; the mesh still ships no resolver of its own.
- **Every node runs systemd-resolved**, through the `systemd-resolved` module. It is installed
everywhere already and enabled nowhere; the mesh still ships no resolver of its own.
- **The runtime's `dns` changes once per node**, which the runtime reads only at start. With
`live-restore` already on, that restart keeps every container running.
- **A LAN loses a resolver it had borrowed.** The router change is an explicit step, done through
+3 -2
View File
@@ -357,8 +357,9 @@ machine — declared or not — is what a nameserver in `resolv.conf` would be f
seat of capacity one, placed on the node every tunnel converges on. It holds one wildcard per node —
`<node>.internal` and everything under it — and listens on the private network only. Every node's
`node-resolver-config` routes the mesh's suffix to it and leaves every other name with public
resolvers; plain `resolv.conf` cannot route by domain, so the asking side is a stub that can
(systemd-resolved, the tunnel's link carrying the suffix as its routing domain). The container runtime
resolvers; plain `resolv.conf` cannot route by domain, so the asking side is a stub that can — a
`systemd-resolved` module claiming `node-resolver-config` in place of `resolv-conf`, routing the suffix
to `mesh-resolver`. The container runtime
cannot use a loopback stub, so its `dns` names `mesh-resolver`, which forwards public names for
containers — the one place a public name passes through the mesh
([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)).