ADR 0194: the mesh has one resolver, and every node asks it for the mesh's names #326
+20
-8
@@ -72,10 +72,22 @@ that it shares the overlay's single point rather than adding one. Which daemon f
|
||||
the module's business, as the connectivity design says.
|
||||
|
||||
**Every node asks it for the mesh's names and nothing else.** `node-resolver-config` routes the
|
||||
mesh's suffix to `mesh-resolver` and leaves every other name with public resolvers. Plain
|
||||
`resolv.conf` cannot route by domain, so the asking side is a stub that can — systemd-resolved, with
|
||||
the tunnel's link carrying the mesh resolver and the suffix as its routing domain. The stub holds no
|
||||
names; it decides only where a question goes.
|
||||
mesh's suffix to `mesh-resolver` and leaves every other name with public resolvers.
|
||||
|
||||
**Why a stub on every node.** `resolv.conf` cannot route by domain: the C library asks the servers it
|
||||
lists in order, for every name, and moves to the next only when one does not answer — an NXDOMAIN
|
||||
from the first is final. Listing `mesh-resolver` first sends every public name through the tunnel
|
||||
(option 2); listing a public resolver first means `.internal` is never asked of the mesh. Something on
|
||||
the node has to look at the name before choosing a server, and that is a stub resolver. Keeping
|
||||
dnsmasq for it would keep a daemon that reads hosts files and can be told to answer a LAN — the two
|
||||
ways copies went wrong. systemd-resolved holds no names of its own, routes by domain natively (a
|
||||
routing domain `~<suffix>` on the server that answers it), and is part of systemd, already installed
|
||||
on every node and enabled on none.
|
||||
|
||||
**So the asking side is a `systemd-resolved` module**, claiming `node-resolver-config` — the same claim
|
||||
as the `resolv-conf` module it replaces, so the mesh refuses both on one node. It enables the service,
|
||||
writes its configuration (the mesh resolver for the suffix, public resolvers for everything else), and
|
||||
writes `/etc/resolv.conf` as a file naming the stub — a file the module owns, not a link to one.
|
||||
|
||||
**A container asks the mesh's resolver directly.** The container runtime cannot use a loopback stub
|
||||
and drops its routing domains, so the runtime's `dns` names `mesh-resolver`, which forwards public
|
||||
@@ -95,8 +107,8 @@ as a LAN's DNS server is pointed elsewhere before that node stops answering.
|
||||
**The order is fixed, because every step before the last leaves a working resolver:**
|
||||
|
||||
1. `mesh-resolver` is assigned and answers on the private network.
|
||||
2. Each node's `node-resolver-config` switches to the routing stub, and the container runtime's `dns`
|
||||
to `mesh-resolver`.
|
||||
2. Each node's `node-resolver-config` moves from `resolv-conf` to `systemd-resolved`, and the container
|
||||
runtime's `dns` to `mesh-resolver`.
|
||||
3. A LAN whose router points at a node's resolver is pointed at its router or a public resolver.
|
||||
4. `node-dns-resolver` is unassigned from every node, and the hosts region is withdrawn.
|
||||
|
||||
@@ -108,8 +120,8 @@ as a LAN's DNS server is pointed elsewhere before that node stops answering.
|
||||
since every tunnel goes through it. Public resolution on every node is unaffected.
|
||||
- **A container's public resolution depends on the mesh's resolver.** Accepted, and named in the
|
||||
decision; a container that must resolve public names with the tunnel down is the case it costs.
|
||||
- **Every node gains systemd-resolved as its asking stub**, which today none runs. It is configured
|
||||
by the module holding `node-resolver-config`; the mesh still ships no resolver of its own.
|
||||
- **Every node runs systemd-resolved**, through the `systemd-resolved` module. It is installed
|
||||
everywhere already and enabled nowhere; the mesh still ships no resolver of its own.
|
||||
- **The runtime's `dns` changes once per node**, which the runtime reads only at start. With
|
||||
`live-restore` already on, that restart keeps every container running.
|
||||
- **A LAN loses a resolver it had borrowed.** The router change is an explicit step, done through
|
||||
|
||||
@@ -357,8 +357,9 @@ machine — declared or not — is what a nameserver in `resolv.conf` would be f
|
||||
seat of capacity one, placed on the node every tunnel converges on. It holds one wildcard per node —
|
||||
`<node>.internal` and everything under it — and listens on the private network only. Every node's
|
||||
`node-resolver-config` routes the mesh's suffix to it and leaves every other name with public
|
||||
resolvers; plain `resolv.conf` cannot route by domain, so the asking side is a stub that can
|
||||
(systemd-resolved, the tunnel's link carrying the suffix as its routing domain). The container runtime
|
||||
resolvers; plain `resolv.conf` cannot route by domain, so the asking side is a stub that can — a
|
||||
`systemd-resolved` module claiming `node-resolver-config` in place of `resolv-conf`, routing the suffix
|
||||
to `mesh-resolver`. The container runtime
|
||||
cannot use a loopback stub, so its `dns` names `mesh-resolver`, which forwards public names for
|
||||
containers — the one place a public name passes through the mesh
|
||||
([ADR 0194](../../02-DECISIONS/0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md)).
|
||||
|
||||
Reference in New Issue
Block a user