ADR 0194 rejected sending every query to the mesh's resolver because a node with its tunnel down would resolve nothing; a public resolver listed second answers exactly then. That drops the systemd-resolved stub and the runtime's dns: containers copy the machine's resolvers. Narrows 0194; amends connectivity §2.
158 lines
10 KiB
Markdown
158 lines
10 KiB
Markdown
---
|
|
topic: the tiers
|
|
status: accepted
|
|
date: 2026-10-03
|
|
deciders: jochen
|
|
reconstructed: false
|
|
supersedes-in-part:
|
|
- 0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md
|
|
extends: 0191-the-meshs-resolver-holds-only-the-meshs-own-names.md
|
|
---
|
|
|
|
# 194. The mesh has one resolver, and every node asks it for the mesh's names
|
|
|
|
> **Narrowed, not replaced — 2026-10-03.** How a node asks is decided again by
|
|
> [ADR 0196](0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md): every
|
|
> node and container asks `mesh-resolver` first and a public resolver only when it is silent. There is
|
|
> no `systemd-resolved` module and no runtime `dns` naming `mesh-resolver`, and step 2 of the migration
|
|
> reads as 0196 states it. Option 2 below was rejected for a laptop with its tunnel down resolving
|
|
> nothing; a public resolver listed second answers exactly then. The one resolver, its placement and
|
|
> the retirement of every per-node copy stand.
|
|
|
|
## Context
|
|
|
|
**Every node runs its own resolver and holds its own copy of the mesh's names.** On the production
|
|
mesh on 2026-10-03, each of the four nodes held `node-dns-resolver` with dnsmasq, fed on every push
|
|
with a zones file (one wildcard per node) and a region of `/etc/hosts` (the machines), and pointed
|
|
its own `/etc/resolv.conf` at itself. The controller computes the names once; four daemons then hold
|
|
four copies, each read in its own way.
|
|
|
|
**Every resolution fault found that day was a copy disagreeing with the truth, not the truth being
|
|
wrong:**
|
|
|
|
- **A copy read once.** dnsmasq reads `/etc/hosts` at start. After the controller stopped publishing
|
|
public names ([ADR 0191](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md)), every node's
|
|
hosts file was right and every resolver still answered the mail server's public name with a
|
|
tunnel address, until each was restarted.
|
|
- **A copy beside other copies.** On the workstation, a name resolved to two addresses in rotation:
|
|
the mesh's region gave the tunnel address, and two lines the operator had written before the mesh
|
|
existed — one in `/etc/hosts`, one in a file the resolver also reads — gave the LAN address. A
|
|
comment beside one of them said to delete it once the mesh took over; nothing made that happen.
|
|
- **A copy that became somebody else's resolver.** The home server's resolver also answers its LAN
|
|
(a listen address added 2026-10-02), and the LAN's router hands that address out as the only DNS
|
|
server. Every phone and television on the LAN resolved through a mesh node's private copy, which is
|
|
how ADR 0191's outage reached them.
|
|
|
|
**And the overlay already has one centre.** Every node has exactly one tunnel peer — the anchor —
|
|
and routes the whole private range through it. Two nodes on the same LAN reach each other through
|
|
the anchor. So a name under `.internal` is only ever useful while the anchor is reachable: a resolver
|
|
anywhere else adds a copy without adding an answer anybody can use.
|
|
|
|
**What a node asks is already a separate role.** The connectivity design split *serving* (answers
|
|
the names) from *asking* (decides what the machine asks), because systemd-resolved cannot answer a
|
|
wildcard and can only route the mesh's suffix to something that can
|
|
([connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md)). ADR 0121 kept them as two seats,
|
|
`node-dns-resolver` and `node-resolver-config`, both at node scope. No node runs systemd-resolved
|
|
today; each writes `/etc/resolv.conf` as a plain file pointing at its own dnsmasq.
|
|
|
|
## Considered Options
|
|
|
|
**1. Keep a resolver on every node, and make the copies more careful.** Restart on every file it
|
|
reads, own every file it reads, refuse to listen on a LAN. Each is a fix for one way a copy goes
|
|
stale, and the next way is not on the list yet. It keeps four answers to one question.
|
|
|
|
**2. One resolver for the mesh, and every node sends it every query.** The simplest asking side —
|
|
`resolv.conf` names the mesh's resolver and nothing else. Rejected: public resolution then depends on
|
|
the tunnel. A laptop whose tunnel is down could resolve nothing at all, and a public name would take
|
|
a detour through the anchor for no reason ADR 0191 left standing.
|
|
|
|
**3. One resolver for the mesh's names; each node asks it for those only.** The mesh's resolver holds
|
|
every node's internal domain. Each node's asking role routes the mesh's suffix to it and every other
|
|
name to public resolvers. Chosen.
|
|
|
|
## Decision
|
|
|
|
**The mesh has one resolver.** It is a module holding a new mesh-scoped seat, **`mesh-resolver`**
|
|
(capacity one). It holds each node's internal domain — `<node>.internal` and everything under it, at
|
|
that node's private address ([ADR 0191](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md))
|
|
— and listens on the private network only. It is placed on the node every tunnel converges on, so
|
|
that it shares the overlay's single point rather than adding one. Which daemon fills the seat stays
|
|
the module's business, as the connectivity design says.
|
|
|
|
**Every node asks it for the mesh's names and nothing else.** `node-resolver-config` routes the
|
|
mesh's suffix to `mesh-resolver` and leaves every other name with public resolvers.
|
|
|
|
**Why a stub on every node.** `resolv.conf` cannot route by domain: the C library asks the servers it
|
|
lists in order, for every name, and moves to the next only when one does not answer — an NXDOMAIN
|
|
from the first is final. Listing `mesh-resolver` first sends every public name through the tunnel
|
|
(option 2); listing a public resolver first means `.internal` is never asked of the mesh. Something on
|
|
the node has to look at the name before choosing a server, and that is a stub resolver. Keeping
|
|
dnsmasq for it would keep a daemon that reads hosts files and can be told to answer a LAN — the two
|
|
ways copies went wrong. systemd-resolved holds no names of its own, routes by domain natively (a
|
|
routing domain `~<suffix>` on the server that answers it), and is part of systemd, already installed
|
|
on every node and enabled on none.
|
|
|
|
**So the asking side is a `systemd-resolved` module**, claiming `node-resolver-config` — the same claim
|
|
as the `resolv-conf` module it replaces, so the mesh refuses both on one node. It enables the service,
|
|
writes its configuration (the mesh resolver for the suffix, public resolvers for everything else), and
|
|
writes `/etc/resolv.conf` as a file naming the stub — a file the module owns, not a link to one.
|
|
|
|
**A container asks the mesh's resolver directly.** The container runtime cannot use a loopback stub
|
|
and drops its routing domains, so the runtime's `dns` names `mesh-resolver`, which forwards public
|
|
names for the containers that ask it. This is the one place a public name passes through the mesh,
|
|
and it is stated rather than hidden.
|
|
|
|
**`node-dns-resolver` is retired**, and with it every per-node copy: the zones file, the mesh's region
|
|
of `/etc/hosts` (the floor connectivity §2 already planned to remove), and the daemon on every node
|
|
but the one holding `mesh-resolver`. This narrows ADR 0121's *"the-dns-port → node-dns-resolver"*:
|
|
the serving role keeps its distinction from the asking role and moves to mesh scope, as ADR 0121 did
|
|
for the private network.
|
|
|
|
**A LAN's resolver is not the mesh's.** No device that is not a member can reach a private address,
|
|
so no member's resolver answers a LAN on the mesh's behalf. A router that hands out a node's address
|
|
as a LAN's DNS server is pointed elsewhere before that node stops answering.
|
|
|
|
**The order is fixed, because every step before the last leaves a working resolver:**
|
|
|
|
1. `mesh-resolver` is assigned and answers on the private network.
|
|
2. Each node's `node-resolver-config` moves from `resolv-conf` to `systemd-resolved`, and the container
|
|
runtime's `dns` to `mesh-resolver`.
|
|
3. A LAN whose router points at a node's resolver is pointed at its router or a public resolver.
|
|
4. `node-dns-resolver` is unassigned from every node, and the hosts region is withdrawn.
|
|
|
|
## Consequences
|
|
|
|
- **One answer per name.** A name is wrong in one place or right everywhere; no node can hold a copy
|
|
that disagrees, and no operator file on a node is read by the mesh's resolver.
|
|
- **The anchor down means no `.internal` names** — which it already meant for `.internal` traffic,
|
|
since every tunnel goes through it. Public resolution on every node is unaffected.
|
|
- **A container's public resolution depends on the mesh's resolver.** Accepted, and named in the
|
|
decision; a container that must resolve public names with the tunnel down is the case it costs.
|
|
- **Every node runs systemd-resolved**, through the `systemd-resolved` module. It is installed
|
|
everywhere already and enabled nowhere; the mesh still ships no resolver of its own.
|
|
- **The runtime's `dns` changes once per node**, which the runtime reads only at start. With
|
|
`live-restore` already on, that restart keeps every container running.
|
|
- **A LAN loses a resolver it had borrowed.** The router change is an explicit step, done through
|
|
the module that manages the router, before the node's resolver goes.
|
|
|
|
**How each is checked:**
|
|
|
|
- **One holder:** the seat has capacity one, so a second assignment is refused by the controller.
|
|
- **Asking:** on each node, `resolvectl` shows the tunnel's link with `mesh-resolver` and the suffix as
|
|
its routing domain; a name under `.internal` is answered by it, and a public name is answered
|
|
without it (its query log shows no public name from a node).
|
|
- **No copies:** no node but the holder answers DNS on a private or LAN address — every other node's
|
|
port 53 is systemd-resolved's loopback stub and nothing else — and no node's `/etc/hosts` carries a
|
|
mesh region.
|
|
- **A LAN:** the router's DHCP DNS option names no node's address.
|
|
|
|
## References
|
|
|
|
- [ADR 0191](0191-the-meshs-resolver-holds-only-the-meshs-own-names.md) — what the mesh's resolver
|
|
holds; this record decides where it runs and how nodes reach it.
|
|
- [ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md) — the two
|
|
resolver seats, and the private network's move to mesh scope this mirrors.
|
|
- [Connectivity §2](../03-DESIGN/01-to-be/08-connectivity.md) — serving and asking as two roles;
|
|
amended alongside this record.
|
|
- [The seats](../03-DESIGN/01-to-be/26-the-seats.md) — the seat table, amended alongside.
|