Files
hq/02-DECISIONS/0196-a-node-asks-the-meshs-resolver-first-and-a-public-one-only-when-it-is-silent.md
T
jschoubben f5d518d256 ADR 0196: a node asks the mesh's resolver first, and a public one only when it is silent
ADR 0194 rejected sending every query to the mesh's resolver because a node with its tunnel down
would resolve nothing; a public resolver listed second answers exactly then. That drops the
systemd-resolved stub and the runtime's dns: containers copy the machine's resolvers. Narrows 0194;
amends connectivity §2.
2026-10-03 21:56:33 +02:00

5.6 KiB

topic, status, date, deciders, reconstructed, supersedes-in-part
topic status date deciders reconstructed supersedes-in-part
the tiers accepted 2026-10-03 jochen false
0194-the-mesh-has-one-resolver-and-every-node-asks-it-for-the-meshs-names.md

196. A node asks the mesh's resolver first, and a public one only when it is silent

Context

ADR 0194 gave the mesh one resolver and had each node ask it for the mesh's names only. Because resolv.conf cannot route by domain, that needed a stub on every node — a systemd-resolved module — and a separate dns for the container runtime, which cannot use a loopback stub. It rejected the simpler shape, every node sending every query to the mesh's resolver, on the grounds that "a laptop whose tunnel is down could resolve nothing at all."

That is true only of a resolv.conf naming the mesh's resolver alone. The C library asks the servers it lists in order and moves to the next when one does not answer within its timeout. A public resolver listed second is asked exactly when the mesh's is unreachable — the anchor down, the tunnel down, a laptop behind a captive portal that has not let the tunnel up — and never otherwise. An answer from the first, including "no such name", is final, so .internal is never asked of a public resolver while the mesh's answers.

And the container runtime copies a machine's resolvers into its containers when they are not loopback addresses. With the mesh's resolver and a public one listed, every container gets both, as they are, with nothing configured for the runtime.

Considered Options

1. Keep ADR 0194's stub. Public names never touch the mesh, and a node with the anchor down resolves public names at full speed. It costs a module and a running service on every node, a second configuration for containers, and the one asymmetry ADR 0194 had to state — containers' public names through the mesh, nodes' not.

2. Every node asks the mesh's resolver for everything, with a public resolver as the silent fallback. One server answers every node and every container; nothing on a node routes, holds names, or runs. Chosen.

Decision

A node's /etc/resolv.conf names the mesh's resolver first and a public resolver second, with a short timeout and a single attempt. It is written by the module holding node-resolver-config — the existing resolv-conf — which now names mesh-resolver's address instead of the machine's own. The mesh's resolver answers the mesh's names from what it holds and forwards every other name, giving the public answer (ADR 0191 is unchanged: no public name gets a private answer).

Containers take the same two resolvers from their machine. The container runtime's own dns setting is not written; the runtime copies the machine's non-loopback resolvers into every container.

This replaces, from ADR 0194: the asking side as a stub ("So the asking side is a systemd-resolved module"), the container runtime's dns naming mesh-resolver, and step 2 of the migration as written. There is no systemd-resolved module. Everything else in ADR 0194 stands — one mesh-resolver, on the node every tunnel converges on, holding each node's internal domain, the retirement of node-dns-resolver and every per-node copy, and a LAN's resolver not being the mesh's.

The migration, as it now reads:

  1. mesh-resolver is assigned and answers on the private network.
  2. Each node's resolv-conf names mesh-resolver first and a public resolver second.
  3. A LAN whose router points at a node's resolver is pointed at its router.
  4. node-dns-resolver is unassigned from every node, and the hosts region is withdrawn.

Consequences

  • Every name a node or container asks goes through the anchor while it is up. A public lookup takes a few milliseconds longer than asking a public resolver directly, and the mesh's resolver sees every name its nodes look up. It is the operator's own server.
  • With the anchor unreachable, each lookup waits out one timeout, then resolves publicly. .internal names fail then — as .internal traffic does, every tunnel going through the anchor.
  • A LAN is unaffected by this choice. Devices that are not members never read a node's resolv.conf; they get their resolver from their router, which step 3 points at itself.
  • Nothing new runs on a node. No stub, no module, no per-node configuration for containers.
  • The runtime's dns key goes with node-dns-resolver. The dnsmasq module wrote it into the runtime's configuration; unassigning that module in step 4 withdraws it, and the runtime reads the change only when it next starts — with live-restore on, that restart keeps every container running.

How each is checked:

  • Order: each node's /etc/resolv.conf lists mesh-resolver's private address first and a public resolver second, and nothing else.
  • Fallback: with mesh-resolver unreachable from a node, a public name still resolves there, after the timeout.
  • Containers: a container started on a node lists the same two resolvers.
  • A LAN: the router's DHCP DNS option names the router, not a node.

References

  • ADR 0194 — the one resolver; this record replaces how nodes and containers ask it.
  • ADR 0191 — what the resolver holds.
  • Connectivity §2, amended alongside this record.