It declared udp/53 only. The daemon listens on tcp/53 as well, and a resolver is asked over tcp whenever an answer will not fit in a datagram — so on every converged machine that port is closed while the service reports itself healthy and the manifest reads as though the resolver were fully declared. The same fault as issue 136 in miniature: the declaration covers part of what the service does, and the gap is silent because nothing compares the two.
86 lines
7.7 KiB
JSON
86 lines
7.7 KiB
JSON
{
|
|
"module": "dnsmasq",
|
|
"version": "1",
|
|
"provides": [
|
|
"wildcard-resolution"
|
|
],
|
|
"requires": [
|
|
"mesh-addressing"
|
|
],
|
|
"emits": [
|
|
"name.added",
|
|
"name.removed"
|
|
],
|
|
"own-secrets": {
|
|
"broker": "/var/lib/mesh/dnsmasq/broker"
|
|
},
|
|
"claims": [
|
|
{
|
|
"name": "node-dns-resolver",
|
|
"scope": "node"
|
|
}
|
|
],
|
|
"listens": [
|
|
{
|
|
"port": 53,
|
|
"protocol": "udp",
|
|
"from": "mesh",
|
|
"why": "every name for this machine and what it runs \u2014 the mesh's own answered here, the rest forwarded",
|
|
"fixed": true
|
|
},
|
|
{
|
|
"port": 53,
|
|
"protocol": "tcp",
|
|
"from": "mesh",
|
|
"why": "the same names over tcp, which a resolver answers on as well and is asked for whenever an answer will not fit in a datagram. Declared because the daemon serves it: a declaration that covers one of the two protocols its own service listens on leaves the other closed while everything reports success",
|
|
"fixed": true
|
|
}
|
|
],
|
|
"resources": [
|
|
{
|
|
"id": "mesh-state",
|
|
"type": "directory",
|
|
"path": "/var/lib/mesh/dnsmasq",
|
|
"mode": "0700"
|
|
},
|
|
{
|
|
"id": "package",
|
|
"type": "package",
|
|
"package": "dnsmasq"
|
|
},
|
|
{
|
|
"id": "config",
|
|
"type": "file",
|
|
"path": "/etc/dnsmasq.conf",
|
|
"mode": "0644",
|
|
"content": "# Managed by the mesh. dnsmasq's own defaults are replaced whole rather than\n# patched, because this module owns the file and a patch would leave whatever\n# was there before to be discovered later.\n\n# What the mesh computed: one wildcard per machine \u2014 its name and everything\n# under it \u2014 and the mesh's own suffix as a local domain, so a name under it is\n# answered here or not at all and is never asked upstream. Rewritten whenever a\n# machine joins or leaves, which is why the service below restarts on it: a\n# reload makes dnsmasq re-read hosts files, not its configuration, and a\n# wildcard is configuration.\nconf-file=/etc/mesh-resolver/nodes.conf\n\n# Where it answers. Both are names the mesh chose, so this file needs to know\n# nothing about this particular machine:\n#\n# mesh0 the private network, so anything on it can ask \u2014 including\n# this machine's containers. This module writes the runtime's\n# `dns` key into its own configuration file, beside whatever the\n# machine had there (novox/hq ADR 0102), naming this address: a\n# container cannot reach the machine's loopback, and a runtime\n# whose host resolves at loopback falls back to a public resolver\n# and never sees a mesh name. The runtime reads that key when it\n# starts and not on a reload, and a restart stops every container\n# on the machine, so this module orders neither: the key holds for\n# every container created after the runtime next starts. On the\n# machine this replaces the predecessor wrote the same value, so\n# nothing there is waiting on it.\n# 127.0.0.1 this machine's own use. The predecessor's resolver answered\n# here, and the resolv.conf it wrote on every machine says so;\n# that file stays in force on an adopted machine until the mesh's\n# module for it is taken, so the resolver has to answer where the\n# machine already asks or the machine loses DNS the moment this\n# module is taken. Not .53 or .54: systemd-resolved holds BOTH \u2014\n# .53 is its stub and .54 its proxy stub \u2014 and neither is .1, so\n# the two coexist on a machine that runs it. This module used to\n# answer on 127.0.0.55 instead: a convention of its own, beside\n# the one every machine already followed. One address, this one,\n# and the modules that point a machine at the mesh name the same.\n#\n# Whatever address it listens on, it takes the machine's DNS port.\n# That is why this module claims `node-dns-resolver`.\n#\n# bind-dynamic rather than bind-interfaces: mesh0 does not exist until the\n# machine is on the private network, and binding an interface that is not there\n# yet fails to start rather than waiting for it.\nbind-dynamic\ninterface=mesh0\nlisten-address=127.0.0.1\n\n# **It must never read resolv.conf to find out where to forward.** Whatever\n# points this machine at the mesh writes this resolver's own address there \u2014 so\n# a resolver that read it for upstreams would find itself, and every query it\n# could not answer locally would loop until its receive queue filled. That is\n# not theoretical: it filled with 15KB of queries and every lookup on the\n# machine hung. no-resolv is what makes that loop impossible: the upstreams are\n# the two lines below, and nothing on the machine can redirect them.\n#\n# It forwards, because it is now asked for everything. The module that points\n# this machine at the mesh names this resolver alone \u2014 as the predecessor's\n# did \u2014 so the host and every container resolve the world through it. The\n# upstreams are the ones the predecessor's module shipped as its defaults. The\n# mesh's own names never reach them: the local= line in the file above stops\n# them here, answered or refused.\nno-resolv\nserver=1.1.1.1\nserver=8.8.8.8\n\n# A name without a dot is never forwarded \u2014 a bare hostname is answered from\n# /etc/hosts or not at all \u2014 and reverse lookups of private ranges are answered\n# here rather than asking the world who 10.x is.\ndomain-needed\nbogus-priv\n\n# The operator's own names have a home the mesh never rewrites (novox/hq issue\n# 122: a workstation's job includes names \u2014 Mediahuis's 13, say \u2014 that are\n# neither a mesh machine nor a routed name). Two homes, because both shapes\n# exist in the wild and neither is the mesh's to own:\n#\n# /etc/dnsmasq.d/*.conf drop-in dnsmasq directives \u2014 an address=, a second\n# upstream for one domain, a cname. HAL's dnsmasq-app\n# carried exactly this line, so it is a proven shape\n# and the files a migrating workstation already has\n# land here untouched.\n# /etc/hosts.local plain `<ip> <name>` lines, the /etc/hosts a person\n# kept \u2014 read as additional hosts, so the generated\n# /etc/hosts (which the mesh owns and rewrites) never\n# has to carry an operator entry to keep it resolving.\n#\n# Both are the operator's: the mesh creates neither and rewrites neither, and a\n# machine with no such file loses nothing. This is what lets mesh-wireguard take\n# /etc/hosts without taking the names a workstation needs down with it \u2014 they\n# were moved here first.\nconf-dir=/etc/dnsmasq.d/,*.conf\naddn-hosts=/etc/hosts.local\n"
|
|
},
|
|
{
|
|
"id": "runtime-dns",
|
|
"type": "file",
|
|
"path": "/etc/docker/daemon.json",
|
|
"mode": "0644",
|
|
"merge": "json",
|
|
"into": "json",
|
|
"content": "{\"dns\": [\"${machine:address}\"]}\n"
|
|
},
|
|
{
|
|
"id": "service",
|
|
"type": "service",
|
|
"unit": "dnsmasq.service",
|
|
"state": "running",
|
|
"boot": "enabled",
|
|
"restart-on": [
|
|
"config",
|
|
"dnsmasq.fact-node-zones"
|
|
]
|
|
}
|
|
],
|
|
"facts": {
|
|
"node-zones": {
|
|
"path": "/etc/mesh-resolver/nodes.conf",
|
|
"template": "# Generated by the mesh. Do not edit \u2014 this file is replaced whenever a machine\n# joins or leaves, and an edit would survive until then and vanish.\n\nlocal=/{{.Suffix}}/\n{{range .Machines}}address=/{{.FQDN}}/{{.Address}}\n{{end}}"
|
|
}
|
|
}
|
|
}
|