Files
hq/04-ISSUES/171-a-modules-own-resolver-knows-no-mesh-name/00-report.md
T
jschoubben af170e3a67 Issue 171: a module that names its own resolver knows no mesh name
Found and fixed the afternoon ADR 0148 landed: mailu-admin lost its
database behind Mailu's own resolver. Two catalogue PRs; an insight on
0148 that a container's dns is a decision, not a preference.
2026-09-30 15:20:26 +02:00

4.5 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
resolved 2026-09-30
mesh-catalog modules/mailu (eight containers name Mailu's own resolver, and one of them binds a mesh name)
mesh-catalog modules/dnsmasq (dropped the DNSSEC bit its upstreams set)
mesh-catalog PR 178 (mailu-admin uses the machine's resolver) and PR 179 (the machine's resolver passes the DNSSEC bit down) — 2026-09-30, the same afternoon

171 — A module that names its own resolver knows no mesh name

What was observed

The afternoon ADR 0148 landed — no container is given the mesh's names any more; it asks the machine's resolver — the mail system's admin container began logging, 523 times in three minutes:

psycopg2.OperationalError: could not translate host name "novox.internal" to address: Name does not resolve

Mail was accepted on every port and the web front answered; the admin and the spam filter beside it were unhealthy, and anything that needed the database — a mailbox change through the API, the spam filter's domain list — failed. Found by the operator asking whether mail was back, forty minutes in.

Mailu ships its own resolver, an unbound in a container, and every other Mailu container is told to use it — the module carries dns: [192.168.203.254] on eight containers. That resolver recurses from the root and knows nothing under .internal. Until that afternoon the admin container had the database's name anyway, because the mesh wrote every name into every container at creation; the copy was the only reason a container behind its own resolver could reach anything by a mesh name, and nobody knew it was load-bearing.

Removing the override was not enough. Given the machine's resolver instead, the admin refused to start: Your DNS resolver at 127.0.0.11 isn't doing DNSSEC validation. Mailu checks, at start, that its resolver returns the Authenticated Data bit for a signed name. The mesh's resolver forwards to two upstreams that validate and set the bit, and dropped it on the way down — dnsmasq does unless told otherwise. Mailu's own unbound has no hook to forward a zone elsewhere, so it could not be taught the mesh's names either.

Why it matters beyond this instance

A container with a resolver of its own has opted out of the machine's, and nothing says so. 0148 made the machine's resolver load-bearing for every container; a dns on a container is a quiet exception to that, and the exception used to be papered over by the copy the record removed. The manifest field reads like a preference and is a decision about whether mesh names exist inside the container.

A resolver that forwards to validating upstreams and hides the fact is less useful than it could be, and the first program to check found out.

The mesh reported nothing. Every container ran; the failing one accepted connections; the report was about bytes. It is issue 145 again, and the check that would have caught it is the same unbuilt one.

What was done

  • The one Mailu container that binds a mesh name — the admin, through the database it is granted — no longer names Mailu's resolver and uses the machine's, like every container without a dns of its own (mesh-catalog PR 178). The other seven keep unbound: the spam filter needs a validating resolver for its blocklist lookups, and none of them asks for a mesh name.
  • The machine's resolver passes the DNSSEC bit down from its upstreams, proxy-dnssec (PR 179). It does not validate itself; the trust is the upstream's and the path to it, as a forwarding resolver's always was, and the configuration says so.

What checks it

The admin container's own start-up check, which is what failed, and the mesh's status once it reads healthy. A container-level check that a mesh name resolves from inside every declared container is the one 110 and 145 both ask for and is not built.

Open questions

  • Should a container's dns be refused, or made to say what it gives up? A module that names its own resolver and binds a mesh name is a contradiction the controller can see at composition — the grant hands it a name its resolver will not answer.
  • Should the machine's resolver validate rather than proxy? It would cost a trust anchor on every machine and make the resolver slower to start; proxying was enough for the one program that asked.