Found and fixed the afternoon ADR 0148 landed: mailu-admin lost its database behind Mailu's own resolver. Two catalogue PRs; an insight on 0148 that a container's dns is a decision, not a preference.
4.5 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design | ||
|---|---|---|---|---|---|---|
| resolved | 2026-09-30 |
|
mesh-catalog PR 178 (mailu-admin uses the machine's resolver) and PR 179 (the machine's resolver passes the DNSSEC bit down) — 2026-09-30, the same afternoon |
171 — A module that names its own resolver knows no mesh name
What was observed
The afternoon ADR 0148 landed — no container is given the mesh's names any more; it asks the machine's resolver — the mail system's admin container began logging, 523 times in three minutes:
psycopg2.OperationalError: could not translate host name "novox.internal" to address: Name does not resolve
Mail was accepted on every port and the web front answered; the admin and the spam filter beside it were unhealthy, and anything that needed the database — a mailbox change through the API, the spam filter's domain list — failed. Found by the operator asking whether mail was back, forty minutes in.
Mailu ships its own resolver, an unbound in a container, and every other Mailu container is told to
use it — the module carries dns: [192.168.203.254] on eight containers. That resolver recurses from the
root and knows nothing under .internal. Until that afternoon the admin container had the database's
name anyway, because the mesh wrote every name into every container at creation; the copy was the only
reason a container behind its own resolver could reach anything by a mesh name, and nobody knew it was
load-bearing.
Removing the override was not enough. Given the machine's resolver instead, the admin refused to
start: Your DNS resolver at 127.0.0.11 isn't doing DNSSEC validation. Mailu checks, at start, that
its resolver returns the Authenticated Data bit for a signed name. The mesh's resolver forwards to two
upstreams that validate and set the bit, and dropped it on the way down — dnsmasq does unless told
otherwise. Mailu's own unbound has no hook to forward a zone elsewhere, so it could not be taught the
mesh's names either.
Why it matters beyond this instance
A container with a resolver of its own has opted out of the machine's, and nothing says so. 0148
made the machine's resolver load-bearing for every container; a dns on a container is a quiet
exception to that, and the exception used to be papered over by the copy the record removed. The
manifest field reads like a preference and is a decision about whether mesh names exist inside the
container.
A resolver that forwards to validating upstreams and hides the fact is less useful than it could be, and the first program to check found out.
The mesh reported nothing. Every container ran; the failing one accepted connections; the report was about bytes. It is issue 145 again, and the check that would have caught it is the same unbuilt one.
What was done
- The one Mailu container that binds a mesh name — the admin, through the database it is granted —
no longer names Mailu's resolver and uses the machine's, like every container without a
dnsof its own (mesh-catalog PR 178). The other seven keep unbound: the spam filter needs a validating resolver for its blocklist lookups, and none of them asks for a mesh name. - The machine's resolver passes the DNSSEC bit down from its upstreams,
proxy-dnssec(PR 179). It does not validate itself; the trust is the upstream's and the path to it, as a forwarding resolver's always was, and the configuration says so.
What checks it
The admin container's own start-up check, which is what failed, and the mesh's status once it reads healthy. A container-level check that a mesh name resolves from inside every declared container is the one 110 and 145 both ask for and is not built.
Open questions
- Should a container's
dnsbe refused, or made to say what it gives up? A module that names its own resolver and binds a mesh name is a contradiction the controller can see at composition — the grant hands it a name its resolver will not answer. - Should the machine's resolver validate rather than proxy? It would cost a trust anchor on every machine and make the resolver slower to start; proxying was enough for the one program that asked.