Configure the resolver rather than fight it
The first attempt wrote /etc/resolv.conf. On these images that is a symlink owned by systemd-resolved, so the file is either reverted or the link is broken — found by reading a running machine instead of assuming the change had worked. The real shape shows in resolvectl: the machine has sensible global fallbacks, and the link carrying the default route has exactly one server, the uplink gateway. resolved will not reach a global fallback while the link has a server of its own, so one unanswered packet is one failed lookup. Three runs have died that way, each long after the egress check passed. The uplink stays first, so the modelled path is still what is used and still what the check proves. Verified on a live machine: three servers on the link, uplink first, resolution intact.
This commit is contained in:
+17
-8
@@ -131,19 +131,28 @@ async function reaches(name: string, waitSeconds: number): Promise<string | null
|
||||
* explicitly and nothing else defaults.
|
||||
*/
|
||||
/**
|
||||
* Keep the uplink as the resolver and give it company.
|
||||
* Give the machine's uplink more than one resolver, through the thing that owns resolvers.
|
||||
*
|
||||
* Appended rather than replacing: the uplink is still asked first and still proves the modelled
|
||||
* path. The fallbacks only answer when it does not, which under a four-machine image pull is a
|
||||
* thing that happens and has ended three runs.
|
||||
* **Not by writing /etc/resolv.conf**, which was the first attempt and was wrong: on these images
|
||||
* that path is a symlink managed by systemd-resolved, so a file written over it is either reverted
|
||||
* or breaks the link. Checked on a running machine rather than assumed.
|
||||
*
|
||||
* The shape of the problem is visible in `resolvectl status`: the machine has sensible global
|
||||
* fallbacks, and the *link* carrying the default route has exactly one server — the uplink gateway.
|
||||
* resolved will not reach for a global fallback while the link it is using has a server of its own,
|
||||
* so one unanswered packet is one failed lookup. Under four machines pulling images at once that
|
||||
* happens, and it has ended three runs long after the egress check passed.
|
||||
*
|
||||
* The uplink stays first, so the modelled path is still the one used and still proven by the check
|
||||
* above. The others answer only when it does not.
|
||||
*/
|
||||
async function resilientResolver(name: string): Promise<void> {
|
||||
await incus([
|
||||
"exec", name, "--", "sh", "-c",
|
||||
`via=$(ip -4 route show default | awk '{print $3}' | head -n1); ` +
|
||||
`{ [ -n "$via" ] && printf 'nameserver %s\\n' "$via"; ` +
|
||||
`printf 'nameserver 1.1.1.1\\nnameserver 8.8.8.8\\n'; ` +
|
||||
`printf 'options timeout:2 attempts:3\\n'; } > /etc/resolv.conf; true`,
|
||||
`link=$(ip -4 route show default | awk '{print $5}' | head -n1); ` +
|
||||
`via=$(ip -4 route show default | awk '{print $3}' | head -n1); ` +
|
||||
`if [ -n "$link" ] && command -v resolvectl >/dev/null 2>&1; then ` +
|
||||
`resolvectl dns "$link" $via 1.1.1.1 8.8.8.8 >/dev/null 2>&1 || true; fi; true`,
|
||||
], 30_000);
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user