Raise machines that are somebody, and a bed that hands over nothing
Every machine in a bed is a clone of one base image, so all of them booted with the same /etc/machine-id. systemd's DHCP client derives its client identifier from that file and dnsmasq keys leases on the identifier rather than the MAC, so four machines with four distinct MACs were handed one address and the host kept one ARP entry for it. Whichever machine last answered an ARP request received everybody's replies. This is the fault behind every run lost to "flaky lab DNS": resolution that works two times in three, pulls that succeed on a retry, and one machine out of four being fine while the rest have no path at all. It survived an earlier diagnosis that blamed resolver ordering, because reordering resolvers on a machine that has just won the ARP race looks exactly like a fix. Each machine is now given its own machine-id before the uplink lease is asked for, and a check after addresses are applied refuses to go on if two machines took the same one — the positive control this never had, since the fault is invisible where it happens and unrecognisable where it surfaces. The egress check also now demands five consecutive lookups rather than one. A single answer is what let a machine resolving one query in three pass and then die twenty minutes later inside a pull. And fresh-mesh: whole-mesh-full's topology with genesis-single's honesty. The four-machine bed loads thirty-four of the mesh's own images onto its machines from the workstation because it does not build them, which is a shape no real installation has and the same fiction the lab removed when it deleted its own registry. This scenario names no images at all. The machines pull what is public, the installer builds the control plane, and the mesh builds the rest — including, last and deliberately, a module on a machine that did not build it. Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
This commit is contained in:
+48
-5
@@ -89,9 +89,40 @@ export async function confirmEgress(
|
||||
// the retry is tightened so a silent server costs seconds rather than the step. A broken uplink
|
||||
// still fails the check above, before any of this.
|
||||
await resilientResolver(name);
|
||||
|
||||
// **And then prove it is STEADY, because one success proved nothing.**
|
||||
//
|
||||
// The check above is satisfied by a single answer, and that is how a machine resolving one
|
||||
// query in three passed it and then killed two installs twenty minutes later. What a run needs
|
||||
// is not "a name resolved once" but "names resolve reliably", and those differ by exactly the
|
||||
// failure that has cost the most time here. Consecutive, because alternating success and
|
||||
// timeout is the observed shape — a total count would pass on the same machine.
|
||||
const steady = await steadilyResolves(name, 5);
|
||||
if (steady < 5) {
|
||||
throw new EgressError(
|
||||
`${machine} resolves ${UPSTREAM} only ${steady} time(s) in five consecutive tries.\n` +
|
||||
` It has egress and an unreliable resolver, which does not fail here — it fails later, ` +
|
||||
`inside a pull or a clone, as "could not resolve host" with the cause long out of view.\n` +
|
||||
` The uplink's own resolver is the usual culprit and is deliberately last in the list; ` +
|
||||
`a machine still failing this has something wrong upstream of the lab.`,
|
||||
);
|
||||
}
|
||||
log(` ${machine} resolves steadily (5/5)`);
|
||||
}
|
||||
}
|
||||
|
||||
/** How many of `tries` consecutive lookups answered. Stops at the first failure. */
|
||||
async function steadilyResolves(name: string, tries: number): Promise<number> {
|
||||
for (let i = 0; i < tries; i++) {
|
||||
const said = await incusOk(
|
||||
["exec", name, "--", "sh", "-c",
|
||||
`timeout 8 getent hosts ${UPSTREAM} >/dev/null && echo yes`], 20_000,
|
||||
);
|
||||
if (said?.trim() !== "yes") return i;
|
||||
}
|
||||
return tries;
|
||||
}
|
||||
|
||||
async function resolves(name: string, waitSeconds: number): Promise<boolean> {
|
||||
const deadline = Date.now() + waitSeconds * 1_000;
|
||||
while (Date.now() < deadline) {
|
||||
@@ -140,11 +171,22 @@ async function reaches(name: string, waitSeconds: number): Promise<string | null
|
||||
* The shape of the problem is visible in `resolvectl status`: the machine has sensible global
|
||||
* fallbacks, and the *link* carrying the default route has exactly one server — the uplink gateway.
|
||||
* resolved will not reach for a global fallback while the link it is using has a server of its own,
|
||||
* so one unanswered packet is one failed lookup. Under four machines pulling images at once that
|
||||
* happens, and it has ended three runs long after the egress check passed.
|
||||
* so one unanswered packet is one failed lookup.
|
||||
*
|
||||
* The uplink stays first, so the modelled path is still the one used and still proven by the check
|
||||
* above. The others answer only when it does not.
|
||||
* **The uplink goes last, and this is a preference, not a fix.** The uplink's resolver is a single
|
||||
* dnsmasq with no unique knowledge any scenario machine needs — machines here address each other by
|
||||
* address, and the mesh writes its own names into /etc/hosts — so there is no reason for anything
|
||||
* to wait on it. It is kept, last, so DHCP-supplied names still answer.
|
||||
*
|
||||
* **It is written down as a preference because it was once mistaken for the cure.** Lookups were
|
||||
* timing out two times in three; the uplink was first in the list; reordering it made five in five;
|
||||
* the conclusion drew itself and was wrong. The machines were sharing ONE DHCP lease (see
|
||||
* `distinguishMachines` in raise.ts), so which machine could resolve anything depended on which had
|
||||
* last won an ARP race — and reordering resolvers on a machine that has just won looks exactly like
|
||||
* a fix. Nothing below this line will save a bed whose machines share an address.
|
||||
*
|
||||
* The caches are flushed afterwards, because resolved remembers the failures it collected while
|
||||
* the bad server was in front.
|
||||
*/
|
||||
async function resilientResolver(name: string): Promise<void> {
|
||||
await incus([
|
||||
@@ -152,7 +194,8 @@ async function resilientResolver(name: string): Promise<void> {
|
||||
`link=$(ip -4 route show default | awk '{print $5}' | head -n1); ` +
|
||||
`via=$(ip -4 route show default | awk '{print $3}' | head -n1); ` +
|
||||
`if [ -n "$link" ] && command -v resolvectl >/dev/null 2>&1; then ` +
|
||||
`resolvectl dns "$link" $via 1.1.1.1 8.8.8.8 >/dev/null 2>&1 || true; fi; true`,
|
||||
`resolvectl dns "$link" 1.1.1.1 8.8.8.8 9.9.9.9 $via >/dev/null 2>&1 || true; ` +
|
||||
`resolvectl flush-caches >/dev/null 2>&1 || true; fi; true`,
|
||||
], 30_000);
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user