Raise machines that are somebody, and a bed that hands over nothing

Every machine in a bed is a clone of one base image, so all of them booted with
the same /etc/machine-id. systemd's DHCP client derives its client identifier
from that file and dnsmasq keys leases on the identifier rather than the MAC, so
four machines with four distinct MACs were handed one address and the host kept
one ARP entry for it. Whichever machine last answered an ARP request received
everybody's replies.

This is the fault behind every run lost to "flaky lab DNS": resolution that works
two times in three, pulls that succeed on a retry, and one machine out of four
being fine while the rest have no path at all. It survived an earlier diagnosis
that blamed resolver ordering, because reordering resolvers on a machine that has
just won the ARP race looks exactly like a fix.

Each machine is now given its own machine-id before the uplink lease is asked
for, and a check after addresses are applied refuses to go on if two machines
took the same one — the positive control this never had, since the fault is
invisible where it happens and unrecognisable where it surfaces.

The egress check also now demands five consecutive lookups rather than one. A
single answer is what let a machine resolving one query in three pass and then
die twenty minutes later inside a pull.

And fresh-mesh: whole-mesh-full's topology with genesis-single's honesty. The
four-machine bed loads thirty-four of the mesh's own images onto its machines
from the workstation because it does not build them, which is a shape no real
installation has and the same fiction the lab removed when it deleted its own
registry. This scenario names no images at all. The machines pull what is public,
the installer builds the control plane, and the mesh builds the rest — including,
last and deliberately, a module on a machine that did not build it.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
This commit is contained in:
2026-09-14 20:41:55 +02:00
parent 7fa1f1f0bb
commit babf08b9f8
4 changed files with 600 additions and 5 deletions
+48 -5
View File
@@ -89,9 +89,40 @@ export async function confirmEgress(
// the retry is tightened so a silent server costs seconds rather than the step. A broken uplink
// still fails the check above, before any of this.
await resilientResolver(name);
// **And then prove it is STEADY, because one success proved nothing.**
//
// The check above is satisfied by a single answer, and that is how a machine resolving one
// query in three passed it and then killed two installs twenty minutes later. What a run needs
// is not "a name resolved once" but "names resolve reliably", and those differ by exactly the
// failure that has cost the most time here. Consecutive, because alternating success and
// timeout is the observed shape — a total count would pass on the same machine.
const steady = await steadilyResolves(name, 5);
if (steady < 5) {
throw new EgressError(
`${machine} resolves ${UPSTREAM} only ${steady} time(s) in five consecutive tries.\n` +
` It has egress and an unreliable resolver, which does not fail here — it fails later, ` +
`inside a pull or a clone, as "could not resolve host" with the cause long out of view.\n` +
` The uplink's own resolver is the usual culprit and is deliberately last in the list; ` +
`a machine still failing this has something wrong upstream of the lab.`,
);
}
log(` ${machine} resolves steadily (5/5)`);
}
}
/** How many of `tries` consecutive lookups answered. Stops at the first failure. */
async function steadilyResolves(name: string, tries: number): Promise<number> {
for (let i = 0; i < tries; i++) {
const said = await incusOk(
["exec", name, "--", "sh", "-c",
`timeout 8 getent hosts ${UPSTREAM} >/dev/null && echo yes`], 20_000,
);
if (said?.trim() !== "yes") return i;
}
return tries;
}
async function resolves(name: string, waitSeconds: number): Promise<boolean> {
const deadline = Date.now() + waitSeconds * 1_000;
while (Date.now() < deadline) {
@@ -140,11 +171,22 @@ async function reaches(name: string, waitSeconds: number): Promise<string | null
* The shape of the problem is visible in `resolvectl status`: the machine has sensible global
* fallbacks, and the *link* carrying the default route has exactly one server — the uplink gateway.
* resolved will not reach for a global fallback while the link it is using has a server of its own,
* so one unanswered packet is one failed lookup. Under four machines pulling images at once that
* happens, and it has ended three runs long after the egress check passed.
* so one unanswered packet is one failed lookup.
*
* The uplink stays first, so the modelled path is still the one used and still proven by the check
* above. The others answer only when it does not.
* **The uplink goes last, and this is a preference, not a fix.** The uplink's resolver is a single
* dnsmasq with no unique knowledge any scenario machine needs — machines here address each other by
* address, and the mesh writes its own names into /etc/hosts — so there is no reason for anything
* to wait on it. It is kept, last, so DHCP-supplied names still answer.
*
* **It is written down as a preference because it was once mistaken for the cure.** Lookups were
* timing out two times in three; the uplink was first in the list; reordering it made five in five;
* the conclusion drew itself and was wrong. The machines were sharing ONE DHCP lease (see
* `distinguishMachines` in raise.ts), so which machine could resolve anything depended on which had
* last won an ARP race — and reordering resolvers on a machine that has just won looks exactly like
* a fix. Nothing below this line will save a bed whose machines share an address.
*
* The caches are flushed afterwards, because resolved remembers the failures it collected while
* the bad server was in front.
*/
async function resilientResolver(name: string): Promise<void> {
await incus([
@@ -152,7 +194,8 @@ async function resilientResolver(name: string): Promise<void> {
`link=$(ip -4 route show default | awk '{print $5}' | head -n1); ` +
`via=$(ip -4 route show default | awk '{print $3}' | head -n1); ` +
`if [ -n "$link" ] && command -v resolvectl >/dev/null 2>&1; then ` +
`resolvectl dns "$link" $via 1.1.1.1 8.8.8.8 >/dev/null 2>&1 || true; fi; true`,
`resolvectl dns "$link" 1.1.1.1 8.8.8.8 9.9.9.9 $via >/dev/null 2>&1 || true; ` +
`resolvectl flush-caches >/dev/null 2>&1 || true; fi; true`,
], 30_000);
}
+91
View File
@@ -307,9 +307,18 @@ export async function raise(
enter("waiting for machines to become usable");
await waitUntilAllUsable(created, readyTimeout, log);
// **Before the uplink lease is asked for, make sure each machine asks as itself.**
enter("giving each machine its own identity");
await distinguishMachines(created, log);
enter("applying declared addresses");
await applyAddresses(scenario, instanceId, byMachine, log);
// The positive control for the step above: if two machines share an uplink address, say so
// HERE, where it is one obvious sentence, rather than letting it surface an hour later as
// intermittent name resolution on some machines and not others.
await noTwoMachinesShareAnAddress(scenario, byMachine, log);
// Transit first: a gateway's default route points at it, so it has to exist.
enter("wiring the public networks together");
const transit = await raiseTransit(scenario, instanceId, log);
@@ -356,3 +365,85 @@ export async function raise(
throw new RaiseError(instanceId, step, cause);
}
}
/**
* Give every machine a machine-id of its own, before any of them asks for an uplink lease.
*
* **Every machine in a bed is a clone of one base image, so they all boot with the SAME
* `/etc/machine-id`.** systemd's DHCP client derives its client identifier from that file, and the
* uplink's dnsmasq keys leases on the client identifier rather than on the MAC — so four machines
* with four distinct MACs were all handed *the same address*, and the host kept one ARP entry for
* it. Whichever machine last answered an ARP request got everybody's replies.
*
* This is the fault behind every "the lab's DNS is flaky" run. It does not present as an address
* conflict; it presents as name resolution that works two times in three, as pulls that succeed on
* a retry, and as one machine out of four being fine — because that machine happened to be holding
* the address. It survived an earlier diagnosis that blamed resolver ordering, because reordering
* resolvers on a machine that has just won the ARP race does appear to fix it.
*
* **Ordering is the whole of it, and the first version got it wrong in the other direction.** That
* version also restarted networkd and waited for a new lease, which is what you would do to repair
* a machine already holding a shared address. Here there is nothing to repair yet: the uplink is
* still down at this point and `applyAddresses` is what asks for the lease. So this writes the
* identity and stops, and the request that follows is made as somebody.
*
* Machines without `systemd-machine-id-setup` are left alone; the step is advisory.
*/
async function distinguishMachines(names: string[], log: (m: string) => void): Promise<void> {
for (const name of names) {
const said = await incusOk([
"exec", name, "--", "sh", "-c",
// Regenerated rather than written: `systemd-machine-id-setup` owns the format, and a
// hand-made value that is not 32 hex characters is rejected by systemd at next boot.
`command -v systemd-machine-id-setup >/dev/null 2>&1 || { echo unchanged; exit 0; }; ` +
`rm -f /etc/machine-id && systemd-machine-id-setup >/dev/null 2>&1; ` +
`cat /etc/machine-id`,
], 60_000);
log(` ${name} is ${said?.trim().slice(0, 12) || "unchanged"}`);
}
}
/**
* Refuse to go on if two machines took the same uplink address.
*
* A check rather than a comment, because the failure it guards is invisible where it happens and
* unrecognisable where it surfaces. Every machine that declares egress is asked what address it
* holds; two the same is a stop, named as what it is.
*/
async function noTwoMachinesShareAnAddress(
scenario: Scenario,
byMachine: Map<string, string>,
log: (m: string) => void,
): Promise<void> {
const held = new Map<string, string[]>();
for (const [machine, spec] of Object.entries(scenario.machines)) {
if (!spec.egress || spec.at === "detached") continue;
const name = byMachine.get(machine);
if (!name) continue;
const said = (await incusOk([
"exec", name, "--", "sh", "-c",
// The `src` of the default route, NOT its last field — the first version of this took
// `$NF` and compared four machines' route METRIC, which is identical by construction and
// made the guard fire on every bed. A check that cannot be wrong is worth less than one
// that is read carefully once.
`ip -4 -o route show default 2>/dev/null | ` +
`awk '{for (i = 1; i <= NF; i++) if ($i == "src") { print $(i + 1); exit }}' | head -n1`,
], 30_000))?.trim();
if (!said) continue;
held.set(said, [...(held.get(said) ?? []), machine]);
}
const shared = [...held.entries()].filter(([, who]) => who.length > 1);
if (shared.length === 0) {
if (held.size > 0) log(` every machine with egress took an address of its own`);
return;
}
throw new Error(
`two machines took the SAME uplink address: ` +
shared.map(([a, who]) => `${a} held by ${who.join(" and ")}`).join("; ") + `.\n` +
` They are clones of one image, so they present one DHCP client identity unless each is ` +
`given its own machine-id before the lease is asked for.\n` +
` This does not fail as an address conflict. It fails later, as name resolution that works ` +
`about two times in three and as pulls that succeed on a retry, because the host holds one ` +
`ARP entry and whichever machine answered last receives the replies.`,
);
}