The registry is addressed the way every other machine is

novox/hq 04-ISSUES/024. The registry machine had its address set with
`ip addr add`; every other machine gets a systemd-networkd unit. That
one difference stalled the lab indefinitely.

An address set by hand leaves networkd waiting to configure a link it
was never told about, so the link sits at `configuring` for ever.
`systemd-networkd-wait-online` has TimeoutStartUSec=infinity, so
`network-online.target` is never reached — and Docker is ordered after
it. `docker load` then blocked on a socket whose daemon was queued
behind a target that would never come.

Measured before and after on the same scenario: stuck with five pending
systemd jobs and `docker` inactive; now `enp5s0 configured`, `docker`
active, no jobs, and the whole raise completes in 87.5s.

The guess in the issue was wrong, and it was wrong in the usual way —
stocking had just been changed, so stocking looked guilty. Stocking
takes 34s and always did.

Two things that made this cost hours rather than minutes are fixed with
it. Placing an image now waits for the container runtime to answer and
refuses after 120s naming what systemd is waiting on, so a stall becomes
a failure that says why instead of three stacked timeouts totalling 35
minutes. And the end-to-end test passes `onProgress`, so a raise says
what step it is on — it printed nothing at all until it finished, which
is why 35 minutes of nothing read as a slow test.
This commit is contained in:
2026-09-01 09:37:16 +02:00
parent c30d71e249
commit 3503ad990b
4 changed files with 100 additions and 10 deletions
+20 -8
View File
@@ -22,6 +22,7 @@ import { spawn } from "node:child_process";
import { incus, incusOk, succeeds } from "../incus/client.ts";
import { macFor, networkName } from "./names.ts";
import { addressLink } from "./address.ts";
import { BASE_IMAGE_ALIAS, placeImage } from "./place.ts";
import { mkdtemp, rm } from "node:fs/promises";
import { tmpdir } from "node:os";
@@ -296,14 +297,25 @@ export async function raiseRegistry(
await succeeds(["start", name], 60_000);
await waitForAgent(name);
// Address it by MAC, never by interface name: a machine with a container runtime has a
// `docker0` that sorts before `enp5s0`, and naive selection configures that instead — which
// then overlaps the segment and breaks routing on the machine.
const mac = macFor(instanceId, "registry", 0);
await incus(["exec", name, "--", "sh", "-c",
`dev=$(ip -o link | awk -F': ' '/${mac}/ {print $2}' | head -1); ` +
`[ -n "$dev" ] && ip addr add ${address}${prefix} dev "$dev" 2>/dev/null; ` +
`[ -n "$dev" ] && ip link set "$dev" up`], 60_000);
// Addressed the way every other machine is: a systemd-networkd unit matching the MAC.
//
// **This used to be `ip addr add`, and it stalled the lab.** An address set by hand leaves
// networkd waiting to configure a link it was never told about, so the link sits at
// `configuring`, `systemd-networkd-wait-online` never returns — its timeout is `infinity` —
// and `network-online.target` is never reached. Docker is ordered after that target, so
// `docker load` two lines below blocked on a socket whose daemon was queued behind a target
// that would never come.
//
// Matching on MAC and not on interface name is still the rule: a machine with a container
// runtime has a `docker0` that sorts before `enp5s0`, and naive selection configures that.
await addressLink(name, {
device: "eth0",
mac: macFor(instanceId, "registry", 0),
addresses: [`${address}${prefix}`],
// The registry takes the segment's default. It carried no MTU before this and still does
// not: what a scenario sets an MTU for is the path under test, and this is scenery.
mtu: undefined,
});
log(` registry on ${segment.name} at ${address}`);