The registry is addressed the way every other machine is

novox/hq 04-ISSUES/024. The registry machine had its address set with
`ip addr add`; every other machine gets a systemd-networkd unit. That
one difference stalled the lab indefinitely.

An address set by hand leaves networkd waiting to configure a link it
was never told about, so the link sits at `configuring` for ever.
`systemd-networkd-wait-online` has TimeoutStartUSec=infinity, so
`network-online.target` is never reached — and Docker is ordered after
it. `docker load` then blocked on a socket whose daemon was queued
behind a target that would never come.

Measured before and after on the same scenario: stuck with five pending
systemd jobs and `docker` inactive; now `enp5s0 configured`, `docker`
active, no jobs, and the whole raise completes in 87.5s.

The guess in the issue was wrong, and it was wrong in the usual way —
stocking had just been changed, so stocking looked guilty. Stocking
takes 34s and always did.

Two things that made this cost hours rather than minutes are fixed with
it. Placing an image now waits for the container runtime to answer and
refuses after 120s naming what systemd is waiting on, so a stall becomes
a failure that says why instead of three stacked timeouts totalling 35
minutes. And the end-to-end test passes `onProgress`, so a raise says
what step it is on — it printed nothing at all until it finished, which
is why 35 minutes of nothing read as a slow test.
This commit is contained in:
2026-09-01 09:37:16 +02:00
parent c30d71e249
commit 3503ad990b
4 changed files with 100 additions and 10 deletions
+30
View File
@@ -47,6 +47,36 @@ function networkUnit(wire: Wire): string {
return lines.join("\n") + "\n";
}
/**
* Give one link a static address through systemd-networkd, and wait until networkd says it is
* configured.
*
* **Not `ip addr add`**, which is what this replaced and what cost a lab that could not finish.
* An address set by hand leaves the link `configuring` for ever, because networkd is still
* waiting to configure something it was never told about. `systemd-networkd-wait-online` then
* never returns — its timeout is `infinity` — so `network-online.target` is never reached, and
* **anything ordered after it never starts**. On these machines that is Docker, which meant
* `docker load` blocked on a socket whose daemon was queued behind a target that would never
* come. The lab stalled for thirty-five minutes with nothing to say.
*
* Every machine already did it this way. The registry did not, and it was the only one that
* needed Docker before anything else ran.
*/
export async function addressLink(
instanceName: string,
wire: Wire,
index = 0,
): Promise<void> {
const unit = networkUnit(wire);
await incus(
["exec", instanceName, "--", "sh", "-c",
`mkdir -p /etc/systemd/network && cat > /etc/systemd/network/10-mlab-${index}.network <<'MLAB'\n${unit}MLAB`],
30_000,
);
await incus(["exec", instanceName, "--", "systemctl", "enable", "--now", "systemd-networkd"], 60_000);
await incus(["exec", instanceName, "--", "systemctl", "restart", "systemd-networkd"], 60_000);
}
export async function applyAddresses(
scenario: Scenario,
instanceId: string,