Containers coming back is not the mesh coming back

E2 asserted that every container was running after a reboot and stopped there.
The runtime restarts containers by itself; what makes a machine part of a mesh is
an agent listening for what it should be. A machine whose containers returned and
whose agent did not looks healthy and cannot be told anything.

The installer is explicit that a host started the way the lab starts it does not
survive a reboot, so this may now fail — and if it does, it is the packaging gap
the design already records under what is not yet true, not a fault in the mesh.
Better a named failure than a pass that means less than it appears to.

Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
This commit is contained in:
2026-09-14 23:44:24 +02:00
parent 9b8b21ac17
commit 3ebf0bb38c
2 changed files with 24 additions and 13 deletions
+12 -1
View File
@@ -901,7 +901,18 @@ before(async () => {
const ps = (await on(CONTROL, `docker ps -a --format '{{.Names}}\t{{.Status}}'`)).out;
assert.equal(missing.length, 0,
`after a reboot these are not running: ${missing.join(", ")}\n\ncontainers:\n${ps}`);
return ps;
// **Containers coming back is not the mesh coming back.** The runtime restarts containers on
// its own; what makes a machine part of a mesh is an agent listening for what it should be. A
// machine whose containers returned and whose agent did not looks healthy and cannot be told
// anything — and the installer is explicit that a host started the way the lab starts it, with
// --host-in-background, does not survive a reboot. So this asks, and a failure here is the
// packaging gap the design already records rather than a fault in the mesh.
const agent = (await on(CONTROL, `pgrep -af '[m]esh-host run' | head -3`)).out.trim();
assert.ok(agent,
`every container came back and no host agent did, so the machine is running the right ` +
`things and can no longer be told anything:\n${ps}`);
return `${ps}\n${agent}`;
});
stateOutcome();