The init is asked for start and restart; a launcher does the rest

ADR 0061. Recovery was the most systemd-specific part of the host, and it is
the part that must work on a machine where nothing else does -- which made
unit-file syntax a poor place for it, because syntax cannot be tested and the
one time it runs is the one time nobody can afford it wrong.

So StartLimitBurst and OnFailure move into a launcher script that init starts
instead of the host. The unit drops to start-at-boot and restart-on-exit, which
OpenRC, runit, s6 and an Android init.rc can all express. Everything 0059
decided is kept: two watchdogs, roll back once, recovery is local, the rollback
shares no code with the host.

The counter is the whole mechanism, so it is what the tests are mostly about.
Three real problems came out of writing them:

A counter file holding "1 2" became "12" -- `tr -d [:space:]` concatenates
rather than rejecting -- which is past the limit, so a HEALTHY node rolled
itself back. Now it reads the first field and insists on a plain integer.

The corrupt-counter test used "not-a-number", which shell arithmetic happens to
evaluate to 0, so it passed with the guard removed and proved nothing. Replaced
with values that discriminate: "5x" errors under set -e and kills the launcher,
and "0x10" is read as HEX 16 -- past the limit, so again a healthy node rolls
back.

And the test harness itself was wrong. With `set -e` and a bare launcher call,
removing a guard killed the script at the first corrupt case and silently
skipped everything after -- reporting a full pass over tests that never ran.
Every launcher call now records its failure instead of aborting. Same class as
the placebo assertion found last time, and the reason to keep injecting faults
rather than trusting green.

Both scripts run in `make check`. 27 launcher tests, 9 rollback tests, all
confirmed to bite.
This commit is contained in:
2026-08-27 23:45:57 +02:00
parent f4143806c2
commit 057f34f924
7 changed files with 222 additions and 28 deletions
+135
View File
@@ -0,0 +1,135 @@
#!/bin/sh
# Tests for nox-mesh-host-launch.
#
# The counter is the whole mechanism and it is the part to get wrong: never cleared and a node
# rolls back on a healthy boot; cleared too eagerly and it never rolls back at all. So the
# counter is what most of these assert.
set -eu
cd "$(dirname "$0")"
LAUNCH="$PWD/nox-mesh-host-launch"
PASS=0; FAIL=0
setup() {
WORK="$(mktemp -d)"
export MESH_HOST_STATE_DIR="$WORK/state"
export MESH_HOST_LIBEXEC="$WORK/libexec"
export MESH_HOST_BIN="$WORK/bin/nox-mesh-host"
export MESH_HOST_START_LIMIT=3
mkdir -p "$MESH_HOST_STATE_DIR" "$MESH_HOST_LIBEXEC" "$WORK/bin"
# A host that records being started. It exits immediately, which is what the launcher's
# exec makes indistinguishable from a host that ran for a week — the launcher is gone by
# then either way.
cat > "$MESH_HOST_BIN" <<'STUB'
#!/bin/sh
echo "$@" >> "$MESH_HOST_STATE_DIR/host.starts"
exit 0
STUB
cat > "$MESH_HOST_LIBEXEC/rollback" <<'STUB'
#!/bin/sh
echo rolled-back >> "$MESH_HOST_STATE_DIR/rollback.calls"
[ -n "${STUB_ROLLBACK_FAILS:-}" ] && exit 1
echo "$(cat "$MESH_HOST_STATE_DIR/known-good" 2>/dev/null)" > "$MESH_HOST_STATE_DIR/rollback-attempted"
exit 0
STUB
chmod +x "$MESH_HOST_BIN" "$MESH_HOST_LIBEXEC/rollback"
unset STUB_ROLLBACK_FAILS || true
}
# `|| true` on every launcher call above: a launcher that exits non-zero is something to
# ASSERT, not something to abort on. With `set -e` and a bare call, removing a guard from the
# launcher killed this script at the first corrupt-counter case and silently skipped the rest —
# reporting a full pass over tests that never ran.
check() { if [ "$3" = "$4" ]; then PASS=$((PASS+1)); printf ' ok %s\n' "$1"
else FAIL=$((FAIL+1)); printf ' FAIL %s\n %s\n got: %s\n expected: %s\n' "$1" "$2" "$3" "$4"; fi; }
count() { cat "$MESH_HOST_STATE_DIR/start-attempts" 2>/dev/null || echo MISSING; }
started() { [ -f "$MESH_HOST_STATE_DIR/host.starts" ] && echo yes || echo no; }
rolled() { [ -f "$MESH_HOST_STATE_DIR/rollback.calls" ] && echo yes || echo no; }
# --- the ordinary start -----------------------------------------------------------------------
setup
"$LAUNCH" >/dev/null 2>&1 || true
check "starts the host" "the common case, every boot" "$(started)" "yes"
check "counts the attempt" "the counter is what decides a rollback later" "$(count)" "1"
check "does not roll back" "a first start is not a failure" "$(rolled)" "no"
# --- failures below the limit -----------------------------------------------------------------
setup
i=1; while [ $i -le 3 ]; do "$LAUNCH" >/dev/null 2>&1 || true; i=$((i+1)); done
check "three starts do not trigger a rollback" "the limit is exceeded, not reached" "$(rolled)" "no"
check "counts them all" "" "$(count)" "3"
# --- past the limit ---------------------------------------------------------------------------
setup
echo "1.4.2" > "$MESH_HOST_STATE_DIR/known-good"
i=1; while [ $i -le 4 ]; do "$LAUNCH" >/dev/null 2>&1 || true; i=$((i+1)); done
check "the fourth start rolls back" "three failures is a binary that does not work" "$(rolled)" "yes"
check "and still starts the host" "the rolled-back version has to be run" "$(started)" "yes"
check "resets the counter after rolling back" "the new version deserves its own attempts, or it halts at once" \
"$(count)" "0"
# --- the host clears the counter on success ----------------------------------------------------
setup
i=1; while [ $i -le 2 ]; do "$LAUNCH" >/dev/null 2>&1 || true; i=$((i+1)); done
printf '0\n' > "$MESH_HOST_STATE_DIR/start-attempts" # what the host does on a completed reconcile
i=1; while [ $i -le 3 ]; do "$LAUNCH" >/dev/null 2>&1 || true; i=$((i+1)); done
check "a cleared counter prevents a rollback" "a node up for months must not roll back on a healthy boot" \
"$(rolled)" "no"
# --- rolled back once already -------------------------------------------------------------------
setup
echo "1.4.2" > "$MESH_HOST_STATE_DIR/known-good"
echo "1.4.2" > "$MESH_HOST_STATE_DIR/rollback-attempted"
i=1; while [ $i -le 4 ]; do "$LAUNCH" >/dev/null 2>&1 || true; i=$((i+1)); done
check "does not roll back twice" "the previous version failing too means the machine, not the binary" \
"$(rolled)" "no"
check "halts instead" "" "$([ -f "$MESH_HOST_STATE_DIR/halted" ] && echo halted || echo running)" "halted"
# --- halted stays halted --------------------------------------------------------------------------
setup
echo "rolled back and still failing" > "$MESH_HOST_STATE_DIR/halted"
"$LAUNCH" >/dev/null 2>&1 || true
check "a halted node does not start the host" "nothing further is tried automatically" "$(started)" "no"
set +e; "$LAUNCH" >/dev/null 2>&1; RC=$?; set -e
check "a halted node exits zero" "a supervisor loop that is slow and visible beats a crash loop" "$RC" "0"
# --- the rollback itself fails ----------------------------------------------------------------------
setup
echo "1.4.2" > "$MESH_HOST_STATE_DIR/known-good"
STUB_ROLLBACK_FAILS=1; export STUB_ROLLBACK_FAILS
i=1; while [ $i -le 4 ]; do "$LAUNCH" >/dev/null 2>&1 || true; i=$((i+1)); done
check "a failed rollback halts" "restarting into the same failure would loop forever" \
"$([ -f "$MESH_HOST_STATE_DIR/halted" ] && echo halted || echo running)" "halted"
# Counted, not "was it ever started": the first three attempts DID start it, correctly, and
# only the fourth must not. An earlier version of this asserted the host was never started and
# failed for that reason rather than for a fault.
check "and does not start it on the halting attempt" "three starts, not four" \
"$(wc -l < "$MESH_HOST_STATE_DIR/host.starts" 2>/dev/null || echo 0)" "3"
# --- a corrupt counter ------------------------------------------------------------------------------
#
# The values here are chosen because they DISCRIMINATE. An earlier version used
# "not-a-number", which shell arithmetic happens to evaluate to 0 — so the test passed with the
# guard removed and proved nothing. These two do not:
#
# 5x shell arithmetic errors, and under `set -e` the launcher dies without starting the host
# 0x10 is read as HEX 16 — past the limit, so a healthy node would roll back for no reason
for corrupt in "5x" "0x10" "1 2" ""; do
setup
echo "1.4.2" > "$MESH_HOST_STATE_DIR/known-good"
printf '%s\n' "$corrupt" > "$MESH_HOST_STATE_DIR/start-attempts"
"$LAUNCH" >/dev/null 2>&1 || true
# The exact number is not the property — "1 2" legitimately recovers a leading 1, while
# "5x" is rejected to 0. What must hold for every one of them is that the launcher
# survives its own state and does not read it as "past the limit".
check "corrupt counter [$corrupt]: starts the host" "the launcher must not die on its own state" \
"$(started)" "yes"
check "corrupt counter [$corrupt]: does not roll back" "a healthy node must not roll back on a bad counter" \
"$(rolled)" "no"
check "corrupt counter [$corrupt]: counter is a sane integer" "it is written back for the next start to read" \
"$(count | grep -cE '^[0-9]+$')" "1"
done
printf '\nlaunch: %d passed, %d failed\n' "$PASS" "$FAIL"
[ "$FAIL" -eq 0 ]