ADR 0061. Recovery was the most systemd-specific part of the host, and it is the part that must work on a machine where nothing else does -- which made unit-file syntax a poor place for it, because syntax cannot be tested and the one time it runs is the one time nobody can afford it wrong. So StartLimitBurst and OnFailure move into a launcher script that init starts instead of the host. The unit drops to start-at-boot and restart-on-exit, which OpenRC, runit, s6 and an Android init.rc can all express. Everything 0059 decided is kept: two watchdogs, roll back once, recovery is local, the rollback shares no code with the host. The counter is the whole mechanism, so it is what the tests are mostly about. Three real problems came out of writing them: A counter file holding "1 2" became "12" -- `tr -d [:space:]` concatenates rather than rejecting -- which is past the limit, so a HEALTHY node rolled itself back. Now it reads the first field and insists on a plain integer. The corrupt-counter test used "not-a-number", which shell arithmetic happens to evaluate to 0, so it passed with the guard removed and proved nothing. Replaced with values that discriminate: "5x" errors under set -e and kills the launcher, and "0x10" is read as HEX 16 -- past the limit, so again a healthy node rolls back. And the test harness itself was wrong. With `set -e` and a bare launcher call, removing a guard killed the script at the first corrupt case and silently skipped everything after -- reporting a full pass over tests that never ran. Every launcher call now records its failure instead of aborting. Same class as the placebo assertion found last time, and the reason to keep injecting faults rather than trusting green. Both scripts run in `make check`. 27 launcher tests, 9 rollback tests, all confirmed to bite.
73 lines
2.7 KiB
Bash
Executable File
73 lines
2.7 KiB
Bash
Executable File
#!/bin/sh
|
|
# Start the host, and decide what to do when it will not start.
|
|
#
|
|
# novox/hq ADR 0061. The init is asked for two things — start this at boot, start it again if it
|
|
# exits — and everything else is here, because this is the one piece that has to work on a
|
|
# machine where the host does not. Unit-file syntax cannot be tested; this can.
|
|
#
|
|
# POSIX sh, no bashisms, nothing that has to be installed.
|
|
set -eu
|
|
|
|
STATE_DIR="${MESH_HOST_STATE_DIR:-/var/lib/mesh-host}"
|
|
LIBEXEC="${MESH_HOST_LIBEXEC:-/usr/lib/nox-mesh-host}"
|
|
HOST="${MESH_HOST_BIN:-/usr/bin/nox-mesh-host}"
|
|
LIMIT="${MESH_HOST_START_LIMIT:-3}"
|
|
|
|
ATTEMPTS="$STATE_DIR/start-attempts"
|
|
HALTED="$STATE_DIR/halted"
|
|
|
|
say() { echo "nox-mesh-host-launch: $*" >&2; }
|
|
|
|
mkdir -p "$STATE_DIR"
|
|
|
|
# Halted: rolled back once and the previous version failed too, so the binary is not the problem.
|
|
# Nothing further is tried automatically. Exit zero — a supervisor restarting this forever is a
|
|
# slow visible loop rather than a crash loop, and the node stays down until a person looks.
|
|
if [ -e "$HALTED" ]; then
|
|
say "halted: $(cat "$HALTED" 2>/dev/null || echo 'reason not recorded')"
|
|
say "not starting the host. this node needs a person."
|
|
exit 0
|
|
fi
|
|
|
|
# Read the FIRST FIELD, then insist it is a plain integer.
|
|
#
|
|
# Stripping whitespace instead concatenates, and that is not a hypothetical: a counter file
|
|
# holding "1 2" became "12", which is past the limit, so a healthy node rolled itself back. An
|
|
# unreadable counter must fail towards "start normally", never towards "give up".
|
|
count=0
|
|
if [ -s "$ATTEMPTS" ]; then
|
|
read -r count _ < "$ATTEMPTS" 2>/dev/null || count=0
|
|
fi
|
|
case "${count:-}" in
|
|
'' | *[!0-9]*) count=0 ;;
|
|
esac
|
|
|
|
count=$((count + 1))
|
|
printf '%s\n' "$count" > "$ATTEMPTS"
|
|
|
|
if [ "$count" -gt "$LIMIT" ]; then
|
|
# The host has failed to get through a reconcile $LIMIT times running. The counter is
|
|
# cleared by the host itself on success, so reaching here means none of those starts
|
|
# worked — not that the machine has been up a long time.
|
|
if [ -e "$STATE_DIR/rollback-attempted" ]; then
|
|
say "the host failed $count times after a rollback. the previous version does not"
|
|
say "start either, so this is the machine and not the binary."
|
|
printf 'rolled back and still failing\n' > "$HALTED"
|
|
exit 0
|
|
fi
|
|
|
|
say "the host failed $count times. rolling back."
|
|
if "$LIBEXEC/rollback"; then
|
|
# Fresh count for the version we just installed: it deserves its own attempts, and
|
|
# without this it inherits a count already over the limit and halts immediately.
|
|
printf '0\n' > "$ATTEMPTS"
|
|
else
|
|
say "rollback failed. halting rather than restarting into the same failure."
|
|
printf 'rollback failed\n' > "$HALTED"
|
|
exit 0
|
|
fi
|
|
fi
|
|
|
|
# exec, so the host is what the supervisor watches and signals reach it directly.
|
|
exec "$HOST" run
|