Build the rollback mechanism, and test it
ADR 0059's recovery path: the pieces that run when the host will not start. internal/upgrade -- two facts, neither of them the host judging its health. Whether the executable this process started from has been replaced on disk, and which version last completed a reconcile. The first design was wrong and the tests caught it, not review. It asked /proc/self/exe whether it was marked deleted. That is Linux procfs behaviour rather than a fact about files, and it catches only unlink -- a binary swapped by rename onto the same path reads as untouched, which is exactly what a package manager does. Now the identity is captured at start and compared later: no procfs, and neither case missed. known-good is one bare line. The reader is a shell script on a machine where the host is failing to start, so it must not need a parser to be present and working. Written only after a clean apply, which is the whole claim -- not health, because a disconnected node is ordinary and a failing resource is the machine's problem rather than the binary's. packaging/ -- the unit, the rollback unit, and the rollback script. The script shares no code with the host and calls none of it: a binary that cannot start cannot be its own recovery. POSIX sh, nothing that has to be installed. The unit carries Restart=always with a comment saying why on-failure would break every upgrade. Both are tested and both sets of tests were confirmed to bite. Injecting five faults broke exactly the intended tests -- except one, and chasing why it did not found a placebo assertion I had written: `check "exits zero" ... "0" "0"` compares a literal to itself and can never fail. Replaced with the real exit code, after which the injection bites. Also caught: an injection that produced a build failure rather than a test failure, which my grep read as "no failure". Re-run so it compiled, and the test did bite. The script test runs in `make check`, so it is a gate rather than something that was run once. Verified against the real binary: known-good is written beside the store after a clean apply and is NOT written after a failed one.
This commit is contained in:
Executable
+66
@@ -0,0 +1,66 @@
|
||||
#!/bin/sh
|
||||
# Put the host back on the last version that worked.
|
||||
#
|
||||
# novox/hq ADR 0059. This runs when nox-mesh-host will not start, so it shares no code with it
|
||||
# and calls none of it: a binary that cannot start cannot be its own recovery. POSIX sh, no
|
||||
# bashisms, nothing that has to be installed.
|
||||
#
|
||||
# It is deliberately dull. Everything it does is one of: read a file, run the package manager,
|
||||
# ask the service manager to try again.
|
||||
set -eu
|
||||
|
||||
STATE_DIR="${MESH_HOST_STATE_DIR:-/var/lib/mesh-host}"
|
||||
PKG_CACHE="${MESH_HOST_PKG_CACHE:-/var/cache/pacman/pkg}"
|
||||
PACKAGE="${MESH_HOST_PACKAGE:-nox-mesh-host}"
|
||||
|
||||
KNOWN_GOOD="$STATE_DIR/known-good"
|
||||
ATTEMPTED="$STATE_DIR/rollback-attempted"
|
||||
|
||||
say() { echo "nox-mesh-host-rollback: $*" >&2; }
|
||||
|
||||
# Roll back once. A second failure is a different diagnosis: the previously working binary also
|
||||
# does not run, so the binary is not the problem — the machine is. Rolling back again would flap
|
||||
# between two versions forever and bury the actual cause under a loop.
|
||||
if [ -e "$ATTEMPTED" ]; then
|
||||
say "already rolled back once, to $(cat "$ATTEMPTED" 2>/dev/null || echo unknown)."
|
||||
say "the previous version also failed to start, so this is the machine and not the binary."
|
||||
say "not rolling back again. this node needs a person."
|
||||
exit 0
|
||||
fi
|
||||
|
||||
# A machine whose host never completed a reconcile has no version to go back to. That is a real
|
||||
# state rather than a fault: the node was never working, so the failure belongs to the
|
||||
# installation. Guessing a version here is how a recovery becomes a second fault.
|
||||
if [ ! -s "$KNOWN_GOOD" ]; then
|
||||
say "no known-good version recorded — this host has never completed a reconcile."
|
||||
say "there is nothing to roll back to. this is an installation failure, not an upgrade one."
|
||||
exit 0
|
||||
fi
|
||||
|
||||
VERSION="$(tr -d '[:space:]' < "$KNOWN_GOOD")"
|
||||
if [ -z "$VERSION" ]; then
|
||||
say "known-good is empty. refusing to guess."
|
||||
exit 0
|
||||
fi
|
||||
|
||||
PKG="$(ls "$PKG_CACHE"/"$PACKAGE"-"$VERSION"-*.pkg.tar.* 2>/dev/null | head -n 1 || true)"
|
||||
if [ -z "$PKG" ]; then
|
||||
say "known-good is $VERSION and no package for it is in $PKG_CACHE."
|
||||
say "the cache was cleaned, or that version was never installed from here."
|
||||
say "cannot roll back. this node needs a person."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
say "rolling back to $VERSION ($PKG)"
|
||||
printf '%s\n' "$VERSION" > "$ATTEMPTED"
|
||||
|
||||
if ! pacman -U --noconfirm "$PKG"; then
|
||||
say "the package manager refused to install $PKG."
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# reset-failed first, or the start limit that brought us here is still in force.
|
||||
systemctl reset-failed "$PACKAGE".service 2>/dev/null || true
|
||||
systemctl start "$PACKAGE".service
|
||||
|
||||
say "rolled back to $VERSION and started it. the node is on the previous version."
|
||||
@@ -0,0 +1,8 @@
|
||||
[Unit]
|
||||
Description=Roll the Novox Mesh node host back to the last version that started
|
||||
# No OnFailure of its own. If the rollback fails there is nothing further to try
|
||||
# automatically, and the node needs a person.
|
||||
|
||||
[Service]
|
||||
Type=oneshot
|
||||
ExecStart=/usr/lib/nox-mesh-host/rollback
|
||||
@@ -0,0 +1,21 @@
|
||||
[Unit]
|
||||
Description=Novox Mesh node host
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
# When the supervisor gives up, recover rather than leaving the node quiet — a host that will
|
||||
# not start looks exactly like a machine somebody switched off (novox/hq ADR 0059).
|
||||
OnFailure=nox-mesh-host-rollback.service
|
||||
|
||||
[Service]
|
||||
Type=notify
|
||||
ExecStart=/usr/bin/nox-mesh-host run
|
||||
# always, NOT on-failure: the host restarts onto a new binary by exiting CLEANLY
|
||||
# (novox/hq ADR 0057), and on-failure would leave an upgraded node stopped.
|
||||
Restart=always
|
||||
RestartSec=5s
|
||||
StartLimitBurst=3
|
||||
StartLimitIntervalSec=120
|
||||
StateDirectory=mesh-host
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
Executable
+98
@@ -0,0 +1,98 @@
|
||||
#!/bin/sh
|
||||
# Tests for nox-mesh-host-rollback.
|
||||
#
|
||||
# It runs on a machine where the host will not start, which is the one moment nobody can afford
|
||||
# it to be wrong — and the one moment it is hardest to debug. So it is tested here, against a
|
||||
# real filesystem, with a stub package manager that records what it was asked to do.
|
||||
set -eu
|
||||
cd "$(dirname "$0")"
|
||||
SCRIPT="$PWD/nox-mesh-host-rollback"
|
||||
PASS=0; FAIL=0
|
||||
|
||||
setup() {
|
||||
WORK="$(mktemp -d)"
|
||||
export MESH_HOST_STATE_DIR="$WORK/state"
|
||||
export MESH_HOST_PKG_CACHE="$WORK/cache"
|
||||
export MESH_HOST_PACKAGE="nox-mesh-host"
|
||||
mkdir -p "$MESH_HOST_STATE_DIR" "$MESH_HOST_PKG_CACHE" "$WORK/bin"
|
||||
|
||||
# Stubs on PATH. Not mocks of the script's own logic — the boundary is real commands, and
|
||||
# these record the calls so a test can assert what the script asked the machine to do.
|
||||
cat > "$WORK/bin/pacman" <<'STUB'
|
||||
#!/bin/sh
|
||||
echo "$@" >> "$MESH_HOST_STATE_DIR/pacman.calls"
|
||||
[ -n "${STUB_PACMAN_FAILS:-}" ] && exit 1
|
||||
exit 0
|
||||
STUB
|
||||
cat > "$WORK/bin/systemctl" <<'STUB'
|
||||
#!/bin/sh
|
||||
echo "$@" >> "$MESH_HOST_STATE_DIR/systemctl.calls"
|
||||
exit 0
|
||||
STUB
|
||||
chmod +x "$WORK/bin/pacman" "$WORK/bin/systemctl"
|
||||
PATH="$WORK/bin:$PATH"; export PATH
|
||||
unset STUB_PACMAN_FAILS || true
|
||||
}
|
||||
|
||||
check() { # name, condition-description, actual, expected
|
||||
if [ "$3" = "$4" ]; then PASS=$((PASS+1)); printf ' ok %s\n' "$1"
|
||||
else FAIL=$((FAIL+1)); printf ' FAIL %s\n %s\n got: %s\n expected: %s\n' "$1" "$2" "$3" "$4"; fi
|
||||
}
|
||||
|
||||
# --- a normal rollback ---------------------------------------------------------------------
|
||||
setup
|
||||
echo "1.4.2" > "$MESH_HOST_STATE_DIR/known-good"
|
||||
touch "$MESH_HOST_PKG_CACHE/nox-mesh-host-1.4.2-1-x86_64.pkg.tar.zst"
|
||||
"$SCRIPT" >/dev/null 2>&1
|
||||
check "installs the known-good version" "pacman is asked to install the cached package" \
|
||||
"$(grep -c 'nox-mesh-host-1.4.2' "$MESH_HOST_STATE_DIR/pacman.calls" 2>/dev/null || echo 0)" "1"
|
||||
check "resets the start limit before starting" "reset-failed precedes start" \
|
||||
"$(head -1 "$MESH_HOST_STATE_DIR/systemctl.calls" | cut -d' ' -f1)" "reset-failed"
|
||||
check "starts the host again" "systemctl start is called" \
|
||||
"$(grep -c '^start ' "$MESH_HOST_STATE_DIR/systemctl.calls" 2>/dev/null || echo 0)" "1"
|
||||
check "records that it rolled back" "the attempted marker holds the version" \
|
||||
"$(cat "$MESH_HOST_STATE_DIR/rollback-attempted" 2>/dev/null || echo MISSING)" "1.4.2"
|
||||
|
||||
# --- it rolls back only once ---------------------------------------------------------------
|
||||
setup
|
||||
echo "1.4.2" > "$MESH_HOST_STATE_DIR/known-good"
|
||||
echo "1.4.2" > "$MESH_HOST_STATE_DIR/rollback-attempted"
|
||||
touch "$MESH_HOST_PKG_CACHE/nox-mesh-host-1.4.2-1-x86_64.pkg.tar.zst"
|
||||
"$SCRIPT" >/dev/null 2>&1
|
||||
check "does not roll back twice" "a second failure is the machine, not the binary" \
|
||||
"$([ -f "$MESH_HOST_STATE_DIR/pacman.calls" ] && echo called || echo not-called)" "not-called"
|
||||
|
||||
# --- nothing to roll back to ---------------------------------------------------------------
|
||||
setup
|
||||
set +e; "$SCRIPT" >/dev/null 2>&1; RC=$?; set -e
|
||||
check "no known-good: does nothing" "a host that never reconciled has no version to return to" \
|
||||
"$([ -f "$MESH_HOST_STATE_DIR/pacman.calls" ] && echo called || echo not-called)" "not-called"
|
||||
# The exit code is asserted from a real run, not from a literal. An earlier version of this
|
||||
# compared "0" to "0" and could not fail — which hid an injected fault that made the script die
|
||||
# here instead of returning cleanly.
|
||||
check "no known-good: exits zero" "an installation failure is not a rollback failure" "$RC" "0"
|
||||
|
||||
setup
|
||||
printf ' \n' > "$MESH_HOST_STATE_DIR/known-good"
|
||||
"$SCRIPT" >/dev/null 2>&1
|
||||
check "blank known-good: refuses to guess" "installing nothing and reporting success is the fault this prevents" \
|
||||
"$([ -f "$MESH_HOST_STATE_DIR/pacman.calls" ] && echo called || echo not-called)" "not-called"
|
||||
|
||||
# --- the cache was cleaned ------------------------------------------------------------------
|
||||
setup
|
||||
echo "1.4.2" > "$MESH_HOST_STATE_DIR/known-good"
|
||||
set +e; "$SCRIPT" >/dev/null 2>&1; RC=$?; set -e
|
||||
check "missing package: fails loudly" "cannot roll back, and says so rather than reporting success" "$RC" "1"
|
||||
|
||||
# --- the package manager refuses -------------------------------------------------------------
|
||||
setup
|
||||
echo "1.4.2" > "$MESH_HOST_STATE_DIR/known-good"
|
||||
touch "$MESH_HOST_PKG_CACHE/nox-mesh-host-1.4.2-1-x86_64.pkg.tar.zst"
|
||||
STUB_PACMAN_FAILS=1 ; export STUB_PACMAN_FAILS
|
||||
set +e; "$SCRIPT" >/dev/null 2>&1; RC=$?; set -e
|
||||
check "pacman fails: does not start the host" "starting the broken binary again would loop" \
|
||||
"$([ -f "$MESH_HOST_STATE_DIR/systemctl.calls" ] && echo started || echo not-started)" "not-started"
|
||||
check "pacman fails: exits non-zero" "a failed rollback is a failure" "$RC" "1"
|
||||
|
||||
printf '\nrollback: %d passed, %d failed\n' "$PASS" "$FAIL"
|
||||
[ "$FAIL" -eq 0 ]
|
||||
Reference in New Issue
Block a user