Files
mesh-host/packaging/nox-mesh-host-rollback
T
jschoubben f4143806c2 Build the rollback mechanism, and test it
ADR 0059's recovery path: the pieces that run when the host will not start.

internal/upgrade -- two facts, neither of them the host judging its health.
Whether the executable this process started from has been replaced on disk, and
which version last completed a reconcile.

The first design was wrong and the tests caught it, not review. It asked
/proc/self/exe whether it was marked deleted. That is Linux procfs behaviour
rather than a fact about files, and it catches only unlink -- a binary swapped
by rename onto the same path reads as untouched, which is exactly what a
package manager does. Now the identity is captured at start and compared later:
no procfs, and neither case missed.

known-good is one bare line. The reader is a shell script on a machine where
the host is failing to start, so it must not need a parser to be present and
working. Written only after a clean apply, which is the whole claim -- not
health, because a disconnected node is ordinary and a failing resource is the
machine's problem rather than the binary's.

packaging/ -- the unit, the rollback unit, and the rollback script. The script
shares no code with the host and calls none of it: a binary that cannot start
cannot be its own recovery. POSIX sh, nothing that has to be installed. The
unit carries Restart=always with a comment saying why on-failure would break
every upgrade.

Both are tested and both sets of tests were confirmed to bite. Injecting five
faults broke exactly the intended tests -- except one, and chasing why it did
not found a placebo assertion I had written: `check "exits zero" ... "0" "0"`
compares a literal to itself and can never fail. Replaced with the real exit
code, after which the injection bites.

Also caught: an injection that produced a build failure rather than a test
failure, which my grep read as "no failure". Re-run so it compiled, and the
test did bite.

The script test runs in `make check`, so it is a gate rather than something
that was run once.

Verified against the real binary: known-good is written beside the store after
a clean apply and is NOT written after a failed one.
2026-08-27 22:24:31 +02:00

67 lines
2.6 KiB
Bash
Executable File

#!/bin/sh
# Put the host back on the last version that worked.
#
# novox/hq ADR 0059. This runs when nox-mesh-host will not start, so it shares no code with it
# and calls none of it: a binary that cannot start cannot be its own recovery. POSIX sh, no
# bashisms, nothing that has to be installed.
#
# It is deliberately dull. Everything it does is one of: read a file, run the package manager,
# ask the service manager to try again.
set -eu
STATE_DIR="${MESH_HOST_STATE_DIR:-/var/lib/mesh-host}"
PKG_CACHE="${MESH_HOST_PKG_CACHE:-/var/cache/pacman/pkg}"
PACKAGE="${MESH_HOST_PACKAGE:-nox-mesh-host}"
KNOWN_GOOD="$STATE_DIR/known-good"
ATTEMPTED="$STATE_DIR/rollback-attempted"
say() { echo "nox-mesh-host-rollback: $*" >&2; }
# Roll back once. A second failure is a different diagnosis: the previously working binary also
# does not run, so the binary is not the problem — the machine is. Rolling back again would flap
# between two versions forever and bury the actual cause under a loop.
if [ -e "$ATTEMPTED" ]; then
say "already rolled back once, to $(cat "$ATTEMPTED" 2>/dev/null || echo unknown)."
say "the previous version also failed to start, so this is the machine and not the binary."
say "not rolling back again. this node needs a person."
exit 0
fi
# A machine whose host never completed a reconcile has no version to go back to. That is a real
# state rather than a fault: the node was never working, so the failure belongs to the
# installation. Guessing a version here is how a recovery becomes a second fault.
if [ ! -s "$KNOWN_GOOD" ]; then
say "no known-good version recorded — this host has never completed a reconcile."
say "there is nothing to roll back to. this is an installation failure, not an upgrade one."
exit 0
fi
VERSION="$(tr -d '[:space:]' < "$KNOWN_GOOD")"
if [ -z "$VERSION" ]; then
say "known-good is empty. refusing to guess."
exit 0
fi
PKG="$(ls "$PKG_CACHE"/"$PACKAGE"-"$VERSION"-*.pkg.tar.* 2>/dev/null | head -n 1 || true)"
if [ -z "$PKG" ]; then
say "known-good is $VERSION and no package for it is in $PKG_CACHE."
say "the cache was cleaned, or that version was never installed from here."
say "cannot roll back. this node needs a person."
exit 1
fi
say "rolling back to $VERSION ($PKG)"
printf '%s\n' "$VERSION" > "$ATTEMPTED"
if ! pacman -U --noconfirm "$PKG"; then
say "the package manager refused to install $PKG."
exit 1
fi
# reset-failed first, or the start limit that brought us here is still in force.
systemctl reset-failed "$PACKAGE".service 2>/dev/null || true
systemctl start "$PACKAGE".service
say "rolled back to $VERSION and started it. the node is on the previous version."