The supervision was already right: a clean exit means the host stood aside, and the launcher's next turn runs what is on disk. Two things made it dead code — nothing told the running host a successor was waiting, and the rollback resolved its known-good version through pacman, which no machine here uses and which two of three operating systems do not have. Keeping a version rather than a path was the clue. Versions now live in directories named for them: - the launcher picks the newest delivered one every time round the loop, or the one a rollback pinned, or the host placed by hand when nothing is delivered; - the running host stands aside between reconciles, never inside one, by exiting cleanly — and returns nil so the launcher does not count it as a crash; - a completed reconcile retires what is older than the predecessor, keeping the predecessor because that is what a rollback starts, and never the running one; - rollback pins the predecessor instead of reinstalling a package: no package manager, no cache anyone may clean, same script on every operating system; - the report says which host version produced it, so 'behind' is answerable. Newest is when it arrived, never how the name sorts: '1.10' orders before '1.9', and ordering by name would start an older host and call it an upgrade. novox/hq ADR 0141. The delivery half — a module carrying the next host — follows; until then nothing delivers a version and every machine takes the fallback, which is what it does today.
76 lines
3.6 KiB
Bash
Executable File
76 lines
3.6 KiB
Bash
Executable File
#!/bin/sh
|
|
# Put the host back on the last version that worked.
|
|
#
|
|
# novox/hq ADR 0005 and ADR 0141. This runs when nox-mesh-host will not start, so it shares no code
|
|
# with it and calls none of it: a binary that cannot start cannot be its own recovery. POSIX sh, no
|
|
# bashisms, nothing that has to be installed.
|
|
#
|
|
# It is deliberately dull. Everything it does is one of: read a file, look at a directory, write a
|
|
# file.
|
|
#
|
|
# **It used to reinstall a package.** It read the known-good version and asked one operating system's
|
|
# package manager for it, out of that package manager's cache. Two things were wrong with that. No
|
|
# machine in this mesh had the host installed as a package, so the recovery could not run on any of
|
|
# them; and the host is built per operating system (ADR 0005), so a recovery written in one package
|
|
# manager's terms could not run on two of the three. Versions now live side by side in directories
|
|
# named for them, so going back is choosing a directory — which is the same on every machine.
|
|
set -eu
|
|
|
|
STATE_DIR="${MESH_HOST_STATE_DIR:-/var/lib/mesh-host}"
|
|
LIBEXEC="${MESH_HOST_LIBEXEC:-/usr/lib/nox-mesh-host}"
|
|
VERSIONS="$LIBEXEC/versions"
|
|
BINARY="nox-mesh-host"
|
|
|
|
KNOWN_GOOD="$STATE_DIR/known-good"
|
|
ATTEMPTED="$STATE_DIR/rollback-attempted"
|
|
PINNED="$STATE_DIR/rollback-pinned"
|
|
|
|
say() { echo "nox-mesh-host-rollback: $*" >&2; }
|
|
|
|
# Roll back once. A second failure is a different diagnosis: the previously working binary also
|
|
# does not run, so the binary is not the problem — the machine is. Rolling back again would flap
|
|
# between two versions forever and bury the actual cause under a loop.
|
|
if [ -e "$ATTEMPTED" ]; then
|
|
say "already rolled back once, to $(cat "$ATTEMPTED" 2>/dev/null || echo unknown)."
|
|
say "the previous version also failed to start, so this is the machine and not the binary."
|
|
say "not rolling back again. this node needs a person."
|
|
exit 0
|
|
fi
|
|
|
|
# A machine whose host never completed a reconcile has no version to go back to. That is a real
|
|
# state rather than a fault: the node was never working, so the failure belongs to the
|
|
# installation. Guessing a version here is how a recovery becomes a second fault.
|
|
if [ ! -s "$KNOWN_GOOD" ]; then
|
|
say "no known-good version recorded — this host has never completed a reconcile."
|
|
say "there is nothing to roll back to. this is an installation failure, not an upgrade one."
|
|
exit 0
|
|
fi
|
|
|
|
VERSION="$(tr -d '[:space:]' < "$KNOWN_GOOD")"
|
|
if [ -z "$VERSION" ]; then
|
|
say "known-good is empty. refusing to guess."
|
|
exit 0
|
|
fi
|
|
|
|
# The version that last worked may be the one that was placed by hand, which is not delivered and has
|
|
# no directory. Nothing to choose, and saying so is better than pinning a version that is not there —
|
|
# the launcher would ignore the pin and start the newest again, which is the binary that is failing.
|
|
if [ ! -x "$VERSIONS/$VERSION/$BINARY" ]; then
|
|
say "known-good is $VERSION and no such version is delivered under $VERSIONS."
|
|
say "it was retired, or that host was placed by hand and never delivered."
|
|
say "cannot roll back. this node needs a person."
|
|
exit 1
|
|
fi
|
|
|
|
say "rolling back to $VERSION ($VERSIONS/$VERSION/$BINARY)"
|
|
printf '%s\n' "$VERSION" > "$ATTEMPTED"
|
|
|
|
# The pin is what stops the launcher starting the newest again. Written last, so a failure above
|
|
# leaves the machine choosing for itself rather than pinned to something this script did not verify.
|
|
printf '%s\n' "$VERSION" > "$PINNED"
|
|
|
|
# Deliberately does NOT start anything. The launcher called this and will run the host next, so
|
|
# starting it here would run two. novox/hq ADR 0005 moved that responsibility; this script chooses a
|
|
# version and says so, and nothing else.
|
|
say "pinned $VERSION. the launcher will start it."
|