Build the rollback mechanism, and test it

ADR 0059's recovery path: the pieces that run when the host will not start.

internal/upgrade -- two facts, neither of them the host judging its health.
Whether the executable this process started from has been replaced on disk, and
which version last completed a reconcile.

The first design was wrong and the tests caught it, not review. It asked
/proc/self/exe whether it was marked deleted. That is Linux procfs behaviour
rather than a fact about files, and it catches only unlink -- a binary swapped
by rename onto the same path reads as untouched, which is exactly what a
package manager does. Now the identity is captured at start and compared later:
no procfs, and neither case missed.

known-good is one bare line. The reader is a shell script on a machine where
the host is failing to start, so it must not need a parser to be present and
working. Written only after a clean apply, which is the whole claim -- not
health, because a disconnected node is ordinary and a failing resource is the
machine's problem rather than the binary's.

packaging/ -- the unit, the rollback unit, and the rollback script. The script
shares no code with the host and calls none of it: a binary that cannot start
cannot be its own recovery. POSIX sh, nothing that has to be installed. The
unit carries Restart=always with a comment saying why on-failure would break
every upgrade.

Both are tested and both sets of tests were confirmed to bite. Injecting five
faults broke exactly the intended tests -- except one, and chasing why it did
not found a placebo assertion I had written: `check "exits zero" ... "0" "0"`
compares a literal to itself and can never fail. Replaced with the real exit
code, after which the injection bites.

Also caught: an injection that produced a build failure rather than a test
failure, which my grep read as "no failure". Re-run so it compiled, and the
test did bite.

The script test runs in `make check`, so it is a gate rather than something
that was run once.

Verified against the real binary: known-good is written beside the store after
a clean apply and is NOT written after a failed one.
This commit is contained in:
2026-08-27 22:24:31 +02:00
parent 9a9937b7e6
commit f4143806c2
8 changed files with 544 additions and 1 deletions
+16
View File
@@ -24,6 +24,7 @@ import (
"github.com/novox/mesh-host/internal/inventory"
"github.com/novox/mesh-host/internal/profile"
"github.com/novox/mesh-host/internal/store"
"github.com/novox/mesh-host/internal/upgrade"
)
// version is stamped at build time. Unset in a development build, and said so rather than
@@ -312,6 +313,21 @@ func runApply(ctx context.Context, opts options, d *declaration.Declaration, sou
return applyErr
}
// Only now, and only after a clean apply: this version got as far as a completed
// reconcile, which is the whole of what "known good" claims (novox/hq ADR 0059). Not
// health — a disconnected node is ordinary, and a resource that fails is the machine's
// problem rather than the binary's.
//
// A failure to record is reported and does not fail the apply. The apply worked; what is
// lost is a rollback's ability to come back here, which is worse to hide than to say.
if version != "" {
if err := upgrade.RecordKnownGood(upgrade.KnownGoodPath(opts.state), version); err != nil {
fmt.Fprintf(os.Stderr,
"mesh-host: applied, but could not record %s as known-good: %v\n"+
" a rollback would have nothing to return to.\n", version, err)
}
}
if opts.json {
return writeJSON(report)
}