Build the rollback mechanism, and test it
ADR 0059's recovery path: the pieces that run when the host will not start. internal/upgrade -- two facts, neither of them the host judging its health. Whether the executable this process started from has been replaced on disk, and which version last completed a reconcile. The first design was wrong and the tests caught it, not review. It asked /proc/self/exe whether it was marked deleted. That is Linux procfs behaviour rather than a fact about files, and it catches only unlink -- a binary swapped by rename onto the same path reads as untouched, which is exactly what a package manager does. Now the identity is captured at start and compared later: no procfs, and neither case missed. known-good is one bare line. The reader is a shell script on a machine where the host is failing to start, so it must not need a parser to be present and working. Written only after a clean apply, which is the whole claim -- not health, because a disconnected node is ordinary and a failing resource is the machine's problem rather than the binary's. packaging/ -- the unit, the rollback unit, and the rollback script. The script shares no code with the host and calls none of it: a binary that cannot start cannot be its own recovery. POSIX sh, nothing that has to be installed. The unit carries Restart=always with a comment saying why on-failure would break every upgrade. Both are tested and both sets of tests were confirmed to bite. Injecting five faults broke exactly the intended tests -- except one, and chasing why it did not found a placebo assertion I had written: `check "exits zero" ... "0" "0"` compares a literal to itself and can never fail. Replaced with the real exit code, after which the injection bites. Also caught: an injection that produced a build failure rather than a test failure, which my grep read as "no failure". Re-run so it compiled, and the test did bite. The script test runs in `make check`, so it is a gate rather than something that was run once. Verified against the real binary: known-good is written beside the store after a clean apply and is NOT written after a failed one.
This commit is contained in:
@@ -0,0 +1,21 @@
|
||||
[Unit]
|
||||
Description=Novox Mesh node host
|
||||
After=network-online.target
|
||||
Wants=network-online.target
|
||||
# When the supervisor gives up, recover rather than leaving the node quiet — a host that will
|
||||
# not start looks exactly like a machine somebody switched off (novox/hq ADR 0059).
|
||||
OnFailure=nox-mesh-host-rollback.service
|
||||
|
||||
[Service]
|
||||
Type=notify
|
||||
ExecStart=/usr/bin/nox-mesh-host run
|
||||
# always, NOT on-failure: the host restarts onto a new binary by exiting CLEANLY
|
||||
# (novox/hq ADR 0057), and on-failure would leave an upgraded node stopped.
|
||||
Restart=always
|
||||
RestartSec=5s
|
||||
StartLimitBurst=3
|
||||
StartLimitIntervalSec=120
|
||||
StateDirectory=mesh-host
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
Reference in New Issue
Block a user