02f1fcc865e27738ca7b8584975bd1002404ff63
3
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5b7b280e3a |
The launcher supervises the host instead of exec'ing it
Jochen: "I thought we did not want to run the host under a systemd/openrc/init loop, but instead had our own host-init program?" -- and that was right. I had moved the give-up logic out of unit files and left RESTART in them, with the launcher exec'ing the host and disappearing. So init still decided when the host came back, which is the arrangement 0061 exists to remove. The launcher now stays and supervises: starts the host as a child, waits, decides. Init is asked for one thing, run this at boot. There is an OpenRC script beside the systemd unit now, four lines each, which is the point -- a second init is transcription rather than a port. The cost of not exec'ing is signals. A supervisor that exits while its child runs leaves the host to be killed rather than to stop, and an apply interrupted that way is the half-configured machine this project is about. So SIGTERM is trapped, passed down, and waited on. Two bugs, both found by the tests rather than by review: A clean exit was counted as a failure. The host exits cleanly to stand aside for a new binary after an upgrade (0057), so a host that upgraded itself three times rolled itself back having worked perfectly every time. The counter now counts CONSECUTIVE FAILURES, incremented after the wait rather than before the start. And when rolling back I reset the counter file but not the variable, so the next failure counted from the old value -- the rolled-back version got one attempt instead of three. Also: the host now clears the counter when it completes a reconcile, at the same moment it records known-good and for the same reason. Without it the count only climbs, and a node up for months rolls itself back on its third ordinary restart -- a healthy machine undone by its own recovery. One test expectation was tightened rather than fixed: "resets the counter after rolling back" asserted exactly 0, which was true only under the old count-before-start semantics. It now asserts the property -- below the limit -- since 1 is correct after a rollback plus one failure. 32 launcher tests, all confirmed to bite. |
||
|
|
057f34f924 |
The init is asked for start and restart; a launcher does the rest
ADR 0061. Recovery was the most systemd-specific part of the host, and it is the part that must work on a machine where nothing else does -- which made unit-file syntax a poor place for it, because syntax cannot be tested and the one time it runs is the one time nobody can afford it wrong. So StartLimitBurst and OnFailure move into a launcher script that init starts instead of the host. The unit drops to start-at-boot and restart-on-exit, which OpenRC, runit, s6 and an Android init.rc can all express. Everything 0059 decided is kept: two watchdogs, roll back once, recovery is local, the rollback shares no code with the host. The counter is the whole mechanism, so it is what the tests are mostly about. Three real problems came out of writing them: A counter file holding "1 2" became "12" -- `tr -d [:space:]` concatenates rather than rejecting -- which is past the limit, so a HEALTHY node rolled itself back. Now it reads the first field and insists on a plain integer. The corrupt-counter test used "not-a-number", which shell arithmetic happens to evaluate to 0, so it passed with the guard removed and proved nothing. Replaced with values that discriminate: "5x" errors under set -e and kills the launcher, and "0x10" is read as HEX 16 -- past the limit, so again a healthy node rolls back. And the test harness itself was wrong. With `set -e` and a bare launcher call, removing a guard killed the script at the first corrupt case and silently skipped everything after -- reporting a full pass over tests that never ran. Every launcher call now records its failure instead of aborting. Same class as the placebo assertion found last time, and the reason to keep injecting faults rather than trusting green. Both scripts run in `make check`. 27 launcher tests, 9 rollback tests, all confirmed to bite. |
||
|
|
f4143806c2 |
Build the rollback mechanism, and test it
ADR 0059's recovery path: the pieces that run when the host will not start. internal/upgrade -- two facts, neither of them the host judging its health. Whether the executable this process started from has been replaced on disk, and which version last completed a reconcile. The first design was wrong and the tests caught it, not review. It asked /proc/self/exe whether it was marked deleted. That is Linux procfs behaviour rather than a fact about files, and it catches only unlink -- a binary swapped by rename onto the same path reads as untouched, which is exactly what a package manager does. Now the identity is captured at start and compared later: no procfs, and neither case missed. known-good is one bare line. The reader is a shell script on a machine where the host is failing to start, so it must not need a parser to be present and working. Written only after a clean apply, which is the whole claim -- not health, because a disconnected node is ordinary and a failing resource is the machine's problem rather than the binary's. packaging/ -- the unit, the rollback unit, and the rollback script. The script shares no code with the host and calls none of it: a binary that cannot start cannot be its own recovery. POSIX sh, nothing that has to be installed. The unit carries Restart=always with a comment saying why on-failure would break every upgrade. Both are tested and both sets of tests were confirmed to bite. Injecting five faults broke exactly the intended tests -- except one, and chasing why it did not found a placebo assertion I had written: `check "exits zero" ... "0" "0"` compares a literal to itself and can never fail. Replaced with the real exit code, after which the injection bites. Also caught: an injection that produced a build failure rather than a test failure, which my grep read as "no failure". Re-run so it compiled, and the test did bite. The script test runs in `make check`, so it is a gate rather than something that was run once. Verified against the real binary: known-good is written beside the store after a clean apply and is NOT written after a failed one. |