Commit Graph
4 Commits
Author SHA1 Message Date
jschoubben 057f34f924 The init is asked for start and restart; a launcher does the rest
ADR 0061. Recovery was the most systemd-specific part of the host, and it is
the part that must work on a machine where nothing else does -- which made
unit-file syntax a poor place for it, because syntax cannot be tested and the
one time it runs is the one time nobody can afford it wrong.

So StartLimitBurst and OnFailure move into a launcher script that init starts
instead of the host. The unit drops to start-at-boot and restart-on-exit, which
OpenRC, runit, s6 and an Android init.rc can all express. Everything 0059
decided is kept: two watchdogs, roll back once, recovery is local, the rollback
shares no code with the host.

The counter is the whole mechanism, so it is what the tests are mostly about.
Three real problems came out of writing them:

A counter file holding "1 2" became "12" -- `tr -d [:space:]` concatenates
rather than rejecting -- which is past the limit, so a HEALTHY node rolled
itself back. Now it reads the first field and insists on a plain integer.

The corrupt-counter test used "not-a-number", which shell arithmetic happens to
evaluate to 0, so it passed with the guard removed and proved nothing. Replaced
with values that discriminate: "5x" errors under set -e and kills the launcher,
and "0x10" is read as HEX 16 -- past the limit, so again a healthy node rolls
back.

And the test harness itself was wrong. With `set -e` and a bare launcher call,
removing a guard killed the script at the first corrupt case and silently
skipped everything after -- reporting a full pass over tests that never ran.
Every launcher call now records its failure instead of aborting. Same class as
the placebo assertion found last time, and the reason to keep injecting faults
rather than trusting green.

Both scripts run in `make check`. 27 launcher tests, 9 rollback tests, all
confirmed to bite.
2026-08-27 23:45:57 +02:00
jschoubben f4143806c2 Build the rollback mechanism, and test it
ADR 0059's recovery path: the pieces that run when the host will not start.

internal/upgrade -- two facts, neither of them the host judging its health.
Whether the executable this process started from has been replaced on disk, and
which version last completed a reconcile.

The first design was wrong and the tests caught it, not review. It asked
/proc/self/exe whether it was marked deleted. That is Linux procfs behaviour
rather than a fact about files, and it catches only unlink -- a binary swapped
by rename onto the same path reads as untouched, which is exactly what a
package manager does. Now the identity is captured at start and compared later:
no procfs, and neither case missed.

known-good is one bare line. The reader is a shell script on a machine where
the host is failing to start, so it must not need a parser to be present and
working. Written only after a clean apply, which is the whole claim -- not
health, because a disconnected node is ordinary and a failing resource is the
machine's problem rather than the binary's.

packaging/ -- the unit, the rollback unit, and the rollback script. The script
shares no code with the host and calls none of it: a binary that cannot start
cannot be its own recovery. POSIX sh, nothing that has to be installed. The
unit carries Restart=always with a comment saying why on-failure would break
every upgrade.

Both are tested and both sets of tests were confirmed to bite. Injecting five
faults broke exactly the intended tests -- except one, and chasing why it did
not found a placebo assertion I had written: `check "exits zero" ... "0" "0"`
compares a literal to itself and can never fail. Replaced with the real exit
code, after which the injection bites.

Also caught: an injection that produced a build failure rather than a test
failure, which my grep read as "no failure". Re-run so it compiled, and the
test did bite.

The script test runs in `make check`, so it is a gate rather than something
that was run once.

Verified against the real binary: known-good is written beside the store after
a clean apply and is NOT written after a failed one.
2026-08-27 22:24:31 +02:00
jschoubben 08a1263a81 Stage 2 — the bundle a host carries
novox/hq ADR 0038: one behaviour, two sources of declaration. This is the source
that does not need a mesh — the first node's path.

The bundle is embedded in the binary rather than shipped beside it, because
"copy it onto a machine and run it is the whole installation" stops being true
the moment a second file has to arrive with it. `make host BUNDLE=...` builds a
host carrying one; `mesh-host reconcile` applies it; `mesh-host bundle` shows it.

A default build carries nothing and REFUSES to reconcile, saying why. A host
that applied nothing and reported success would look exactly like one that
raised a first node, and the difference would surface later as a mesh that never
came up with nothing to point at.

Proved on a sealed machine: no route out, no name resolution, one binary copied
on, and it configured itself from what it carried. Idempotent on the second run.

One bug found by running rather than reasoning, and it is a shape worth naming:
`mesh-host bundle` validated the carried bundle through a path that strips
comments, while `reconcile` handed the raw bytes to the parser. So the command
whose whole job is to check the bundle said yes, and the command that uses it
said no. Two paths to one artefact, disagreeing. There is one path now, and a
test asserts that what validates is what is applied.

What this does NOT prove is stated in the README rather than left implied: the
claim under stage 2 is that one host can raise the substrate alone, and the
substrate is four container services. There is no container type, because a
container needs an image and where images come from is open; what belongs in a
substrate is not known, because the closure for a one-node mesh is what research
011 and 012 exist to answer; and the machine used to test this cannot install a
container runtime through a sealed network.

The mechanism is finished. The claim is not, and shipping a host that claimed a
substrate it has never raised would be the fault this whole project is about.

65 tests.
2026-08-26 22:06:54 +02:00
jschoubben 73c010e7ef Stage 1 — the host reports what a machine is and can do
Tier 0's first slice, per novox/hq 03-DESIGN/01-to-be/05-the-node-host.md. It
applies nothing, connects to nothing, listens on nothing. 2.9 MB, static, no
dynamic dependencies: copy it onto a machine and run it is the whole install,
which is the property ADR 0041 rests on.

A capability is detected, never assumed. Every detector runs something that only
succeeds if the thing FUNCTIONS — the daemon is asked for its version, the
package database is queried, the firewall is asked to list a ruleset, which
needs the privilege as well as the tool. 04-ISSUES/007 is the fault this
prevents: a client on disk with its daemon down looks exactly like a working
runtime, and a node assigned work on that basis fails when the work arrives.

Every verdict carries the reason and the method. A capability reported absent
with no reason is the same fault in a new place: something nobody can act on.

Two bugs found by running rather than reasoning, both silent:

systemctl is-system-running exits non-zero for every state except `running` —
including `degraded`, which means units failed and the init is emphatically
there. Reading the exit code reported NO service manager on a machine whose init
it was. That is 007 in the mirror, and both directions place work wrongly. A
verdict now reads what a tool says about itself, not only how it exited.

And `mesh-host inventory --json` printed text: the standard library stops
parsing at the first non-flag argument, so the flag sat unread and the command
exited 0 having ignored what was asked. The parser now takes the subcommand off
the front, and a stray or mistyped argument is refused rather than dropped.

Detection deliberately does NOT follow ADR 0008. That rule governs applying
state, where a failed step means the machine is not what was asked for. A failed
probe is a finding — "absent, because the probe failed" — and aborting would
replace one legible absence with total ignorance of the rest.

25 tests: structure and logic with a fake runner, and the same detectors against
this machine, because a test that fakes the system under detection asserts only
that the fake behaves as expected.
2026-08-26 00:25:08 +02:00