Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24

Merged
jschoubben merged 177 commits from reconcile-init-into-main into main 2026-09-05 10:27:11 +00:00
Showing only changes of commit 66df0eb53e - Show all commits
@@ -38,23 +38,41 @@ is a container the runtime restarts. Nothing in it declares a `service`.
## Decision
**An init is asked for two things: start this at boot, and start it again if it exits.** Both are
expressible in systemd, OpenRC, runit, s6 and an Android `init.rc`.
**An init is asked for one thing: start this at boot.**
An earlier version of this record asked for two — start, and restart on exit — and left the
restart in the unit file while moving the give-up logic out. That was half a change: it kept the
init deciding *when the host comes back*, which is the thing being removed. **The launcher does
not exec the host; it supervises it**, so restarting is ours as well.
The cost of not exec'ing is signals, and it is the reason people reach for a service manager in
the first place. A supervisor that exits while its child is still running leaves the host to be
*killed* rather than to *stop*, and an apply interrupted that way is the half-configured machine
this project is about. So the launcher traps the shutdown signal, passes it to the host, and
waits.
**Everything else moves into a launcher**, which is what the init actually starts:
```
init ──► nox-mesh-host-launch ──► nox-mesh-host
init ──► nox-mesh-host-launch ──► nox-mesh-host (a child, not an exec)
│
├─ halted? say so and stop. a person has to look
├─ count this start attempt
├─ too many, and not yet rolled back? roll back, then start
├─ too many, and already rolled back? halt — the machine is the problem
└─ otherwise start the host
└─ loop:
halted? say so and stop; a person has to look
too many failures? roll back once, then halt
start the host, and wait
exited 0 it upgraded itself — start the new binary,
and do NOT count it
crashed count it, back off, loop
shutting down pass the signal down, wait, exit
host, on a completed reconcile ──► clears the counter, records known-good
```
**The counter counts consecutive failures, not starts**, and the difference is not cosmetic.
Counting starts meant a host that upgraded itself three times rolled itself back, having worked
perfectly every time — the clean exit *is* the upgrade path
([ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)).
**This keeps everything ADR 0059 decided and changes only where it lives.** Two watchdogs still,
and neither substitutes for the other: the mesh stages a host rollout and stops when nodes go
quiet; the node recovers itself. Recovery is still local, because nothing dials a node and a host
@@ -85,10 +103,13 @@ host, because a binary that will not start cannot be its own recovery.
## Consequences
- **`Restart=always` stays load-bearing and stays subtle.** The host restarts onto a new binary
by exiting cleanly ([ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)), so
whatever supervises must restart on a zero exit. This caught out an earlier draft of 0059,
which specified `on-failure` and would have left every upgraded node stopped.
- **A clean exit is the upgrade path, and it is the easiest thing to get wrong.** Twice now:
ADR 0059 specified `on-failure`, which would have left every upgraded node stopped; and the
first supervising loop counted a clean exit as a failure, which would have rolled back a host
that upgraded itself three times. Anything touching restart has to ask what a zero exit means
here.
- **`Restart=` in the unit file becomes a backstop, not the mechanism.** It brings the launcher
back if the launcher itself is killed. It no longer decides anything about the host.
- **The unit file becomes trivial**, which is the point: start, restart, a state directory.
Nothing in it encodes policy, so porting it is transcription rather than design.
- **A halted node is silent**, unchanged from 0059 and still the last gap. What notices is the