0061: the launcher supervises; init is asked for one thing
The record said an init is asked for two things -- start at boot and restart on exit -- which was half a change. It moved the give-up logic out of unit files and left the restart in one, so init still decided when the host came back. The launcher no longer execs the host. It supervises it, so restarting is ours too, and init is asked only to run it at boot. There is an OpenRC script beside the systemd unit now. Records the cost honestly: not exec'ing means the launcher must trap the shutdown signal and pass it down, because a supervisor that exits while its child runs leaves the host to be killed rather than to stop. And records what the implementation found: the counter counts consecutive FAILURES, not starts. Counting starts meant a host that upgraded itself three times rolled itself back, having worked perfectly every time -- because a clean exit IS the upgrade path. That is now the second time a clean exit has been mishandled, so it is called out as the thing to check.
This commit is contained in:
@@ -38,23 +38,41 @@ is a container the runtime restarts. Nothing in it declares a `service`.
|
||||
|
||||
## Decision
|
||||
|
||||
**An init is asked for two things: start this at boot, and start it again if it exits.** Both are
|
||||
expressible in systemd, OpenRC, runit, s6 and an Android `init.rc`.
|
||||
**An init is asked for one thing: start this at boot.**
|
||||
|
||||
An earlier version of this record asked for two — start, and restart on exit — and left the
|
||||
restart in the unit file while moving the give-up logic out. That was half a change: it kept the
|
||||
init deciding *when the host comes back*, which is the thing being removed. **The launcher does
|
||||
not exec the host; it supervises it**, so restarting is ours as well.
|
||||
|
||||
The cost of not exec'ing is signals, and it is the reason people reach for a service manager in
|
||||
the first place. A supervisor that exits while its child is still running leaves the host to be
|
||||
*killed* rather than to *stop*, and an apply interrupted that way is the half-configured machine
|
||||
this project is about. So the launcher traps the shutdown signal, passes it to the host, and
|
||||
waits.
|
||||
|
||||
**Everything else moves into a launcher**, which is what the init actually starts:
|
||||
|
||||
```
|
||||
init ──► nox-mesh-host-launch ──► nox-mesh-host
|
||||
init ──► nox-mesh-host-launch ──► nox-mesh-host (a child, not an exec)
|
||||
│
|
||||
├─ halted? say so and stop. a person has to look
|
||||
├─ count this start attempt
|
||||
├─ too many, and not yet rolled back? roll back, then start
|
||||
├─ too many, and already rolled back? halt — the machine is the problem
|
||||
└─ otherwise start the host
|
||||
└─ loop:
|
||||
halted? say so and stop; a person has to look
|
||||
too many failures? roll back once, then halt
|
||||
start the host, and wait
|
||||
exited 0 it upgraded itself — start the new binary,
|
||||
and do NOT count it
|
||||
crashed count it, back off, loop
|
||||
shutting down pass the signal down, wait, exit
|
||||
|
||||
host, on a completed reconcile ──► clears the counter, records known-good
|
||||
```
|
||||
|
||||
**The counter counts consecutive failures, not starts**, and the difference is not cosmetic.
|
||||
Counting starts meant a host that upgraded itself three times rolled itself back, having worked
|
||||
perfectly every time — the clean exit *is* the upgrade path
|
||||
([ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)).
|
||||
|
||||
**This keeps everything ADR 0059 decided and changes only where it lives.** Two watchdogs still,
|
||||
and neither substitutes for the other: the mesh stages a host rollout and stops when nodes go
|
||||
quiet; the node recovers itself. Recovery is still local, because nothing dials a node and a host
|
||||
@@ -85,10 +103,13 @@ host, because a binary that will not start cannot be its own recovery.
|
||||
|
||||
## Consequences
|
||||
|
||||
- **`Restart=always` stays load-bearing and stays subtle.** The host restarts onto a new binary
|
||||
by exiting cleanly ([ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)), so
|
||||
whatever supervises must restart on a zero exit. This caught out an earlier draft of 0059,
|
||||
which specified `on-failure` and would have left every upgraded node stopped.
|
||||
- **A clean exit is the upgrade path, and it is the easiest thing to get wrong.** Twice now:
|
||||
ADR 0059 specified `on-failure`, which would have left every upgraded node stopped; and the
|
||||
first supervising loop counted a clean exit as a failure, which would have rolled back a host
|
||||
that upgraded itself three times. Anything touching restart has to ask what a zero exit means
|
||||
here.
|
||||
- **`Restart=` in the unit file becomes a backstop, not the mechanism.** It brings the launcher
|
||||
back if the launcher itself is killed. It no longer decides anything about the host.
|
||||
- **The unit file becomes trivial**, which is the point: start, restart, a state directory.
|
||||
Nothing in it encodes policy, so porting it is transcription rather than design.
|
||||
- **A halted node is silent**, unchanged from 0059 and still the last gap. What notices is the
|
||||
|
||||
Reference in New Issue
Block a user