Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24
@@ -38,23 +38,41 @@ is a container the runtime restarts. Nothing in it declares a `service`.
|
|||||||
|
|
||||||
## Decision
|
## Decision
|
||||||
|
|
||||||
**An init is asked for two things: start this at boot, and start it again if it exits.** Both are
|
**An init is asked for one thing: start this at boot.**
|
||||||
expressible in systemd, OpenRC, runit, s6 and an Android `init.rc`.
|
|
||||||
|
An earlier version of this record asked for two — start, and restart on exit — and left the
|
||||||
|
restart in the unit file while moving the give-up logic out. That was half a change: it kept the
|
||||||
|
init deciding *when the host comes back*, which is the thing being removed. **The launcher does
|
||||||
|
not exec the host; it supervises it**, so restarting is ours as well.
|
||||||
|
|
||||||
|
The cost of not exec'ing is signals, and it is the reason people reach for a service manager in
|
||||||
|
the first place. A supervisor that exits while its child is still running leaves the host to be
|
||||||
|
*killed* rather than to *stop*, and an apply interrupted that way is the half-configured machine
|
||||||
|
this project is about. So the launcher traps the shutdown signal, passes it to the host, and
|
||||||
|
waits.
|
||||||
|
|
||||||
**Everything else moves into a launcher**, which is what the init actually starts:
|
**Everything else moves into a launcher**, which is what the init actually starts:
|
||||||
|
|
||||||
```
|
```
|
||||||
init ──► nox-mesh-host-launch ──► nox-mesh-host
|
init ──► nox-mesh-host-launch ──► nox-mesh-host (a child, not an exec)
|
||||||
│
|
│
|
||||||
├─ halted? say so and stop. a person has to look
|
└─ loop:
|
||||||
├─ count this start attempt
|
halted? say so and stop; a person has to look
|
||||||
├─ too many, and not yet rolled back? roll back, then start
|
too many failures? roll back once, then halt
|
||||||
├─ too many, and already rolled back? halt — the machine is the problem
|
start the host, and wait
|
||||||
└─ otherwise start the host
|
exited 0 it upgraded itself — start the new binary,
|
||||||
|
and do NOT count it
|
||||||
|
crashed count it, back off, loop
|
||||||
|
shutting down pass the signal down, wait, exit
|
||||||
|
|
||||||
host, on a completed reconcile ──► clears the counter, records known-good
|
host, on a completed reconcile ──► clears the counter, records known-good
|
||||||
```
|
```
|
||||||
|
|
||||||
|
**The counter counts consecutive failures, not starts**, and the difference is not cosmetic.
|
||||||
|
Counting starts meant a host that upgraded itself three times rolled itself back, having worked
|
||||||
|
perfectly every time — the clean exit *is* the upgrade path
|
||||||
|
([ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)).
|
||||||
|
|
||||||
**This keeps everything ADR 0059 decided and changes only where it lives.** Two watchdogs still,
|
**This keeps everything ADR 0059 decided and changes only where it lives.** Two watchdogs still,
|
||||||
and neither substitutes for the other: the mesh stages a host rollout and stops when nodes go
|
and neither substitutes for the other: the mesh stages a host rollout and stops when nodes go
|
||||||
quiet; the node recovers itself. Recovery is still local, because nothing dials a node and a host
|
quiet; the node recovers itself. Recovery is still local, because nothing dials a node and a host
|
||||||
@@ -85,10 +103,13 @@ host, because a binary that will not start cannot be its own recovery.
|
|||||||
|
|
||||||
## Consequences
|
## Consequences
|
||||||
|
|
||||||
- **`Restart=always` stays load-bearing and stays subtle.** The host restarts onto a new binary
|
- **A clean exit is the upgrade path, and it is the easiest thing to get wrong.** Twice now:
|
||||||
by exiting cleanly ([ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)), so
|
ADR 0059 specified `on-failure`, which would have left every upgraded node stopped; and the
|
||||||
whatever supervises must restart on a zero exit. This caught out an earlier draft of 0059,
|
first supervising loop counted a clean exit as a failure, which would have rolled back a host
|
||||||
which specified `on-failure` and would have left every upgraded node stopped.
|
that upgraded itself three times. Anything touching restart has to ask what a zero exit means
|
||||||
|
here.
|
||||||
|
- **`Restart=` in the unit file becomes a backstop, not the mechanism.** It brings the launcher
|
||||||
|
back if the launcher itself is killed. It no longer decides anything about the host.
|
||||||
- **The unit file becomes trivial**, which is the point: start, restart, a state directory.
|
- **The unit file becomes trivial**, which is the point: start, restart, a state directory.
|
||||||
Nothing in it encodes policy, so porting it is transcription rather than design.
|
Nothing in it encodes policy, so porting it is transcription rather than design.
|
||||||
- **A halted node is silent**, unchanged from 0059 and still the last gap. What notices is the
|
- **A halted node is silent**, unchanged from 0059 and still the last gap. What notices is the
|
||||||
|
|||||||
Reference in New Issue
Block a user