Files
hq/02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md
T
jschoubben 66df0eb53e 0061: the launcher supervises; init is asked for one thing
The record said an init is asked for two things -- start at boot and restart on
exit -- which was half a change. It moved the give-up logic out of unit files
and left the restart in one, so init still decided when the host came back.

The launcher no longer execs the host. It supervises it, so restarting is ours
too, and init is asked only to run it at boot. There is an OpenRC script beside
the systemd unit now.

Records the cost honestly: not exec'ing means the launcher must trap the
shutdown signal and pass it down, because a supervisor that exits while its
child runs leaves the host to be killed rather than to stop.

And records what the implementation found: the counter counts consecutive
FAILURES, not starts. Counting starts meant a host that upgraded itself three
times rolled itself back, having worked perfectly every time -- because a clean
exit IS the upgrade path. That is now the second time a clean exit has been
mishandled, so it is called out as the thing to check.
2026-08-28 00:38:18 +02:00

131 lines
7.1 KiB
Markdown

---
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
supersedes: 0059-a-host-that-cannot-start-rolls-itself-back.md
extends: 0060-the-host-is-built-per-operating-system.md
---
# 61. The host asks an init for start and restart, and nothing else
## Context
[ADR 0059](0059-a-host-that-cannot-start-rolls-itself-back.md) built the node's recovery out of
systemd's own features: `StartLimitBurst` to decide a binary is broken, `OnFailure` to run a
rollback unit. It works, and it makes recovery — the thing that matters most when a node is
stuck — the most systemd-specific part of the whole host.
[ADR 0060](0060-the-host-is-built-per-operating-system.md) makes the host per operating system,
which raises the obvious question: how much of an init does the host actually need?
Counted honestly, three things, and only one of them is special:
| | any init? |
|---|---|
| start at boot | **yes** |
| restart it when it exits | **yes** |
| give up after N failures and run something else | **no** — that is systemd's `StartLimitBurst` and `OnFailure` |
So the recovery mechanism is the only reason the host needs *this* init rather than *an* init.
And it is the part that must work on a machine where the host does not, which makes "it is
expressed in unit-file syntax" a poor place for it: unit syntax is not something we can test, and
the one time it runs is the one time nobody can afford it to be wrong.
**The substrate does not need systemd either**, which is what makes this worth doing rather than
merely tidy. The bootstrap is package → container → action → container, and every mesh workload
is a container the runtime restarts. Nothing in it declares a `service`.
## Decision
**An init is asked for one thing: start this at boot.**
An earlier version of this record asked for two — start, and restart on exit — and left the
restart in the unit file while moving the give-up logic out. That was half a change: it kept the
init deciding *when the host comes back*, which is the thing being removed. **The launcher does
not exec the host; it supervises it**, so restarting is ours as well.
The cost of not exec'ing is signals, and it is the reason people reach for a service manager in
the first place. A supervisor that exits while its child is still running leaves the host to be
*killed* rather than to *stop*, and an apply interrupted that way is the half-configured machine
this project is about. So the launcher traps the shutdown signal, passes it to the host, and
waits.
**Everything else moves into a launcher**, which is what the init actually starts:
```
init ──► nox-mesh-host-launch ──► nox-mesh-host (a child, not an exec)
│
└─ loop:
halted? say so and stop; a person has to look
too many failures? roll back once, then halt
start the host, and wait
exited 0 it upgraded itself — start the new binary,
and do NOT count it
crashed count it, back off, loop
shutting down pass the signal down, wait, exit
host, on a completed reconcile ──► clears the counter, records known-good
```
**The counter counts consecutive failures, not starts**, and the difference is not cosmetic.
Counting starts meant a host that upgraded itself three times rolled itself back, having worked
perfectly every time — the clean exit *is* the upgrade path
([ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)).
**This keeps everything ADR 0059 decided and changes only where it lives.** Two watchdogs still,
and neither substitutes for the other: the mesh stages a host rollout and stops when nodes go
quiet; the node recovers itself. Recovery is still local, because nothing dials a node and a host
that cannot start cannot report. It rolls back once, because a second failure of a
previously-working binary is a different diagnosis. The rollback still shares no code with the
host, because a binary that will not start cannot be its own recovery.
### What is gained by moving it
- **It becomes testable.** A shell script with a counter can be run against a stub package
manager and asserted, which is how the rollback script is already tested. `OnFailure=` can be
read and hoped for.
- **The host becomes runnable under any init**, which is what makes an Alpine or Android host
possible later rather than blocked on porting the recovery.
- **The give-up policy stops being configuration and becomes code we own.** Three attempts is a
decision, and it should live where decisions are read and tested.
### What it costs
- **One more process in the chain**, and it runs before the host on every start.
- **The counter is the whole mechanism, and it is the part to get right.** Never cleared, and
the node rolls back on a healthy boot; cleared too eagerly, and it never rolls back at all.
It is cleared by the host on a **completed reconcile** — the same event that records
known-good, for the same reason.
- **The launcher is a thing that can itself be broken**, and nothing recovers it. That is one
turtle down from where we were, not zero: the alternative was unit syntax, which also cannot
recover itself and additionally cannot be tested.
## Consequences
- **A clean exit is the upgrade path, and it is the easiest thing to get wrong.** Twice now:
ADR 0059 specified `on-failure`, which would have left every upgraded node stopped; and the
first supervising loop counted a clean exit as a failure, which would have rolled back a host
that upgraded itself three times. Anything touching restart has to ask what a zero exit means
here.
- **`Restart=` in the unit file becomes a backstop, not the mechanism.** It brings the launcher
back if the launcher itself is killed. It no longer decides anything about the host.
- **The unit file becomes trivial**, which is the point: start, restart, a state directory.
Nothing in it encodes policy, so porting it is transcription rather than design.
- **A halted node is silent**, unchanged from 0059 and still the last gap. What notices is the
mesh seeing a node it has not heard from.
- **The rollback script grows into a launcher** rather than being replaced. Its tested behaviour —
roll back once, refuse to guess with no known-good, fail loudly when the package is not
cached — carries over and is where the new counter logic joins it.
- **`service` survives as a shape**, and this record does not remove it. Almost nothing in the
design declares one, but adoption takes over machines already in use whose units somebody
chose, and saying the mesh may never manage those is a larger decision than this one.
## References
- [ADR 0059](0059-a-host-that-cannot-start-rolls-itself-back.md) — superseded; its reasoning
about two watchdogs, rolling back once, and recovery being local is kept in full.
- [ADR 0060](0060-the-host-is-built-per-operating-system.md) — why this question was asked.
- [ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md) — restart-by-exiting,
which constrains what the init must do.