Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24

Merged
jschoubben merged 177 commits from reconcile-init-into-main into main 2026-09-05 10:27:11 +00:00
2 changed files with 91 additions and 5 deletions
Showing only changes of commit 0a37d751e2 - Show all commits
@@ -1,8 +1,11 @@
---
status: active
status: graduated
initiated: 2026-08-22
touches: [03-DESIGN/00-as-is/05-runtime-and-installation.md]
became: []
became:
- 02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md
- 02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md
- 03-DESIGN/01-to-be/05-the-node-host.md
---
# 003 — Who supervises a service
@@ -47,7 +50,31 @@ This effort answers the cost half. It does not choose.
Detail and costs in [`analysis.md`](analysis.md).
## Decision needed
## What it became
Which supervision model the mesh adopts, recorded in a decision record before Phase 0 builds
anything. The options and their costs are in `analysis.md` under "Options".
*Closed 2026-08-28.* The decision this effort asked for was taken — and taken without citing it,
which is why the effort sat `active` for five days after being answered. Recorded here because
finding that is the point of a sweep.
**The third option is what the mesh adopted.** `Docker is the supervisor for everything` is
[ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md): the
host is a plain process on the machine and everything above tier 0 is a container. The substrate
bootstrap declares no service at all — it is package, container, action, container — so the
44-of-44 restart policies this effort counted are the supervision, exactly as it argued.
**Fate-sharing was the hard part, and it is solved the way this effort predicted.** It said any
mesh-native supervisor inherits the problem *unless it sits outside the mesh's own process
tree*. [ADR 0061](../../02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md) puts
the launcher there: it supervises the host as a child and shares no code with it, so a host that
cannot start is still recovered.
**It is not all-or-nothing, as this effort insisted.** A human agent acts through a shell and a
desktop, and those are on the host. So is the host itself — the one thing an init starts.
## What is not closed
**The automatic node rescue the documentation describes does not exist.** No unit declares
`OnFailure=`, and nothing calls the rescue script on a timer. That is a documented behaviour
which never happens, and it outlives this effort — filed as
[`04-ISSUES/008`](../../04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md)
rather than closed with it.
@@ -0,0 +1,59 @@
---
status: open
opened: 2026-08-28
located-in: [hal]
fixed-by:
amended-design:
---
# 008 — The documented automatic node rescue does not exist
## Symptom
The mesh's documentation describes an automatic node rescue: a node that fails is recovered
without anybody intervening. **Nothing implements it.**
Found incidentally while investigating supervision
([research 003](../../01-RESEARCH/003-service-supervision/00-overview.md)), which counted what
actually supervises what:
- **no unit declares `OnFailure=`**, so nothing runs when a unit gives up;
- **nothing calls the rescue script on a timer**, so it runs only when a person runs it.
The script exists. The thing that would invoke it does not.
## Why this is worse than having no rescue
A rescue nobody wrote is a gap somebody can see. A rescue that is *documented* and absent is a
gap nobody looks for, because the documentation says it is covered — and it is read exactly when
a node has failed and somebody is deciding whether to intervene.
This is `how-we-build` §5 in its most expensive form: *an unenforced rule is indistinguishable
from a wrong one, and costs more, because people believe it.* Here the belief is that a failed
node recovers itself.
## Scope
**The as-is only.** The design being built has a different answer:
[ADR 0061](../../02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md) puts recovery
in a launcher that supervises the host, and that recovery is tested — 32 assertions, each
confirmed to fail when the behaviour is removed.
So this issue is about the mesh that runs **now**, and it has two possible resolutions rather
than one:
1. **Implement it** — an `OnFailure=` and a timer — if node rescue is wanted before the new host
reaches the fleet.
2. **Delete the documentation** — and say plainly that a failed node needs a person, which is
what is true today.
**Either is honest. Leaving it as it is, is not.** The choice turns on how far away the new host
is, which is a scheduling question rather than a technical one.
## What it would take to be sure
Read back rather than assumed
([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)): list every unit on a
node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above
came from reading the repository, and confirming it against a running node is the difference
between *no unit declares this* and *no unit in the source declares this*.