diff --git a/01-RESEARCH/003-service-supervision/00-overview.md b/01-RESEARCH/003-service-supervision/00-overview.md index 4736fc6..2d8beac 100644 --- a/01-RESEARCH/003-service-supervision/00-overview.md +++ b/01-RESEARCH/003-service-supervision/00-overview.md @@ -1,8 +1,11 @@ --- -status: active +status: graduated initiated: 2026-08-22 touches: [03-DESIGN/00-as-is/05-runtime-and-installation.md] -became: [] +became: + - 02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md + - 02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md + - 03-DESIGN/01-to-be/05-the-node-host.md --- # 003 — Who supervises a service @@ -47,7 +50,31 @@ This effort answers the cost half. It does not choose. Detail and costs in [`analysis.md`](analysis.md). -## Decision needed +## What it became -Which supervision model the mesh adopts, recorded in a decision record before Phase 0 builds -anything. The options and their costs are in `analysis.md` under "Options". +*Closed 2026-08-28.* The decision this effort asked for was taken — and taken without citing it, +which is why the effort sat `active` for five days after being answered. Recorded here because +finding that is the point of a sweep. + +**The third option is what the mesh adopted.** `Docker is the supervisor for everything` is +[ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md): the +host is a plain process on the machine and everything above tier 0 is a container. The substrate +bootstrap declares no service at all — it is package, container, action, container — so the +44-of-44 restart policies this effort counted are the supervision, exactly as it argued. + +**Fate-sharing was the hard part, and it is solved the way this effort predicted.** It said any +mesh-native supervisor inherits the problem *unless it sits outside the mesh's own process +tree*. [ADR 0061](../../02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md) puts +the launcher there: it supervises the host as a child and shares no code with it, so a host that +cannot start is still recovered. + +**It is not all-or-nothing, as this effort insisted.** A human agent acts through a shell and a +desktop, and those are on the host. So is the host itself — the one thing an init starts. + +## What is not closed + +**The automatic node rescue the documentation describes does not exist.** No unit declares +`OnFailure=`, and nothing calls the rescue script on a timer. That is a documented behaviour +which never happens, and it outlives this effort — filed as +[`04-ISSUES/008`](../../04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md) +rather than closed with it. diff --git a/04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md b/04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md new file mode 100644 index 0000000..ad5af50 --- /dev/null +++ b/04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md @@ -0,0 +1,59 @@ +--- +status: open +opened: 2026-08-28 +located-in: [hal] +fixed-by: +amended-design: +--- + +# 008 — The documented automatic node rescue does not exist + +## Symptom + +The mesh's documentation describes an automatic node rescue: a node that fails is recovered +without anybody intervening. **Nothing implements it.** + +Found incidentally while investigating supervision +([research 003](../../01-RESEARCH/003-service-supervision/00-overview.md)), which counted what +actually supervises what: + +- **no unit declares `OnFailure=`**, so nothing runs when a unit gives up; +- **nothing calls the rescue script on a timer**, so it runs only when a person runs it. + +The script exists. The thing that would invoke it does not. + +## Why this is worse than having no rescue + +A rescue nobody wrote is a gap somebody can see. A rescue that is *documented* and absent is a +gap nobody looks for, because the documentation says it is covered — and it is read exactly when +a node has failed and somebody is deciding whether to intervene. + +This is `how-we-build` §5 in its most expensive form: *an unenforced rule is indistinguishable +from a wrong one, and costs more, because people believe it.* Here the belief is that a failed +node recovers itself. + +## Scope + +**The as-is only.** The design being built has a different answer: +[ADR 0061](../../02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md) puts recovery +in a launcher that supervises the host, and that recovery is tested — 32 assertions, each +confirmed to fail when the behaviour is removed. + +So this issue is about the mesh that runs **now**, and it has two possible resolutions rather +than one: + +1. **Implement it** — an `OnFailure=` and a timer — if node rescue is wanted before the new host + reaches the fleet. +2. **Delete the documentation** — and say plainly that a failed node needs a person, which is + what is true today. + +**Either is honest. Leaving it as it is, is not.** The choice turns on how far away the new host +is, which is a scheduling question rather than a technical one. + +## What it would take to be sure + +Read back rather than assumed +([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)): list every unit on a +node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above +came from reading the repository, and confirming it against a running node is the difference +between *no unit declares this* and *no unit in the source declares this*.