Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24
@@ -1,8 +1,11 @@
|
||||
---
|
||||
status: active
|
||||
status: graduated
|
||||
initiated: 2026-08-22
|
||||
touches: [03-DESIGN/00-as-is/05-runtime-and-installation.md]
|
||||
became: []
|
||||
became:
|
||||
- 02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md
|
||||
- 02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md
|
||||
- 03-DESIGN/01-to-be/05-the-node-host.md
|
||||
---
|
||||
|
||||
# 003 — Who supervises a service
|
||||
@@ -47,7 +50,31 @@ This effort answers the cost half. It does not choose.
|
||||
|
||||
Detail and costs in [`analysis.md`](analysis.md).
|
||||
|
||||
## Decision needed
|
||||
## What it became
|
||||
|
||||
Which supervision model the mesh adopts, recorded in a decision record before Phase 0 builds
|
||||
anything. The options and their costs are in `analysis.md` under "Options".
|
||||
*Closed 2026-08-28.* The decision this effort asked for was taken — and taken without citing it,
|
||||
which is why the effort sat `active` for five days after being answered. Recorded here because
|
||||
finding that is the point of a sweep.
|
||||
|
||||
**The third option is what the mesh adopted.** `Docker is the supervisor for everything` is
|
||||
[ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md): the
|
||||
host is a plain process on the machine and everything above tier 0 is a container. The substrate
|
||||
bootstrap declares no service at all — it is package, container, action, container — so the
|
||||
44-of-44 restart policies this effort counted are the supervision, exactly as it argued.
|
||||
|
||||
**Fate-sharing was the hard part, and it is solved the way this effort predicted.** It said any
|
||||
mesh-native supervisor inherits the problem *unless it sits outside the mesh's own process
|
||||
tree*. [ADR 0061](../../02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md) puts
|
||||
the launcher there: it supervises the host as a child and shares no code with it, so a host that
|
||||
cannot start is still recovered.
|
||||
|
||||
**It is not all-or-nothing, as this effort insisted.** A human agent acts through a shell and a
|
||||
desktop, and those are on the host. So is the host itself — the one thing an init starts.
|
||||
|
||||
## What is not closed
|
||||
|
||||
**The automatic node rescue the documentation describes does not exist.** No unit declares
|
||||
`OnFailure=`, and nothing calls the rescue script on a timer. That is a documented behaviour
|
||||
which never happens, and it outlives this effort — filed as
|
||||
[`04-ISSUES/008`](../../04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md)
|
||||
rather than closed with it.
|
||||
|
||||
@@ -0,0 +1,59 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-08-28
|
||||
located-in: [hal]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 008 — The documented automatic node rescue does not exist
|
||||
|
||||
## Symptom
|
||||
|
||||
The mesh's documentation describes an automatic node rescue: a node that fails is recovered
|
||||
without anybody intervening. **Nothing implements it.**
|
||||
|
||||
Found incidentally while investigating supervision
|
||||
([research 003](../../01-RESEARCH/003-service-supervision/00-overview.md)), which counted what
|
||||
actually supervises what:
|
||||
|
||||
- **no unit declares `OnFailure=`**, so nothing runs when a unit gives up;
|
||||
- **nothing calls the rescue script on a timer**, so it runs only when a person runs it.
|
||||
|
||||
The script exists. The thing that would invoke it does not.
|
||||
|
||||
## Why this is worse than having no rescue
|
||||
|
||||
A rescue nobody wrote is a gap somebody can see. A rescue that is *documented* and absent is a
|
||||
gap nobody looks for, because the documentation says it is covered — and it is read exactly when
|
||||
a node has failed and somebody is deciding whether to intervene.
|
||||
|
||||
This is `how-we-build` §5 in its most expensive form: *an unenforced rule is indistinguishable
|
||||
from a wrong one, and costs more, because people believe it.* Here the belief is that a failed
|
||||
node recovers itself.
|
||||
|
||||
## Scope
|
||||
|
||||
**The as-is only.** The design being built has a different answer:
|
||||
[ADR 0061](../../02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md) puts recovery
|
||||
in a launcher that supervises the host, and that recovery is tested — 32 assertions, each
|
||||
confirmed to fail when the behaviour is removed.
|
||||
|
||||
So this issue is about the mesh that runs **now**, and it has two possible resolutions rather
|
||||
than one:
|
||||
|
||||
1. **Implement it** — an `OnFailure=` and a timer — if node rescue is wanted before the new host
|
||||
reaches the fleet.
|
||||
2. **Delete the documentation** — and say plainly that a failed node needs a person, which is
|
||||
what is true today.
|
||||
|
||||
**Either is honest. Leaving it as it is, is not.** The choice turns on how far away the new host
|
||||
is, which is a scheduling question rather than a technical one.
|
||||
|
||||
## What it would take to be sure
|
||||
|
||||
Read back rather than assumed
|
||||
([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)): list every unit on a
|
||||
node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above
|
||||
came from reading the repository, and confirming it against a running node is the difference
|
||||
between *no unit declares this* and *no unit in the source declares this*.
|
||||
Reference in New Issue
Block a user