Graduate research 003; file the rescue that does not exist
A sweep of the nine active research efforts. 003 was answered five days ago and nobody closed it -- the decision it asked for was taken without citing it, which is how an effort stays `active` after being resolved. Its recommendation is what the mesh adopted, and the match is exact rather than approximate. "Run the daemons as containers, making Docker the supervisor for everything" is ADR 0057. Its warning that a mesh-native supervisor inherits fate-sharing "unless it sits outside the mesh's own process tree" is where ADR 0061 put the launcher. And its insistence that it cannot be all-or-nothing is why the host itself is the one thing an init starts. Its incidental finding does not graduate with it, so it is now issue 008: the automatic node rescue the documentation describes does not exist. No unit declares OnFailure=, nothing calls the rescue script on a timer. That is worse than having no rescue. A rescue nobody wrote is a gap somebody can see; a documented one that is absent is a gap nobody looks for, and the documentation is read exactly when a node has failed and somebody is deciding whether to intervene. The issue names two honest resolutions -- implement it, or delete the documentation and say a failed node needs a person -- and says the choice is scheduling rather than technical, since the new host's recovery is built and tested. It also says what would make the finding certain: it came from reading the repository, and confirming it on a running node is the difference between "no unit declares this" and "no unit in the source declares this".
This commit is contained in:
@@ -1,8 +1,11 @@
|
|||||||
---
|
---
|
||||||
status: active
|
status: graduated
|
||||||
initiated: 2026-08-22
|
initiated: 2026-08-22
|
||||||
touches: [03-DESIGN/00-as-is/05-runtime-and-installation.md]
|
touches: [03-DESIGN/00-as-is/05-runtime-and-installation.md]
|
||||||
became: []
|
became:
|
||||||
|
- 02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md
|
||||||
|
- 02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md
|
||||||
|
- 03-DESIGN/01-to-be/05-the-node-host.md
|
||||||
---
|
---
|
||||||
|
|
||||||
# 003 — Who supervises a service
|
# 003 — Who supervises a service
|
||||||
@@ -47,7 +50,31 @@ This effort answers the cost half. It does not choose.
|
|||||||
|
|
||||||
Detail and costs in [`analysis.md`](analysis.md).
|
Detail and costs in [`analysis.md`](analysis.md).
|
||||||
|
|
||||||
## Decision needed
|
## What it became
|
||||||
|
|
||||||
Which supervision model the mesh adopts, recorded in a decision record before Phase 0 builds
|
*Closed 2026-08-28.* The decision this effort asked for was taken — and taken without citing it,
|
||||||
anything. The options and their costs are in `analysis.md` under "Options".
|
which is why the effort sat `active` for five days after being answered. Recorded here because
|
||||||
|
finding that is the point of a sweep.
|
||||||
|
|
||||||
|
**The third option is what the mesh adopted.** `Docker is the supervisor for everything` is
|
||||||
|
[ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md): the
|
||||||
|
host is a plain process on the machine and everything above tier 0 is a container. The substrate
|
||||||
|
bootstrap declares no service at all — it is package, container, action, container — so the
|
||||||
|
44-of-44 restart policies this effort counted are the supervision, exactly as it argued.
|
||||||
|
|
||||||
|
**Fate-sharing was the hard part, and it is solved the way this effort predicted.** It said any
|
||||||
|
mesh-native supervisor inherits the problem *unless it sits outside the mesh's own process
|
||||||
|
tree*. [ADR 0061](../../02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md) puts
|
||||||
|
the launcher there: it supervises the host as a child and shares no code with it, so a host that
|
||||||
|
cannot start is still recovered.
|
||||||
|
|
||||||
|
**It is not all-or-nothing, as this effort insisted.** A human agent acts through a shell and a
|
||||||
|
desktop, and those are on the host. So is the host itself — the one thing an init starts.
|
||||||
|
|
||||||
|
## What is not closed
|
||||||
|
|
||||||
|
**The automatic node rescue the documentation describes does not exist.** No unit declares
|
||||||
|
`OnFailure=`, and nothing calls the rescue script on a timer. That is a documented behaviour
|
||||||
|
which never happens, and it outlives this effort — filed as
|
||||||
|
[`04-ISSUES/008`](../../04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md)
|
||||||
|
rather than closed with it.
|
||||||
|
|||||||
@@ -0,0 +1,59 @@
|
|||||||
|
---
|
||||||
|
status: open
|
||||||
|
opened: 2026-08-28
|
||||||
|
located-in: [hal]
|
||||||
|
fixed-by:
|
||||||
|
amended-design:
|
||||||
|
---
|
||||||
|
|
||||||
|
# 008 — The documented automatic node rescue does not exist
|
||||||
|
|
||||||
|
## Symptom
|
||||||
|
|
||||||
|
The mesh's documentation describes an automatic node rescue: a node that fails is recovered
|
||||||
|
without anybody intervening. **Nothing implements it.**
|
||||||
|
|
||||||
|
Found incidentally while investigating supervision
|
||||||
|
([research 003](../../01-RESEARCH/003-service-supervision/00-overview.md)), which counted what
|
||||||
|
actually supervises what:
|
||||||
|
|
||||||
|
- **no unit declares `OnFailure=`**, so nothing runs when a unit gives up;
|
||||||
|
- **nothing calls the rescue script on a timer**, so it runs only when a person runs it.
|
||||||
|
|
||||||
|
The script exists. The thing that would invoke it does not.
|
||||||
|
|
||||||
|
## Why this is worse than having no rescue
|
||||||
|
|
||||||
|
A rescue nobody wrote is a gap somebody can see. A rescue that is *documented* and absent is a
|
||||||
|
gap nobody looks for, because the documentation says it is covered — and it is read exactly when
|
||||||
|
a node has failed and somebody is deciding whether to intervene.
|
||||||
|
|
||||||
|
This is `how-we-build` §5 in its most expensive form: *an unenforced rule is indistinguishable
|
||||||
|
from a wrong one, and costs more, because people believe it.* Here the belief is that a failed
|
||||||
|
node recovers itself.
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
**The as-is only.** The design being built has a different answer:
|
||||||
|
[ADR 0061](../../02-DECISIONS/0061-the-host-asks-an-init-for-start-and-restart.md) puts recovery
|
||||||
|
in a launcher that supervises the host, and that recovery is tested — 32 assertions, each
|
||||||
|
confirmed to fail when the behaviour is removed.
|
||||||
|
|
||||||
|
So this issue is about the mesh that runs **now**, and it has two possible resolutions rather
|
||||||
|
than one:
|
||||||
|
|
||||||
|
1. **Implement it** — an `OnFailure=` and a timer — if node rescue is wanted before the new host
|
||||||
|
reaches the fleet.
|
||||||
|
2. **Delete the documentation** — and say plainly that a failed node needs a person, which is
|
||||||
|
what is true today.
|
||||||
|
|
||||||
|
**Either is honest. Leaving it as it is, is not.** The choice turns on how far away the new host
|
||||||
|
is, which is a scheduling question rather than a technical one.
|
||||||
|
|
||||||
|
## What it would take to be sure
|
||||||
|
|
||||||
|
Read back rather than assumed
|
||||||
|
([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)): list every unit on a
|
||||||
|
node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above
|
||||||
|
came from reading the repository, and confirming it against a running node is the difference
|
||||||
|
between *no unit declares this* and *no unit in the source declares this*.
|
||||||
Reference in New Issue
Block a user