Files
hq/01-RESEARCH/003-service-supervision/00-overview.md
T
jschoubben becae7ba51 ADR 0080: the development cycle is checked, not trusted
The flow the process overview draws — idea/symptom -> decision -> to-be design -> code ->
as-is — was enforced by nothing. cycle.py now refuses a to-be design naming no decision, an
in-progress/implemented design naming no owning code, a located/fixed issue with no owner,
a fixed/resolved issue with no fix, and a graduated research overview that does not say what
it became. AGENTS.md carries the cycle and a where-to-look table so a fresh session (or a
cleared context) finds the chain in frontmatter instead of assuming it. Grounding the check
surfaced two real gaps, fixed here: the work-ahead design named no owning code, and research
003 listed one became target twice.

https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
2026-09-17 22:28:03 +02:00

80 lines
4.1 KiB
Markdown

---
status: graduated
initiated: 2026-08-22
touches: [03-DESIGN/00-as-is/05-runtime-and-installation.md]
became:
- 02-DECISIONS/0005-the-node-host.md
- 03-DESIGN/01-to-be/05-the-node-host.md
---
# 003 — Who supervises a service
**No longer blocks Phase 0** (see below).
- **Initiated by:** jochen, 2026-08-22, in response to
[`002-local-mesh`](../002-local-mesh/analysis.md) open question 1
- **Areas touched:** every module shipping a `systemd/` directory (14), the
`hal-module@` template, `hal/sdk` feature handlers, `dev_up`, `log_tail` /
`systemd_journal`, the bootstrap scripts.
## The question
`002` asked how a containerised node runs a module service, given that a module service is
defined today as a systemd unit running `docker compose` against `/services/`. The response
was the better question:
> If it's possible to run systemd inside a container, that's the way to go I think. However,
> what would the cost be to step away from systemd to run our services and set it up in a
> different way? More hal mesh approach.
This effort answers the cost half. It does not choose.
## Summary of findings
- **The question splits in two**, and the halves have opposite answers. Supervising
`docker compose` stacks through systemd is largely **redundant** — 44 of 44 module compose
files already declare a restart policy, so Docker is already the supervisor. Supervising
HAL's **9 long-running Node daemons** is not redundant: `Restart=always` is currently the
only thing between a crash and a dead node.
- **The hard part is fate-sharing, not systemd.** meshware already cannot restart itself and
carries a documented workaround for it. Any mesh-native supervisor inherits that problem
recursively unless it sits outside the mesh's own process tree — at which point it is an
OS-level supervisor again, just reinvented.
- **There is a third option neither of us named**, and it is the one that also solves Phase 0:
run HAL's own daemons as containers, making Docker the supervisor for everything. Local and
production then have the same shape rather than a translation layer between them.
- **It cannot be all-or-nothing**, and ADR 0001 already says why: a human agent acts through a
shell and a desktop. Those parts are on the host by definition.
- One incidental finding: the automatic node rescue that documentation describes **does not
exist**. No unit declares `OnFailure=`, and nothing calls `hal-rescue.sh` on a timer.
Detail and costs in [`analysis.md`](analysis.md).
## What it became
*Closed 2026-08-28.* The decision this effort asked for was taken — and taken without citing it,
which is why the effort sat `active` for five days after being answered. Recorded here because
finding that is the point of a sweep.
**The third option is what the mesh adopted.** `Docker is the supervisor for everything` is
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md): the
host is a plain process on the machine and everything above tier 0 is a container. The substrate
bootstrap declares no service at all — it is package, container, action, container — so the
44-of-44 restart policies this effort counted are the supervision, exactly as it argued.
**Fate-sharing was the hard part, and it is solved the way this effort predicted.** It said any
mesh-native supervisor inherits the problem *unless it sits outside the mesh's own process
tree*. [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) puts
the launcher there: it supervises the host as a child and shares no code with it, so a host that
cannot start is still recovered.
**It is not all-or-nothing, as this effort insisted.** A human agent acts through a shell and a
desktop, and those are on the host. So is the host itself — the one thing an init starts.
## What is not closed
**The automatic node rescue the documentation describes does not exist.** No unit declares
`OnFailure=`, and nothing calls the rescue script on a timer. That is a documented behaviour
which never happens, and it outlives this effort — filed as
[`04-ISSUES/008`](../../04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md)
rather than closed with it.