The flow the process overview draws — idea/symptom -> decision -> to-be design -> code -> as-is — was enforced by nothing. cycle.py now refuses a to-be design naming no decision, an in-progress/implemented design naming no owning code, a located/fixed issue with no owner, a fixed/resolved issue with no fix, and a graduated research overview that does not say what it became. AGENTS.md carries the cycle and a where-to-look table so a fresh session (or a cleared context) finds the chain in frontmatter instead of assuming it. Grounding the check surfaced two real gaps, fixed here: the work-ahead design named no owning code, and research 003 listed one became target twice. https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
4.1 KiB
status, initiated, touches, became
| status | initiated | touches | became | |||
|---|---|---|---|---|---|---|
| graduated | 2026-08-22 |
|
|
003 — Who supervises a service
No longer blocks Phase 0 (see below).
- Initiated by: jochen, 2026-08-22, in response to
002-local-meshopen question 1 - Areas touched: every module shipping a
systemd/directory (14), thehal-module@template,hal/sdkfeature handlers,dev_up,log_tail/systemd_journal, the bootstrap scripts.
The question
002 asked how a containerised node runs a module service, given that a module service is
defined today as a systemd unit running docker compose against /services/. The response
was the better question:
If it's possible to run systemd inside a container, that's the way to go I think. However, what would the cost be to step away from systemd to run our services and set it up in a different way? More hal mesh approach.
This effort answers the cost half. It does not choose.
Summary of findings
- The question splits in two, and the halves have opposite answers. Supervising
docker composestacks through systemd is largely redundant — 44 of 44 module compose files already declare a restart policy, so Docker is already the supervisor. Supervising HAL's 9 long-running Node daemons is not redundant:Restart=alwaysis currently the only thing between a crash and a dead node. - The hard part is fate-sharing, not systemd. meshware already cannot restart itself and carries a documented workaround for it. Any mesh-native supervisor inherits that problem recursively unless it sits outside the mesh's own process tree — at which point it is an OS-level supervisor again, just reinvented.
- There is a third option neither of us named, and it is the one that also solves Phase 0: run HAL's own daemons as containers, making Docker the supervisor for everything. Local and production then have the same shape rather than a translation layer between them.
- It cannot be all-or-nothing, and ADR 0001 already says why: a human agent acts through a shell and a desktop. Those parts are on the host by definition.
- One incidental finding: the automatic node rescue that documentation describes does not
exist. No unit declares
OnFailure=, and nothing callshal-rescue.shon a timer.
Detail and costs in analysis.md.
What it became
Closed 2026-08-28. The decision this effort asked for was taken — and taken without citing it,
which is why the effort sat active for five days after being answered. Recorded here because
finding that is the point of a sweep.
The third option is what the mesh adopted. Docker is the supervisor for everything is
ADR 0005: the
host is a plain process on the machine and everything above tier 0 is a container. The substrate
bootstrap declares no service at all — it is package, container, action, container — so the
44-of-44 restart policies this effort counted are the supervision, exactly as it argued.
Fate-sharing was the hard part, and it is solved the way this effort predicted. It said any mesh-native supervisor inherits the problem unless it sits outside the mesh's own process tree. ADR 0005 puts the launcher there: it supervises the host as a child and shares no code with it, so a host that cannot start is still recovered.
It is not all-or-nothing, as this effort insisted. A human agent acts through a shell and a desktop, and those are on the host. So is the host itself — the one thing an init starts.
What is not closed
The automatic node rescue the documentation describes does not exist. No unit declares
OnFailure=, and nothing calls the rescue script on a timer. That is a documented behaviour
which never happens, and it outlives this effort — filed as
04-ISSUES/008
rather than closed with it.