The flow the process overview draws — idea/symptom -> decision -> to-be design -> code -> as-is — was enforced by nothing. cycle.py now refuses a to-be design naming no decision, an in-progress/implemented design naming no owning code, a located/fixed issue with no owner, a fixed/resolved issue with no fix, and a graduated research overview that does not say what it became. AGENTS.md carries the cycle and a where-to-look table so a fresh session (or a cleared context) finds the chain in frontmatter instead of assuming it. Grounding the check surfaced two real gaps, fixed here: the work-ahead design named no owning code, and research 003 listed one became target twice. https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
80 lines
4.1 KiB
Markdown
80 lines
4.1 KiB
Markdown
---
|
|
status: graduated
|
|
initiated: 2026-08-22
|
|
touches: [03-DESIGN/00-as-is/05-runtime-and-installation.md]
|
|
became:
|
|
- 02-DECISIONS/0005-the-node-host.md
|
|
- 03-DESIGN/01-to-be/05-the-node-host.md
|
|
---
|
|
|
|
# 003 — Who supervises a service
|
|
|
|
**No longer blocks Phase 0** (see below).
|
|
- **Initiated by:** jochen, 2026-08-22, in response to
|
|
[`002-local-mesh`](../002-local-mesh/analysis.md) open question 1
|
|
- **Areas touched:** every module shipping a `systemd/` directory (14), the
|
|
`hal-module@` template, `hal/sdk` feature handlers, `dev_up`, `log_tail` /
|
|
`systemd_journal`, the bootstrap scripts.
|
|
|
|
## The question
|
|
|
|
`002` asked how a containerised node runs a module service, given that a module service is
|
|
defined today as a systemd unit running `docker compose` against `/services/`. The response
|
|
was the better question:
|
|
|
|
> If it's possible to run systemd inside a container, that's the way to go I think. However,
|
|
> what would the cost be to step away from systemd to run our services and set it up in a
|
|
> different way? More hal mesh approach.
|
|
|
|
This effort answers the cost half. It does not choose.
|
|
|
|
## Summary of findings
|
|
|
|
- **The question splits in two**, and the halves have opposite answers. Supervising
|
|
`docker compose` stacks through systemd is largely **redundant** — 44 of 44 module compose
|
|
files already declare a restart policy, so Docker is already the supervisor. Supervising
|
|
HAL's **9 long-running Node daemons** is not redundant: `Restart=always` is currently the
|
|
only thing between a crash and a dead node.
|
|
- **The hard part is fate-sharing, not systemd.** meshware already cannot restart itself and
|
|
carries a documented workaround for it. Any mesh-native supervisor inherits that problem
|
|
recursively unless it sits outside the mesh's own process tree — at which point it is an
|
|
OS-level supervisor again, just reinvented.
|
|
- **There is a third option neither of us named**, and it is the one that also solves Phase 0:
|
|
run HAL's own daemons as containers, making Docker the supervisor for everything. Local and
|
|
production then have the same shape rather than a translation layer between them.
|
|
- **It cannot be all-or-nothing**, and ADR 0001 already says why: a human agent acts through a
|
|
shell and a desktop. Those parts are on the host by definition.
|
|
- One incidental finding: the automatic node rescue that documentation describes **does not
|
|
exist**. No unit declares `OnFailure=`, and nothing calls `hal-rescue.sh` on a timer.
|
|
|
|
Detail and costs in [`analysis.md`](analysis.md).
|
|
|
|
## What it became
|
|
|
|
*Closed 2026-08-28.* The decision this effort asked for was taken — and taken without citing it,
|
|
which is why the effort sat `active` for five days after being answered. Recorded here because
|
|
finding that is the point of a sweep.
|
|
|
|
**The third option is what the mesh adopted.** `Docker is the supervisor for everything` is
|
|
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md): the
|
|
host is a plain process on the machine and everything above tier 0 is a container. The substrate
|
|
bootstrap declares no service at all — it is package, container, action, container — so the
|
|
44-of-44 restart policies this effort counted are the supervision, exactly as it argued.
|
|
|
|
**Fate-sharing was the hard part, and it is solved the way this effort predicted.** It said any
|
|
mesh-native supervisor inherits the problem *unless it sits outside the mesh's own process
|
|
tree*. [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) puts
|
|
the launcher there: it supervises the host as a child and shares no code with it, so a host that
|
|
cannot start is still recovered.
|
|
|
|
**It is not all-or-nothing, as this effort insisted.** A human agent acts through a shell and a
|
|
desktop, and those are on the host. So is the host itself — the one thing an init starts.
|
|
|
|
## What is not closed
|
|
|
|
**The automatic node rescue the documentation describes does not exist.** No unit declares
|
|
`OnFailure=`, and nothing calls the rescue script on a timer. That is a documented behaviour
|
|
which never happens, and it outlives this effort — filed as
|
|
[`04-ISSUES/008`](../../04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md)
|
|
rather than closed with it.
|