Files
hq/01-RESEARCH/003-service-supervision/00-overview.md
T
jschoubben c0b35652d0 The numbering is the flow: decisions are 02, design is 03
papa-hq reads 01 research -> 03 decision -> 02 design. The order is a
scar, not a choice: 02-DESIGN existed from its initial commit, and when
adr/ was finally promoted on 2026-07-13 it took the next free number
rather than its place in the sequence. By then design was too settled to
renumber.

hal-hq was three commits old, so it is not. adr/ becomes 02-DECISIONS and
02-DESIGN becomes 03-DESIGN, and following the folder numbers now walks
the process in the order it happens: research produces a decision, the
decision authorises a design.

00-GENESIS becomes 00-META, matching papa's rename from the same
restructure.

Every path reference rewritten across documents, frontmatter, playbooks
and skills. All links resolve; all 58 frontmatter blocks parse and their
path fields still point at files that exist.
2026-08-23 18:05:11 +02:00

54 lines
2.6 KiB
Markdown

---
status: active
initiated: 2026-08-22
touches: [03-DESIGN/00-as-is/05-runtime-and-installation.md]
became: []
---
# 003 — Who supervises a service
**No longer blocks Phase 0** (see below).
- **Initiated by:** jochen, 2026-08-22, in response to
[`002-local-mesh`](../002-local-mesh/analysis.md) open question 1
- **Areas touched:** every module shipping a `systemd/` directory (14), the
`hal-module@` template, `hal/sdk` feature handlers, `dev_up`, `log_tail` /
`systemd_journal`, the bootstrap scripts.
## The question
`002` asked how a containerised node runs a module service, given that a module service is
defined today as a systemd unit running `docker compose` against `/services/`. The response
was the better question:
> If it's possible to run systemd inside a container, that's the way to go I think. However,
> what would the cost be to step away from systemd to run our services and set it up in a
> different way? More hal mesh approach.
This effort answers the cost half. It does not choose.
## Summary of findings
- **The question splits in two**, and the halves have opposite answers. Supervising
`docker compose` stacks through systemd is largely **redundant** — 44 of 44 module compose
files already declare a restart policy, so Docker is already the supervisor. Supervising
HAL's **9 long-running Node daemons** is not redundant: `Restart=always` is currently the
only thing between a crash and a dead node.
- **The hard part is fate-sharing, not systemd.** meshware already cannot restart itself and
carries a documented workaround for it. Any mesh-native supervisor inherits that problem
recursively unless it sits outside the mesh's own process tree — at which point it is an
OS-level supervisor again, just reinvented.
- **There is a third option neither of us named**, and it is the one that also solves Phase 0:
run HAL's own daemons as containers, making Docker the supervisor for everything. Local and
production then have the same shape rather than a translation layer between them.
- **It cannot be all-or-nothing**, and ADR 0015 already says why: a human agent acts through a
shell and a desktop. Those parts are on the host by definition.
- One incidental finding: the automatic node rescue that documentation describes **does not
exist**. No unit declares `OnFailure=`, and nothing calls `hal-rescue.sh` on a timer.
Detail and costs in [`analysis.md`](analysis.md).
## Decision needed
Which supervision model the mesh adopts, recorded in a decision record before Phase 0 builds
anything. The options and their costs are in `analysis.md` under "Options".