--- status: ongoing initiated: 2026-08-22 touches: [02-DESIGN/00-as-is/05-runtime-and-installation.md] became: [] --- # 003 — Who supervises a service **No longer blocks Phase 0** (see below). - **Initiated by:** jochen, 2026-08-22, in response to [`002-local-mesh`](../002-local-mesh/analysis.md) open question 1 - **Areas touched:** every module shipping a `systemd/` directory (14), the `hal-module@` template, `hal/sdk` feature handlers, `dev_up`, `log_tail` / `systemd_journal`, the bootstrap scripts. ## The question `002` asked how a containerised node runs a module service, given that a module service is defined today as a systemd unit running `docker compose` against `/services/`. The response was the better question: > If it's possible to run systemd inside a container, that's the way to go I think. However, > what would the cost be to step away from systemd to run our services and set it up in a > different way? More hal mesh approach. This effort answers the cost half. It does not choose. ## Summary of findings - **The question splits in two**, and the halves have opposite answers. Supervising `docker compose` stacks through systemd is largely **redundant** — 44 of 44 module compose files already declare a restart policy, so Docker is already the supervisor. Supervising HAL's **9 long-running Node daemons** is not redundant: `Restart=always` is currently the only thing between a crash and a dead node. - **The hard part is fate-sharing, not systemd.** meshware already cannot restart itself and carries a documented workaround for it. Any mesh-native supervisor inherits that problem recursively unless it sits outside the mesh's own process tree — at which point it is an OS-level supervisor again, just reinvented. - **There is a third option neither of us named**, and it is the one that also solves Phase 0: run HAL's own daemons as containers, making Docker the supervisor for everything. Local and production then have the same shape rather than a translation layer between them. - **It cannot be all-or-nothing**, and ADR 0015 already says why: a human agent acts through a shell and a desktop. Those parts are on the host by definition. - One incidental finding: the automatic node rescue that documentation describes **does not exist**. No unit declares `OnFailure=`, and nothing calls `hal-rescue.sh` on a timer. Detail and costs in [`analysis.md`](analysis.md). ## Decision needed Which supervision model the mesh adopts, recorded in a decision record before Phase 0 builds anything. The options and their costs are in `analysis.md` under "Options".