Files
hq/01-RESEARCH/003-service-supervision/00-overview.md
T
jschoubben f05e4a0dce Follow papa-hq's research convention; the mesh links nothing
Research efforts move from status.md to 00-overview.md with active /
graduated / abandoned, matching papa-hq so the two repositories read the
same way. Playbooks, skills, README and the ledger follow.

Reverses yesterday's withdrawal of the symlink note in GENESIS. The note
was right and the withdrawal was wrong: the intent is that the mesh
creates no symlinks at all, so a founding document listing "symlinks, not
copies" as a design principle does point the opposite way from where this
is going, and that is a contradiction rather than a stale detail.

ADR 0018 records the position, proposed. ADR 0011 stays as it is — it is
the historical decision and the incident behind it is why anyone believes
either record — and is superseded in intent, not edited. Its one
editorial line, which called the wider reading false, is corrected to
state what is actually true: centralising who may link narrowed the
incident class without closing it, because a link the installer makes
resolves exactly like one made by hand.

The argument that kept linking was staleness. ADR 0004 removed it: every
managed file is already derived and reconciled, so a copy is the natural
form and a pointer into source is the shape the mesh's own model forbids
everywhere else. What is not settled, and is marked open, is how
staleness gets detected — which is the decision that makes or breaks it.
2026-08-23 09:29:09 +02:00

2.6 KiB

status, initiated, touches, became
status initiated touches became
active 2026-08-22
02-DESIGN/00-as-is/05-runtime-and-installation.md

003 — Who supervises a service

No longer blocks Phase 0 (see below).

  • Initiated by: jochen, 2026-08-22, in response to 002-local-mesh open question 1
  • Areas touched: every module shipping a systemd/ directory (14), the hal-module@ template, hal/sdk feature handlers, dev_up, log_tail / systemd_journal, the bootstrap scripts.

The question

002 asked how a containerised node runs a module service, given that a module service is defined today as a systemd unit running docker compose against /services/. The response was the better question:

If it's possible to run systemd inside a container, that's the way to go I think. However, what would the cost be to step away from systemd to run our services and set it up in a different way? More hal mesh approach.

This effort answers the cost half. It does not choose.

Summary of findings

  • The question splits in two, and the halves have opposite answers. Supervising docker compose stacks through systemd is largely redundant — 44 of 44 module compose files already declare a restart policy, so Docker is already the supervisor. Supervising HAL's 9 long-running Node daemons is not redundant: Restart=always is currently the only thing between a crash and a dead node.
  • The hard part is fate-sharing, not systemd. meshware already cannot restart itself and carries a documented workaround for it. Any mesh-native supervisor inherits that problem recursively unless it sits outside the mesh's own process tree — at which point it is an OS-level supervisor again, just reinvented.
  • There is a third option neither of us named, and it is the one that also solves Phase 0: run HAL's own daemons as containers, making Docker the supervisor for everything. Local and production then have the same shape rather than a translation layer between them.
  • It cannot be all-or-nothing, and ADR 0015 already says why: a human agent acts through a shell and a desktop. Those parts are on the host by definition.
  • One incidental finding: the automatic node rescue that documentation describes does not exist. No unit declares OnFailure=, and nothing calls hal-rescue.sh on a timer.

Detail and costs in analysis.md.

Decision needed

Which supervision model the mesh adopts, recorded in a decision record before Phase 0 builds anything. The options and their costs are in analysis.md under "Options".