What the mesh is, what it is becoming, and why. Implementation lives in the code repositories; the reasoning lives here. 00-GENESIS mission, engineering context, effect, and the rules that hold 01-RESEARCH investigations, before they harden into design 02-DESIGN the authoritative specification adr numbered decisions — what was chosen, and what was rejected DECISIONS.md the ledger: every decision, in the order it was taken Written for a reader who is not its author and has no access to the mesh it describes. Addresses use the documentation ranges of RFC 5737 and RFC 1918; nodes are named by role. Single initial commit by intent. The prior history came from a private repository and carried operational detail — a routable address identified as a VPN hub, real domain names, a hosting provider — which sanitising a tip commit would not have removed from the log.
2.5 KiB
2.5 KiB
003 — Who supervises a service
- Status: ONGOING — evidence gathered, options costed, decision open. No longer blocks Phase 0 (see below).
- Initiated by: jochen, 2026-08-22, in response to
002-local-meshopen question 1 - Areas touched: every module shipping a
systemd/directory (14), thehal-module@template,hal/sdkfeature handlers,dev_up,log_tail/systemd_journal, the bootstrap scripts.
The question
002 asked how a containerised node runs a module service, given that a module service is
defined today as a systemd unit running docker compose against /services/. The response
was the better question:
If it's possible to run systemd inside a container, that's the way to go I think. However, what would the cost be to step away from systemd to run our services and set it up in a different way? More hal mesh approach.
This effort answers the cost half. It does not choose.
Summary of findings
- The question splits in two, and the halves have opposite answers. Supervising
docker composestacks through systemd is largely redundant — 44 of 44 module compose files already declare a restart policy, so Docker is already the supervisor. Supervising HAL's 9 long-running Node daemons is not redundant:Restart=alwaysis currently the only thing between a crash and a dead node. - The hard part is fate-sharing, not systemd. meshware already cannot restart itself and carries a documented workaround for it. Any mesh-native supervisor inherits that problem recursively unless it sits outside the mesh's own process tree — at which point it is an OS-level supervisor again, just reinvented.
- There is a third option neither of us named, and it is the one that also solves Phase 0: run HAL's own daemons as containers, making Docker the supervisor for everything. Local and production then have the same shape rather than a translation layer between them.
- It cannot be all-or-nothing, and ADR 0001 already says why: a human agent acts through a shell and a desktop. Those parts are on the host by definition.
- One incidental finding: the automatic node rescue that documentation describes does not
exist. No unit declares
OnFailure=, and nothing callshal-rescue.shon a timer.
Detail and costs in analysis.md.
Decision needed
Which supervision model the mesh adopts, recorded in ADR 0002 before Phase 0 builds
anything. The options and their costs are in analysis.md under "Options".