What the mesh is, what it is becoming, and why. Implementation lives in the code repositories; the reasoning lives here. 00-GENESIS mission, engineering context, effect, and the rules that hold 01-RESEARCH investigations, before they harden into design 02-DESIGN the authoritative specification adr numbered decisions — what was chosen, and what was rejected DECISIONS.md the ledger: every decision, in the order it was taken Written for a reader who is not its author and has no access to the mesh it describes. Addresses use the documentation ranges of RFC 5737 and RFC 1918; nodes are named by role. Single initial commit by intent. The prior history came from a private repository and carried operational detail — a routable address identified as a VPN hub, real domain names, a hosting provider — which sanitising a tip commit would not have removed from the log.
12 KiB
Who supervises a service — the cost of leaving systemd
Measured 2026-08-22. Every claim is a file location or a count.
1. What systemd actually does for HAL
14 modules ship a systemd/ directory. The units divide cleanly into three kinds:
| kind | count | what it is |
|---|---|---|
| long-running daemons | 9 | hal/brain, hal/meshware, hal/coordinator, hal/cortex, hal/env-sync, hal/file-share, hal/thoughts, noxflow/runtime, desktop_notifications — all Restart=always, RestartSec=10 |
| scheduled one-shots | 6 timers | docker-prune, docker-registry maintenance, claude-code sessions + usage, health, mailu cert-sync |
| the template | 1 | hal-module@.service — the per-module Docker lifecycle |
Only mailu is system-scope. Everything else is a user unit under
~/.config/systemd/user, auto-detected from directory presence — the module author names
the file and that filename is the unit name
(modules/hal/sdk/src/feature-handlers/systemd.ts:31-42).
The template is the piece 002 tripped over
(modules/hal/meshware/systemd/hal-module@.service):
ExecStart=docker compose --project-directory /services/%i --project-name %i up --remove-orphans --pull always
ExecStop=docker compose --project-directory /services/%i --project-name %i down
Restart=on-failure
StartLimitIntervalSec=120
StartLimitBurst=5
Systemd's roles, then: restart-on-failure, start at login, env-file loading, ordering,
per-module lifecycle, timers, and — via journald — the only log store the 9 Node daemons
have (modules/hal/sdk/src/tools/log-tail.ts:50-53).
2. The question splits in two, and the halves disagree
For docker compose stacks, systemd is mostly redundant
44 of 44 module docker-compose.yml files declare a container-level restart: policy —
always or unless-stopped, none missing. Docker's own daemon already restarts crashed
containers, independently of systemd.
hal-module@.service does not supervise the containers. It supervises the docker compose up foreground process. It is a second layer on top of a restart policy that already
works. What it genuinely adds is narrower than it looks:
- one uniform verb for every module type —
systemctl --user start hal-module@Xrather than remembering each module's compose invocation, relied on across a dozen call sites inmeshware.ts:124-154,dev-env.ts:142-163,installer-core.ts,infra.ts:166 - recovery when the
docker compose up --pull alwaysprocess itself dies — a failed image pull, not a crashed container - rate-limited restart (
StartLimitBurst=5) so a broken stack does not spin
That is real, but it is a convenience layer, not a safety layer. This half could go at moderate cost.
For the 9 Node daemons, systemd is load-bearing
There is no alternative supervisor anywhere in the repo. No PM2, no forever, no nodemon, no
watchdog loop — all checked, zero hits. Restart=always is the only thing standing between
a crashed daemon and a dead node.
This half is the actual question.
3. The hard part is fate-sharing
The difficulty is not systemd. It is that a supervisor must not share fate with what it supervises, and the codebase already has a scar from exactly this.
meshware cannot restart itself mid-request: killing the process before it ACKs the AMQP
message loses the message. So it defers its own restart by two seconds after closing the
connection (modules/hal/meshware/daemon/src/cerebellum.ts:815-828) — a commented
workaround for a problem that only exists because the thing being restarted is the thing
doing the restarting.
Any mesh-native supervisor inherits this recursively. Something has to be the outermost always-alive process, and if it is written by the mesh, the mesh must supervise it, and so on. The recursion only terminates at a process the mesh does not own. Today that is systemd. A "more mesh" supervisor that is itself a mesh process is not a smaller problem; it is the same problem with a new name.
This is the strongest argument for the status quo, and it is worth stating plainly before looking at alternatives.
4. The option neither of us named
There is a third answer that terminates the recursion in something that is not systemd and not written by us: run HAL's own daemons as containers.
Docker is already the outermost supervisor for 44 of 44 module stacks. It does not share
fate with the mesh. It already has restart policies, backoff, and a log store that
log_tail already speaks (log-tail.ts:43-47 reads docker logs for Docker modules
today). Extending it from "the things the mesh runs" to "the mesh itself" is not new
machinery — it is applying machinery the mesh already trusts to one more case.
What that buys, beyond supervision:
- Phase 0 stops being a translation. A containerised node becomes the same shape as a
production node, rather than a local approximation with a systemd-shaped hole in it. The
divergence that
002open question 3 worries about largely disappears. /services/and~/.config/hal/stop being special. Mounts, not host paths.- The GENESIS "dogfood everything" value gets easier, not harder: the mesh's own components would ship and run exactly like everything else it carries.
What it costs, honestly:
- journald → docker logs for the 9 daemons.
log_tailalready handles both, butsystemd_journaland the health checks that read unit state (modules/hal/mesh/health.sh:34-40,modules/hal/health/hal-health.sh:335-395) would need a container-aware path. - Boot start becomes Docker's
restart: alwaysplus the Docker daemon being enabled at boot — which is still one systemd unit, but the OS's own, not ours. - A container needs the host to be reachable for anything that touches the node itself. Some of these daemons exist precisely to write host files.
5. Why it cannot be all-or-nothing — and ADR 0001 already says so
Some of what runs under systemd today cannot be containerised, and the reason is already in the domain model. ADR 0001:
a non-human agent acts through a spawned session — a human agent acts through a shell or desktop
hal/brain has two modes (modules/hal/brain/daemon/src/brainstem.ts:7-13): a daemon mode
that is an AMQP relay, and a cortex mode that is an MCP server over stdio for an
interactive Claude Code session. The second is a human agent's modality. It runs in the
human's shell, on the human's node, against the human's ~/.claude. Containerising it is
not a hard engineering problem, it is a category error.
The same holds for desktop_notifications and everything hal/desktop-environment touches.
So the line is not "systemd or not". It is:
| belongs where | |
|---|---|
| mesh daemons — meshware, coordinator, env-sync, thoughts, file-share, brain in daemon mode | supervisable by Docker; candidates to containerise |
| human-modality surfaces — brain in cortex mode, desktop notifications, desktop environment | on the host, by definition |
| module stacks | already Docker; systemd layer is the redundant part |
This split is not a compromise between the options. It is what the domain model implies, and it is a decent sign that the model is doing work.
6. Options, with costs
| option | cost | what it buys | |
|---|---|---|---|
| A | Keep systemd; systemd-in-container for Phase 0 | privileged containers, heavy images, slow iteration; local mesh keeps a shape production does not have | nothing changes in production; smallest change to the mesh |
| B | Drop the hal-module@ layer only — let Docker's restart policies supervise stacks directly |
reimplement uniform start/stop across ~12 call sites; lose rate-limited restart and pull-failure recovery | removes the redundant layer; does not solve Phase 0 on its own, since the daemons still need supervising |
| C | Containerise the mesh daemons; Docker supervises | container-aware systemd_journal/health; host access for daemons that write host files; the human-modality surfaces stay on the host regardless |
terminates the fate-sharing recursion without writing a supervisor; makes local and production the same shape; Phase 0 becomes much less of a special case |
| D | Write a mesh-native supervisor | the fate-sharing recursion (§3), plus matching systemd's maturity — backoff, resource limits, clean SIGTERM (noxflow/runtime already depends on TimeoutStopSec=60) — on machines that are somebody's daily driver |
most "mesh"; least justified by the evidence |
D is the option the phrasing "more hal mesh approach" points at, and the evidence argues against it. Supervision is not a domain concern the mesh is better placed to solve than the OS; the mesh's distinguishing feature is brokering capabilities, not restarting processes. C gets the benefit D is reaching for — the mesh not depending on host-specific init — without the recursion.
C and B compose. C is the one that pays for Phase 0.
Worth noting what is not in this table: the dozens of systemctl call sites that
configure the host OS's own units — NetworkManager, resolved, oomd, zram, docker.service,
fail2ban, sshd, zfs, across modules/asusd, g14-power, wireguard, sshd, zfs and
others. Those are not HAL supervising itself; they are HAL configuring an Arch box. They
are out of scope for every option above and do not go away under any of them.
7. An incidental finding
Documentation describes an automatic node rescue: hal-rescue.sh:20 states it is "triggered
automatically by hal-health.timer when hal-meshware is failed".
It is not. hal-health.sh contains no call to hal-rescue.sh (checked, zero matches),
and no unit in the repository declares OnFailure= (checked, zero matches). The only
real triggers are the manual rescue_node tool
(modules/hal/sdk/src/tools/deploy-rescue.ts:77-92) and running install.d/rescue.sh by
hand.
This is worth recording for two reasons. It weakens any argument that systemd-adjacent
self-healing is already wired — it is not. And it is another instance of the pattern this
whole refactor is about: a documented mechanism that does not exist, believed because it
was written down. 00-GENESIS/how-we-build.md calls this out as a rule; here it is again,
found by grep.
8. Open question
Which supervision model does the mesh adopt? A, B, C, D or a combination.
This no longer gates Phase 0. When this was written, the local mesh was assumed to be
built from application containers, which forced the question — there is no natural way to
run an init system inside one. The decision of 2026-08-22 to build development nodes as
system containers (see 02-DESIGN/01-end-to-end-testing.md)
removes that pressure entirely: a system container runs a real init, so the existing model
works unmodified and the lab needs no answer here to exist.
What remains is the question on its own merits, which is worth keeping open because the evidence above still holds: option B removes a layer that 44 of 44 module stacks have already made redundant, and option C would let the mesh stop depending on host-specific init. Neither is urgent. Both are now cheap to try, because there is somewhere to try them.
Option D — a mesh-written supervisor — remains the one the evidence argues against, for the fate-sharing reason in §3.
References
002-local-mesh— the effort this came out ofadr/0001— agent modality, which decides what cannot leave the hostmodules/hal/meshware/daemon/src/cerebellum.ts:815-828— the self-restart workaroundmodules/hal/meshware/systemd/hal-module@.service— the per-module Docker lifecyclemodules/hal/sdk/src/feature-handlers/systemd.ts— detection, install, start