A sweep of the nine active research efforts. 003 was answered five days ago and nobody closed it -- the decision it asked for was taken without citing it, which is how an effort stays `active` after being resolved. Its recommendation is what the mesh adopted, and the match is exact rather than approximate. "Run the daemons as containers, making Docker the supervisor for everything" is ADR 0057. Its warning that a mesh-native supervisor inherits fate-sharing "unless it sits outside the mesh's own process tree" is where ADR 0061 put the launcher. And its insistence that it cannot be all-or-nothing is why the host itself is the one thing an init starts. Its incidental finding does not graduate with it, so it is now issue 008: the automatic node rescue the documentation describes does not exist. No unit declares OnFailure=, nothing calls the rescue script on a timer. That is worse than having no rescue. A rescue nobody wrote is a gap somebody can see; a documented one that is absent is a gap nobody looks for, and the documentation is read exactly when a node has failed and somebody is deciding whether to intervene. The issue names two honest resolutions -- implement it, or delete the documentation and say a failed node needs a person -- and says the choice is scheduling rather than technical, since the new host's recovery is built and tested. It also says what would make the finding certain: it came from reading the repository, and confirming it on a running node is the difference between "no unit declares this" and "no unit in the source declares this".
4.2 KiB
status, initiated, touches, became
| status | initiated | touches | became | ||||
|---|---|---|---|---|---|---|---|
| graduated | 2026-08-22 |
|
|
003 — Who supervises a service
No longer blocks Phase 0 (see below).
- Initiated by: jochen, 2026-08-22, in response to
002-local-meshopen question 1 - Areas touched: every module shipping a
systemd/directory (14), thehal-module@template,hal/sdkfeature handlers,dev_up,log_tail/systemd_journal, the bootstrap scripts.
The question
002 asked how a containerised node runs a module service, given that a module service is
defined today as a systemd unit running docker compose against /services/. The response
was the better question:
If it's possible to run systemd inside a container, that's the way to go I think. However, what would the cost be to step away from systemd to run our services and set it up in a different way? More hal mesh approach.
This effort answers the cost half. It does not choose.
Summary of findings
- The question splits in two, and the halves have opposite answers. Supervising
docker composestacks through systemd is largely redundant — 44 of 44 module compose files already declare a restart policy, so Docker is already the supervisor. Supervising HAL's 9 long-running Node daemons is not redundant:Restart=alwaysis currently the only thing between a crash and a dead node. - The hard part is fate-sharing, not systemd. meshware already cannot restart itself and carries a documented workaround for it. Any mesh-native supervisor inherits that problem recursively unless it sits outside the mesh's own process tree — at which point it is an OS-level supervisor again, just reinvented.
- There is a third option neither of us named, and it is the one that also solves Phase 0: run HAL's own daemons as containers, making Docker the supervisor for everything. Local and production then have the same shape rather than a translation layer between them.
- It cannot be all-or-nothing, and ADR 0015 already says why: a human agent acts through a shell and a desktop. Those parts are on the host by definition.
- One incidental finding: the automatic node rescue that documentation describes does not
exist. No unit declares
OnFailure=, and nothing callshal-rescue.shon a timer.
Detail and costs in analysis.md.
What it became
Closed 2026-08-28. The decision this effort asked for was taken — and taken without citing it,
which is why the effort sat active for five days after being answered. Recorded here because
finding that is the point of a sweep.
The third option is what the mesh adopted. Docker is the supervisor for everything is
ADR 0057: the
host is a plain process on the machine and everything above tier 0 is a container. The substrate
bootstrap declares no service at all — it is package, container, action, container — so the
44-of-44 restart policies this effort counted are the supervision, exactly as it argued.
Fate-sharing was the hard part, and it is solved the way this effort predicted. It said any mesh-native supervisor inherits the problem unless it sits outside the mesh's own process tree. ADR 0061 puts the launcher there: it supervises the host as a child and shares no code with it, so a host that cannot start is still recovered.
It is not all-or-nothing, as this effort insisted. A human agent acts through a shell and a desktop, and those are on the host. So is the host itself — the one thing an init starts.
What is not closed
The automatic node rescue the documentation describes does not exist. No unit declares
OnFailure=, and nothing calls the rescue script on a timer. That is a documented behaviour
which never happens, and it outlives this effort — filed as
04-ISSUES/008
rather than closed with it.