Files
hq/01-RESEARCH/003-service-supervision/00-overview.md
T
jschoubben 77f3a4cea7 Consolidate: 65 decision records to 23
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.

  the node host          8 -> 1    applies not decides, depends on nothing,
                                   per operating system, root service, the
                                   launcher, episodic, what a declaration is,
                                   actions from the bundle only
  a node and how it joins 4 -> 1   what a node is, joining, the link as
                                   security boundary, the enrolment token
  modules and the graph   7 -> 1   everything is a module, no domain modules,
                                   three edges, provisioning, the core library
  substrate and control   6 -> 1   the test, seven contexts, one control plane,
    plane                          the authority is not a database, the named
                                   products, the pinned bundle
  connectivity            3 -> 1   a route is a grant, reachability declared,
                                   filter rules
  delivery                5 -> 1   reconciliation not a pipeline, artifacts,
                                   the three silos, a failed step, the verdict
  the lab                 5 -> 1   (earlier)
  how this repository     10 -> 1  (earlier)
    works

Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.

The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.

The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
2026-08-28 20:03:24 +02:00

4.1 KiB

status, initiated, touches, became
status initiated touches became
graduated 2026-08-22
03-DESIGN/00-as-is/05-runtime-and-installation.md
02-DECISIONS/0037-the-node-host.md
02-DECISIONS/0037-the-node-host.md
03-DESIGN/01-to-be/05-the-node-host.md

003 — Who supervises a service

No longer blocks Phase 0 (see below).

  • Initiated by: jochen, 2026-08-22, in response to 002-local-mesh open question 1
  • Areas touched: every module shipping a systemd/ directory (14), the hal-module@ template, hal/sdk feature handlers, dev_up, log_tail / systemd_journal, the bootstrap scripts.

The question

002 asked how a containerised node runs a module service, given that a module service is defined today as a systemd unit running docker compose against /services/. The response was the better question:

If it's possible to run systemd inside a container, that's the way to go I think. However, what would the cost be to step away from systemd to run our services and set it up in a different way? More hal mesh approach.

This effort answers the cost half. It does not choose.

Summary of findings

  • The question splits in two, and the halves have opposite answers. Supervising docker compose stacks through systemd is largely redundant — 44 of 44 module compose files already declare a restart policy, so Docker is already the supervisor. Supervising HAL's 9 long-running Node daemons is not redundant: Restart=always is currently the only thing between a crash and a dead node.
  • The hard part is fate-sharing, not systemd. meshware already cannot restart itself and carries a documented workaround for it. Any mesh-native supervisor inherits that problem recursively unless it sits outside the mesh's own process tree — at which point it is an OS-level supervisor again, just reinvented.
  • There is a third option neither of us named, and it is the one that also solves Phase 0: run HAL's own daemons as containers, making Docker the supervisor for everything. Local and production then have the same shape rather than a translation layer between them.
  • It cannot be all-or-nothing, and ADR 0015 already says why: a human agent acts through a shell and a desktop. Those parts are on the host by definition.
  • One incidental finding: the automatic node rescue that documentation describes does not exist. No unit declares OnFailure=, and nothing calls hal-rescue.sh on a timer.

Detail and costs in analysis.md.

What it became

Closed 2026-08-28. The decision this effort asked for was taken — and taken without citing it, which is why the effort sat active for five days after being answered. Recorded here because finding that is the point of a sweep.

The third option is what the mesh adopted. Docker is the supervisor for everything is ADR 0037: the host is a plain process on the machine and everything above tier 0 is a container. The substrate bootstrap declares no service at all — it is package, container, action, container — so the 44-of-44 restart policies this effort counted are the supervision, exactly as it argued.

Fate-sharing was the hard part, and it is solved the way this effort predicted. It said any mesh-native supervisor inherits the problem unless it sits outside the mesh's own process tree. ADR 0037 puts the launcher there: it supervises the host as a child and shares no code with it, so a host that cannot start is still recovered.

It is not all-or-nothing, as this effort insisted. A human agent acts through a shell and a desktop, and those are on the host. So is the host itself — the one thing an init starts.

What is not closed

The automatic node rescue the documentation describes does not exist. No unit declares OnFailure=, and nothing calls the rescue script on a timer. That is a documented behaviour which never happens, and it outlives this effort — filed as 04-ISSUES/008 rather than closed with it.