Files
hq/01-RESEARCH/003-service-supervision/00-overview.md
T
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00

4.1 KiB

status, initiated, touches, became
status initiated touches became
graduated 2026-08-22
03-DESIGN/00-as-is/05-runtime-and-installation.md
02-DECISIONS/0005-the-node-host.md
02-DECISIONS/0005-the-node-host.md
03-DESIGN/01-to-be/05-the-node-host.md

003 — Who supervises a service

No longer blocks Phase 0 (see below).

  • Initiated by: jochen, 2026-08-22, in response to 002-local-mesh open question 1
  • Areas touched: every module shipping a systemd/ directory (14), the hal-module@ template, hal/sdk feature handlers, dev_up, log_tail / systemd_journal, the bootstrap scripts.

The question

002 asked how a containerised node runs a module service, given that a module service is defined today as a systemd unit running docker compose against /services/. The response was the better question:

If it's possible to run systemd inside a container, that's the way to go I think. However, what would the cost be to step away from systemd to run our services and set it up in a different way? More hal mesh approach.

This effort answers the cost half. It does not choose.

Summary of findings

  • The question splits in two, and the halves have opposite answers. Supervising docker compose stacks through systemd is largely redundant — 44 of 44 module compose files already declare a restart policy, so Docker is already the supervisor. Supervising HAL's 9 long-running Node daemons is not redundant: Restart=always is currently the only thing between a crash and a dead node.
  • The hard part is fate-sharing, not systemd. meshware already cannot restart itself and carries a documented workaround for it. Any mesh-native supervisor inherits that problem recursively unless it sits outside the mesh's own process tree — at which point it is an OS-level supervisor again, just reinvented.
  • There is a third option neither of us named, and it is the one that also solves Phase 0: run HAL's own daemons as containers, making Docker the supervisor for everything. Local and production then have the same shape rather than a translation layer between them.
  • It cannot be all-or-nothing, and ADR 0001 already says why: a human agent acts through a shell and a desktop. Those parts are on the host by definition.
  • One incidental finding: the automatic node rescue that documentation describes does not exist. No unit declares OnFailure=, and nothing calls hal-rescue.sh on a timer.

Detail and costs in analysis.md.

What it became

Closed 2026-08-28. The decision this effort asked for was taken — and taken without citing it, which is why the effort sat active for five days after being answered. Recorded here because finding that is the point of a sweep.

The third option is what the mesh adopted. Docker is the supervisor for everything is ADR 0005: the host is a plain process on the machine and everything above tier 0 is a container. The substrate bootstrap declares no service at all — it is package, container, action, container — so the 44-of-44 restart policies this effort counted are the supervision, exactly as it argued.

Fate-sharing was the hard part, and it is solved the way this effort predicted. It said any mesh-native supervisor inherits the problem unless it sits outside the mesh's own process tree. ADR 0005 puts the launcher there: it supervises the host as a child and shares no code with it, so a host that cannot start is still recovered.

It is not all-or-nothing, as this effort insisted. A human agent acts through a shell and a desktop, and those are on the host. So is the host itself — the one thing an init starts.

What is not closed

The automatic node rescue the documentation describes does not exist. No unit declares OnFailure=, and nothing calls the rescue script on a timer. That is a documented behaviour which never happens, and it outlives this effort — filed as 04-ISSUES/008 rather than closed with it.