Files
hq/01-RESEARCH/003-service-supervision/00-overview.md
T
jschoubben e1febe8e0f Renumber the records 1 to 23
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18,
19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only
the archaeology of what used to be there.

Renumbered contiguously. Renames run in ascending order, so every target number
is already free and no two files ever collide.

The reference rewrite is one simultaneous pass rather than a sequence of
replacements. Numbers moved into slots other numbers were vacating -- the node
host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time
would have cascaded and silently pointed things at the wrong record.

Seven plain-text references survived the merges as prose rather than links,
naming records that no longer existed: the enrolment token, the link boundary,
what a declaration is, reachability, the repository structure. Each mapped to
the consolidated record that now holds it.

Verified rather than assumed: every [ADR NNNN](path) link now has matching text
and target, checked across the whole repository, and the checker passes.

Frontmatter `consolidates:` lists dropped -- they named records that are gone,
and each consolidated record already says in prose what it absorbed.
2026-08-28 23:28:34 +02:00

81 lines
4.1 KiB
Markdown

---
status: graduated
initiated: 2026-08-22
touches: [03-DESIGN/00-as-is/05-runtime-and-installation.md]
became:
- 02-DECISIONS/0016-the-node-host.md
- 02-DECISIONS/0016-the-node-host.md
- 03-DESIGN/01-to-be/05-the-node-host.md
---
# 003 — Who supervises a service
**No longer blocks Phase 0** (see below).
- **Initiated by:** jochen, 2026-08-22, in response to
[`002-local-mesh`](../002-local-mesh/analysis.md) open question 1
- **Areas touched:** every module shipping a `systemd/` directory (14), the
`hal-module@` template, `hal/sdk` feature handlers, `dev_up`, `log_tail` /
`systemd_journal`, the bootstrap scripts.
## The question
`002` asked how a containerised node runs a module service, given that a module service is
defined today as a systemd unit running `docker compose` against `/services/`. The response
was the better question:
> If it's possible to run systemd inside a container, that's the way to go I think. However,
> what would the cost be to step away from systemd to run our services and set it up in a
> different way? More hal mesh approach.
This effort answers the cost half. It does not choose.
## Summary of findings
- **The question splits in two**, and the halves have opposite answers. Supervising
`docker compose` stacks through systemd is largely **redundant** — 44 of 44 module compose
files already declare a restart policy, so Docker is already the supervisor. Supervising
HAL's **9 long-running Node daemons** is not redundant: `Restart=always` is currently the
only thing between a crash and a dead node.
- **The hard part is fate-sharing, not systemd.** meshware already cannot restart itself and
carries a documented workaround for it. Any mesh-native supervisor inherits that problem
recursively unless it sits outside the mesh's own process tree — at which point it is an
OS-level supervisor again, just reinvented.
- **There is a third option neither of us named**, and it is the one that also solves Phase 0:
run HAL's own daemons as containers, making Docker the supervisor for everything. Local and
production then have the same shape rather than a translation layer between them.
- **It cannot be all-or-nothing**, and ADR 0008 already says why: a human agent acts through a
shell and a desktop. Those parts are on the host by definition.
- One incidental finding: the automatic node rescue that documentation describes **does not
exist**. No unit declares `OnFailure=`, and nothing calls `hal-rescue.sh` on a timer.
Detail and costs in [`analysis.md`](analysis.md).
## What it became
*Closed 2026-08-28.* The decision this effort asked for was taken — and taken without citing it,
which is why the effort sat `active` for five days after being answered. Recorded here because
finding that is the point of a sweep.
**The third option is what the mesh adopted.** `Docker is the supervisor for everything` is
[ADR 0016](../../02-DECISIONS/0016-the-node-host.md): the
host is a plain process on the machine and everything above tier 0 is a container. The substrate
bootstrap declares no service at all — it is package, container, action, container — so the
44-of-44 restart policies this effort counted are the supervision, exactly as it argued.
**Fate-sharing was the hard part, and it is solved the way this effort predicted.** It said any
mesh-native supervisor inherits the problem *unless it sits outside the mesh's own process
tree*. [ADR 0016](../../02-DECISIONS/0016-the-node-host.md) puts
the launcher there: it supervises the host as a child and shares no code with it, so a host that
cannot start is still recovered.
**It is not all-or-nothing, as this effort insisted.** A human agent acts through a shell and a
desktop, and those are on the host. So is the host itself — the one thing an init starts.
## What is not closed
**The automatic node rescue the documentation describes does not exist.** No unit declares
`OnFailure=`, and nothing calls the rescue script on a timer. That is a documented behaviour
which never happens, and it outlives this effort — filed as
[`04-ISSUES/008`](../../04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md)
rather than closed with it.