Jochen asked whether the order made sense. It did not -- it followed when things happened to be decided, which after consolidation is fictional anyway since record 5 alone folds decisions taken across a week. Concretely wrong before: the domain statement sat at 8, after five engineering rules; the constitution was scattered across 5, 12 and 17; the tiers landed at 15, 16, 21 and 22 with process records in between. Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what runs on them and how it gets there (9-10), how it is built (11-16), how it is checked (17-18), how we work (19-23). Two things made this safe rather than free. It is a permutation, not a compaction, so the renames go through temporary names -- otherwise two files want one slot and one is lost. And the reference rewrite is a single simultaneous pass, because almost every number moved into a slot another number was vacating; replacing one at a time would have cascaded and pointed things at the wrong record while still resolving. Verified: 284 [ADR NNNN](path) links across the repository, all with matching text and target. The ordering principle is now stated in 19 rather than left implicit -- the repository already said "the numbering is the flow" about its folders, and there was no reason for the records to be the exception.
81 lines
4.1 KiB
Markdown
81 lines
4.1 KiB
Markdown
---
|
|
status: graduated
|
|
initiated: 2026-08-22
|
|
touches: [03-DESIGN/00-as-is/05-runtime-and-installation.md]
|
|
became:
|
|
- 02-DECISIONS/0005-the-node-host.md
|
|
- 02-DECISIONS/0005-the-node-host.md
|
|
- 03-DESIGN/01-to-be/05-the-node-host.md
|
|
---
|
|
|
|
# 003 — Who supervises a service
|
|
|
|
**No longer blocks Phase 0** (see below).
|
|
- **Initiated by:** jochen, 2026-08-22, in response to
|
|
[`002-local-mesh`](../002-local-mesh/analysis.md) open question 1
|
|
- **Areas touched:** every module shipping a `systemd/` directory (14), the
|
|
`hal-module@` template, `hal/sdk` feature handlers, `dev_up`, `log_tail` /
|
|
`systemd_journal`, the bootstrap scripts.
|
|
|
|
## The question
|
|
|
|
`002` asked how a containerised node runs a module service, given that a module service is
|
|
defined today as a systemd unit running `docker compose` against `/services/`. The response
|
|
was the better question:
|
|
|
|
> If it's possible to run systemd inside a container, that's the way to go I think. However,
|
|
> what would the cost be to step away from systemd to run our services and set it up in a
|
|
> different way? More hal mesh approach.
|
|
|
|
This effort answers the cost half. It does not choose.
|
|
|
|
## Summary of findings
|
|
|
|
- **The question splits in two**, and the halves have opposite answers. Supervising
|
|
`docker compose` stacks through systemd is largely **redundant** — 44 of 44 module compose
|
|
files already declare a restart policy, so Docker is already the supervisor. Supervising
|
|
HAL's **9 long-running Node daemons** is not redundant: `Restart=always` is currently the
|
|
only thing between a crash and a dead node.
|
|
- **The hard part is fate-sharing, not systemd.** meshware already cannot restart itself and
|
|
carries a documented workaround for it. Any mesh-native supervisor inherits that problem
|
|
recursively unless it sits outside the mesh's own process tree — at which point it is an
|
|
OS-level supervisor again, just reinvented.
|
|
- **There is a third option neither of us named**, and it is the one that also solves Phase 0:
|
|
run HAL's own daemons as containers, making Docker the supervisor for everything. Local and
|
|
production then have the same shape rather than a translation layer between them.
|
|
- **It cannot be all-or-nothing**, and ADR 0001 already says why: a human agent acts through a
|
|
shell and a desktop. Those parts are on the host by definition.
|
|
- One incidental finding: the automatic node rescue that documentation describes **does not
|
|
exist**. No unit declares `OnFailure=`, and nothing calls `hal-rescue.sh` on a timer.
|
|
|
|
Detail and costs in [`analysis.md`](analysis.md).
|
|
|
|
## What it became
|
|
|
|
*Closed 2026-08-28.* The decision this effort asked for was taken — and taken without citing it,
|
|
which is why the effort sat `active` for five days after being answered. Recorded here because
|
|
finding that is the point of a sweep.
|
|
|
|
**The third option is what the mesh adopted.** `Docker is the supervisor for everything` is
|
|
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md): the
|
|
host is a plain process on the machine and everything above tier 0 is a container. The substrate
|
|
bootstrap declares no service at all — it is package, container, action, container — so the
|
|
44-of-44 restart policies this effort counted are the supervision, exactly as it argued.
|
|
|
|
**Fate-sharing was the hard part, and it is solved the way this effort predicted.** It said any
|
|
mesh-native supervisor inherits the problem *unless it sits outside the mesh's own process
|
|
tree*. [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) puts
|
|
the launcher there: it supervises the host as a child and shares no code with it, so a host that
|
|
cannot start is still recovered.
|
|
|
|
**It is not all-or-nothing, as this effort insisted.** A human agent acts through a shell and a
|
|
desktop, and those are on the host. So is the host itself — the one thing an init starts.
|
|
|
|
## What is not closed
|
|
|
|
**The automatic node rescue the documentation describes does not exist.** No unit declares
|
|
`OnFailure=`, and nothing calls the rescue script on a timer. That is a documented behaviour
|
|
which never happens, and it outlives this effort — filed as
|
|
[`04-ISSUES/008`](../../04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md)
|
|
rather than closed with it.
|