Files
hq/03-DESIGN/00-as-is/05-runtime-and-installation.md
T
jschoubben 77f3a4cea7 Consolidate: 65 decision records to 23
Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.

  the node host          8 -> 1    applies not decides, depends on nothing,
                                   per operating system, root service, the
                                   launcher, episodic, what a declaration is,
                                   actions from the bundle only
  a node and how it joins 4 -> 1   what a node is, joining, the link as
                                   security boundary, the enrolment token
  modules and the graph   7 -> 1   everything is a module, no domain modules,
                                   three edges, provisioning, the core library
  substrate and control   6 -> 1   the test, seven contexts, one control plane,
    plane                          the authority is not a database, the named
                                   products, the pinned bundle
  connectivity            3 -> 1   a route is a grant, reachability declared,
                                   filter rules
  delivery                5 -> 1   reconciliation not a pipeline, artifacts,
                                   the three silos, a failed step, the verdict
  the lab                 5 -> 1   (earlier)
  how this repository     10 -> 1  (earlier)
    works

Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.

The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.

The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
2026-08-28 20:03:24 +02:00

108 lines
5.0 KiB
Markdown

---
layer: as-is
status: implemented
code: [hal]
updated: 2026-08-23
decisions:
- 02-DECISIONS/0044-modules-and-the-graph.md
- 02-DECISIONS/0018-the-mesh-creates-no-symlinks.md
---
# The node runtime, and how a node comes into being
Every node runs the same runtime. What differs is its assignment.
## Two modes, coexisting
The runtime runs in two modes at once on any node that needs both.
**Daemon** mode is a headless consumer: it connects to the broker, consumes the node's request
queue, routes each request to a local capability, and emits the node's lifecycle events. This
is what makes a node a participant — it is reachable whether or not anyone is logged in.
**Interactive** mode exposes the node's capabilities to a session on that machine over a local
protocol. Its capability surface is larger, because it includes stand-ins for every capability
discovered on peers.
The two are the same code with the same catalogue. A capability is written once and is
available to both.
## Anatomy naming, and what it obscures
The runtime's components are named after brain anatomy: an entry point that bootstraps, a
headless listener, an interactive surface, an installer daemon, and a provisioner daemon.
This is the mesh's most-cited naming problem and it belongs in the as-is layer because it is
what a reader will actually encounter. The names are evocative and describe nothing: the most
suggestive word in the system names the node runtime, and the component whose manifest says
"mesh messaging" is documented elsewhere as the interactive runtime. Anatomy makes attractive
names and poor boundaries.
[ADR 0015](../../02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md) replaces this with names
taken from what each part owns. Until then, this is the vocabulary in the code.
## Starting a module
For each assigned module the runtime, at startup:
1. Reads the manifest, if there is one — a flag module has nothing to read.
2. Skips a module's capabilities if a variable they require is unset. This is quiet by design
and hard to distinguish from a module that has no capabilities.
3. Loads the capabilities the module carries.
4. For a service module, ensures it is installed and running.
Installation is idempotent and does the unglamorous work: ensure the runtime directory,
reconcile the link from the catalogue's definition into it, generate the environment, run any
outstanding local migrations, create data directories with the right ownership, then start the
service under supervision.
**The installer is the only thing that creates a link** ([ADR
0011](../../02-DECISIONS/0018-the-mesh-creates-no-symlinks.md)). It reconciles rather than assumes: a
missing link is created, a stale one repointed, and a real file found where a link belongs is
adopted into the node's override area and replaced. Nothing else — not a hook, not a fix, not a
person debugging — creates one.
That is the as-is. The intent is to remove linking altogether and derive a real file instead,
which the reconciliation machinery already makes possible
([ADR 0018](../../02-DECISIONS/0018-the-mesh-creates-no-symlinks.md), proposed). What is described above
is what runs today.
## Supervision
Services run under the host's init system via a templated unit, one instance per module. It is
a thin layer: the unit starts and stops a container group.
Whether the mesh keeps this, drops the per-module layer, containerises the daemons, or writes
its own supervisor is **open** — costed in research effort 003 and deliberately undecided. It
no longer gates anything.
Two failures worth knowing about, both in the shape of "the change did not apply". A
per-instance copy of the unit template shadows the template, so edits to the template do
nothing. And a session-scoped one-shot job loses the environment it was given, because the
import is one-time and not persisted.
## How a node comes into being
Three bootstrap scripts, and which one runs depends on the situation:
- **First node.** Nothing exists yet, so the script stands up the database the rest of the
mesh reads from, publishes the catalogue, and starts the mesh. This resolves the
circularity of a mesh whose source of truth is itself a provisioned module.
- **Joining.** The node registers, takes its assignment from the database, and syncs.
- **Rescue.** A node that cannot reach the mesh is brought back far enough to.
Bringing a node into being is therefore a **database operation with a script attached**, not a
checkout. There is no per-node content in the repository to copy.
Adoption of a pre-existing machine's configuration was the original path and is now a legacy
one, explicitly out of scope for the lab
([`01-to-be/01-end-to-end-testing.md`](../01-to-be/01-end-to-end-testing.md)).
## Node identity
Each node carries an identity text in its own record, which the daemon writes onto the node at
startup so that a session on that machine knows which node it is on and how that node
presents itself. It carries identity only; shared rules live separately.
Like everything else derived onto a node, it is generated and not edited there.