Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24

Merged
jschoubben merged 177 commits from reconcile-init-into-main into main 2026-09-05 10:27:11 +00:00
6 changed files with 218 additions and 22 deletions
Showing only changes of commit 19997d56c3 - Show all commits
@@ -1,5 +1,5 @@
---
status: proposed
status: accepted
date: 2026-08-27
deciders: jochen
reconstructed: false
@@ -40,6 +40,42 @@ could apply almost nothing, and the almost is where the confusion would live.
reboot to hold it. The command-line entry points remain — they are how a person inspects and
rescues a machine — but the ordinary case is a unit that is always up.
### It cannot run in a container, and the reason is the bootstrap
Worth stating because everything else the mesh runs *is* a container, which makes the host look
like an exception somebody forgot to fix.
**Step 0 of the substrate bootstrap is installing the container runtime**
([`07-the-substrate.md`](../03-DESIGN/01-to-be/07-the-substrate.md)). A host that ran inside a
container could not perform it — it would need the thing it is there to install. On a machine
with no runtime, nothing would ever start.
That is not the only reason, but it is the sufficient one:
- **It would break [ADR 0041](0041-the-host-depends-on-nothing.md).** *Copy it onto a machine and
run it* stops being true when the machine must already have a container runtime.
- **The isolation would be fiction.** To write `/etc`, install packages, manage units and run
containers, it would need the host's mount, PID and network namespaces plus the runtime's own
socket. A container with all of those is a process with extra steps.
**So the host is a plain process on the machine, and everything above tier 0 is a container.**
That split is the tier boundary made concrete rather than an inconsistency.
### What it needs from an init, and why that is not a dependency
The host needs four things from whatever supervises it: start at boot, restart when it exits,
give up after repeated failures, and run something else when it gives up.
**Every machine the mesh targets already has systemd**, and the host already treats the service
manager as a detected capability rather than an assumption. This is not a dependency in
[ADR 0041](0041-the-host-depends-on-nothing.md)'s sense — 0041 is about what must be *installed
before the host works*, and an init is not installed, it is what the machine already is.
**Abstracting over init systems is not done**, because there is no second one to abstract over.
The unit file is the only systemd-specific artefact, it belongs to the package rather than the
binary, and a machine with a different supervisor would ship a different package — which is
where that difference belongs.
### It never manages its own unit
**The host's own service file is not a resource the host applies.** The temptation is obvious —
@@ -1,5 +1,5 @@
---
status: proposed
status: accepted
date: 2026-08-27
deciders: jochen
reconstructed: false
@@ -1,12 +1,12 @@
---
status: proposed
status: accepted
date: 2026-08-27
deciders: jochen
reconstructed: false
extends: 0057-the-host-is-a-root-service-installed-as-a-package.md
---
# 59. A host that cannot start is rolled back by the supervisor
# 59. Two watchdogs: the mesh stages the rollout, the supervisor recovers the node
## Context
@@ -18,21 +18,31 @@ It was left because automatic recovery looked like *the host judging its own hea
the self-reference the rest of that document avoids.
**That objection does not survive being asked properly.** A keepalive is not the host judging
itself — it is something else judging the host. The question is only *what*, and that has one
answer.
itself — it is something else judging the host.
### Why the watchdog must be local
### There are two watchdogs, and they cannot do each other's job
The obvious candidate is the control plane, and it cannot be:
A first draft of this record concluded the watchdog must be local, and stopped there. That was
half an answer: it is true that recovery must be local, and false that the mesh has no part.
| | can see | can act |
|---|---|---|
| **the service manager**, on the node | that *this* process keeps dying | **yes** — restart it, replace it |
| **the mesh** | that *eleven of twelve nodes* went quiet after one declaration | **no** — nothing dials a node |
**Recovery must be local**, and that half stands:
- **Nothing dials a node.** [ADR 0039](0039-the-link-is-the-security-boundary.md) makes the link
outbound and node-initiated, and a node has no listening control surface. The mesh has no way
to reach in and act.
outbound and node-initiated, with no listening control surface. The mesh has no way to reach
in and act.
- **The failure removes the reporting path.** A host that cannot start cannot link, so the mesh
learns nothing to act on.
learns nothing *from that node* to act on.
So the watchdog is local, and the only local thing that is always present, already depended on,
and is not the host is the **service manager**.
**But detection is the mesh's**, and it is the half a local watchdog structurally cannot do. A
node's supervisor sees one process failing and has no idea whether that is a broken machine or a
broken release. **Only something watching every node can tell those apart** — and telling them
apart is what decides whether the right response is *fix this machine* or *stop shipping this
version immediately*.
### The failure this actually prevents
@@ -51,7 +61,35 @@ nothing else.
## Decision
**The service manager rolls the host back to the last version that started.**
**The mesh stages the rollout and stops when nodes go quiet. The service manager recovers the
node it is on.** Prevention and recovery, and neither substitutes for the other.
### The mesh stages a host rollout
A host version does not reach every node at once. The delivery context updates a few nodes'
declarations, **waits for those nodes to heartbeat on the new version**, and only then continues.
```
update 2 nodes ─► heard from both, running the new version ─► continue
└► silence past the window ─► STOP. Report.
```
**Silence is the signal, and it is available because of the heartbeat**
([ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)). A node that upgraded
and cannot start stops reporting; that is indistinguishable from a switched-off machine *for one
node*, and completely distinguishable across a batch that was all told the same thing at the
same time.
**This is what keeps a bad release from becoming a fleet outage.** Local rollback repairs a node
after the fact; staging means most nodes never receive the bad version at all. A stopped rollout
is two broken nodes and a report, rather than every node quiet at once.
**It does not replace local recovery**, for two reasons. The canary nodes still break, and
somebody has to be able to fix them. And a node that was offline during the staged rollout gets
the declaration when it reconnects, with no batch around it and nothing watching — so it must be
able to recover alone.
### On the node: the service manager rolls back
Four parts, and each one is chosen so it works when the host does not:
@@ -78,7 +116,7 @@ import it.
```ini
[Service]
Restart=on-failure
Restart=always
StartLimitBurst=3
StartLimitIntervalSec=120
@@ -86,10 +124,27 @@ StartLimitIntervalSec=120
OnFailure=nox-mesh-host-rollback.service
```
Three failures in two minutes is a binary that does not work, not a transient. The supervisor
stops trying and runs the rollback unit, which downgrades and starts the host again.
**`Restart=always`, not `on-failure`**, and the difference is load-bearing rather than a
preference. [ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md) has the host
restart onto a new binary by **exiting cleanly** — and `on-failure` does not restart a process
that exited zero. An earlier draft of this record specified `on-failure` and would have left
every upgraded node stopped, having successfully upgraded. Caught by reading the two records
against each other rather than by either alone.
### 4 — It rolls back once
Three failures in two minutes is a binary that does not work, not a transient. The supervisor
stops trying, the unit enters a failed state, and `OnFailure` runs the rollback unit — which
downgrades and starts the host again.
### 4 — With nothing to roll back to, it does not try
A machine whose host has *never* completed a reconcile has no `known-good`. The rollback unit
finds nothing, does nothing, and says so.
That is the right outcome: there is no previous version, so the node was never working, and the
failure belongs to the installation rather than to an upgrade. Attempting a rollback here would
mean guessing at a version, which is how a recovery mechanism becomes a second fault.
### 5 — It rolls back once
The rollback unit records that it fired. If the rolled-back version *also* fails to start, it
does **not** fire again — the node stops, loudly, in a failed state.
+11 -4
View File
@@ -12,6 +12,8 @@ decisions:
- 02-DECISIONS/0008-a-failed-step-fails-the-job.md
- 02-DECISIONS/0041-the-host-depends-on-nothing.md
- 02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md
- 02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md
- 02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md
---
# The node host
@@ -115,10 +117,7 @@ the link; never asked downward.
## The process
> **Proposed, not yet decided** —
> [ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md) is
> awaiting review. This section is written against it and moves into the document's `decisions:`
> when the record is accepted.
[ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md).
**One `root` service on every node, plus command-line entry points for a person.** A machine
without one is not a node ([ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)) —
@@ -167,12 +166,20 @@ Type=notify
ExecStart=/usr/bin/nox-mesh-host run
Restart=always
RestartSec=5s
StartLimitBurst=3
StartLimitIntervalSec=120
StateDirectory=mesh-host
[Install]
WantedBy=multi-user.target
```
**`Restart=always` and not `on-failure`**: the host restarts onto a new binary by exiting
*cleanly*, and `on-failure` would not restart it — an upgraded node would be left stopped,
having successfully upgraded. The start limit is what
[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md) uses to decide
a binary is broken rather than unlucky.
**The package owns this file. The host never does.** It manages `service` resources, and its own
unit is a service — the temptation is obvious and it ends with a host stopping itself half way
through an apply, leaving a machine with nothing running to fix it. A declaration naming the
@@ -76,6 +76,59 @@ its own store, and they integrate through the record rather than by reading one
constraint is load-bearing: a single surface can compose them only while there is one interface
in front of them.
## Nothing outside a context touches its store
The question this answers: **can a node write to the registry database?** No — and not "only
through one node", which is the weaker arrangement it might be mistaken for.
> **No node holds a credential to any control-plane store, for writing or for reading.**
That is not a new rule here; it is four already taken, and it is worth seeing them together
because each one alone reads like a detail:
| | |
|---|---|
| [ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md) | the host never queries the mesh database |
| [ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md) | a node holds its own identity **and nothing else** — the shared database credential every node carries today is the exposure this exists to remove |
| [ADR 0045](../../02-DECISIONS/0045-a-context-owns-its-store.md) | a context is granted only what it **exclusively** owns: no shared writes, no read-only roles |
| [ADR 0056](../../02-DECISIONS/0056-the-authority-is-the-control-plane-not-a-database.md) | there is no single mesh database, and nothing reads one |
### So how does anything get in
**Over the broker, as a message; the owning context writes.**
```
node ──event/report──► broker ──► the context that owns that data ──► its own store
```
A node reports what it applied, what it holds, and that it is alive. It **states**; it does not
**write**. The difference is the whole security boundary: a node that can write cannot be
prevented from writing anything, and a node that can only state has its blast radius bounded by
what the message vocabulary can say.
Reads work the same way in reverse — a node is *told*, in declarations. It never asks.
### On volume, which is the real worry underneath
**Most high-frequency writes are not registry writes, and that is the first thing to check
before designing for throughput.** The registry is `inventory`'s store: nodes, modules,
assignments, versions. Those change when somebody changes something.
**Logs, metrics and health checks belong to `observability`**, which owns a different store
([ADR 0045](../../02-DECISIONS/0045-a-context-owns-its-store.md)). Sending them to the registry
would be exactly the shared-schema mistake 0045 exists to stop, arriving through the back door
marked *performance*.
That leaves one genuine funnel: every context's writes go through the process that owns it, and
[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md) says there is one of it.
For a mesh of a handful of machines this is not a scaling problem, and **it should not be solved
by giving nodes database credentials** — that trades a bounded problem for an unbounded one. If
it ever binds, the answers are at the consumer: batch, apply backpressure, or move the highest
volume stream out of a relational store entirely.
**What observability actually stores its data in is not decided**, and it is the one place where
volume genuinely argues against a relational store.
## What it is not
- **Not the thing that changes machines.** It decides; the host applies. It never reaches into a
@@ -10,6 +10,9 @@ decisions:
- 02-DECISIONS/0039-the-link-is-the-security-boundary.md
- 02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md
- 02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md
- 02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md
- 02-DECISIONS/0058-delivery-ends-in-a-declaration.md
- 02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md
---
# The node lifecycle
@@ -97,6 +100,48 @@ derived centrally and pushed down
([ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md),
[`08-connectivity.md`](08-connectivity.md)).
### The first declaration is the overlay, and nothing else
**The mesh makes a node reachable before it makes it useful.** Step 5 is not one declaration
carrying everything the node will ever run. It is two, in order:
```
first the overlay — this node's address, its keys, its peers, its names
then everything else — packages, containers, services, files
```
Three reasons, and the third is the one that matters when something goes wrong:
- **It is forced.** A node cannot join the overlay before contacting the mesh, because its
address and peer set are *assigned* — it generates a keypair, publishes the public half, and
receives the rest ([`08-connectivity.md`](08-connectivity.md)). So the overlay is the first
thing the mesh can give it, and it should be.
- **It is what [ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md) already
says:** *a joining node does the minimum to be reachable, and nothing else.*
- **It is the way back in.** Once the overlay is up, the node is reachable over it — by SSH, by
anything. If a later declaration breaks the machine, there is a route to it that does not
depend on the mesh's control path working. **Sending a large first declaration risks a node
that is broken and unreachable at the same time**, and those two failures are much worse
together than separately.
### Reachable is not the same as having a control surface
Worth stating plainly, because the two rules read as a contradiction and are not.
| | |
|---|---|
| **every node reaches every other node** | over the overlay — SSH, services, ordinary traffic. This is the point of having one |
| **every node consumes from the broker** | its own queue, over its own outbound connection ([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md)) |
| **nothing dials a node to control it** | the host has no inbound control surface ([ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md)) |
**[ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md) is about the control
channel, not about network reachability.** What it forbids is a listening thing that accepts
instructions and changes the machine. A node being reachable on the overlay — the whole purpose
of the overlay — is untouched by it, and so is a person opening a shell on it.
The distinction is *who can tell this machine what to be*: only the control plane, only over the
link the node opened, only in declarations of known shape.
---
## hosted → enrolled: the first node