Approve 0057-0059, with four corrections from review

Not approved as drafted -- four things came out of checking them against each
other, and one was a bug that would have broken every upgrade.

The bug: 0059 specified Restart=on-failure while 0057 has the host restart onto
a new binary by exiting CLEANLY. on-failure does not restart a process that
exited zero, so every upgraded node would have been left stopped, having
successfully upgraded. Found by reading the two records against each other
rather than by either alone. Now Restart=always in all three places that
mention it.

The host cannot run in a container, and the reason is decisive rather than
stylistic: step 0 of the substrate bootstrap installs the container runtime, so
a host inside a container would need the thing it exists to install. It would
also break 0041 -- copy it onto a machine and run it stops being true when the
machine must already have a runtime. Everything above tier 0 is a container;
the host is not. That split is the tier boundary, not an inconsistency.

systemd is named rather than abstracted. An init is not a dependency in 0041's
sense: 0041 is about what must be installed before the host works, and an init
is not installed, it is what the machine already is. The unit file is the only
systemd-specific artefact and it belongs to the package, so a machine with a
different supervisor ships a different package.

The mesh is a watchdog, and my first draft was half an answer. Recovery must be
local -- nothing dials a node, and a host that cannot start cannot report. But
detection is the mesh's, and a local supervisor structurally cannot do it: it
sees one process failing and cannot tell a broken machine from a broken
release. Only something watching every node can, and that distinction decides
whether the response is "fix this machine" or "stop shipping this version". So
a host rollout is staged -- a few nodes, wait for heartbeats, continue or stop
on silence. Local rollback still needed, because the canary nodes break and
because a node offline during the rollout gets the declaration later with no
batch around it.

The first declaration is the overlay and nothing else. Forced, because a node's
address and peers are assigned rather than chosen. But also the way back in: a
node reachable over the overlay can be fixed by hand if a later declaration
breaks it, and a large first declaration risks a node that is broken and
unreachable at once.

Also stated plainly, because it reads as a contradiction: nodes reach each
other over the overlay and every node consumes from the broker; what 0039
forbids is an inbound CONTROL surface, not reachability.

And in 06: no node holds a credential to any control-plane store, for reads or
writes. Four ADRs already say this separately and none of them said it in one
place. Nodes state over the broker; the owning context writes. With a note that
most high-frequency writes are observability's, not the registry's -- routing
logs into the registry would be the shared-schema mistake arriving through a
door marked performance.
This commit is contained in:
2026-08-27 22:12:12 +02:00
parent 605c9fd441
commit 19997d56c3
6 changed files with 218 additions and 22 deletions
@@ -1,5 +1,5 @@
---
status: proposed
status: accepted
date: 2026-08-27
deciders: jochen
reconstructed: false
@@ -40,6 +40,42 @@ could apply almost nothing, and the almost is where the confusion would live.
reboot to hold it. The command-line entry points remain — they are how a person inspects and
rescues a machine — but the ordinary case is a unit that is always up.
### It cannot run in a container, and the reason is the bootstrap
Worth stating because everything else the mesh runs *is* a container, which makes the host look
like an exception somebody forgot to fix.
**Step 0 of the substrate bootstrap is installing the container runtime**
([`07-the-substrate.md`](../03-DESIGN/01-to-be/07-the-substrate.md)). A host that ran inside a
container could not perform it — it would need the thing it is there to install. On a machine
with no runtime, nothing would ever start.
That is not the only reason, but it is the sufficient one:
- **It would break [ADR 0041](0041-the-host-depends-on-nothing.md).** *Copy it onto a machine and
run it* stops being true when the machine must already have a container runtime.
- **The isolation would be fiction.** To write `/etc`, install packages, manage units and run
containers, it would need the host's mount, PID and network namespaces plus the runtime's own
socket. A container with all of those is a process with extra steps.
**So the host is a plain process on the machine, and everything above tier 0 is a container.**
That split is the tier boundary made concrete rather than an inconsistency.
### What it needs from an init, and why that is not a dependency
The host needs four things from whatever supervises it: start at boot, restart when it exits,
give up after repeated failures, and run something else when it gives up.
**Every machine the mesh targets already has systemd**, and the host already treats the service
manager as a detected capability rather than an assumption. This is not a dependency in
[ADR 0041](0041-the-host-depends-on-nothing.md)'s sense — 0041 is about what must be *installed
before the host works*, and an init is not installed, it is what the machine already is.
**Abstracting over init systems is not done**, because there is no second one to abstract over.
The unit file is the only systemd-specific artefact, it belongs to the package rather than the
binary, and a machine with a different supervisor would ship a different package — which is
where that difference belongs.
### It never manages its own unit
**The host's own service file is not a resource the host applies.** The temptation is obvious —
@@ -1,5 +1,5 @@
---
status: proposed
status: accepted
date: 2026-08-27
deciders: jochen
reconstructed: false
@@ -1,12 +1,12 @@
---
status: proposed
status: accepted
date: 2026-08-27
deciders: jochen
reconstructed: false
extends: 0057-the-host-is-a-root-service-installed-as-a-package.md
---
# 59. A host that cannot start is rolled back by the supervisor
# 59. Two watchdogs: the mesh stages the rollout, the supervisor recovers the node
## Context
@@ -18,21 +18,31 @@ It was left because automatic recovery looked like *the host judging its own hea
the self-reference the rest of that document avoids.
**That objection does not survive being asked properly.** A keepalive is not the host judging
itself — it is something else judging the host. The question is only *what*, and that has one
answer.
itself — it is something else judging the host.
### Why the watchdog must be local
### There are two watchdogs, and they cannot do each other's job
The obvious candidate is the control plane, and it cannot be:
A first draft of this record concluded the watchdog must be local, and stopped there. That was
half an answer: it is true that recovery must be local, and false that the mesh has no part.
| | can see | can act |
|---|---|---|
| **the service manager**, on the node | that *this* process keeps dying | **yes** — restart it, replace it |
| **the mesh** | that *eleven of twelve nodes* went quiet after one declaration | **no** — nothing dials a node |
**Recovery must be local**, and that half stands:
- **Nothing dials a node.** [ADR 0039](0039-the-link-is-the-security-boundary.md) makes the link
outbound and node-initiated, and a node has no listening control surface. The mesh has no way
to reach in and act.
outbound and node-initiated, with no listening control surface. The mesh has no way to reach
in and act.
- **The failure removes the reporting path.** A host that cannot start cannot link, so the mesh
learns nothing to act on.
learns nothing *from that node* to act on.
So the watchdog is local, and the only local thing that is always present, already depended on,
and is not the host is the **service manager**.
**But detection is the mesh's**, and it is the half a local watchdog structurally cannot do. A
node's supervisor sees one process failing and has no idea whether that is a broken machine or a
broken release. **Only something watching every node can tell those apart** — and telling them
apart is what decides whether the right response is *fix this machine* or *stop shipping this
version immediately*.
### The failure this actually prevents
@@ -51,7 +61,35 @@ nothing else.
## Decision
**The service manager rolls the host back to the last version that started.**
**The mesh stages the rollout and stops when nodes go quiet. The service manager recovers the
node it is on.** Prevention and recovery, and neither substitutes for the other.
### The mesh stages a host rollout
A host version does not reach every node at once. The delivery context updates a few nodes'
declarations, **waits for those nodes to heartbeat on the new version**, and only then continues.
```
update 2 nodes ─► heard from both, running the new version ─► continue
└► silence past the window ─► STOP. Report.
```
**Silence is the signal, and it is available because of the heartbeat**
([ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)). A node that upgraded
and cannot start stops reporting; that is indistinguishable from a switched-off machine *for one
node*, and completely distinguishable across a batch that was all told the same thing at the
same time.
**This is what keeps a bad release from becoming a fleet outage.** Local rollback repairs a node
after the fact; staging means most nodes never receive the bad version at all. A stopped rollout
is two broken nodes and a report, rather than every node quiet at once.
**It does not replace local recovery**, for two reasons. The canary nodes still break, and
somebody has to be able to fix them. And a node that was offline during the staged rollout gets
the declaration when it reconnects, with no batch around it and nothing watching — so it must be
able to recover alone.
### On the node: the service manager rolls back
Four parts, and each one is chosen so it works when the host does not:
@@ -78,7 +116,7 @@ import it.
```ini
[Service]
Restart=on-failure
Restart=always
StartLimitBurst=3
StartLimitIntervalSec=120
@@ -86,10 +124,27 @@ StartLimitIntervalSec=120
OnFailure=nox-mesh-host-rollback.service
```
Three failures in two minutes is a binary that does not work, not a transient. The supervisor
stops trying and runs the rollback unit, which downgrades and starts the host again.
**`Restart=always`, not `on-failure`**, and the difference is load-bearing rather than a
preference. [ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md) has the host
restart onto a new binary by **exiting cleanly** — and `on-failure` does not restart a process
that exited zero. An earlier draft of this record specified `on-failure` and would have left
every upgraded node stopped, having successfully upgraded. Caught by reading the two records
against each other rather than by either alone.
### 4 — It rolls back once
Three failures in two minutes is a binary that does not work, not a transient. The supervisor
stops trying, the unit enters a failed state, and `OnFailure` runs the rollback unit — which
downgrades and starts the host again.
### 4 — With nothing to roll back to, it does not try
A machine whose host has *never* completed a reconcile has no `known-good`. The rollback unit
finds nothing, does nothing, and says so.
That is the right outcome: there is no previous version, so the node was never working, and the
failure belongs to the installation rather than to an upgrade. Attempting a rollback here would
mean guessing at a version, which is how a recovery mechanism becomes a second fault.
### 5 — It rolls back once
The rollback unit records that it fired. If the rolled-back version *also* fails to start, it
does **not** fire again — the node stops, loudly, in a failed state.
+11 -4
View File
@@ -12,6 +12,8 @@ decisions:
- 02-DECISIONS/0008-a-failed-step-fails-the-job.md
- 02-DECISIONS/0041-the-host-depends-on-nothing.md
- 02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md
- 02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md
- 02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md
---
# The node host
@@ -115,10 +117,7 @@ the link; never asked downward.
## The process
> **Proposed, not yet decided** —
> [ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md) is
> awaiting review. This section is written against it and moves into the document's `decisions:`
> when the record is accepted.
[ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md).
**One `root` service on every node, plus command-line entry points for a person.** A machine
without one is not a node ([ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)) —
@@ -167,12 +166,20 @@ Type=notify
ExecStart=/usr/bin/nox-mesh-host run
Restart=always
RestartSec=5s
StartLimitBurst=3
StartLimitIntervalSec=120
StateDirectory=mesh-host
[Install]
WantedBy=multi-user.target
```
**`Restart=always` and not `on-failure`**: the host restarts onto a new binary by exiting
*cleanly*, and `on-failure` would not restart it — an upgraded node would be left stopped,
having successfully upgraded. The start limit is what
[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md) uses to decide
a binary is broken rather than unlucky.
**The package owns this file. The host never does.** It manages `service` resources, and its own
unit is a service — the temptation is obvious and it ends with a host stopping itself half way
through an apply, leaving a machine with nothing running to fix it. A declaration naming the
@@ -76,6 +76,59 @@ its own store, and they integrate through the record rather than by reading one
constraint is load-bearing: a single surface can compose them only while there is one interface
in front of them.
## Nothing outside a context touches its store
The question this answers: **can a node write to the registry database?** No — and not "only
through one node", which is the weaker arrangement it might be mistaken for.
> **No node holds a credential to any control-plane store, for writing or for reading.**
That is not a new rule here; it is four already taken, and it is worth seeing them together
because each one alone reads like a detail:
| | |
|---|---|
| [ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md) | the host never queries the mesh database |
| [ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md) | a node holds its own identity **and nothing else** — the shared database credential every node carries today is the exposure this exists to remove |
| [ADR 0045](../../02-DECISIONS/0045-a-context-owns-its-store.md) | a context is granted only what it **exclusively** owns: no shared writes, no read-only roles |
| [ADR 0056](../../02-DECISIONS/0056-the-authority-is-the-control-plane-not-a-database.md) | there is no single mesh database, and nothing reads one |
### So how does anything get in
**Over the broker, as a message; the owning context writes.**
```
node ──event/report──► broker ──► the context that owns that data ──► its own store
```
A node reports what it applied, what it holds, and that it is alive. It **states**; it does not
**write**. The difference is the whole security boundary: a node that can write cannot be
prevented from writing anything, and a node that can only state has its blast radius bounded by
what the message vocabulary can say.
Reads work the same way in reverse — a node is *told*, in declarations. It never asks.
### On volume, which is the real worry underneath
**Most high-frequency writes are not registry writes, and that is the first thing to check
before designing for throughput.** The registry is `inventory`'s store: nodes, modules,
assignments, versions. Those change when somebody changes something.
**Logs, metrics and health checks belong to `observability`**, which owns a different store
([ADR 0045](../../02-DECISIONS/0045-a-context-owns-its-store.md)). Sending them to the registry
would be exactly the shared-schema mistake 0045 exists to stop, arriving through the back door
marked *performance*.
That leaves one genuine funnel: every context's writes go through the process that owns it, and
[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md) says there is one of it.
For a mesh of a handful of machines this is not a scaling problem, and **it should not be solved
by giving nodes database credentials** — that trades a bounded problem for an unbounded one. If
it ever binds, the answers are at the consumer: batch, apply backpressure, or move the highest
volume stream out of a relational store entirely.
**What observability actually stores its data in is not decided**, and it is the one place where
volume genuinely argues against a relational store.
## What it is not
- **Not the thing that changes machines.** It decides; the host applies. It never reaches into a
@@ -10,6 +10,9 @@ decisions:
- 02-DECISIONS/0039-the-link-is-the-security-boundary.md
- 02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md
- 02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md
- 02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md
- 02-DECISIONS/0058-delivery-ends-in-a-declaration.md
- 02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md
---
# The node lifecycle
@@ -97,6 +100,48 @@ derived centrally and pushed down
([ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md),
[`08-connectivity.md`](08-connectivity.md)).
### The first declaration is the overlay, and nothing else
**The mesh makes a node reachable before it makes it useful.** Step 5 is not one declaration
carrying everything the node will ever run. It is two, in order:
```
first the overlay — this node's address, its keys, its peers, its names
then everything else — packages, containers, services, files
```
Three reasons, and the third is the one that matters when something goes wrong:
- **It is forced.** A node cannot join the overlay before contacting the mesh, because its
address and peer set are *assigned* — it generates a keypair, publishes the public half, and
receives the rest ([`08-connectivity.md`](08-connectivity.md)). So the overlay is the first
thing the mesh can give it, and it should be.
- **It is what [ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md) already
says:** *a joining node does the minimum to be reachable, and nothing else.*
- **It is the way back in.** Once the overlay is up, the node is reachable over it — by SSH, by
anything. If a later declaration breaks the machine, there is a route to it that does not
depend on the mesh's control path working. **Sending a large first declaration risks a node
that is broken and unreachable at the same time**, and those two failures are much worse
together than separately.
### Reachable is not the same as having a control surface
Worth stating plainly, because the two rules read as a contradiction and are not.
| | |
|---|---|
| **every node reaches every other node** | over the overlay — SSH, services, ordinary traffic. This is the point of having one |
| **every node consumes from the broker** | its own queue, over its own outbound connection ([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md)) |
| **nothing dials a node to control it** | the host has no inbound control surface ([ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md)) |
**[ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md) is about the control
channel, not about network reachability.** What it forbids is a listening thing that accepts
instructions and changes the machine. A node being reachable on the overlay — the whole purpose
of the overlay — is untouched by it, and so is a person opening a shell on it.
The distinction is *who can tell this machine what to be*: only the control plane, only over the
link the node opened, only in declarations of known shape.
---
## hosted → enrolled: the first node