Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24
@@ -1,5 +1,5 @@
|
|||||||
---
|
---
|
||||||
status: proposed
|
status: accepted
|
||||||
date: 2026-08-27
|
date: 2026-08-27
|
||||||
deciders: jochen
|
deciders: jochen
|
||||||
reconstructed: false
|
reconstructed: false
|
||||||
@@ -40,6 +40,42 @@ could apply almost nothing, and the almost is where the confusion would live.
|
|||||||
reboot to hold it. The command-line entry points remain — they are how a person inspects and
|
reboot to hold it. The command-line entry points remain — they are how a person inspects and
|
||||||
rescues a machine — but the ordinary case is a unit that is always up.
|
rescues a machine — but the ordinary case is a unit that is always up.
|
||||||
|
|
||||||
|
### It cannot run in a container, and the reason is the bootstrap
|
||||||
|
|
||||||
|
Worth stating because everything else the mesh runs *is* a container, which makes the host look
|
||||||
|
like an exception somebody forgot to fix.
|
||||||
|
|
||||||
|
**Step 0 of the substrate bootstrap is installing the container runtime**
|
||||||
|
([`07-the-substrate.md`](../03-DESIGN/01-to-be/07-the-substrate.md)). A host that ran inside a
|
||||||
|
container could not perform it — it would need the thing it is there to install. On a machine
|
||||||
|
with no runtime, nothing would ever start.
|
||||||
|
|
||||||
|
That is not the only reason, but it is the sufficient one:
|
||||||
|
|
||||||
|
- **It would break [ADR 0041](0041-the-host-depends-on-nothing.md).** *Copy it onto a machine and
|
||||||
|
run it* stops being true when the machine must already have a container runtime.
|
||||||
|
- **The isolation would be fiction.** To write `/etc`, install packages, manage units and run
|
||||||
|
containers, it would need the host's mount, PID and network namespaces plus the runtime's own
|
||||||
|
socket. A container with all of those is a process with extra steps.
|
||||||
|
|
||||||
|
**So the host is a plain process on the machine, and everything above tier 0 is a container.**
|
||||||
|
That split is the tier boundary made concrete rather than an inconsistency.
|
||||||
|
|
||||||
|
### What it needs from an init, and why that is not a dependency
|
||||||
|
|
||||||
|
The host needs four things from whatever supervises it: start at boot, restart when it exits,
|
||||||
|
give up after repeated failures, and run something else when it gives up.
|
||||||
|
|
||||||
|
**Every machine the mesh targets already has systemd**, and the host already treats the service
|
||||||
|
manager as a detected capability rather than an assumption. This is not a dependency in
|
||||||
|
[ADR 0041](0041-the-host-depends-on-nothing.md)'s sense — 0041 is about what must be *installed
|
||||||
|
before the host works*, and an init is not installed, it is what the machine already is.
|
||||||
|
|
||||||
|
**Abstracting over init systems is not done**, because there is no second one to abstract over.
|
||||||
|
The unit file is the only systemd-specific artefact, it belongs to the package rather than the
|
||||||
|
binary, and a machine with a different supervisor would ship a different package — which is
|
||||||
|
where that difference belongs.
|
||||||
|
|
||||||
### It never manages its own unit
|
### It never manages its own unit
|
||||||
|
|
||||||
**The host's own service file is not a resource the host applies.** The temptation is obvious —
|
**The host's own service file is not a resource the host applies.** The temptation is obvious —
|
||||||
|
|||||||
@@ -1,5 +1,5 @@
|
|||||||
---
|
---
|
||||||
status: proposed
|
status: accepted
|
||||||
date: 2026-08-27
|
date: 2026-08-27
|
||||||
deciders: jochen
|
deciders: jochen
|
||||||
reconstructed: false
|
reconstructed: false
|
||||||
|
|||||||
@@ -1,12 +1,12 @@
|
|||||||
---
|
---
|
||||||
status: proposed
|
status: accepted
|
||||||
date: 2026-08-27
|
date: 2026-08-27
|
||||||
deciders: jochen
|
deciders: jochen
|
||||||
reconstructed: false
|
reconstructed: false
|
||||||
extends: 0057-the-host-is-a-root-service-installed-as-a-package.md
|
extends: 0057-the-host-is-a-root-service-installed-as-a-package.md
|
||||||
---
|
---
|
||||||
|
|
||||||
# 59. A host that cannot start is rolled back by the supervisor
|
# 59. Two watchdogs: the mesh stages the rollout, the supervisor recovers the node
|
||||||
|
|
||||||
## Context
|
## Context
|
||||||
|
|
||||||
@@ -18,21 +18,31 @@ It was left because automatic recovery looked like *the host judging its own hea
|
|||||||
the self-reference the rest of that document avoids.
|
the self-reference the rest of that document avoids.
|
||||||
|
|
||||||
**That objection does not survive being asked properly.** A keepalive is not the host judging
|
**That objection does not survive being asked properly.** A keepalive is not the host judging
|
||||||
itself — it is something else judging the host. The question is only *what*, and that has one
|
itself — it is something else judging the host.
|
||||||
answer.
|
|
||||||
|
|
||||||
### Why the watchdog must be local
|
### There are two watchdogs, and they cannot do each other's job
|
||||||
|
|
||||||
The obvious candidate is the control plane, and it cannot be:
|
A first draft of this record concluded the watchdog must be local, and stopped there. That was
|
||||||
|
half an answer: it is true that recovery must be local, and false that the mesh has no part.
|
||||||
|
|
||||||
|
| | can see | can act |
|
||||||
|
|---|---|---|
|
||||||
|
| **the service manager**, on the node | that *this* process keeps dying | **yes** — restart it, replace it |
|
||||||
|
| **the mesh** | that *eleven of twelve nodes* went quiet after one declaration | **no** — nothing dials a node |
|
||||||
|
|
||||||
|
**Recovery must be local**, and that half stands:
|
||||||
|
|
||||||
- **Nothing dials a node.** [ADR 0039](0039-the-link-is-the-security-boundary.md) makes the link
|
- **Nothing dials a node.** [ADR 0039](0039-the-link-is-the-security-boundary.md) makes the link
|
||||||
outbound and node-initiated, and a node has no listening control surface. The mesh has no way
|
outbound and node-initiated, with no listening control surface. The mesh has no way to reach
|
||||||
to reach in and act.
|
in and act.
|
||||||
- **The failure removes the reporting path.** A host that cannot start cannot link, so the mesh
|
- **The failure removes the reporting path.** A host that cannot start cannot link, so the mesh
|
||||||
learns nothing to act on.
|
learns nothing *from that node* to act on.
|
||||||
|
|
||||||
So the watchdog is local, and the only local thing that is always present, already depended on,
|
**But detection is the mesh's**, and it is the half a local watchdog structurally cannot do. A
|
||||||
and is not the host is the **service manager**.
|
node's supervisor sees one process failing and has no idea whether that is a broken machine or a
|
||||||
|
broken release. **Only something watching every node can tell those apart** — and telling them
|
||||||
|
apart is what decides whether the right response is *fix this machine* or *stop shipping this
|
||||||
|
version immediately*.
|
||||||
|
|
||||||
### The failure this actually prevents
|
### The failure this actually prevents
|
||||||
|
|
||||||
@@ -51,7 +61,35 @@ nothing else.
|
|||||||
|
|
||||||
## Decision
|
## Decision
|
||||||
|
|
||||||
**The service manager rolls the host back to the last version that started.**
|
**The mesh stages the rollout and stops when nodes go quiet. The service manager recovers the
|
||||||
|
node it is on.** Prevention and recovery, and neither substitutes for the other.
|
||||||
|
|
||||||
|
### The mesh stages a host rollout
|
||||||
|
|
||||||
|
A host version does not reach every node at once. The delivery context updates a few nodes'
|
||||||
|
declarations, **waits for those nodes to heartbeat on the new version**, and only then continues.
|
||||||
|
|
||||||
|
```
|
||||||
|
update 2 nodes ─► heard from both, running the new version ─► continue
|
||||||
|
└► silence past the window ─► STOP. Report.
|
||||||
|
```
|
||||||
|
|
||||||
|
**Silence is the signal, and it is available because of the heartbeat**
|
||||||
|
([ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)). A node that upgraded
|
||||||
|
and cannot start stops reporting; that is indistinguishable from a switched-off machine *for one
|
||||||
|
node*, and completely distinguishable across a batch that was all told the same thing at the
|
||||||
|
same time.
|
||||||
|
|
||||||
|
**This is what keeps a bad release from becoming a fleet outage.** Local rollback repairs a node
|
||||||
|
after the fact; staging means most nodes never receive the bad version at all. A stopped rollout
|
||||||
|
is two broken nodes and a report, rather than every node quiet at once.
|
||||||
|
|
||||||
|
**It does not replace local recovery**, for two reasons. The canary nodes still break, and
|
||||||
|
somebody has to be able to fix them. And a node that was offline during the staged rollout gets
|
||||||
|
the declaration when it reconnects, with no batch around it and nothing watching — so it must be
|
||||||
|
able to recover alone.
|
||||||
|
|
||||||
|
### On the node: the service manager rolls back
|
||||||
|
|
||||||
Four parts, and each one is chosen so it works when the host does not:
|
Four parts, and each one is chosen so it works when the host does not:
|
||||||
|
|
||||||
@@ -78,7 +116,7 @@ import it.
|
|||||||
|
|
||||||
```ini
|
```ini
|
||||||
[Service]
|
[Service]
|
||||||
Restart=on-failure
|
Restart=always
|
||||||
StartLimitBurst=3
|
StartLimitBurst=3
|
||||||
StartLimitIntervalSec=120
|
StartLimitIntervalSec=120
|
||||||
|
|
||||||
@@ -86,10 +124,27 @@ StartLimitIntervalSec=120
|
|||||||
OnFailure=nox-mesh-host-rollback.service
|
OnFailure=nox-mesh-host-rollback.service
|
||||||
```
|
```
|
||||||
|
|
||||||
Three failures in two minutes is a binary that does not work, not a transient. The supervisor
|
**`Restart=always`, not `on-failure`**, and the difference is load-bearing rather than a
|
||||||
stops trying and runs the rollback unit, which downgrades and starts the host again.
|
preference. [ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md) has the host
|
||||||
|
restart onto a new binary by **exiting cleanly** — and `on-failure` does not restart a process
|
||||||
|
that exited zero. An earlier draft of this record specified `on-failure` and would have left
|
||||||
|
every upgraded node stopped, having successfully upgraded. Caught by reading the two records
|
||||||
|
against each other rather than by either alone.
|
||||||
|
|
||||||
### 4 — It rolls back once
|
Three failures in two minutes is a binary that does not work, not a transient. The supervisor
|
||||||
|
stops trying, the unit enters a failed state, and `OnFailure` runs the rollback unit — which
|
||||||
|
downgrades and starts the host again.
|
||||||
|
|
||||||
|
### 4 — With nothing to roll back to, it does not try
|
||||||
|
|
||||||
|
A machine whose host has *never* completed a reconcile has no `known-good`. The rollback unit
|
||||||
|
finds nothing, does nothing, and says so.
|
||||||
|
|
||||||
|
That is the right outcome: there is no previous version, so the node was never working, and the
|
||||||
|
failure belongs to the installation rather than to an upgrade. Attempting a rollback here would
|
||||||
|
mean guessing at a version, which is how a recovery mechanism becomes a second fault.
|
||||||
|
|
||||||
|
### 5 — It rolls back once
|
||||||
|
|
||||||
The rollback unit records that it fired. If the rolled-back version *also* fails to start, it
|
The rollback unit records that it fired. If the rolled-back version *also* fails to start, it
|
||||||
does **not** fire again — the node stops, loudly, in a failed state.
|
does **not** fire again — the node stops, loudly, in a failed state.
|
||||||
|
|||||||
@@ -12,6 +12,8 @@ decisions:
|
|||||||
- 02-DECISIONS/0008-a-failed-step-fails-the-job.md
|
- 02-DECISIONS/0008-a-failed-step-fails-the-job.md
|
||||||
- 02-DECISIONS/0041-the-host-depends-on-nothing.md
|
- 02-DECISIONS/0041-the-host-depends-on-nothing.md
|
||||||
- 02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md
|
- 02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md
|
||||||
|
- 02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md
|
||||||
|
- 02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md
|
||||||
---
|
---
|
||||||
|
|
||||||
# The node host
|
# The node host
|
||||||
@@ -115,10 +117,7 @@ the link; never asked downward.
|
|||||||
|
|
||||||
## The process
|
## The process
|
||||||
|
|
||||||
> **Proposed, not yet decided** —
|
[ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md).
|
||||||
> [ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md) is
|
|
||||||
> awaiting review. This section is written against it and moves into the document's `decisions:`
|
|
||||||
> when the record is accepted.
|
|
||||||
|
|
||||||
**One `root` service on every node, plus command-line entry points for a person.** A machine
|
**One `root` service on every node, plus command-line entry points for a person.** A machine
|
||||||
without one is not a node ([ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)) —
|
without one is not a node ([ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)) —
|
||||||
@@ -167,12 +166,20 @@ Type=notify
|
|||||||
ExecStart=/usr/bin/nox-mesh-host run
|
ExecStart=/usr/bin/nox-mesh-host run
|
||||||
Restart=always
|
Restart=always
|
||||||
RestartSec=5s
|
RestartSec=5s
|
||||||
|
StartLimitBurst=3
|
||||||
|
StartLimitIntervalSec=120
|
||||||
StateDirectory=mesh-host
|
StateDirectory=mesh-host
|
||||||
|
|
||||||
[Install]
|
[Install]
|
||||||
WantedBy=multi-user.target
|
WantedBy=multi-user.target
|
||||||
```
|
```
|
||||||
|
|
||||||
|
**`Restart=always` and not `on-failure`**: the host restarts onto a new binary by exiting
|
||||||
|
*cleanly*, and `on-failure` would not restart it — an upgraded node would be left stopped,
|
||||||
|
having successfully upgraded. The start limit is what
|
||||||
|
[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md) uses to decide
|
||||||
|
a binary is broken rather than unlucky.
|
||||||
|
|
||||||
**The package owns this file. The host never does.** It manages `service` resources, and its own
|
**The package owns this file. The host never does.** It manages `service` resources, and its own
|
||||||
unit is a service — the temptation is obvious and it ends with a host stopping itself half way
|
unit is a service — the temptation is obvious and it ends with a host stopping itself half way
|
||||||
through an apply, leaving a machine with nothing running to fix it. A declaration naming the
|
through an apply, leaving a machine with nothing running to fix it. A declaration naming the
|
||||||
|
|||||||
@@ -76,6 +76,59 @@ its own store, and they integrate through the record rather than by reading one
|
|||||||
constraint is load-bearing: a single surface can compose them only while there is one interface
|
constraint is load-bearing: a single surface can compose them only while there is one interface
|
||||||
in front of them.
|
in front of them.
|
||||||
|
|
||||||
|
## Nothing outside a context touches its store
|
||||||
|
|
||||||
|
The question this answers: **can a node write to the registry database?** No — and not "only
|
||||||
|
through one node", which is the weaker arrangement it might be mistaken for.
|
||||||
|
|
||||||
|
> **No node holds a credential to any control-plane store, for writing or for reading.**
|
||||||
|
|
||||||
|
That is not a new rule here; it is four already taken, and it is worth seeing them together
|
||||||
|
because each one alone reads like a detail:
|
||||||
|
|
||||||
|
| | |
|
||||||
|
|---|---|
|
||||||
|
| [ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md) | the host never queries the mesh database |
|
||||||
|
| [ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md) | a node holds its own identity **and nothing else** — the shared database credential every node carries today is the exposure this exists to remove |
|
||||||
|
| [ADR 0045](../../02-DECISIONS/0045-a-context-owns-its-store.md) | a context is granted only what it **exclusively** owns: no shared writes, no read-only roles |
|
||||||
|
| [ADR 0056](../../02-DECISIONS/0056-the-authority-is-the-control-plane-not-a-database.md) | there is no single mesh database, and nothing reads one |
|
||||||
|
|
||||||
|
### So how does anything get in
|
||||||
|
|
||||||
|
**Over the broker, as a message; the owning context writes.**
|
||||||
|
|
||||||
|
```
|
||||||
|
node ──event/report──► broker ──► the context that owns that data ──► its own store
|
||||||
|
```
|
||||||
|
|
||||||
|
A node reports what it applied, what it holds, and that it is alive. It **states**; it does not
|
||||||
|
**write**. The difference is the whole security boundary: a node that can write cannot be
|
||||||
|
prevented from writing anything, and a node that can only state has its blast radius bounded by
|
||||||
|
what the message vocabulary can say.
|
||||||
|
|
||||||
|
Reads work the same way in reverse — a node is *told*, in declarations. It never asks.
|
||||||
|
|
||||||
|
### On volume, which is the real worry underneath
|
||||||
|
|
||||||
|
**Most high-frequency writes are not registry writes, and that is the first thing to check
|
||||||
|
before designing for throughput.** The registry is `inventory`'s store: nodes, modules,
|
||||||
|
assignments, versions. Those change when somebody changes something.
|
||||||
|
|
||||||
|
**Logs, metrics and health checks belong to `observability`**, which owns a different store
|
||||||
|
([ADR 0045](../../02-DECISIONS/0045-a-context-owns-its-store.md)). Sending them to the registry
|
||||||
|
would be exactly the shared-schema mistake 0045 exists to stop, arriving through the back door
|
||||||
|
marked *performance*.
|
||||||
|
|
||||||
|
That leaves one genuine funnel: every context's writes go through the process that owns it, and
|
||||||
|
[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md) says there is one of it.
|
||||||
|
For a mesh of a handful of machines this is not a scaling problem, and **it should not be solved
|
||||||
|
by giving nodes database credentials** — that trades a bounded problem for an unbounded one. If
|
||||||
|
it ever binds, the answers are at the consumer: batch, apply backpressure, or move the highest
|
||||||
|
volume stream out of a relational store entirely.
|
||||||
|
|
||||||
|
**What observability actually stores its data in is not decided**, and it is the one place where
|
||||||
|
volume genuinely argues against a relational store.
|
||||||
|
|
||||||
## What it is not
|
## What it is not
|
||||||
|
|
||||||
- **Not the thing that changes machines.** It decides; the host applies. It never reaches into a
|
- **Not the thing that changes machines.** It decides; the host applies. It never reaches into a
|
||||||
|
|||||||
@@ -10,6 +10,9 @@ decisions:
|
|||||||
- 02-DECISIONS/0039-the-link-is-the-security-boundary.md
|
- 02-DECISIONS/0039-the-link-is-the-security-boundary.md
|
||||||
- 02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md
|
- 02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md
|
||||||
- 02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md
|
- 02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md
|
||||||
|
- 02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md
|
||||||
|
- 02-DECISIONS/0058-delivery-ends-in-a-declaration.md
|
||||||
|
- 02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md
|
||||||
---
|
---
|
||||||
|
|
||||||
# The node lifecycle
|
# The node lifecycle
|
||||||
@@ -97,6 +100,48 @@ derived centrally and pushed down
|
|||||||
([ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md),
|
([ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md),
|
||||||
[`08-connectivity.md`](08-connectivity.md)).
|
[`08-connectivity.md`](08-connectivity.md)).
|
||||||
|
|
||||||
|
### The first declaration is the overlay, and nothing else
|
||||||
|
|
||||||
|
**The mesh makes a node reachable before it makes it useful.** Step 5 is not one declaration
|
||||||
|
carrying everything the node will ever run. It is two, in order:
|
||||||
|
|
||||||
|
```
|
||||||
|
first the overlay — this node's address, its keys, its peers, its names
|
||||||
|
then everything else — packages, containers, services, files
|
||||||
|
```
|
||||||
|
|
||||||
|
Three reasons, and the third is the one that matters when something goes wrong:
|
||||||
|
|
||||||
|
- **It is forced.** A node cannot join the overlay before contacting the mesh, because its
|
||||||
|
address and peer set are *assigned* — it generates a keypair, publishes the public half, and
|
||||||
|
receives the rest ([`08-connectivity.md`](08-connectivity.md)). So the overlay is the first
|
||||||
|
thing the mesh can give it, and it should be.
|
||||||
|
- **It is what [ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md) already
|
||||||
|
says:** *a joining node does the minimum to be reachable, and nothing else.*
|
||||||
|
- **It is the way back in.** Once the overlay is up, the node is reachable over it — by SSH, by
|
||||||
|
anything. If a later declaration breaks the machine, there is a route to it that does not
|
||||||
|
depend on the mesh's control path working. **Sending a large first declaration risks a node
|
||||||
|
that is broken and unreachable at the same time**, and those two failures are much worse
|
||||||
|
together than separately.
|
||||||
|
|
||||||
|
### Reachable is not the same as having a control surface
|
||||||
|
|
||||||
|
Worth stating plainly, because the two rules read as a contradiction and are not.
|
||||||
|
|
||||||
|
| | |
|
||||||
|
|---|---|
|
||||||
|
| **every node reaches every other node** | over the overlay — SSH, services, ordinary traffic. This is the point of having one |
|
||||||
|
| **every node consumes from the broker** | its own queue, over its own outbound connection ([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md)) |
|
||||||
|
| **nothing dials a node to control it** | the host has no inbound control surface ([ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md)) |
|
||||||
|
|
||||||
|
**[ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md) is about the control
|
||||||
|
channel, not about network reachability.** What it forbids is a listening thing that accepts
|
||||||
|
instructions and changes the machine. A node being reachable on the overlay — the whole purpose
|
||||||
|
of the overlay — is untouched by it, and so is a person opening a shell on it.
|
||||||
|
|
||||||
|
The distinction is *who can tell this machine what to be*: only the control plane, only over the
|
||||||
|
link the node opened, only in declarations of known shape.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## hosted → enrolled: the first node
|
## hosted → enrolled: the first node
|
||||||
|
|||||||
Reference in New Issue
Block a user