Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24
@@ -1,5 +1,5 @@
|
||||
---
|
||||
status: proposed
|
||||
status: accepted
|
||||
date: 2026-08-27
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
@@ -40,6 +40,42 @@ could apply almost nothing, and the almost is where the confusion would live.
|
||||
reboot to hold it. The command-line entry points remain — they are how a person inspects and
|
||||
rescues a machine — but the ordinary case is a unit that is always up.
|
||||
|
||||
### It cannot run in a container, and the reason is the bootstrap
|
||||
|
||||
Worth stating because everything else the mesh runs *is* a container, which makes the host look
|
||||
like an exception somebody forgot to fix.
|
||||
|
||||
**Step 0 of the substrate bootstrap is installing the container runtime**
|
||||
([`07-the-substrate.md`](../03-DESIGN/01-to-be/07-the-substrate.md)). A host that ran inside a
|
||||
container could not perform it — it would need the thing it is there to install. On a machine
|
||||
with no runtime, nothing would ever start.
|
||||
|
||||
That is not the only reason, but it is the sufficient one:
|
||||
|
||||
- **It would break [ADR 0041](0041-the-host-depends-on-nothing.md).** *Copy it onto a machine and
|
||||
run it* stops being true when the machine must already have a container runtime.
|
||||
- **The isolation would be fiction.** To write `/etc`, install packages, manage units and run
|
||||
containers, it would need the host's mount, PID and network namespaces plus the runtime's own
|
||||
socket. A container with all of those is a process with extra steps.
|
||||
|
||||
**So the host is a plain process on the machine, and everything above tier 0 is a container.**
|
||||
That split is the tier boundary made concrete rather than an inconsistency.
|
||||
|
||||
### What it needs from an init, and why that is not a dependency
|
||||
|
||||
The host needs four things from whatever supervises it: start at boot, restart when it exits,
|
||||
give up after repeated failures, and run something else when it gives up.
|
||||
|
||||
**Every machine the mesh targets already has systemd**, and the host already treats the service
|
||||
manager as a detected capability rather than an assumption. This is not a dependency in
|
||||
[ADR 0041](0041-the-host-depends-on-nothing.md)'s sense — 0041 is about what must be *installed
|
||||
before the host works*, and an init is not installed, it is what the machine already is.
|
||||
|
||||
**Abstracting over init systems is not done**, because there is no second one to abstract over.
|
||||
The unit file is the only systemd-specific artefact, it belongs to the package rather than the
|
||||
binary, and a machine with a different supervisor would ship a different package — which is
|
||||
where that difference belongs.
|
||||
|
||||
### It never manages its own unit
|
||||
|
||||
**The host's own service file is not a resource the host applies.** The temptation is obvious —
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
---
|
||||
status: proposed
|
||||
status: accepted
|
||||
date: 2026-08-27
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
|
||||
@@ -1,12 +1,12 @@
|
||||
---
|
||||
status: proposed
|
||||
status: accepted
|
||||
date: 2026-08-27
|
||||
deciders: jochen
|
||||
reconstructed: false
|
||||
extends: 0057-the-host-is-a-root-service-installed-as-a-package.md
|
||||
---
|
||||
|
||||
# 59. A host that cannot start is rolled back by the supervisor
|
||||
# 59. Two watchdogs: the mesh stages the rollout, the supervisor recovers the node
|
||||
|
||||
## Context
|
||||
|
||||
@@ -18,21 +18,31 @@ It was left because automatic recovery looked like *the host judging its own hea
|
||||
the self-reference the rest of that document avoids.
|
||||
|
||||
**That objection does not survive being asked properly.** A keepalive is not the host judging
|
||||
itself — it is something else judging the host. The question is only *what*, and that has one
|
||||
answer.
|
||||
itself — it is something else judging the host.
|
||||
|
||||
### Why the watchdog must be local
|
||||
### There are two watchdogs, and they cannot do each other's job
|
||||
|
||||
The obvious candidate is the control plane, and it cannot be:
|
||||
A first draft of this record concluded the watchdog must be local, and stopped there. That was
|
||||
half an answer: it is true that recovery must be local, and false that the mesh has no part.
|
||||
|
||||
| | can see | can act |
|
||||
|---|---|---|
|
||||
| **the service manager**, on the node | that *this* process keeps dying | **yes** — restart it, replace it |
|
||||
| **the mesh** | that *eleven of twelve nodes* went quiet after one declaration | **no** — nothing dials a node |
|
||||
|
||||
**Recovery must be local**, and that half stands:
|
||||
|
||||
- **Nothing dials a node.** [ADR 0039](0039-the-link-is-the-security-boundary.md) makes the link
|
||||
outbound and node-initiated, and a node has no listening control surface. The mesh has no way
|
||||
to reach in and act.
|
||||
outbound and node-initiated, with no listening control surface. The mesh has no way to reach
|
||||
in and act.
|
||||
- **The failure removes the reporting path.** A host that cannot start cannot link, so the mesh
|
||||
learns nothing to act on.
|
||||
learns nothing *from that node* to act on.
|
||||
|
||||
So the watchdog is local, and the only local thing that is always present, already depended on,
|
||||
and is not the host is the **service manager**.
|
||||
**But detection is the mesh's**, and it is the half a local watchdog structurally cannot do. A
|
||||
node's supervisor sees one process failing and has no idea whether that is a broken machine or a
|
||||
broken release. **Only something watching every node can tell those apart** — and telling them
|
||||
apart is what decides whether the right response is *fix this machine* or *stop shipping this
|
||||
version immediately*.
|
||||
|
||||
### The failure this actually prevents
|
||||
|
||||
@@ -51,7 +61,35 @@ nothing else.
|
||||
|
||||
## Decision
|
||||
|
||||
**The service manager rolls the host back to the last version that started.**
|
||||
**The mesh stages the rollout and stops when nodes go quiet. The service manager recovers the
|
||||
node it is on.** Prevention and recovery, and neither substitutes for the other.
|
||||
|
||||
### The mesh stages a host rollout
|
||||
|
||||
A host version does not reach every node at once. The delivery context updates a few nodes'
|
||||
declarations, **waits for those nodes to heartbeat on the new version**, and only then continues.
|
||||
|
||||
```
|
||||
update 2 nodes ─► heard from both, running the new version ─► continue
|
||||
└► silence past the window ─► STOP. Report.
|
||||
```
|
||||
|
||||
**Silence is the signal, and it is available because of the heartbeat**
|
||||
([ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)). A node that upgraded
|
||||
and cannot start stops reporting; that is indistinguishable from a switched-off machine *for one
|
||||
node*, and completely distinguishable across a batch that was all told the same thing at the
|
||||
same time.
|
||||
|
||||
**This is what keeps a bad release from becoming a fleet outage.** Local rollback repairs a node
|
||||
after the fact; staging means most nodes never receive the bad version at all. A stopped rollout
|
||||
is two broken nodes and a report, rather than every node quiet at once.
|
||||
|
||||
**It does not replace local recovery**, for two reasons. The canary nodes still break, and
|
||||
somebody has to be able to fix them. And a node that was offline during the staged rollout gets
|
||||
the declaration when it reconnects, with no batch around it and nothing watching — so it must be
|
||||
able to recover alone.
|
||||
|
||||
### On the node: the service manager rolls back
|
||||
|
||||
Four parts, and each one is chosen so it works when the host does not:
|
||||
|
||||
@@ -78,7 +116,7 @@ import it.
|
||||
|
||||
```ini
|
||||
[Service]
|
||||
Restart=on-failure
|
||||
Restart=always
|
||||
StartLimitBurst=3
|
||||
StartLimitIntervalSec=120
|
||||
|
||||
@@ -86,10 +124,27 @@ StartLimitIntervalSec=120
|
||||
OnFailure=nox-mesh-host-rollback.service
|
||||
```
|
||||
|
||||
Three failures in two minutes is a binary that does not work, not a transient. The supervisor
|
||||
stops trying and runs the rollback unit, which downgrades and starts the host again.
|
||||
**`Restart=always`, not `on-failure`**, and the difference is load-bearing rather than a
|
||||
preference. [ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md) has the host
|
||||
restart onto a new binary by **exiting cleanly** — and `on-failure` does not restart a process
|
||||
that exited zero. An earlier draft of this record specified `on-failure` and would have left
|
||||
every upgraded node stopped, having successfully upgraded. Caught by reading the two records
|
||||
against each other rather than by either alone.
|
||||
|
||||
### 4 — It rolls back once
|
||||
Three failures in two minutes is a binary that does not work, not a transient. The supervisor
|
||||
stops trying, the unit enters a failed state, and `OnFailure` runs the rollback unit — which
|
||||
downgrades and starts the host again.
|
||||
|
||||
### 4 — With nothing to roll back to, it does not try
|
||||
|
||||
A machine whose host has *never* completed a reconcile has no `known-good`. The rollback unit
|
||||
finds nothing, does nothing, and says so.
|
||||
|
||||
That is the right outcome: there is no previous version, so the node was never working, and the
|
||||
failure belongs to the installation rather than to an upgrade. Attempting a rollback here would
|
||||
mean guessing at a version, which is how a recovery mechanism becomes a second fault.
|
||||
|
||||
### 5 — It rolls back once
|
||||
|
||||
The rollback unit records that it fired. If the rolled-back version *also* fails to start, it
|
||||
does **not** fire again — the node stops, loudly, in a failed state.
|
||||
|
||||
@@ -12,6 +12,8 @@ decisions:
|
||||
- 02-DECISIONS/0008-a-failed-step-fails-the-job.md
|
||||
- 02-DECISIONS/0041-the-host-depends-on-nothing.md
|
||||
- 02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md
|
||||
- 02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md
|
||||
- 02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md
|
||||
---
|
||||
|
||||
# The node host
|
||||
@@ -115,10 +117,7 @@ the link; never asked downward.
|
||||
|
||||
## The process
|
||||
|
||||
> **Proposed, not yet decided** —
|
||||
> [ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md) is
|
||||
> awaiting review. This section is written against it and moves into the document's `decisions:`
|
||||
> when the record is accepted.
|
||||
[ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md).
|
||||
|
||||
**One `root` service on every node, plus command-line entry points for a person.** A machine
|
||||
without one is not a node ([ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)) —
|
||||
@@ -167,12 +166,20 @@ Type=notify
|
||||
ExecStart=/usr/bin/nox-mesh-host run
|
||||
Restart=always
|
||||
RestartSec=5s
|
||||
StartLimitBurst=3
|
||||
StartLimitIntervalSec=120
|
||||
StateDirectory=mesh-host
|
||||
|
||||
[Install]
|
||||
WantedBy=multi-user.target
|
||||
```
|
||||
|
||||
**`Restart=always` and not `on-failure`**: the host restarts onto a new binary by exiting
|
||||
*cleanly*, and `on-failure` would not restart it — an upgraded node would be left stopped,
|
||||
having successfully upgraded. The start limit is what
|
||||
[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md) uses to decide
|
||||
a binary is broken rather than unlucky.
|
||||
|
||||
**The package owns this file. The host never does.** It manages `service` resources, and its own
|
||||
unit is a service — the temptation is obvious and it ends with a host stopping itself half way
|
||||
through an apply, leaving a machine with nothing running to fix it. A declaration naming the
|
||||
|
||||
@@ -76,6 +76,59 @@ its own store, and they integrate through the record rather than by reading one
|
||||
constraint is load-bearing: a single surface can compose them only while there is one interface
|
||||
in front of them.
|
||||
|
||||
## Nothing outside a context touches its store
|
||||
|
||||
The question this answers: **can a node write to the registry database?** No — and not "only
|
||||
through one node", which is the weaker arrangement it might be mistaken for.
|
||||
|
||||
> **No node holds a credential to any control-plane store, for writing or for reading.**
|
||||
|
||||
That is not a new rule here; it is four already taken, and it is worth seeing them together
|
||||
because each one alone reads like a detail:
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| [ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md) | the host never queries the mesh database |
|
||||
| [ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md) | a node holds its own identity **and nothing else** — the shared database credential every node carries today is the exposure this exists to remove |
|
||||
| [ADR 0045](../../02-DECISIONS/0045-a-context-owns-its-store.md) | a context is granted only what it **exclusively** owns: no shared writes, no read-only roles |
|
||||
| [ADR 0056](../../02-DECISIONS/0056-the-authority-is-the-control-plane-not-a-database.md) | there is no single mesh database, and nothing reads one |
|
||||
|
||||
### So how does anything get in
|
||||
|
||||
**Over the broker, as a message; the owning context writes.**
|
||||
|
||||
```
|
||||
node ──event/report──► broker ──► the context that owns that data ──► its own store
|
||||
```
|
||||
|
||||
A node reports what it applied, what it holds, and that it is alive. It **states**; it does not
|
||||
**write**. The difference is the whole security boundary: a node that can write cannot be
|
||||
prevented from writing anything, and a node that can only state has its blast radius bounded by
|
||||
what the message vocabulary can say.
|
||||
|
||||
Reads work the same way in reverse — a node is *told*, in declarations. It never asks.
|
||||
|
||||
### On volume, which is the real worry underneath
|
||||
|
||||
**Most high-frequency writes are not registry writes, and that is the first thing to check
|
||||
before designing for throughput.** The registry is `inventory`'s store: nodes, modules,
|
||||
assignments, versions. Those change when somebody changes something.
|
||||
|
||||
**Logs, metrics and health checks belong to `observability`**, which owns a different store
|
||||
([ADR 0045](../../02-DECISIONS/0045-a-context-owns-its-store.md)). Sending them to the registry
|
||||
would be exactly the shared-schema mistake 0045 exists to stop, arriving through the back door
|
||||
marked *performance*.
|
||||
|
||||
That leaves one genuine funnel: every context's writes go through the process that owns it, and
|
||||
[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md) says there is one of it.
|
||||
For a mesh of a handful of machines this is not a scaling problem, and **it should not be solved
|
||||
by giving nodes database credentials** — that trades a bounded problem for an unbounded one. If
|
||||
it ever binds, the answers are at the consumer: batch, apply backpressure, or move the highest
|
||||
volume stream out of a relational store entirely.
|
||||
|
||||
**What observability actually stores its data in is not decided**, and it is the one place where
|
||||
volume genuinely argues against a relational store.
|
||||
|
||||
## What it is not
|
||||
|
||||
- **Not the thing that changes machines.** It decides; the host applies. It never reaches into a
|
||||
|
||||
@@ -10,6 +10,9 @@ decisions:
|
||||
- 02-DECISIONS/0039-the-link-is-the-security-boundary.md
|
||||
- 02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md
|
||||
- 02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md
|
||||
- 02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md
|
||||
- 02-DECISIONS/0058-delivery-ends-in-a-declaration.md
|
||||
- 02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md
|
||||
---
|
||||
|
||||
# The node lifecycle
|
||||
@@ -97,6 +100,48 @@ derived centrally and pushed down
|
||||
([ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md),
|
||||
[`08-connectivity.md`](08-connectivity.md)).
|
||||
|
||||
### The first declaration is the overlay, and nothing else
|
||||
|
||||
**The mesh makes a node reachable before it makes it useful.** Step 5 is not one declaration
|
||||
carrying everything the node will ever run. It is two, in order:
|
||||
|
||||
```
|
||||
first the overlay — this node's address, its keys, its peers, its names
|
||||
then everything else — packages, containers, services, files
|
||||
```
|
||||
|
||||
Three reasons, and the third is the one that matters when something goes wrong:
|
||||
|
||||
- **It is forced.** A node cannot join the overlay before contacting the mesh, because its
|
||||
address and peer set are *assigned* — it generates a keypair, publishes the public half, and
|
||||
receives the rest ([`08-connectivity.md`](08-connectivity.md)). So the overlay is the first
|
||||
thing the mesh can give it, and it should be.
|
||||
- **It is what [ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md) already
|
||||
says:** *a joining node does the minimum to be reachable, and nothing else.*
|
||||
- **It is the way back in.** Once the overlay is up, the node is reachable over it — by SSH, by
|
||||
anything. If a later declaration breaks the machine, there is a route to it that does not
|
||||
depend on the mesh's control path working. **Sending a large first declaration risks a node
|
||||
that is broken and unreachable at the same time**, and those two failures are much worse
|
||||
together than separately.
|
||||
|
||||
### Reachable is not the same as having a control surface
|
||||
|
||||
Worth stating plainly, because the two rules read as a contradiction and are not.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **every node reaches every other node** | over the overlay — SSH, services, ordinary traffic. This is the point of having one |
|
||||
| **every node consumes from the broker** | its own queue, over its own outbound connection ([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md)) |
|
||||
| **nothing dials a node to control it** | the host has no inbound control surface ([ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md)) |
|
||||
|
||||
**[ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md) is about the control
|
||||
channel, not about network reachability.** What it forbids is a listening thing that accepts
|
||||
instructions and changes the machine. A node being reachable on the overlay — the whole purpose
|
||||
of the overlay — is untouched by it, and so is a person opening a shell on it.
|
||||
|
||||
The distinction is *who can tell this machine what to be*: only the control plane, only over the
|
||||
link the node opened, only in declarations of known shape.
|
||||
|
||||
---
|
||||
|
||||
## hosted → enrolled: the first node
|
||||
|
||||
Reference in New Issue
Block a user