diff --git a/02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md b/02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md index e42ab8c..dde2109 100644 --- a/02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md +++ b/02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md @@ -1,5 +1,5 @@ --- -status: proposed +status: accepted date: 2026-08-27 deciders: jochen reconstructed: false @@ -40,6 +40,42 @@ could apply almost nothing, and the almost is where the confusion would live. reboot to hold it. The command-line entry points remain — they are how a person inspects and rescues a machine — but the ordinary case is a unit that is always up. +### It cannot run in a container, and the reason is the bootstrap + +Worth stating because everything else the mesh runs *is* a container, which makes the host look +like an exception somebody forgot to fix. + +**Step 0 of the substrate bootstrap is installing the container runtime** +([`07-the-substrate.md`](../03-DESIGN/01-to-be/07-the-substrate.md)). A host that ran inside a +container could not perform it — it would need the thing it is there to install. On a machine +with no runtime, nothing would ever start. + +That is not the only reason, but it is the sufficient one: + +- **It would break [ADR 0041](0041-the-host-depends-on-nothing.md).** *Copy it onto a machine and + run it* stops being true when the machine must already have a container runtime. +- **The isolation would be fiction.** To write `/etc`, install packages, manage units and run + containers, it would need the host's mount, PID and network namespaces plus the runtime's own + socket. A container with all of those is a process with extra steps. + +**So the host is a plain process on the machine, and everything above tier 0 is a container.** +That split is the tier boundary made concrete rather than an inconsistency. + +### What it needs from an init, and why that is not a dependency + +The host needs four things from whatever supervises it: start at boot, restart when it exits, +give up after repeated failures, and run something else when it gives up. + +**Every machine the mesh targets already has systemd**, and the host already treats the service +manager as a detected capability rather than an assumption. This is not a dependency in +[ADR 0041](0041-the-host-depends-on-nothing.md)'s sense — 0041 is about what must be *installed +before the host works*, and an init is not installed, it is what the machine already is. + +**Abstracting over init systems is not done**, because there is no second one to abstract over. +The unit file is the only systemd-specific artefact, it belongs to the package rather than the +binary, and a machine with a different supervisor would ship a different package — which is +where that difference belongs. + ### It never manages its own unit **The host's own service file is not a resource the host applies.** The temptation is obvious — diff --git a/02-DECISIONS/0058-delivery-ends-in-a-declaration.md b/02-DECISIONS/0058-delivery-ends-in-a-declaration.md index 75ad3e7..b424d57 100644 --- a/02-DECISIONS/0058-delivery-ends-in-a-declaration.md +++ b/02-DECISIONS/0058-delivery-ends-in-a-declaration.md @@ -1,5 +1,5 @@ --- -status: proposed +status: accepted date: 2026-08-27 deciders: jochen reconstructed: false diff --git a/02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md b/02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md index d92a9c8..a957f12 100644 --- a/02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md +++ b/02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md @@ -1,12 +1,12 @@ --- -status: proposed +status: accepted date: 2026-08-27 deciders: jochen reconstructed: false extends: 0057-the-host-is-a-root-service-installed-as-a-package.md --- -# 59. A host that cannot start is rolled back by the supervisor +# 59. Two watchdogs: the mesh stages the rollout, the supervisor recovers the node ## Context @@ -18,21 +18,31 @@ It was left because automatic recovery looked like *the host judging its own hea the self-reference the rest of that document avoids. **That objection does not survive being asked properly.** A keepalive is not the host judging -itself — it is something else judging the host. The question is only *what*, and that has one -answer. +itself — it is something else judging the host. -### Why the watchdog must be local +### There are two watchdogs, and they cannot do each other's job -The obvious candidate is the control plane, and it cannot be: +A first draft of this record concluded the watchdog must be local, and stopped there. That was +half an answer: it is true that recovery must be local, and false that the mesh has no part. + +| | can see | can act | +|---|---|---| +| **the service manager**, on the node | that *this* process keeps dying | **yes** — restart it, replace it | +| **the mesh** | that *eleven of twelve nodes* went quiet after one declaration | **no** — nothing dials a node | + +**Recovery must be local**, and that half stands: - **Nothing dials a node.** [ADR 0039](0039-the-link-is-the-security-boundary.md) makes the link - outbound and node-initiated, and a node has no listening control surface. The mesh has no way - to reach in and act. + outbound and node-initiated, with no listening control surface. The mesh has no way to reach + in and act. - **The failure removes the reporting path.** A host that cannot start cannot link, so the mesh - learns nothing to act on. + learns nothing *from that node* to act on. -So the watchdog is local, and the only local thing that is always present, already depended on, -and is not the host is the **service manager**. +**But detection is the mesh's**, and it is the half a local watchdog structurally cannot do. A +node's supervisor sees one process failing and has no idea whether that is a broken machine or a +broken release. **Only something watching every node can tell those apart** — and telling them +apart is what decides whether the right response is *fix this machine* or *stop shipping this +version immediately*. ### The failure this actually prevents @@ -51,7 +61,35 @@ nothing else. ## Decision -**The service manager rolls the host back to the last version that started.** +**The mesh stages the rollout and stops when nodes go quiet. The service manager recovers the +node it is on.** Prevention and recovery, and neither substitutes for the other. + +### The mesh stages a host rollout + +A host version does not reach every node at once. The delivery context updates a few nodes' +declarations, **waits for those nodes to heartbeat on the new version**, and only then continues. + +``` +update 2 nodes ─► heard from both, running the new version ─► continue + └► silence past the window ─► STOP. Report. +``` + +**Silence is the signal, and it is available because of the heartbeat** +([ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md)). A node that upgraded +and cannot start stops reporting; that is indistinguishable from a switched-off machine *for one +node*, and completely distinguishable across a batch that was all told the same thing at the +same time. + +**This is what keeps a bad release from becoming a fleet outage.** Local rollback repairs a node +after the fact; staging means most nodes never receive the bad version at all. A stopped rollout +is two broken nodes and a report, rather than every node quiet at once. + +**It does not replace local recovery**, for two reasons. The canary nodes still break, and +somebody has to be able to fix them. And a node that was offline during the staged rollout gets +the declaration when it reconnects, with no batch around it and nothing watching — so it must be +able to recover alone. + +### On the node: the service manager rolls back Four parts, and each one is chosen so it works when the host does not: @@ -78,7 +116,7 @@ import it. ```ini [Service] -Restart=on-failure +Restart=always StartLimitBurst=3 StartLimitIntervalSec=120 @@ -86,10 +124,27 @@ StartLimitIntervalSec=120 OnFailure=nox-mesh-host-rollback.service ``` -Three failures in two minutes is a binary that does not work, not a transient. The supervisor -stops trying and runs the rollback unit, which downgrades and starts the host again. +**`Restart=always`, not `on-failure`**, and the difference is load-bearing rather than a +preference. [ADR 0057](0057-the-host-is-a-root-service-installed-as-a-package.md) has the host +restart onto a new binary by **exiting cleanly** — and `on-failure` does not restart a process +that exited zero. An earlier draft of this record specified `on-failure` and would have left +every upgraded node stopped, having successfully upgraded. Caught by reading the two records +against each other rather than by either alone. -### 4 — It rolls back once +Three failures in two minutes is a binary that does not work, not a transient. The supervisor +stops trying, the unit enters a failed state, and `OnFailure` runs the rollback unit — which +downgrades and starts the host again. + +### 4 — With nothing to roll back to, it does not try + +A machine whose host has *never* completed a reconcile has no `known-good`. The rollback unit +finds nothing, does nothing, and says so. + +That is the right outcome: there is no previous version, so the node was never working, and the +failure belongs to the installation rather than to an upgrade. Attempting a rollback here would +mean guessing at a version, which is how a recovery mechanism becomes a second fault. + +### 5 — It rolls back once The rollback unit records that it fired. If the rolled-back version *also* fails to start, it does **not** fire again — the node stops, loudly, in a failed state. diff --git a/03-DESIGN/01-to-be/05-the-node-host.md b/03-DESIGN/01-to-be/05-the-node-host.md index 15e8ed5..c31cba0 100644 --- a/03-DESIGN/01-to-be/05-the-node-host.md +++ b/03-DESIGN/01-to-be/05-the-node-host.md @@ -12,6 +12,8 @@ decisions: - 02-DECISIONS/0008-a-failed-step-fails-the-job.md - 02-DECISIONS/0041-the-host-depends-on-nothing.md - 02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md + - 02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md + - 02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md --- # The node host @@ -115,10 +117,7 @@ the link; never asked downward. ## The process -> **Proposed, not yet decided** — -> [ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md) is -> awaiting review. This section is written against it and moves into the document's `decisions:` -> when the record is accepted. +[ADR 0057](../../02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md). **One `root` service on every node, plus command-line entry points for a person.** A machine without one is not a node ([ADR 0036](../../02-DECISIONS/0036-a-node-is-a-managed-machine.md)) — @@ -167,12 +166,20 @@ Type=notify ExecStart=/usr/bin/nox-mesh-host run Restart=always RestartSec=5s +StartLimitBurst=3 +StartLimitIntervalSec=120 StateDirectory=mesh-host [Install] WantedBy=multi-user.target ``` +**`Restart=always` and not `on-failure`**: the host restarts onto a new binary by exiting +*cleanly*, and `on-failure` would not restart it — an upgraded node would be left stopped, +having successfully upgraded. The start limit is what +[ADR 0059](../../02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md) uses to decide +a binary is broken rather than unlucky. + **The package owns this file. The host never does.** It manages `service` resources, and its own unit is a service — the temptation is obvious and it ends with a host stopping itself half way through an apply, leaving a machine with nothing running to fix it. A declaration naming the diff --git a/03-DESIGN/01-to-be/06-the-control-plane.md b/03-DESIGN/01-to-be/06-the-control-plane.md index 1947f8f..c074a74 100644 --- a/03-DESIGN/01-to-be/06-the-control-plane.md +++ b/03-DESIGN/01-to-be/06-the-control-plane.md @@ -76,6 +76,59 @@ its own store, and they integrate through the record rather than by reading one constraint is load-bearing: a single surface can compose them only while there is one interface in front of them. +## Nothing outside a context touches its store + +The question this answers: **can a node write to the registry database?** No — and not "only +through one node", which is the weaker arrangement it might be mistaken for. + +> **No node holds a credential to any control-plane store, for writing or for reading.** + +That is not a new rule here; it is four already taken, and it is worth seeing them together +because each one alone reads like a detail: + +| | | +|---|---| +| [ADR 0037](../../02-DECISIONS/0037-the-host-applies-it-does-not-decide.md) | the host never queries the mesh database | +| [ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md) | a node holds its own identity **and nothing else** — the shared database credential every node carries today is the exposure this exists to remove | +| [ADR 0045](../../02-DECISIONS/0045-a-context-owns-its-store.md) | a context is granted only what it **exclusively** owns: no shared writes, no read-only roles | +| [ADR 0056](../../02-DECISIONS/0056-the-authority-is-the-control-plane-not-a-database.md) | there is no single mesh database, and nothing reads one | + +### So how does anything get in + +**Over the broker, as a message; the owning context writes.** + +``` +node ──event/report──► broker ──► the context that owns that data ──► its own store +``` + +A node reports what it applied, what it holds, and that it is alive. It **states**; it does not +**write**. The difference is the whole security boundary: a node that can write cannot be +prevented from writing anything, and a node that can only state has its blast radius bounded by +what the message vocabulary can say. + +Reads work the same way in reverse — a node is *told*, in declarations. It never asks. + +### On volume, which is the real worry underneath + +**Most high-frequency writes are not registry writes, and that is the first thing to check +before designing for throughput.** The registry is `inventory`'s store: nodes, modules, +assignments, versions. Those change when somebody changes something. + +**Logs, metrics and health checks belong to `observability`**, which owns a different store +([ADR 0045](../../02-DECISIONS/0045-a-context-owns-its-store.md)). Sending them to the registry +would be exactly the shared-schema mistake 0045 exists to stop, arriving through the back door +marked *performance*. + +That leaves one genuine funnel: every context's writes go through the process that owns it, and +[ADR 0053](../../02-DECISIONS/0053-one-control-plane-and-no-failover.md) says there is one of it. +For a mesh of a handful of machines this is not a scaling problem, and **it should not be solved +by giving nodes database credentials** — that trades a bounded problem for an unbounded one. If +it ever binds, the answers are at the consumer: batch, apply backpressure, or move the highest +volume stream out of a relational store entirely. + +**What observability actually stores its data in is not decided**, and it is the one place where +volume genuinely argues against a relational store. + ## What it is not - **Not the thing that changes machines.** It decides; the host applies. It never reaches into a diff --git a/03-DESIGN/01-to-be/09-the-node-lifecycle.md b/03-DESIGN/01-to-be/09-the-node-lifecycle.md index f8eaaac..1fc32ab 100644 --- a/03-DESIGN/01-to-be/09-the-node-lifecycle.md +++ b/03-DESIGN/01-to-be/09-the-node-lifecycle.md @@ -10,6 +10,9 @@ decisions: - 02-DECISIONS/0039-the-link-is-the-security-boundary.md - 02-DECISIONS/0043-a-declaration-is-an-ordered-list-of-owned-resources.md - 02-DECISIONS/0051-the-enrolment-token-carries-the-mesh.md + - 02-DECISIONS/0057-the-host-is-a-root-service-installed-as-a-package.md + - 02-DECISIONS/0058-delivery-ends-in-a-declaration.md + - 02-DECISIONS/0059-a-host-that-cannot-start-rolls-itself-back.md --- # The node lifecycle @@ -97,6 +100,48 @@ derived centrally and pushed down ([ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md), [`08-connectivity.md`](08-connectivity.md)). +### The first declaration is the overlay, and nothing else + +**The mesh makes a node reachable before it makes it useful.** Step 5 is not one declaration +carrying everything the node will ever run. It is two, in order: + +``` +first the overlay — this node's address, its keys, its peers, its names +then everything else — packages, containers, services, files +``` + +Three reasons, and the third is the one that matters when something goes wrong: + +- **It is forced.** A node cannot join the overlay before contacting the mesh, because its + address and peer set are *assigned* — it generates a keypair, publishes the public half, and + receives the rest ([`08-connectivity.md`](08-connectivity.md)). So the overlay is the first + thing the mesh can give it, and it should be. +- **It is what [ADR 0038](../../02-DECISIONS/0038-a-node-joins-by-linking-first.md) already + says:** *a joining node does the minimum to be reachable, and nothing else.* +- **It is the way back in.** Once the overlay is up, the node is reachable over it — by SSH, by + anything. If a later declaration breaks the machine, there is a route to it that does not + depend on the mesh's control path working. **Sending a large first declaration risks a node + that is broken and unreachable at the same time**, and those two failures are much worse + together than separately. + +### Reachable is not the same as having a control surface + +Worth stating plainly, because the two rules read as a contradiction and are not. + +| | | +|---|---| +| **every node reaches every other node** | over the overlay — SSH, services, ordinary traffic. This is the point of having one | +| **every node consumes from the broker** | its own queue, over its own outbound connection ([ADR 0001](../../02-DECISIONS/0001-nodes-communicate-over-a-broker.md)) | +| **nothing dials a node to control it** | the host has no inbound control surface ([ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md)) | + +**[ADR 0039](../../02-DECISIONS/0039-the-link-is-the-security-boundary.md) is about the control +channel, not about network reachability.** What it forbids is a listening thing that accepts +instructions and changes the machine. A node being reachable on the overlay — the whole purpose +of the overlay — is untouched by it, and so is a person opening a shell on it. + +The distinction is *who can tell this machine what to be*: only the control plane, only over the +link the node opened, only in declarations of known shape. + --- ## hosted → enrolled: the first node