Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.
the node host 8 -> 1 applies not decides, depends on nothing,
per operating system, root service, the
launcher, episodic, what a declaration is,
actions from the bundle only
a node and how it joins 4 -> 1 what a node is, joining, the link as
security boundary, the enrolment token
modules and the graph 7 -> 1 everything is a module, no domain modules,
three edges, provisioning, the core library
substrate and control 6 -> 1 the test, seven contexts, one control plane,
plane the authority is not a database, the named
products, the pinned bundle
connectivity 3 -> 1 a route is a grant, reachability declared,
filter rules
delivery 5 -> 1 reconciliation not a pipeline, artifacts,
the three silos, a failed step, the verdict
the lab 5 -> 1 (earlier)
how this repository 10 -> 1 (earlier)
works
Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.
The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.
The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
60 lines
2.4 KiB
Markdown
60 lines
2.4 KiB
Markdown
---
|
|
status: open
|
|
opened: 2026-08-28
|
|
located-in: [hal]
|
|
fixed-by:
|
|
amended-design:
|
|
---
|
|
|
|
# 008 — The documented automatic node rescue does not exist
|
|
|
|
## Symptom
|
|
|
|
The mesh's documentation describes an automatic node rescue: a node that fails is recovered
|
|
without anybody intervening. **Nothing implements it.**
|
|
|
|
Found incidentally while investigating supervision
|
|
([research 003](../../01-RESEARCH/003-service-supervision/00-overview.md)), which counted what
|
|
actually supervises what:
|
|
|
|
- **no unit declares `OnFailure=`**, so nothing runs when a unit gives up;
|
|
- **nothing calls the rescue script on a timer**, so it runs only when a person runs it.
|
|
|
|
The script exists. The thing that would invoke it does not.
|
|
|
|
## Why this is worse than having no rescue
|
|
|
|
A rescue nobody wrote is a gap somebody can see. A rescue that is *documented* and absent is a
|
|
gap nobody looks for, because the documentation says it is covered — and it is read exactly when
|
|
a node has failed and somebody is deciding whether to intervene.
|
|
|
|
This is `how-we-build` §5 in its most expensive form: *an unenforced rule is indistinguishable
|
|
from a wrong one, and costs more, because people believe it.* Here the belief is that a failed
|
|
node recovers itself.
|
|
|
|
## Scope
|
|
|
|
**The as-is only.** The design being built has a different answer:
|
|
[ADR 0037](../../02-DECISIONS/0037-the-node-host.md) puts recovery
|
|
in a launcher that supervises the host, and that recovery is tested — 32 assertions, each
|
|
confirmed to fail when the behaviour is removed.
|
|
|
|
So this issue is about the mesh that runs **now**, and it has two possible resolutions rather
|
|
than one:
|
|
|
|
1. **Implement it** — an `OnFailure=` and a timer — if node rescue is wanted before the new host
|
|
reaches the fleet.
|
|
2. **Delete the documentation** — and say plainly that a failed node needs a person, which is
|
|
what is true today.
|
|
|
|
**Either is honest. Leaving it as it is, is not.** The choice turns on how far away the new host
|
|
is, which is a scheduling question rather than a technical one.
|
|
|
|
## What it would take to be sure
|
|
|
|
Read back rather than assumed
|
|
([ADR 0035](../../02-DECISIONS/0035-a-picture-is-read-from-what-runs.md)): list every unit on a
|
|
node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above
|
|
came from reading the repository, and confirming it against a running node is the difference
|
|
between *no unit declares this* and *no unit in the source declares this*.
|