Jochen asked whether the order made sense. It did not -- it followed when things happened to be decided, which after consolidation is fictional anyway since record 5 alone folds decisions taken across a week. Concretely wrong before: the domain statement sat at 8, after five engineering rules; the constitution was scattered across 5, 12 and 17; the tiers landed at 15, 16, 21 and 22 with process records in between. Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what runs on them and how it gets there (9-10), how it is built (11-16), how it is checked (17-18), how we work (19-23). Two things made this safe rather than free. It is a permutation, not a compaction, so the renames go through temporary names -- otherwise two files want one slot and one is lost. And the reference rewrite is a single simultaneous pass, because almost every number moved into a slot another number was vacating; replacing one at a time would have cascaded and pointed things at the wrong record while still resolving. Verified: 284 [ADR NNNN](path) links across the repository, all with matching text and target. The ordering principle is now stated in 19 rather than left implicit -- the repository already said "the numbering is the flow" about its folders, and there was no reason for the records to be the exception.
2.4 KiB
status, opened, located-in, fixed-by, amended-design
| status | opened | located-in | fixed-by | amended-design | |
|---|---|---|---|---|---|
| open | 2026-08-28 |
|
008 — The documented automatic node rescue does not exist
Symptom
The mesh's documentation describes an automatic node rescue: a node that fails is recovered without anybody intervening. Nothing implements it.
Found incidentally while investigating supervision (research 003), which counted what actually supervises what:
- no unit declares
OnFailure=, so nothing runs when a unit gives up; - nothing calls the rescue script on a timer, so it runs only when a person runs it.
The script exists. The thing that would invoke it does not.
Why this is worse than having no rescue
A rescue nobody wrote is a gap somebody can see. A rescue that is documented and absent is a gap nobody looks for, because the documentation says it is covered — and it is read exactly when a node has failed and somebody is deciding whether to intervene.
This is how-we-build §5 in its most expensive form: an unenforced rule is indistinguishable
from a wrong one, and costs more, because people believe it. Here the belief is that a failed
node recovers itself.
Scope
The as-is only. The design being built has a different answer: ADR 0005 puts recovery in a launcher that supervises the host, and that recovery is tested — 32 assertions, each confirmed to fail when the behaviour is removed.
So this issue is about the mesh that runs now, and it has two possible resolutions rather than one:
- Implement it — an
OnFailure=and a timer — if node rescue is wanted before the new host reaches the fleet. - Delete the documentation — and say plainly that a failed node needs a person, which is what is true today.
Either is honest. Leaving it as it is, is not. The choice turns on how far away the new host is, which is a scheduling question rather than a technical one.
What it would take to be sure
Read back rather than assumed
(ADR 0018): list every unit on a
node and grep for OnFailure=; list every timer and check what each one calls. The finding above
came from reading the repository, and confirming it against a running node is the difference
between no unit declares this and no unit in the source declares this.