Jochen asked whether the order made sense. It did not -- it followed when things happened to be decided, which after consolidation is fictional anyway since record 5 alone folds decisions taken across a week. Concretely wrong before: the domain statement sat at 8, after five engineering rules; the constitution was scattered across 5, 12 and 17; the tiers landed at 15, 16, 21 and 22 with process records in between. Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what runs on them and how it gets there (9-10), how it is built (11-16), how it is checked (17-18), how we work (19-23). Two things made this safe rather than free. It is a permutation, not a compaction, so the renames go through temporary names -- otherwise two files want one slot and one is lost. And the reference rewrite is a single simultaneous pass, because almost every number moved into a slot another number was vacating; replacing one at a time would have cascaded and pointed things at the wrong record while still resolving. Verified: 284 [ADR NNNN](path) links across the repository, all with matching text and target. The ordering principle is now stated in 19 rather than left implicit -- the repository already said "the numbering is the flow" about its folders, and there was no reason for the records to be the exception.
60 lines
2.4 KiB
Markdown
60 lines
2.4 KiB
Markdown
---
|
|
status: open
|
|
opened: 2026-08-28
|
|
located-in: [hal]
|
|
fixed-by:
|
|
amended-design:
|
|
---
|
|
|
|
# 008 — The documented automatic node rescue does not exist
|
|
|
|
## Symptom
|
|
|
|
The mesh's documentation describes an automatic node rescue: a node that fails is recovered
|
|
without anybody intervening. **Nothing implements it.**
|
|
|
|
Found incidentally while investigating supervision
|
|
([research 003](../../01-RESEARCH/003-service-supervision/00-overview.md)), which counted what
|
|
actually supervises what:
|
|
|
|
- **no unit declares `OnFailure=`**, so nothing runs when a unit gives up;
|
|
- **nothing calls the rescue script on a timer**, so it runs only when a person runs it.
|
|
|
|
The script exists. The thing that would invoke it does not.
|
|
|
|
## Why this is worse than having no rescue
|
|
|
|
A rescue nobody wrote is a gap somebody can see. A rescue that is *documented* and absent is a
|
|
gap nobody looks for, because the documentation says it is covered — and it is read exactly when
|
|
a node has failed and somebody is deciding whether to intervene.
|
|
|
|
This is `how-we-build` §5 in its most expensive form: *an unenforced rule is indistinguishable
|
|
from a wrong one, and costs more, because people believe it.* Here the belief is that a failed
|
|
node recovers itself.
|
|
|
|
## Scope
|
|
|
|
**The as-is only.** The design being built has a different answer:
|
|
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) puts recovery
|
|
in a launcher that supervises the host, and that recovery is tested — 32 assertions, each
|
|
confirmed to fail when the behaviour is removed.
|
|
|
|
So this issue is about the mesh that runs **now**, and it has two possible resolutions rather
|
|
than one:
|
|
|
|
1. **Implement it** — an `OnFailure=` and a timer — if node rescue is wanted before the new host
|
|
reaches the fleet.
|
|
2. **Delete the documentation** — and say plainly that a failed node needs a person, which is
|
|
what is true today.
|
|
|
|
**Either is honest. Leaving it as it is, is not.** The choice turns on how far away the new host
|
|
is, which is a scheduling question rather than a technical one.
|
|
|
|
## What it would take to be sure
|
|
|
|
Read back rather than assumed
|
|
([ADR 0018](../../02-DECISIONS/0018-a-picture-is-read-from-what-runs.md)): list every unit on a
|
|
node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above
|
|
came from reading the repository, and confirming it against a running node is the difference
|
|
between *no unit declares this* and *no unit in the source declares this*.
|