Files
hq/04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md
T
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00

60 lines
2.4 KiB
Markdown

---
status: open
opened: 2026-08-28
located-in: [hal]
fixed-by:
amended-design:
---
# 008 — The documented automatic node rescue does not exist
## Symptom
The mesh's documentation describes an automatic node rescue: a node that fails is recovered
without anybody intervening. **Nothing implements it.**
Found incidentally while investigating supervision
([research 003](../../01-RESEARCH/003-service-supervision/00-overview.md)), which counted what
actually supervises what:
- **no unit declares `OnFailure=`**, so nothing runs when a unit gives up;
- **nothing calls the rescue script on a timer**, so it runs only when a person runs it.
The script exists. The thing that would invoke it does not.
## Why this is worse than having no rescue
A rescue nobody wrote is a gap somebody can see. A rescue that is *documented* and absent is a
gap nobody looks for, because the documentation says it is covered — and it is read exactly when
a node has failed and somebody is deciding whether to intervene.
This is `how-we-build` §5 in its most expensive form: *an unenforced rule is indistinguishable
from a wrong one, and costs more, because people believe it.* Here the belief is that a failed
node recovers itself.
## Scope
**The as-is only.** The design being built has a different answer:
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) puts recovery
in a launcher that supervises the host, and that recovery is tested — 32 assertions, each
confirmed to fail when the behaviour is removed.
So this issue is about the mesh that runs **now**, and it has two possible resolutions rather
than one:
1. **Implement it** — an `OnFailure=` and a timer — if node rescue is wanted before the new host
reaches the fleet.
2. **Delete the documentation** — and say plainly that a failed node needs a person, which is
what is true today.
**Either is honest. Leaving it as it is, is not.** The choice turns on how far away the new host
is, which is a scheduling question rather than a technical one.
## What it would take to be sure
Read back rather than assumed
([ADR 0018](../../02-DECISIONS/0018-a-picture-is-read-from-what-runs.md)): list every unit on a
node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above
came from reading the repository, and confirming it against a running node is the difference
between *no unit declares this* and *no unit in the source declares this*.