Files
hq/04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md
T
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00

2.4 KiB

status, opened, located-in, fixed-by, amended-design
status opened located-in fixed-by amended-design
open 2026-08-28
hal

008 — The documented automatic node rescue does not exist

Symptom

The mesh's documentation describes an automatic node rescue: a node that fails is recovered without anybody intervening. Nothing implements it.

Found incidentally while investigating supervision (research 003), which counted what actually supervises what:

  • no unit declares OnFailure=, so nothing runs when a unit gives up;
  • nothing calls the rescue script on a timer, so it runs only when a person runs it.

The script exists. The thing that would invoke it does not.

Why this is worse than having no rescue

A rescue nobody wrote is a gap somebody can see. A rescue that is documented and absent is a gap nobody looks for, because the documentation says it is covered — and it is read exactly when a node has failed and somebody is deciding whether to intervene.

This is how-we-build §5 in its most expensive form: an unenforced rule is indistinguishable from a wrong one, and costs more, because people believe it. Here the belief is that a failed node recovers itself.

Scope

The as-is only. The design being built has a different answer: ADR 0005 puts recovery in a launcher that supervises the host, and that recovery is tested — 32 assertions, each confirmed to fail when the behaviour is removed.

So this issue is about the mesh that runs now, and it has two possible resolutions rather than one:

  1. Implement it — an OnFailure= and a timer — if node rescue is wanted before the new host reaches the fleet.
  2. Delete the documentation — and say plainly that a failed node needs a person, which is what is true today.

Either is honest. Leaving it as it is, is not. The choice turns on how far away the new host is, which is a scheduling question rather than a technical one.

What it would take to be sure

Read back rather than assumed (ADR 0018): list every unit on a node and grep for OnFailure=; list every timer and check what each one calls. The finding above came from reading the repository, and confirming it against a running node is the difference between no unit declares this and no unit in the source declares this.