Files
hq/04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md
T
jschoubben e1febe8e0f Renumber the records 1 to 23
The consolidation left a sparse sequence -- 1, 4, 6, 7, 9, 10, 12, 15, 16, 18,
19, 25, 34, 35, 36, 37, 40, 42, 44, 45, 48, 49, 58 -- where the gaps were only
the archaeology of what used to be there.

Renumbered contiguously. Renames run in ascending order, so every target number
is already free and no two files ever collide.

The reference rewrite is one simultaneous pass rather than a sequence of
replacements. Numbers moved into slots other numbers were vacating -- the node
host went 37 to 16 while the lab went 16 to 9 -- so replacing one at a time
would have cascaded and silently pointed things at the wrong record.

Seven plain-text references survived the merges as prose rather than links,
naming records that no longer existed: the enrolment token, the link boundary,
what a declaration is, reachability, the repository structure. Each mapped to
the consolidated record that now holds it.

Verified rather than assumed: every [ADR NNNN](path) link now has matching text
and target, checked across the whole repository, and the checker passes.

Frontmatter `consolidates:` lists dropped -- they named records that are gone,
and each consolidated record already says in prose what it absorbed.
2026-08-28 23:28:34 +02:00

60 lines
2.4 KiB
Markdown

---
status: open
opened: 2026-08-28
located-in: [hal]
fixed-by:
amended-design:
---
# 008 — The documented automatic node rescue does not exist
## Symptom
The mesh's documentation describes an automatic node rescue: a node that fails is recovered
without anybody intervening. **Nothing implements it.**
Found incidentally while investigating supervision
([research 003](../../01-RESEARCH/003-service-supervision/00-overview.md)), which counted what
actually supervises what:
- **no unit declares `OnFailure=`**, so nothing runs when a unit gives up;
- **nothing calls the rescue script on a timer**, so it runs only when a person runs it.
The script exists. The thing that would invoke it does not.
## Why this is worse than having no rescue
A rescue nobody wrote is a gap somebody can see. A rescue that is *documented* and absent is a
gap nobody looks for, because the documentation says it is covered — and it is read exactly when
a node has failed and somebody is deciding whether to intervene.
This is `how-we-build` §5 in its most expensive form: *an unenforced rule is indistinguishable
from a wrong one, and costs more, because people believe it.* Here the belief is that a failed
node recovers itself.
## Scope
**The as-is only.** The design being built has a different answer:
[ADR 0016](../../02-DECISIONS/0016-the-node-host.md) puts recovery
in a launcher that supervises the host, and that recovery is tested — 32 assertions, each
confirmed to fail when the behaviour is removed.
So this issue is about the mesh that runs **now**, and it has two possible resolutions rather
than one:
1. **Implement it** — an `OnFailure=` and a timer — if node rescue is wanted before the new host
reaches the fleet.
2. **Delete the documentation** — and say plainly that a failed node needs a person, which is
what is true today.
**Either is honest. Leaving it as it is, is not.** The choice turns on how far away the new host
is, which is a scheduling question rather than a technical one.
## What it would take to be sure
Read back rather than assumed
([ADR 0014](../../02-DECISIONS/0014-a-picture-is-read-from-what-runs.md)): list every unit on a
node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above
came from reading the repository, and confirming it against a running node is the difference
between *no unit declares this* and *no unit in the source declares this*.