Files
hq/04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md
T
jschoubben c192810fba 004 and 008 resolved; 006 narrowed to the decision it actually needs
**004 — certificate issuance.** The resolver declared no authority at
all, so the client fell to its built-in production default: there was no
setting set wrongly, there was no setting. It is now a node property
defaulting to staging, which answers the first open question. Staging by
default rather than production-with-an-override, because the alternative
leaves the safe path depending on remembering to opt out of it — 005's
lesson, in a second place. The rollout is ordered and the order is the
dangerous part; recorded, not performed.

**008 — node rescue.** Read back from running nodes as the report asked,
and one of its own claims was wrong in a way that matters: the health
timer does exist and does fire. It simply never calls the rescue script.
A trigger that exists and does not do what the script claims survives a
halfway check, which makes it worse than the absence the report
described. Resolved by making the documentation true, not by
implementing rescue — the replacement host already supervises recovery,
and wiring unattended restart into the fleet being retired is a
deliberate decision rather than a tidy-up. Two "self-healing" claims
narrowed to what they actually do.

**006 — deliberately not closed.** Re-checked today: the indexing still
does not exist. What is gone is the reason it was an issue — the claim
is no longer load-bearing, because the README names the gap and the
decision's reasoning never invoked indexing. A signpost now points here
from the knowledge base, and was measured rather than assumed: it is
reachable, it is not surfacing. Closing it while the indexing does not
exist would be this repository's own named failure, one folder from
where it names it.
2026-08-31 15:20:11 +02:00

103 lines
4.8 KiB
Markdown

---
status: resolved
opened: 2026-08-28
located-in: [hal]
fixed-by: hal — the script says it is manual, because it is
amended-design:
---
# 008 — The documented automatic node rescue does not exist
## Symptom
The mesh's documentation describes an automatic node rescue: a node that fails is recovered
without anybody intervening. **Nothing implements it.**
Found incidentally while investigating supervision
([research 003](../../01-RESEARCH/003-service-supervision/00-overview.md)), which counted what
actually supervises what:
- **no unit declares `OnFailure=`**, so nothing runs when a unit gives up;
- **nothing calls the rescue script on a timer**, so it runs only when a person runs it.
The script exists. The thing that would invoke it does not.
## Why this is worse than having no rescue
A rescue nobody wrote is a gap somebody can see. A rescue that is *documented* and absent is a
gap nobody looks for, because the documentation says it is covered — and it is read exactly when
a node has failed and somebody is deciding whether to intervene.
This is `how-we-build` §5 in its most expensive form: *an unenforced rule is indistinguishable
from a wrong one, and costs more, because people believe it.* Here the belief is that a failed
node recovers itself.
## Scope
**The as-is only.** The design being built has a different answer:
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) puts recovery
in a launcher that supervises the host, and that recovery is tested — 32 assertions, each
confirmed to fail when the behaviour is removed.
So this issue is about the mesh that runs **now**, and it has two possible resolutions rather
than one:
1. **Implement it** — an `OnFailure=` and a timer — if node rescue is wanted before the new host
reaches the fleet.
2. **Delete the documentation** — and say plainly that a failed node needs a person, which is
what is true today.
**Either is honest. Leaving it as it is, is not.** The choice turns on how far away the new host
is, which is a scheduling question rather than a technical one.
## What it would take to be sure
Read back rather than assumed
([ADR 0018](../../02-DECISIONS/0018-a-picture-is-read-from-what-runs.md)): list every unit on a
node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above
came from reading the repository, and confirming it against a running node is the difference
between *no unit declares this* and *no unit in the source declares this*.
## Resolution
*2026-08-31.* **Resolved the second way: the documentation now says what is true.**
### Read back from running nodes, and the finding sharpened
This report was written from the repository. Checked against three running nodes, as the section
above asks — and one of its own claims was wrong in a way that matters:
| Claim | Verified |
|---|---|
| no unit declares `OnFailure=` | **true**, and in *both* scopes — of 604 system units, the only two are stock (`snapd`, `local-fs.target`); across ~30 user units, none |
| nothing calls the rescue script on a timer | **the timer exists.** `hal-health.timer` fires every 30 minutes |
| the rescue is not actually triggered | **true** — the *deployed* `hal-health.sh`, not only the repository's copy, contains no call to it |
**The trigger exists and does not do the thing the script says it does.** That is worse than the
absence this report described, because it survives a halfway check: somebody verifying "is there a
health timer?" finds one, and stops.
The precise falsehood was a single line in the rescue script — *triggered automatically by
`hal-health.timer` when the mesh service is failed.* It is now the opposite sentence, and says
what does trigger it.
**Two further claims were found and narrowed.** Documentation in two places called the mesh
*self-healing* — once for peers refreshing a tool cache, once for nodes auto-updating on a repo
change. Both behaviours are real and neither is healing. **A phrase that overstates by a category
is read in the crisis it describes**, and "the mesh is self-healing" is exactly the sentence that
stops somebody intervening.
### Why not the first way
Implementing it was the other honest option, and it was not taken. The replacement host already
supervises recovery and that recovery is tested; adding unattended download-and-restart behaviour
to the fleet it is replacing is a change with real failure modes of its own — a rescue loop that
thrashes is worse than a node that waits.
**This is the scheduling judgement this report said the choice turned on, and it is recorded
rather than assumed.** If rescue is wanted on the current fleet before the new host reaches it,
that is a decision to take deliberately: an `OnFailure=` and a call from the health script, both
small, neither free.
Until then the documentation is true, which is the part that was costing something.