--- status: resolved opened: 2026-08-28 located-in: [hal] fixed-by: hal — the script says it is manual, because it is amended-design: --- # 008 — The documented automatic node rescue does not exist ## Symptom The mesh's documentation describes an automatic node rescue: a node that fails is recovered without anybody intervening. **Nothing implements it.** Found incidentally while investigating supervision ([research 003](../../01-RESEARCH/003-service-supervision/00-overview.md)), which counted what actually supervises what: - **no unit declares `OnFailure=`**, so nothing runs when a unit gives up; - **nothing calls the rescue script on a timer**, so it runs only when a person runs it. The script exists. The thing that would invoke it does not. ## Why this is worse than having no rescue A rescue nobody wrote is a gap somebody can see. A rescue that is *documented* and absent is a gap nobody looks for, because the documentation says it is covered — and it is read exactly when a node has failed and somebody is deciding whether to intervene. This is `how-we-build` §5 in its most expensive form: *an unenforced rule is indistinguishable from a wrong one, and costs more, because people believe it.* Here the belief is that a failed node recovers itself. ## Scope **The as-is only.** The design being built has a different answer: [ADR 0005](../../02-DECISIONS/0005-the-node-host.md) puts recovery in a launcher that supervises the host, and that recovery is tested — 32 assertions, each confirmed to fail when the behaviour is removed. So this issue is about the mesh that runs **now**, and it has two possible resolutions rather than one: 1. **Implement it** — an `OnFailure=` and a timer — if node rescue is wanted before the new host reaches the fleet. 2. **Delete the documentation** — and say plainly that a failed node needs a person, which is what is true today. **Either is honest. Leaving it as it is, is not.** The choice turns on how far away the new host is, which is a scheduling question rather than a technical one. ## What it would take to be sure Read back rather than assumed ([ADR 0018](../../02-DECISIONS/0018-a-picture-is-read-from-what-runs.md)): list every unit on a node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above came from reading the repository, and confirming it against a running node is the difference between *no unit declares this* and *no unit in the source declares this*. ## Resolution *2026-08-31.* **Resolved the second way: the documentation now says what is true.** ### Read back from running nodes, and the finding sharpened This report was written from the repository. Checked against three running nodes, as the section above asks — and one of its own claims was wrong in a way that matters: | Claim | Verified | |---|---| | no unit declares `OnFailure=` | **true**, and in *both* scopes — of 604 system units, the only two are stock (`snapd`, `local-fs.target`); across ~30 user units, none | | nothing calls the rescue script on a timer | **the timer exists.** `hal-health.timer` fires every 30 minutes | | the rescue is not actually triggered | **true** — the *deployed* `hal-health.sh`, not only the repository's copy, contains no call to it | **The trigger exists and does not do the thing the script says it does.** That is worse than the absence this report described, because it survives a halfway check: somebody verifying "is there a health timer?" finds one, and stops. The precise falsehood was a single line in the rescue script — *triggered automatically by `hal-health.timer` when the mesh service is failed.* It is now the opposite sentence, and says what does trigger it. **Two further claims were found and narrowed.** Documentation in two places called the mesh *self-healing* — once for peers refreshing a tool cache, once for nodes auto-updating on a repo change. Both behaviours are real and neither is healing. **A phrase that overstates by a category is read in the crisis it describes**, and "the mesh is self-healing" is exactly the sentence that stops somebody intervening. ### Why not the first way Implementing it was the other honest option, and it was not taken. The replacement host already supervises recovery and that recovery is tested; adding unattended download-and-restart behaviour to the fleet it is replacing is a change with real failure modes of its own — a rescue loop that thrashes is worse than a node that waits. **This is the scheduling judgement this report said the choice turned on, and it is recorded rather than assumed.** If rescue is wanted on the current fleet before the new host reaches it, that is a decision to take deliberately: an `OnFailure=` and a call from the health script, both small, neither free. Until then the documentation is true, which is the part that was costing something.