**004 — certificate issuance.** The resolver declared no authority at all, so the client fell to its built-in production default: there was no setting set wrongly, there was no setting. It is now a node property defaulting to staging, which answers the first open question. Staging by default rather than production-with-an-override, because the alternative leaves the safe path depending on remembering to opt out of it — 005's lesson, in a second place. The rollout is ordered and the order is the dangerous part; recorded, not performed. **008 — node rescue.** Read back from running nodes as the report asked, and one of its own claims was wrong in a way that matters: the health timer does exist and does fire. It simply never calls the rescue script. A trigger that exists and does not do what the script claims survives a halfway check, which makes it worse than the absence the report described. Resolved by making the documentation true, not by implementing rescue — the replacement host already supervises recovery, and wiring unattended restart into the fleet being retired is a deliberate decision rather than a tidy-up. Two "self-healing" claims narrowed to what they actually do. **006 — deliberately not closed.** Re-checked today: the indexing still does not exist. What is gone is the reason it was an issue — the claim is no longer load-bearing, because the README names the gap and the decision's reasoning never invoked indexing. A signpost now points here from the knowledge base, and was measured rather than assumed: it is reachable, it is not surfacing. Closing it while the indexing does not exist would be this repository's own named failure, one folder from where it names it.
103 lines
4.8 KiB
Markdown
103 lines
4.8 KiB
Markdown
---
|
|
status: resolved
|
|
opened: 2026-08-28
|
|
located-in: [hal]
|
|
fixed-by: hal — the script says it is manual, because it is
|
|
amended-design:
|
|
---
|
|
|
|
# 008 — The documented automatic node rescue does not exist
|
|
|
|
## Symptom
|
|
|
|
The mesh's documentation describes an automatic node rescue: a node that fails is recovered
|
|
without anybody intervening. **Nothing implements it.**
|
|
|
|
Found incidentally while investigating supervision
|
|
([research 003](../../01-RESEARCH/003-service-supervision/00-overview.md)), which counted what
|
|
actually supervises what:
|
|
|
|
- **no unit declares `OnFailure=`**, so nothing runs when a unit gives up;
|
|
- **nothing calls the rescue script on a timer**, so it runs only when a person runs it.
|
|
|
|
The script exists. The thing that would invoke it does not.
|
|
|
|
## Why this is worse than having no rescue
|
|
|
|
A rescue nobody wrote is a gap somebody can see. A rescue that is *documented* and absent is a
|
|
gap nobody looks for, because the documentation says it is covered — and it is read exactly when
|
|
a node has failed and somebody is deciding whether to intervene.
|
|
|
|
This is `how-we-build` §5 in its most expensive form: *an unenforced rule is indistinguishable
|
|
from a wrong one, and costs more, because people believe it.* Here the belief is that a failed
|
|
node recovers itself.
|
|
|
|
## Scope
|
|
|
|
**The as-is only.** The design being built has a different answer:
|
|
[ADR 0005](../../02-DECISIONS/0005-the-node-host.md) puts recovery
|
|
in a launcher that supervises the host, and that recovery is tested — 32 assertions, each
|
|
confirmed to fail when the behaviour is removed.
|
|
|
|
So this issue is about the mesh that runs **now**, and it has two possible resolutions rather
|
|
than one:
|
|
|
|
1. **Implement it** — an `OnFailure=` and a timer — if node rescue is wanted before the new host
|
|
reaches the fleet.
|
|
2. **Delete the documentation** — and say plainly that a failed node needs a person, which is
|
|
what is true today.
|
|
|
|
**Either is honest. Leaving it as it is, is not.** The choice turns on how far away the new host
|
|
is, which is a scheduling question rather than a technical one.
|
|
|
|
## What it would take to be sure
|
|
|
|
Read back rather than assumed
|
|
([ADR 0018](../../02-DECISIONS/0018-a-picture-is-read-from-what-runs.md)): list every unit on a
|
|
node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above
|
|
came from reading the repository, and confirming it against a running node is the difference
|
|
between *no unit declares this* and *no unit in the source declares this*.
|
|
|
|
## Resolution
|
|
|
|
*2026-08-31.* **Resolved the second way: the documentation now says what is true.**
|
|
|
|
### Read back from running nodes, and the finding sharpened
|
|
|
|
This report was written from the repository. Checked against three running nodes, as the section
|
|
above asks — and one of its own claims was wrong in a way that matters:
|
|
|
|
| Claim | Verified |
|
|
|---|---|
|
|
| no unit declares `OnFailure=` | **true**, and in *both* scopes — of 604 system units, the only two are stock (`snapd`, `local-fs.target`); across ~30 user units, none |
|
|
| nothing calls the rescue script on a timer | **the timer exists.** `hal-health.timer` fires every 30 minutes |
|
|
| the rescue is not actually triggered | **true** — the *deployed* `hal-health.sh`, not only the repository's copy, contains no call to it |
|
|
|
|
**The trigger exists and does not do the thing the script says it does.** That is worse than the
|
|
absence this report described, because it survives a halfway check: somebody verifying "is there a
|
|
health timer?" finds one, and stops.
|
|
|
|
The precise falsehood was a single line in the rescue script — *triggered automatically by
|
|
`hal-health.timer` when the mesh service is failed.* It is now the opposite sentence, and says
|
|
what does trigger it.
|
|
|
|
**Two further claims were found and narrowed.** Documentation in two places called the mesh
|
|
*self-healing* — once for peers refreshing a tool cache, once for nodes auto-updating on a repo
|
|
change. Both behaviours are real and neither is healing. **A phrase that overstates by a category
|
|
is read in the crisis it describes**, and "the mesh is self-healing" is exactly the sentence that
|
|
stops somebody intervening.
|
|
|
|
### Why not the first way
|
|
|
|
Implementing it was the other honest option, and it was not taken. The replacement host already
|
|
supervises recovery and that recovery is tested; adding unattended download-and-restart behaviour
|
|
to the fleet it is replacing is a change with real failure modes of its own — a rescue loop that
|
|
thrashes is worse than a node that waits.
|
|
|
|
**This is the scheduling judgement this report said the choice turned on, and it is recorded
|
|
rather than assumed.** If rescue is wanted on the current fleet before the new host reaches it,
|
|
that is a decision to take deliberately: an `OnFailure=` and a call from the health script, both
|
|
small, neither free.
|
|
|
|
Until then the documentation is true, which is the part that was costing something.
|