004 and 008 resolved; 006 narrowed to the decision it actually needs
**004 — certificate issuance.** The resolver declared no authority at all, so the client fell to its built-in production default: there was no setting set wrongly, there was no setting. It is now a node property defaulting to staging, which answers the first open question. Staging by default rather than production-with-an-override, because the alternative leaves the safe path depending on remembering to opt out of it — 005's lesson, in a second place. The rollout is ordered and the order is the dangerous part; recorded, not performed. **008 — node rescue.** Read back from running nodes as the report asked, and one of its own claims was wrong in a way that matters: the health timer does exist and does fire. It simply never calls the rescue script. A trigger that exists and does not do what the script claims survives a halfway check, which makes it worse than the absence the report described. Resolved by making the documentation true, not by implementing rescue — the replacement host already supervises recovery, and wiring unattended restart into the fleet being retired is a deliberate decision rather than a tidy-up. Two "self-healing" claims narrowed to what they actually do. **006 — deliberately not closed.** Re-checked today: the indexing still does not exist. What is gone is the reason it was an issue — the claim is no longer load-bearing, because the README names the gap and the decision's reasoning never invoked indexing. A signpost now points here from the knowledge base, and was measured rather than assumed: it is reachable, it is not surfacing. Closing it while the indexing does not exist would be this repository's own named failure, one folder from where it names it.
This commit is contained in:
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: open
|
||||
status: resolved
|
||||
opened: 2026-08-28
|
||||
located-in: [hal]
|
||||
fixed-by:
|
||||
fixed-by: hal — the script says it is manual, because it is
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -57,3 +57,46 @@ Read back rather than assumed
|
||||
node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above
|
||||
came from reading the repository, and confirming it against a running node is the difference
|
||||
between *no unit declares this* and *no unit in the source declares this*.
|
||||
|
||||
## Resolution
|
||||
|
||||
*2026-08-31.* **Resolved the second way: the documentation now says what is true.**
|
||||
|
||||
### Read back from running nodes, and the finding sharpened
|
||||
|
||||
This report was written from the repository. Checked against three running nodes, as the section
|
||||
above asks — and one of its own claims was wrong in a way that matters:
|
||||
|
||||
| Claim | Verified |
|
||||
|---|---|
|
||||
| no unit declares `OnFailure=` | **true**, and in *both* scopes — of 604 system units, the only two are stock (`snapd`, `local-fs.target`); across ~30 user units, none |
|
||||
| nothing calls the rescue script on a timer | **the timer exists.** `hal-health.timer` fires every 30 minutes |
|
||||
| the rescue is not actually triggered | **true** — the *deployed* `hal-health.sh`, not only the repository's copy, contains no call to it |
|
||||
|
||||
**The trigger exists and does not do the thing the script says it does.** That is worse than the
|
||||
absence this report described, because it survives a halfway check: somebody verifying "is there a
|
||||
health timer?" finds one, and stops.
|
||||
|
||||
The precise falsehood was a single line in the rescue script — *triggered automatically by
|
||||
`hal-health.timer` when the mesh service is failed.* It is now the opposite sentence, and says
|
||||
what does trigger it.
|
||||
|
||||
**Two further claims were found and narrowed.** Documentation in two places called the mesh
|
||||
*self-healing* — once for peers refreshing a tool cache, once for nodes auto-updating on a repo
|
||||
change. Both behaviours are real and neither is healing. **A phrase that overstates by a category
|
||||
is read in the crisis it describes**, and "the mesh is self-healing" is exactly the sentence that
|
||||
stops somebody intervening.
|
||||
|
||||
### Why not the first way
|
||||
|
||||
Implementing it was the other honest option, and it was not taken. The replacement host already
|
||||
supervises recovery and that recovery is tested; adding unattended download-and-restart behaviour
|
||||
to the fleet it is replacing is a change with real failure modes of its own — a rescue loop that
|
||||
thrashes is worse than a node that waits.
|
||||
|
||||
**This is the scheduling judgement this report said the choice turned on, and it is recorded
|
||||
rather than assumed.** If rescue is wanted on the current fleet before the new host reaches it,
|
||||
that is a decision to take deliberately: an `OnFailure=` and a call from the health script, both
|
||||
small, neither free.
|
||||
|
||||
Until then the documentation is true, which is the part that was costing something.
|
||||
|
||||
Reference in New Issue
Block a user