004 and 008 resolved; 006 narrowed to the decision it actually needs

**004 — certificate issuance.** The resolver declared no authority at
all, so the client fell to its built-in production default: there was no
setting set wrongly, there was no setting. It is now a node property
defaulting to staging, which answers the first open question. Staging by
default rather than production-with-an-override, because the alternative
leaves the safe path depending on remembering to opt out of it — 005's
lesson, in a second place. The rollout is ordered and the order is the
dangerous part; recorded, not performed.

**008 — node rescue.** Read back from running nodes as the report asked,
and one of its own claims was wrong in a way that matters: the health
timer does exist and does fire. It simply never calls the rescue script.
A trigger that exists and does not do what the script claims survives a
halfway check, which makes it worse than the absence the report
described. Resolved by making the documentation true, not by
implementing rescue — the replacement host already supervises recovery,
and wiring unattended restart into the fleet being retired is a
deliberate decision rather than a tidy-up. Two "self-healing" claims
narrowed to what they actually do.

**006 — deliberately not closed.** Re-checked today: the indexing still
does not exist. What is gone is the reason it was an issue — the claim
is no longer load-bearing, because the README names the gap and the
decision's reasoning never invoked indexing. A signpost now points here
from the knowledge base, and was measured rather than assumed: it is
reachable, it is not surfacing. Closing it while the indexing does not
exist would be this repository's own named failure, one folder from
where it names it.
This commit is contained in:
2026-08-31 15:20:11 +02:00
parent 345bbe0552
commit c192810fba
3 changed files with 135 additions and 6 deletions
@@ -1,8 +1,8 @@
---
status: open
status: resolved
opened: 2026-08-22
located-in: []
fixed-by:
located-in: [hal]
fixed-by: hal — ACME_CA_SERVER selects the authority, and defaults to staging
amended-design:
---
@@ -35,3 +35,41 @@ node, changing how names resolve, or testing the lab's certificate authority spl
- The lab issues its own certificates and so does not consume public quota at all. Does that
make this a problem only for experiments run outside the lab, and therefore an argument for
running them inside it?
## Resolution
*2026-08-31.* **The authority is now selectable, and the default is staging.**
Confirmed first: the resolver declared no `caServer` at all, so the client fell to its built-in
default of the public authority's production endpoint. There was no setting to change — not a
setting set wrongly.
`ACME_CA_SERVER` now names the authority, and **defaults to staging**. That answers the first
open question in the affirmative: it is a node property, and the node that serves real traffic is
the one that states so.
**Why staging is the default rather than the safe-looking alternative.** Defaulting to production
and documenting the override would leave the safe path depending on somebody remembering to opt
out of it — on exactly the work most likely to iterate. That is the same fault as
[005](../005-pipeline-test-harness-unbuildable/00-report.md): *remembering is not a mechanism.*
A staging certificate is trusted by no browser, so the mistake announces itself in the first
request rather than a fortnight later when the quota is gone. **The failure that is loud and
immediate is the cheaper one**, and quota exhaustion is neither.
**The second open question is answered too, and it is not the whole answer.** The lab issues its
own certificates and consumes no public quota, so experiments belong there. But "run it in the
lab" is advice, and the nodes this issue is about are the ones outside it — the default is what
protects those.
### The rollout is ordered, and the order is the dangerous part
Both public-serving nodes were checked: neither set the variable. Applying the change without
pinning them first would re-issue their public certificates from an untrusted authority and break
TLS for every hosted service — **this issue's own failure mode, arriving through its fix.** Pin
first, then merge. The commit carries the exact commands.
## Deliberately not done
**Nothing was changed on a running node.** Pinning the public nodes and regenerating their
environment restarts the reverse proxy that fronts every hosted service, and that is an operator's
decision rather than a fix's.
@@ -1,5 +1,5 @@
---
status: open
status: located
opened: 2026-08-23
located-in: [hal, hq]
fixed-by:
@@ -88,3 +88,51 @@ consults Nox — the answer is yes and the original promise holds.
That is a design question for Nox, not a defect in this repository, and it should be settled
before ADR 0019 is treated as answered.
## Where this stands
*2026-08-31. Re-checked, and deliberately not closed.*
**The indexing still does not exist.** Two searches today, against both the symptom-indexed
memory and the structured archive, using a decision record's full title and a distinctive phrase
from a design document: no results, no partial match, no stale copy. The symptom in this report
is unchanged.
**But the part that made it an issue is gone.** This report's argument was that the claim was
*load-bearing* — that a decision rested on a mechanism nobody had checked. It no longer rests on
it. The README now names the gap in the place the claim used to sit, and says it is left standing
rather than quietly reworded. The decision record that separates this repository does not invoke
indexing at all; its reasoning is cadence, reviewers, and scope, none of which depend on it.
So what remains is not a false claim. It is an unbuilt capability and an open design question,
and those are different things.
### What was done
**A signpost, in the knowledge base, pointing here** — what lives in this repository, which
folders hold what, and when to come looking rather than search there. Explicitly a pointer and
not a copy: a derived copy drifts, and the enforced copy wins while the reasoned one quietly
stops being true.
**It was tested, and it half works.** A search for *design records, decisions, repository* returns
it. A search phrased the way somebody would actually ask — *why is the mesh built this way* —
returns nothing, because the store matches terms rather than meaning.
That is this report's own distinction, confirmed by measurement rather than argued: **a signpost
is reachable, it is not surfacing.** Someone who suspects the answer exists will now find it.
Someone debugging an error, with no reason to think this repository knows anything about their
symptom, still will not.
### Why it stays open
The question this report narrows to is unchanged and unanswered:
> When a symptom is searched and the answer happens to live in a design document or a decision
> record here, does the searcher find it without already suspecting it exists?
Today: **no.** Closing this means choosing between a one-way sync into the knowledge base and an
agent that reads this repository and contributes to a symptom search — and that is a decision
about how the knowledge system works, not a defect to be fixed quietly.
**Marking it resolved while the indexing does not exist would be the failure this repository was
created to name**, one folder away from where it names it.
@@ -1,8 +1,8 @@
---
status: open
status: resolved
opened: 2026-08-28
located-in: [hal]
fixed-by:
fixed-by: hal — the script says it is manual, because it is
amended-design:
---
@@ -57,3 +57,46 @@ Read back rather than assumed
node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above
came from reading the repository, and confirming it against a running node is the difference
between *no unit declares this* and *no unit in the source declares this*.
## Resolution
*2026-08-31.* **Resolved the second way: the documentation now says what is true.**
### Read back from running nodes, and the finding sharpened
This report was written from the repository. Checked against three running nodes, as the section
above asks — and one of its own claims was wrong in a way that matters:
| Claim | Verified |
|---|---|
| no unit declares `OnFailure=` | **true**, and in *both* scopes — of 604 system units, the only two are stock (`snapd`, `local-fs.target`); across ~30 user units, none |
| nothing calls the rescue script on a timer | **the timer exists.** `hal-health.timer` fires every 30 minutes |
| the rescue is not actually triggered | **true** — the *deployed* `hal-health.sh`, not only the repository's copy, contains no call to it |
**The trigger exists and does not do the thing the script says it does.** That is worse than the
absence this report described, because it survives a halfway check: somebody verifying "is there a
health timer?" finds one, and stops.
The precise falsehood was a single line in the rescue script — *triggered automatically by
`hal-health.timer` when the mesh service is failed.* It is now the opposite sentence, and says
what does trigger it.
**Two further claims were found and narrowed.** Documentation in two places called the mesh
*self-healing* — once for peers refreshing a tool cache, once for nodes auto-updating on a repo
change. Both behaviours are real and neither is healing. **A phrase that overstates by a category
is read in the crisis it describes**, and "the mesh is self-healing" is exactly the sentence that
stops somebody intervening.
### Why not the first way
Implementing it was the other honest option, and it was not taken. The replacement host already
supervises recovery and that recovery is tested; adding unattended download-and-restart behaviour
to the fleet it is replacing is a change with real failure modes of its own — a rescue loop that
thrashes is worse than a node that waits.
**This is the scheduling judgement this report said the choice turned on, and it is recorded
rather than assumed.** If rescue is wanted on the current fleet before the new host reaches it,
that is a decision to take deliberately: an `OnFailure=` and a call from the health script, both
small, neither free.
Until then the documentation is true, which is the part that was costing something.