Reconcile: adopt initialization's consolidated HQ as canonical, re-home this session's new work #24
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: open
|
||||
status: resolved
|
||||
opened: 2026-08-22
|
||||
located-in: []
|
||||
fixed-by:
|
||||
located-in: [hal]
|
||||
fixed-by: hal — ACME_CA_SERVER selects the authority, and defaults to staging
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -35,3 +35,41 @@ node, changing how names resolve, or testing the lab's certificate authority spl
|
||||
- The lab issues its own certificates and so does not consume public quota at all. Does that
|
||||
make this a problem only for experiments run outside the lab, and therefore an argument for
|
||||
running them inside it?
|
||||
|
||||
## Resolution
|
||||
|
||||
*2026-08-31.* **The authority is now selectable, and the default is staging.**
|
||||
|
||||
Confirmed first: the resolver declared no `caServer` at all, so the client fell to its built-in
|
||||
default of the public authority's production endpoint. There was no setting to change — not a
|
||||
setting set wrongly.
|
||||
|
||||
`ACME_CA_SERVER` now names the authority, and **defaults to staging**. That answers the first
|
||||
open question in the affirmative: it is a node property, and the node that serves real traffic is
|
||||
the one that states so.
|
||||
|
||||
**Why staging is the default rather than the safe-looking alternative.** Defaulting to production
|
||||
and documenting the override would leave the safe path depending on somebody remembering to opt
|
||||
out of it — on exactly the work most likely to iterate. That is the same fault as
|
||||
[005](../005-pipeline-test-harness-unbuildable/00-report.md): *remembering is not a mechanism.*
|
||||
A staging certificate is trusted by no browser, so the mistake announces itself in the first
|
||||
request rather than a fortnight later when the quota is gone. **The failure that is loud and
|
||||
immediate is the cheaper one**, and quota exhaustion is neither.
|
||||
|
||||
**The second open question is answered too, and it is not the whole answer.** The lab issues its
|
||||
own certificates and consumes no public quota, so experiments belong there. But "run it in the
|
||||
lab" is advice, and the nodes this issue is about are the ones outside it — the default is what
|
||||
protects those.
|
||||
|
||||
### The rollout is ordered, and the order is the dangerous part
|
||||
|
||||
Both public-serving nodes were checked: neither set the variable. Applying the change without
|
||||
pinning them first would re-issue their public certificates from an untrusted authority and break
|
||||
TLS for every hosted service — **this issue's own failure mode, arriving through its fix.** Pin
|
||||
first, then merge. The commit carries the exact commands.
|
||||
|
||||
## Deliberately not done
|
||||
|
||||
**Nothing was changed on a running node.** Pinning the public nodes and regenerating their
|
||||
environment restarts the reverse proxy that fronts every hosted service, and that is an operator's
|
||||
decision rather than a fix's.
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
---
|
||||
status: open
|
||||
status: located
|
||||
opened: 2026-08-23
|
||||
located-in: [hal, hq]
|
||||
fixed-by:
|
||||
@@ -88,3 +88,51 @@ consults Nox — the answer is yes and the original promise holds.
|
||||
|
||||
That is a design question for Nox, not a defect in this repository, and it should be settled
|
||||
before ADR 0019 is treated as answered.
|
||||
|
||||
## Where this stands
|
||||
|
||||
*2026-08-31. Re-checked, and deliberately not closed.*
|
||||
|
||||
**The indexing still does not exist.** Two searches today, against both the symptom-indexed
|
||||
memory and the structured archive, using a decision record's full title and a distinctive phrase
|
||||
from a design document: no results, no partial match, no stale copy. The symptom in this report
|
||||
is unchanged.
|
||||
|
||||
**But the part that made it an issue is gone.** This report's argument was that the claim was
|
||||
*load-bearing* — that a decision rested on a mechanism nobody had checked. It no longer rests on
|
||||
it. The README now names the gap in the place the claim used to sit, and says it is left standing
|
||||
rather than quietly reworded. The decision record that separates this repository does not invoke
|
||||
indexing at all; its reasoning is cadence, reviewers, and scope, none of which depend on it.
|
||||
|
||||
So what remains is not a false claim. It is an unbuilt capability and an open design question,
|
||||
and those are different things.
|
||||
|
||||
### What was done
|
||||
|
||||
**A signpost, in the knowledge base, pointing here** — what lives in this repository, which
|
||||
folders hold what, and when to come looking rather than search there. Explicitly a pointer and
|
||||
not a copy: a derived copy drifts, and the enforced copy wins while the reasoned one quietly
|
||||
stops being true.
|
||||
|
||||
**It was tested, and it half works.** A search for *design records, decisions, repository* returns
|
||||
it. A search phrased the way somebody would actually ask — *why is the mesh built this way* —
|
||||
returns nothing, because the store matches terms rather than meaning.
|
||||
|
||||
That is this report's own distinction, confirmed by measurement rather than argued: **a signpost
|
||||
is reachable, it is not surfacing.** Someone who suspects the answer exists will now find it.
|
||||
Someone debugging an error, with no reason to think this repository knows anything about their
|
||||
symptom, still will not.
|
||||
|
||||
### Why it stays open
|
||||
|
||||
The question this report narrows to is unchanged and unanswered:
|
||||
|
||||
> When a symptom is searched and the answer happens to live in a design document or a decision
|
||||
> record here, does the searcher find it without already suspecting it exists?
|
||||
|
||||
Today: **no.** Closing this means choosing between a one-way sync into the knowledge base and an
|
||||
agent that reads this repository and contributes to a symptom search — and that is a decision
|
||||
about how the knowledge system works, not a defect to be fixed quietly.
|
||||
|
||||
**Marking it resolved while the indexing does not exist would be the failure this repository was
|
||||
created to name**, one folder away from where it names it.
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: open
|
||||
status: resolved
|
||||
opened: 2026-08-28
|
||||
located-in: [hal]
|
||||
fixed-by:
|
||||
fixed-by: hal — the script says it is manual, because it is
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -57,3 +57,46 @@ Read back rather than assumed
|
||||
node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above
|
||||
came from reading the repository, and confirming it against a running node is the difference
|
||||
between *no unit declares this* and *no unit in the source declares this*.
|
||||
|
||||
## Resolution
|
||||
|
||||
*2026-08-31.* **Resolved the second way: the documentation now says what is true.**
|
||||
|
||||
### Read back from running nodes, and the finding sharpened
|
||||
|
||||
This report was written from the repository. Checked against three running nodes, as the section
|
||||
above asks — and one of its own claims was wrong in a way that matters:
|
||||
|
||||
| Claim | Verified |
|
||||
|---|---|
|
||||
| no unit declares `OnFailure=` | **true**, and in *both* scopes — of 604 system units, the only two are stock (`snapd`, `local-fs.target`); across ~30 user units, none |
|
||||
| nothing calls the rescue script on a timer | **the timer exists.** `hal-health.timer` fires every 30 minutes |
|
||||
| the rescue is not actually triggered | **true** — the *deployed* `hal-health.sh`, not only the repository's copy, contains no call to it |
|
||||
|
||||
**The trigger exists and does not do the thing the script says it does.** That is worse than the
|
||||
absence this report described, because it survives a halfway check: somebody verifying "is there a
|
||||
health timer?" finds one, and stops.
|
||||
|
||||
The precise falsehood was a single line in the rescue script — *triggered automatically by
|
||||
`hal-health.timer` when the mesh service is failed.* It is now the opposite sentence, and says
|
||||
what does trigger it.
|
||||
|
||||
**Two further claims were found and narrowed.** Documentation in two places called the mesh
|
||||
*self-healing* — once for peers refreshing a tool cache, once for nodes auto-updating on a repo
|
||||
change. Both behaviours are real and neither is healing. **A phrase that overstates by a category
|
||||
is read in the crisis it describes**, and "the mesh is self-healing" is exactly the sentence that
|
||||
stops somebody intervening.
|
||||
|
||||
### Why not the first way
|
||||
|
||||
Implementing it was the other honest option, and it was not taken. The replacement host already
|
||||
supervises recovery and that recovery is tested; adding unattended download-and-restart behaviour
|
||||
to the fleet it is replacing is a change with real failure modes of its own — a rescue loop that
|
||||
thrashes is worse than a node that waits.
|
||||
|
||||
**This is the scheduling judgement this report said the choice turned on, and it is recorded
|
||||
rather than assumed.** If rescue is wanted on the current fleet before the new host reaches it,
|
||||
that is a decision to take deliberately: an `OnFailure=` and a call from the health script, both
|
||||
small, neither free.
|
||||
|
||||
Until then the documentation is true, which is the part that was costing something.
|
||||
|
||||
Reference in New Issue
Block a user