From c192810fba91c4466c913cdee9956dcd0a9ecdb9 Mon Sep 17 00:00:00 2001 From: jochen Date: Mon, 31 Aug 2026 15:20:11 +0200 Subject: [PATCH] 004 and 008 resolved; 006 narrowed to the decision it actually needs MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit **004 — certificate issuance.** The resolver declared no authority at all, so the client fell to its built-in production default: there was no setting set wrongly, there was no setting. It is now a node property defaulting to staging, which answers the first open question. Staging by default rather than production-with-an-override, because the alternative leaves the safe path depending on remembering to opt out of it — 005's lesson, in a second place. The rollout is ordered and the order is the dangerous part; recorded, not performed. **008 — node rescue.** Read back from running nodes as the report asked, and one of its own claims was wrong in a way that matters: the health timer does exist and does fire. It simply never calls the rescue script. A trigger that exists and does not do what the script claims survives a halfway check, which makes it worse than the absence the report described. Resolved by making the documentation true, not by implementing rescue — the replacement host already supervises recovery, and wiring unattended restart into the fleet being retired is a deliberate decision rather than a tidy-up. Two "self-healing" claims narrowed to what they actually do. **006 — deliberately not closed.** Re-checked today: the indexing still does not exist. What is gone is the reason it was an issue — the claim is no longer load-bearing, because the README names the gap and the decision's reasoning never invoked indexing. A signpost now points here from the knowledge base, and was measured rather than assumed: it is reachable, it is not surfacing. Closing it while the indexing does not exist would be this repository's own named failure, one folder from where it names it. --- .../00-report.md | 44 ++++++++++++++-- .../00-report.md | 50 ++++++++++++++++++- .../00-report.md | 47 ++++++++++++++++- 3 files changed, 135 insertions(+), 6 deletions(-) diff --git a/04-ISSUES/004-certificate-issuance-targets-production/00-report.md b/04-ISSUES/004-certificate-issuance-targets-production/00-report.md index 72317bd..f1dc315 100644 --- a/04-ISSUES/004-certificate-issuance-targets-production/00-report.md +++ b/04-ISSUES/004-certificate-issuance-targets-production/00-report.md @@ -1,8 +1,8 @@ --- -status: open +status: resolved opened: 2026-08-22 -located-in: [] -fixed-by: +located-in: [hal] +fixed-by: hal — ACME_CA_SERVER selects the authority, and defaults to staging amended-design: --- @@ -35,3 +35,41 @@ node, changing how names resolve, or testing the lab's certificate authority spl - The lab issues its own certificates and so does not consume public quota at all. Does that make this a problem only for experiments run outside the lab, and therefore an argument for running them inside it? + +## Resolution + +*2026-08-31.* **The authority is now selectable, and the default is staging.** + +Confirmed first: the resolver declared no `caServer` at all, so the client fell to its built-in +default of the public authority's production endpoint. There was no setting to change — not a +setting set wrongly. + +`ACME_CA_SERVER` now names the authority, and **defaults to staging**. That answers the first +open question in the affirmative: it is a node property, and the node that serves real traffic is +the one that states so. + +**Why staging is the default rather than the safe-looking alternative.** Defaulting to production +and documenting the override would leave the safe path depending on somebody remembering to opt +out of it — on exactly the work most likely to iterate. That is the same fault as +[005](../005-pipeline-test-harness-unbuildable/00-report.md): *remembering is not a mechanism.* +A staging certificate is trusted by no browser, so the mistake announces itself in the first +request rather than a fortnight later when the quota is gone. **The failure that is loud and +immediate is the cheaper one**, and quota exhaustion is neither. + +**The second open question is answered too, and it is not the whole answer.** The lab issues its +own certificates and consumes no public quota, so experiments belong there. But "run it in the +lab" is advice, and the nodes this issue is about are the ones outside it — the default is what +protects those. + +### The rollout is ordered, and the order is the dangerous part + +Both public-serving nodes were checked: neither set the variable. Applying the change without +pinning them first would re-issue their public certificates from an untrusted authority and break +TLS for every hosted service — **this issue's own failure mode, arriving through its fix.** Pin +first, then merge. The commit carries the exact commands. + +## Deliberately not done + +**Nothing was changed on a running node.** Pinning the public nodes and regenerating their +environment restarts the reverse proxy that fronts every hosted service, and that is an operator's +decision rather than a fix's. diff --git a/04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md b/04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md index a770566..56c75be 100644 --- a/04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md +++ b/04-ISSUES/006-hq-is-not-indexed-into-the-knowledge-base/00-report.md @@ -1,5 +1,5 @@ --- -status: open +status: located opened: 2026-08-23 located-in: [hal, hq] fixed-by: @@ -88,3 +88,51 @@ consults Nox — the answer is yes and the original promise holds. That is a design question for Nox, not a defect in this repository, and it should be settled before ADR 0019 is treated as answered. + +## Where this stands + +*2026-08-31. Re-checked, and deliberately not closed.* + +**The indexing still does not exist.** Two searches today, against both the symptom-indexed +memory and the structured archive, using a decision record's full title and a distinctive phrase +from a design document: no results, no partial match, no stale copy. The symptom in this report +is unchanged. + +**But the part that made it an issue is gone.** This report's argument was that the claim was +*load-bearing* — that a decision rested on a mechanism nobody had checked. It no longer rests on +it. The README now names the gap in the place the claim used to sit, and says it is left standing +rather than quietly reworded. The decision record that separates this repository does not invoke +indexing at all; its reasoning is cadence, reviewers, and scope, none of which depend on it. + +So what remains is not a false claim. It is an unbuilt capability and an open design question, +and those are different things. + +### What was done + +**A signpost, in the knowledge base, pointing here** — what lives in this repository, which +folders hold what, and when to come looking rather than search there. Explicitly a pointer and +not a copy: a derived copy drifts, and the enforced copy wins while the reasoned one quietly +stops being true. + +**It was tested, and it half works.** A search for *design records, decisions, repository* returns +it. A search phrased the way somebody would actually ask — *why is the mesh built this way* — +returns nothing, because the store matches terms rather than meaning. + +That is this report's own distinction, confirmed by measurement rather than argued: **a signpost +is reachable, it is not surfacing.** Someone who suspects the answer exists will now find it. +Someone debugging an error, with no reason to think this repository knows anything about their +symptom, still will not. + +### Why it stays open + +The question this report narrows to is unchanged and unanswered: + +> When a symptom is searched and the answer happens to live in a design document or a decision +> record here, does the searcher find it without already suspecting it exists? + +Today: **no.** Closing this means choosing between a one-way sync into the knowledge base and an +agent that reads this repository and contributes to a symptom search — and that is a decision +about how the knowledge system works, not a defect to be fixed quietly. + +**Marking it resolved while the indexing does not exist would be the failure this repository was +created to name**, one folder away from where it names it. diff --git a/04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md b/04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md index 51e3a10..8839275 100644 --- a/04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md +++ b/04-ISSUES/008-the-documented-node-rescue-does-not-exist/00-report.md @@ -1,8 +1,8 @@ --- -status: open +status: resolved opened: 2026-08-28 located-in: [hal] -fixed-by: +fixed-by: hal — the script says it is manual, because it is amended-design: --- @@ -57,3 +57,46 @@ Read back rather than assumed node and grep for `OnFailure=`; list every timer and check what each one calls. The finding above came from reading the repository, and confirming it against a running node is the difference between *no unit declares this* and *no unit in the source declares this*. + +## Resolution + +*2026-08-31.* **Resolved the second way: the documentation now says what is true.** + +### Read back from running nodes, and the finding sharpened + +This report was written from the repository. Checked against three running nodes, as the section +above asks — and one of its own claims was wrong in a way that matters: + +| Claim | Verified | +|---|---| +| no unit declares `OnFailure=` | **true**, and in *both* scopes — of 604 system units, the only two are stock (`snapd`, `local-fs.target`); across ~30 user units, none | +| nothing calls the rescue script on a timer | **the timer exists.** `hal-health.timer` fires every 30 minutes | +| the rescue is not actually triggered | **true** — the *deployed* `hal-health.sh`, not only the repository's copy, contains no call to it | + +**The trigger exists and does not do the thing the script says it does.** That is worse than the +absence this report described, because it survives a halfway check: somebody verifying "is there a +health timer?" finds one, and stops. + +The precise falsehood was a single line in the rescue script — *triggered automatically by +`hal-health.timer` when the mesh service is failed.* It is now the opposite sentence, and says +what does trigger it. + +**Two further claims were found and narrowed.** Documentation in two places called the mesh +*self-healing* — once for peers refreshing a tool cache, once for nodes auto-updating on a repo +change. Both behaviours are real and neither is healing. **A phrase that overstates by a category +is read in the crisis it describes**, and "the mesh is self-healing" is exactly the sentence that +stops somebody intervening. + +### Why not the first way + +Implementing it was the other honest option, and it was not taken. The replacement host already +supervises recovery and that recovery is tested; adding unattended download-and-restart behaviour +to the fleet it is replacing is a change with real failure modes of its own — a rescue loop that +thrashes is worse than a node that waits. + +**This is the scheduling judgement this report said the choice turned on, and it is recorded +rather than assumed.** If rescue is wanted on the current fleet before the new host reaches it, +that is a decision to take deliberately: an `OnFailure=` and a call from the health script, both +small, neither free. + +Until then the documentation is true, which is the part that was costing something.