004 and 008 resolved; 006 narrowed to the decision it actually needs
**004 — certificate issuance.** The resolver declared no authority at all, so the client fell to its built-in production default: there was no setting set wrongly, there was no setting. It is now a node property defaulting to staging, which answers the first open question. Staging by default rather than production-with-an-override, because the alternative leaves the safe path depending on remembering to opt out of it — 005's lesson, in a second place. The rollout is ordered and the order is the dangerous part; recorded, not performed. **008 — node rescue.** Read back from running nodes as the report asked, and one of its own claims was wrong in a way that matters: the health timer does exist and does fire. It simply never calls the rescue script. A trigger that exists and does not do what the script claims survives a halfway check, which makes it worse than the absence the report described. Resolved by making the documentation true, not by implementing rescue — the replacement host already supervises recovery, and wiring unattended restart into the fleet being retired is a deliberate decision rather than a tidy-up. Two "self-healing" claims narrowed to what they actually do. **006 — deliberately not closed.** Re-checked today: the indexing still does not exist. What is gone is the reason it was an issue — the claim is no longer load-bearing, because the README names the gap and the decision's reasoning never invoked indexing. A signpost now points here from the knowledge base, and was measured rather than assumed: it is reachable, it is not surfacing. Closing it while the indexing does not exist would be this repository's own named failure, one folder from where it names it.
This commit is contained in:
@@ -1,8 +1,8 @@
|
||||
---
|
||||
status: open
|
||||
status: resolved
|
||||
opened: 2026-08-22
|
||||
located-in: []
|
||||
fixed-by:
|
||||
located-in: [hal]
|
||||
fixed-by: hal — ACME_CA_SERVER selects the authority, and defaults to staging
|
||||
amended-design:
|
||||
---
|
||||
|
||||
@@ -35,3 +35,41 @@ node, changing how names resolve, or testing the lab's certificate authority spl
|
||||
- The lab issues its own certificates and so does not consume public quota at all. Does that
|
||||
make this a problem only for experiments run outside the lab, and therefore an argument for
|
||||
running them inside it?
|
||||
|
||||
## Resolution
|
||||
|
||||
*2026-08-31.* **The authority is now selectable, and the default is staging.**
|
||||
|
||||
Confirmed first: the resolver declared no `caServer` at all, so the client fell to its built-in
|
||||
default of the public authority's production endpoint. There was no setting to change — not a
|
||||
setting set wrongly.
|
||||
|
||||
`ACME_CA_SERVER` now names the authority, and **defaults to staging**. That answers the first
|
||||
open question in the affirmative: it is a node property, and the node that serves real traffic is
|
||||
the one that states so.
|
||||
|
||||
**Why staging is the default rather than the safe-looking alternative.** Defaulting to production
|
||||
and documenting the override would leave the safe path depending on somebody remembering to opt
|
||||
out of it — on exactly the work most likely to iterate. That is the same fault as
|
||||
[005](../005-pipeline-test-harness-unbuildable/00-report.md): *remembering is not a mechanism.*
|
||||
A staging certificate is trusted by no browser, so the mistake announces itself in the first
|
||||
request rather than a fortnight later when the quota is gone. **The failure that is loud and
|
||||
immediate is the cheaper one**, and quota exhaustion is neither.
|
||||
|
||||
**The second open question is answered too, and it is not the whole answer.** The lab issues its
|
||||
own certificates and consumes no public quota, so experiments belong there. But "run it in the
|
||||
lab" is advice, and the nodes this issue is about are the ones outside it — the default is what
|
||||
protects those.
|
||||
|
||||
### The rollout is ordered, and the order is the dangerous part
|
||||
|
||||
Both public-serving nodes were checked: neither set the variable. Applying the change without
|
||||
pinning them first would re-issue their public certificates from an untrusted authority and break
|
||||
TLS for every hosted service — **this issue's own failure mode, arriving through its fix.** Pin
|
||||
first, then merge. The commit carries the exact commands.
|
||||
|
||||
## Deliberately not done
|
||||
|
||||
**Nothing was changed on a running node.** Pinning the public nodes and regenerating their
|
||||
environment restarts the reverse proxy that fronts every hosted service, and that is an operator's
|
||||
decision rather than a fix's.
|
||||
|
||||
Reference in New Issue
Block a user