Files
hq/04-ISSUES/004-certificate-issuance-targets-production/00-report.md
jschoubben c192810fba 004 and 008 resolved; 006 narrowed to the decision it actually needs
**004 — certificate issuance.** The resolver declared no authority at
all, so the client fell to its built-in production default: there was no
setting set wrongly, there was no setting. It is now a node property
defaulting to staging, which answers the first open question. Staging by
default rather than production-with-an-override, because the alternative
leaves the safe path depending on remembering to opt out of it — 005's
lesson, in a second place. The rollout is ordered and the order is the
dangerous part; recorded, not performed.

**008 — node rescue.** Read back from running nodes as the report asked,
and one of its own claims was wrong in a way that matters: the health
timer does exist and does fire. It simply never calls the rescue script.
A trigger that exists and does not do what the script claims survives a
halfway check, which makes it worse than the absence the report
described. Resolved by making the documentation true, not by
implementing rescue — the replacement host already supervises recovery,
and wiring unattended restart into the fleet being retired is a
deliberate decision rather than a tidy-up. Two "self-healing" claims
narrowed to what they actually do.

**006 — deliberately not closed.** Re-checked today: the indexing still
does not exist. What is gone is the reason it was an issue — the claim
is no longer load-bearing, because the README names the gap and the
decision's reasoning never invoked indexing. A signpost now points here
from the knowledge base, and was measured rather than assumed: it is
reachable, it is not surfacing. Closing it while the indexing does not
exist would be this repository's own named failure, one folder from
where it names it.
2026-08-31 15:20:11 +02:00

76 lines
3.5 KiB
Markdown

---
status: resolved
opened: 2026-08-22
located-in: [hal]
fixed-by: hal — ACME_CA_SERVER selects the authority, and defaults to staging
amended-design:
---
# 004 — Certificate issuance always targets the authority's production endpoint
## Symptom
The reverse proxy sets no staging endpoint for its certificate resolver. Issuance therefore
goes to the public authority's production endpoint in every case, including experiments.
## Why this matters
Production issuance is rate-limited per domain and per account. Every certificate experiment on
a real node consumes quota that is not replenished quickly, and exhausting it is not
recoverable by retrying — it removes the ability to issue a certificate anyone actually needs.
The consequence lands hardest on exactly the work most likely to iterate: standing up a new
node, changing how names resolve, or testing the lab's certificate authority split
([ADR 0016](../../02-DECISIONS/0016-the-lab.md)).
## Evidence
- The resolver configuration declares no staging endpoint.
- Observed 2026-08-22.
## Open questions
- Should the endpoint be a node property — production for nodes serving real traffic, staging
everywhere else — rather than a fixed proxy setting?
- The lab issues its own certificates and so does not consume public quota at all. Does that
make this a problem only for experiments run outside the lab, and therefore an argument for
running them inside it?
## Resolution
*2026-08-31.* **The authority is now selectable, and the default is staging.**
Confirmed first: the resolver declared no `caServer` at all, so the client fell to its built-in
default of the public authority's production endpoint. There was no setting to change — not a
setting set wrongly.
`ACME_CA_SERVER` now names the authority, and **defaults to staging**. That answers the first
open question in the affirmative: it is a node property, and the node that serves real traffic is
the one that states so.
**Why staging is the default rather than the safe-looking alternative.** Defaulting to production
and documenting the override would leave the safe path depending on somebody remembering to opt
out of it — on exactly the work most likely to iterate. That is the same fault as
[005](../005-pipeline-test-harness-unbuildable/00-report.md): *remembering is not a mechanism.*
A staging certificate is trusted by no browser, so the mistake announces itself in the first
request rather than a fortnight later when the quota is gone. **The failure that is loud and
immediate is the cheaper one**, and quota exhaustion is neither.
**The second open question is answered too, and it is not the whole answer.** The lab issues its
own certificates and consumes no public quota, so experiments belong there. But "run it in the
lab" is advice, and the nodes this issue is about are the ones outside it — the default is what
protects those.
### The rollout is ordered, and the order is the dangerous part
Both public-serving nodes were checked: neither set the variable. Applying the change without
pinning them first would re-issue their public certificates from an untrusted authority and break
TLS for every hosted service — **this issue's own failure mode, arriving through its fix.** Pin
first, then merge. The commit carries the exact commands.
## Deliberately not done
**Nothing was changed on a running node.** Pinning the public nodes and regenerating their
environment restarts the reverse proxy that fronts every hosted service, and that is an operator's
decision rather than a fix's.