Merge pull request 'Issue 129 is live and reproduced, and needs three steps rather than one' (#198) from issue/129-and-what-reproducing-it-found into main
This commit was merged in pull request #198.
This commit is contained in:
@@ -1,5 +1,5 @@
|
||||
---
|
||||
status: located
|
||||
status: diagnosing
|
||||
opened: 2026-09-26
|
||||
located-in: [mesh-catalog ca-trust]
|
||||
---
|
||||
|
||||
@@ -1,45 +1,82 @@
|
||||
# Diagnosis
|
||||
# 129 — diagnosis
|
||||
|
||||
*2026-09-29.*
|
||||
*2026-09-30, from the workstation the issue was opened on.*
|
||||
|
||||
## What was ruled out
|
||||
## Still live, and reproduced exactly
|
||||
|
||||
**That something already carries the root and it is only misplaced.** It does not. The authority
|
||||
serves its root at a path beside its ACME directory, and the one thing that fetches it — the route
|
||||
proxy — puts it in a directory of its own and hands it to one program. Nothing has ever written
|
||||
into a machine's trust store. Measured on three converged machines: the anchors present are the
|
||||
predecessor's authority and a developer tool's local root, and on the machines where the
|
||||
predecessor's was deliberately removed, every internal name fails verification.
|
||||
The certificate is genuine, the authority is the mesh's, and nothing on the machine trusts it:
|
||||
|
||||
**That the private network could carry it, the way it carries the registry's trust.** That is what
|
||||
the report proposed, and it was rejected on consideration rather than on difficulty
|
||||
([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md), option 1): being on
|
||||
the network is what makes the registry *reachable* and is therefore the right trigger there, while
|
||||
trusting an authority is a separate fact from being able to reach it. The anchor's directory and
|
||||
the command that refreshes the extracted bundles are also one operating system's difference, which
|
||||
is the host's half of the mesh and not the controller's.
|
||||
```
|
||||
$ openssl s_client -connect keycloak.novox.internal:443 -servername keycloak.novox.internal
|
||||
subject=CN=keycloak.novox.internal
|
||||
issuer=O=Mesh Internal CA, CN=Mesh Internal CA Intermediate CA
|
||||
Verify return code: 20 (unable to get local issuer certificate)
|
||||
|
||||
**That it needs a new host resource type.** It does not, today. A file and a service say the whole
|
||||
of it, which the packet filter already proves. The primitive becomes the right answer when a second
|
||||
operating system is in play, and not before.
|
||||
$ curl https://keycloak.novox.internal/
|
||||
curl: (60) SSL certificate OpenSSL verify result: unable to get local issuer certificate (20)
|
||||
```
|
||||
|
||||
## Where it belongs
|
||||
`trust list` holds no entry for the mesh. The anchors present are two `mkcert` development roots and
|
||||
the predecessor's lab root — the report's account of the trust store is unchanged.
|
||||
|
||||
A module in the catalogue: it requires `internal-acme-ca`, fetches the root over the mesh's own
|
||||
network, installs it as a trust anchor, refreshes the machine's bundles, and — because being
|
||||
unassigned stops its unit, and stopping the unit is what undoes it — takes both away again.
|
||||
The public name on the same proxy verifies cleanly (`CN=keycloak.novox.be`, Let's Encrypt, return code
|
||||
0), which places the fault exactly where the report puts it: not in the proxy, not in the authority,
|
||||
and not in the certificate.
|
||||
|
||||
The owner is therefore `mesh-catalog`, module `ca-trust`, and nothing in the control plane.
|
||||
**The name matters, and the report's "every HTTPS name the mesh serves internally" is too broad.** The
|
||||
served internal name is `<label>.<node>.internal`. The hosts file also carries
|
||||
`<label>.<public-domain>.internal`, which nothing serves and which fails differently — that is
|
||||
[issue 157](../157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md), found while
|
||||
reproducing this, and it cost the first several minutes of this diagnosis.
|
||||
|
||||
## The module exists, and this stays open until a machine holds it
|
||||
## The authority serves what the module needs
|
||||
|
||||
*2026-09-29.* `ca-trust` is in the catalogue and merged
|
||||
([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md)), and what it renders
|
||||
is checked in the control plane's own suite: the script fetches from the authority it was bound to,
|
||||
and the unit runs it both ways.
|
||||
`step-ca` is up and healthy, and publishes `roots: /roots.pem` for both `acme-ca` and
|
||||
`internal-acme-ca`. That endpoint returns PEM:
|
||||
|
||||
**No machine has been assigned it, and nothing has verified a name because of it.** The bed written
|
||||
for that cannot run ([issue 146](../146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md)),
|
||||
and the live mesh has not been given the module. So the symptom this record opened on — every
|
||||
internal name failing verification on every machine — is still true everywhere, and the record stays
|
||||
`located` until it is not. Closing it on a module that exists would be closing it on an intention.
|
||||
```
|
||||
$ curl -sk https://127.0.0.1:9000/roots.pem
|
||||
-----BEGIN CERTIFICATE-----
|
||||
MIIBvzCCAWWgAwIBAgIQYa2CkdJk16JyG/dVy2qoEzAKBggqhkjOPQQDAjA+…
|
||||
```
|
||||
|
||||
So `${bound:internal-acme-ca:roots}` in the `ca-trust` module composes to a URL that returns a
|
||||
certificate, and the module's own check — refuse a body that is not one — is checking the right thing.
|
||||
|
||||
Worth recording because it was nearly filed as a defect: step-ca *also* serves `/roots`, which returns
|
||||
`{"crts":["-----BEGIN CERTIFICATE-----\n…"]}`. That body contains the literal text the module greps
|
||||
for, so had the module used `/roots` it would have installed JSON into the anchors directory and
|
||||
reported success. It does not use it. The guard is sound only because the published path is the PEM
|
||||
one, which is worth knowing before anybody changes either.
|
||||
|
||||
## What is actually in the way
|
||||
|
||||
**The module is not registered.** The report and the work plan both say it exists and is merged, which
|
||||
it does — `mesh-catalog modules/ca-trust`, on `main`. But the mesh has never been told about it:
|
||||
|
||||
```
|
||||
$ mesh-controller module list | grep -iE 'ca-trust|step-ca'
|
||||
step-ca 1 built 67f5f4cf on novox
|
||||
```
|
||||
|
||||
39 of the catalogue's 76 manifests are registered. `ca-trust` is one of the 37 that are not, so it
|
||||
cannot be assigned to anything — "assign it to one machine" has no module to name.
|
||||
|
||||
A dry run confirms it registers cleanly and needs no artifact built: it declares a directory, a script,
|
||||
a unit and a service, and no image.
|
||||
|
||||
```
|
||||
$ mesh-controller build <catalogue> --path modules/ca-trust --dry-run
|
||||
… the manifest, parsed and validated
|
||||
```
|
||||
|
||||
## So the remaining work is three steps, not one
|
||||
|
||||
1. **Register it** — build it from the catalogue, which pins nothing because it has no artifacts.
|
||||
2. **Assign it** to a machine. The workstation this was observed on is the honest first choice: it is
|
||||
where a person meets the fault, and it is where the check can be made with a plain client.
|
||||
3. **Verify** `curl https://<label>.<node>.internal/` with no flags, and `trust list` naming the mesh.
|
||||
|
||||
Then removal, which the module declares and nothing has exercised: unassigning must take the anchor
|
||||
away and refresh the bundles ([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md)),
|
||||
and that is the half most likely to be wrong, because it is the half nobody reaches by accident.
|
||||
|
||||
@@ -0,0 +1,63 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-09-30
|
||||
located-in:
|
||||
- mesh-controller internal/catalogue/roster.go (the roster's entries for a routed name)
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 157 — A routed name is published with an `.internal` alias that nothing serves
|
||||
|
||||
## What was observed
|
||||
|
||||
Every machine's hosts file carries two entries for each routed name — the name, and the name with
|
||||
`.internal` appended:
|
||||
|
||||
```
|
||||
10.10.0.1 keycloak.novox.be.internal keycloak.novox.be
|
||||
10.10.0.1 drive.novox.be.internal drive.novox.be
|
||||
10.10.0.1 umami.novox.be.internal umami.novox.be
|
||||
```
|
||||
|
||||
The suffixed one resolves and is served by nothing. The proxy refuses it during the handshake, and
|
||||
says so exactly:
|
||||
|
||||
```
|
||||
http: TLS handshake error from 10.10.0.3:33480:
|
||||
no public route for "keycloak.novox.be.internal" in this mesh, so no certificate is asked for
|
||||
```
|
||||
|
||||
A client sees `curl: (35) TLS connect error ... tlsv1 alert internal error` and no peer certificate —
|
||||
a server-side refusal, with nothing in it to say the name was never real.
|
||||
|
||||
The name the proxy does serve is `<label>.<node>.internal` — `keycloak.novox.internal` — which is what
|
||||
[issue 139](../139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md) describes. So
|
||||
there are two internal shapes for one service, one of them composed by appending the suffix to a name
|
||||
that already has a domain.
|
||||
|
||||
## Why it matters
|
||||
|
||||
**It sends a reader to the wrong diagnosis.** Reproducing
|
||||
[issue 129](../129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md) on 2026-09-30, the
|
||||
first three names tried came from the hosts file, all failed with a TLS alert rather than the
|
||||
verification error 129 reports, and the evidence pointed at the proxy having lost its internal
|
||||
certificates — a regression that had not happened. The correct name reproduces 129 exactly. Several
|
||||
minutes went into a fault that did not exist, and the only thing that distinguished the two was reading
|
||||
the proxy's own log.
|
||||
|
||||
It is also a name in every machine's hosts file, and in every container's, that cannot be reached: the
|
||||
shape the mesh is otherwise careful about — writing a name that resolves to something that does not
|
||||
answer is worse than not writing it, because a connection to an address that does not answer hangs
|
||||
where a name that does not resolve fails at once ([ADR 0007](../../02-DECISIONS/0007-connectivity.md),
|
||||
and the same reasoning in `namesInTheMesh`).
|
||||
|
||||
## Where to look
|
||||
|
||||
The roster template renders one entry per name as `{{.FQDN}} {{.Name}}`, and `FQDN` is composed by
|
||||
appending the mesh suffix to the bare name. For a node that is right — `novox` becomes
|
||||
`novox.internal`. For a routed name the bare name is already a fully qualified public name, so the
|
||||
composition produces `keycloak.novox.be.internal`, which is not a name anything was told to serve.
|
||||
|
||||
The fix is a judgement about what a routed name's internal form is, and 139 is the record that asks it;
|
||||
this one is the evidence that the current answer publishes a third thing that is neither.
|
||||
+49
@@ -0,0 +1,49 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-30
|
||||
located-in: []
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 158 — The proxy re-reads and re-logs every route it serves, every two seconds
|
||||
|
||||
## What was observed
|
||||
|
||||
The route proxy on the control node logs the whole of what it serves about every two seconds —
|
||||
measured 2026-09-30, **31 times in sixty seconds**, each line naming all 52 routes:
|
||||
|
||||
```
|
||||
23:27:19 serving 52 route(s): autoconfig.novox.be, autoconfig.novox.internal, … www.praktijkdespiegel.be
|
||||
23:27:21 serving 52 route(s): autoconfig.novox.be, autoconfig.novox.internal, … www.praktijkdespiegel.be
|
||||
23:27:23 serving 52 route(s): autoconfig.novox.be, autoconfig.novox.internal, … www.praktijkdespiegel.be
|
||||
```
|
||||
|
||||
Nothing is changing. The route set is identical every time.
|
||||
|
||||
## Why it matters
|
||||
|
||||
**It buries the only line that matters.** Between two of those entries sits the one error that explained
|
||||
a failing name:
|
||||
|
||||
```
|
||||
23:27:23 http: TLS handshake error from 10.10.0.3:33480:
|
||||
no public route for "keycloak.novox.be.internal" in this mesh, so no certificate is asked for
|
||||
```
|
||||
|
||||
One line of signal to roughly 5,000 characters of repetition, and the diagnosis it belonged to
|
||||
([issue 157](../157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md)) was found by
|
||||
grepping past it. A log that says the same true thing every two seconds is a log nobody reads, which is
|
||||
the same failure as one that says nothing — and it is the mesh's own accuracy rule pointed the other
|
||||
way: a report that is loud about the unchanging is not reporting.
|
||||
|
||||
Whether the *re-read* is also wasteful is secondary and unmeasured — it may be a cheap file stat. The
|
||||
logging is not in question.
|
||||
|
||||
## Where to look
|
||||
|
||||
Not localised. The proxy is `mesh-controller examples/route-proxy`; whether it re-reads on a timer or on
|
||||
a file watch, and whether it logs unconditionally or only on change, is the first thing to read. Saying
|
||||
what changed — or saying nothing — is the behaviour wanted, and the mesh already has the rule written
|
||||
down for its own reports: a log that is quiet on success and loud on failure reads as broken when it is
|
||||
working, and one that is loud always reads as nothing.
|
||||
Reference in New Issue
Block a user