Merge pull request 'Issue 129 is live and reproduced, and needs three steps rather than one' (#198) from issue/129-and-what-reproducing-it-found into main

This commit was merged in pull request #198.
This commit is contained in:
2026-09-29 23:30:48 +00:00
4 changed files with 184 additions and 35 deletions
@@ -1,5 +1,5 @@
---
status: located
status: diagnosing
opened: 2026-09-26
located-in: [mesh-catalog ca-trust]
---
@@ -1,45 +1,82 @@
# Diagnosis
# 129 — diagnosis
*2026-09-29.*
*2026-09-30, from the workstation the issue was opened on.*
## What was ruled out
## Still live, and reproduced exactly
**That something already carries the root and it is only misplaced.** It does not. The authority
serves its root at a path beside its ACME directory, and the one thing that fetches it — the route
proxy — puts it in a directory of its own and hands it to one program. Nothing has ever written
into a machine's trust store. Measured on three converged machines: the anchors present are the
predecessor's authority and a developer tool's local root, and on the machines where the
predecessor's was deliberately removed, every internal name fails verification.
The certificate is genuine, the authority is the mesh's, and nothing on the machine trusts it:
**That the private network could carry it, the way it carries the registry's trust.** That is what
the report proposed, and it was rejected on consideration rather than on difficulty
([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md), option 1): being on
the network is what makes the registry *reachable* and is therefore the right trigger there, while
trusting an authority is a separate fact from being able to reach it. The anchor's directory and
the command that refreshes the extracted bundles are also one operating system's difference, which
is the host's half of the mesh and not the controller's.
```
$ openssl s_client -connect keycloak.novox.internal:443 -servername keycloak.novox.internal
subject=CN=keycloak.novox.internal
issuer=O=Mesh Internal CA, CN=Mesh Internal CA Intermediate CA
Verify return code: 20 (unable to get local issuer certificate)
**That it needs a new host resource type.** It does not, today. A file and a service say the whole
of it, which the packet filter already proves. The primitive becomes the right answer when a second
operating system is in play, and not before.
$ curl https://keycloak.novox.internal/
curl: (60) SSL certificate OpenSSL verify result: unable to get local issuer certificate (20)
```
## Where it belongs
`trust list` holds no entry for the mesh. The anchors present are two `mkcert` development roots and
the predecessor's lab root — the report's account of the trust store is unchanged.
A module in the catalogue: it requires `internal-acme-ca`, fetches the root over the mesh's own
network, installs it as a trust anchor, refreshes the machine's bundles, and — because being
unassigned stops its unit, and stopping the unit is what undoes it — takes both away again.
The public name on the same proxy verifies cleanly (`CN=keycloak.novox.be`, Let's Encrypt, return code
0), which places the fault exactly where the report puts it: not in the proxy, not in the authority,
and not in the certificate.
The owner is therefore `mesh-catalog`, module `ca-trust`, and nothing in the control plane.
**The name matters, and the report's "every HTTPS name the mesh serves internally" is too broad.** The
served internal name is `<label>.<node>.internal`. The hosts file also carries
`<label>.<public-domain>.internal`, which nothing serves and which fails differently — that is
[issue 157](../157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md), found while
reproducing this, and it cost the first several minutes of this diagnosis.
## The module exists, and this stays open until a machine holds it
## The authority serves what the module needs
*2026-09-29.* `ca-trust` is in the catalogue and merged
([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md)), and what it renders
is checked in the control plane's own suite: the script fetches from the authority it was bound to,
and the unit runs it both ways.
`step-ca` is up and healthy, and publishes `roots: /roots.pem` for both `acme-ca` and
`internal-acme-ca`. That endpoint returns PEM:
**No machine has been assigned it, and nothing has verified a name because of it.** The bed written
for that cannot run ([issue 146](../146-the-foundation-cannot-be-raised-on-the-bus-the-mesh-runs-on/00-report.md)),
and the live mesh has not been given the module. So the symptom this record opened on — every
internal name failing verification on every machine — is still true everywhere, and the record stays
`located` until it is not. Closing it on a module that exists would be closing it on an intention.
```
$ curl -sk https://127.0.0.1:9000/roots.pem
-----BEGIN CERTIFICATE-----
MIIBvzCCAWWgAwIBAgIQYa2CkdJk16JyG/dVy2qoEzAKBggqhkjOPQQDAjA+…
```
So `${bound:internal-acme-ca:roots}` in the `ca-trust` module composes to a URL that returns a
certificate, and the module's own check — refuse a body that is not one — is checking the right thing.
Worth recording because it was nearly filed as a defect: step-ca *also* serves `/roots`, which returns
`{"crts":["-----BEGIN CERTIFICATE-----\n…"]}`. That body contains the literal text the module greps
for, so had the module used `/roots` it would have installed JSON into the anchors directory and
reported success. It does not use it. The guard is sound only because the published path is the PEM
one, which is worth knowing before anybody changes either.
## What is actually in the way
**The module is not registered.** The report and the work plan both say it exists and is merged, which
it does — `mesh-catalog modules/ca-trust`, on `main`. But the mesh has never been told about it:
```
$ mesh-controller module list | grep -iE 'ca-trust|step-ca'
step-ca 1 built 67f5f4cf on novox
```
39 of the catalogue's 76 manifests are registered. `ca-trust` is one of the 37 that are not, so it
cannot be assigned to anything — "assign it to one machine" has no module to name.
A dry run confirms it registers cleanly and needs no artifact built: it declares a directory, a script,
a unit and a service, and no image.
```
$ mesh-controller build <catalogue> --path modules/ca-trust --dry-run
… the manifest, parsed and validated
```
## So the remaining work is three steps, not one
1. **Register it** — build it from the catalogue, which pins nothing because it has no artifacts.
2. **Assign it** to a machine. The workstation this was observed on is the honest first choice: it is
where a person meets the fault, and it is where the check can be made with a plain client.
3. **Verify** `curl https://<label>.<node>.internal/` with no flags, and `trust list` naming the mesh.
Then removal, which the module declares and nothing has exercised: unassigning must take the anchor
away and refresh the bundles ([ADR 0147](../../02-DECISIONS/0147-a-module-anchors-the-meshs-authority.md)),
and that is the half most likely to be wrong, because it is the half nobody reaches by accident.
@@ -0,0 +1,63 @@
---
status: located
opened: 2026-09-30
located-in:
- mesh-controller internal/catalogue/roster.go (the roster's entries for a routed name)
fixed-by:
amended-design:
---
# 157 — A routed name is published with an `.internal` alias that nothing serves
## What was observed
Every machine's hosts file carries two entries for each routed name — the name, and the name with
`.internal` appended:
```
10.10.0.1 keycloak.novox.be.internal keycloak.novox.be
10.10.0.1 drive.novox.be.internal drive.novox.be
10.10.0.1 umami.novox.be.internal umami.novox.be
```
The suffixed one resolves and is served by nothing. The proxy refuses it during the handshake, and
says so exactly:
```
http: TLS handshake error from 10.10.0.3:33480:
no public route for "keycloak.novox.be.internal" in this mesh, so no certificate is asked for
```
A client sees `curl: (35) TLS connect error ... tlsv1 alert internal error` and no peer certificate —
a server-side refusal, with nothing in it to say the name was never real.
The name the proxy does serve is `<label>.<node>.internal` — `keycloak.novox.internal` — which is what
[issue 139](../139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md) describes. So
there are two internal shapes for one service, one of them composed by appending the suffix to a name
that already has a domain.
## Why it matters
**It sends a reader to the wrong diagnosis.** Reproducing
[issue 129](../129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md) on 2026-09-30, the
first three names tried came from the hosts file, all failed with a TLS alert rather than the
verification error 129 reports, and the evidence pointed at the proxy having lost its internal
certificates — a regression that had not happened. The correct name reproduces 129 exactly. Several
minutes went into a fault that did not exist, and the only thing that distinguished the two was reading
the proxy's own log.
It is also a name in every machine's hosts file, and in every container's, that cannot be reached: the
shape the mesh is otherwise careful about — writing a name that resolves to something that does not
answer is worse than not writing it, because a connection to an address that does not answer hangs
where a name that does not resolve fails at once ([ADR 0007](../../02-DECISIONS/0007-connectivity.md),
and the same reasoning in `namesInTheMesh`).
## Where to look
The roster template renders one entry per name as `{{.FQDN}} {{.Name}}`, and `FQDN` is composed by
appending the mesh suffix to the bare name. For a node that is right — `novox` becomes
`novox.internal`. For a routed name the bare name is already a fully qualified public name, so the
composition produces `keycloak.novox.be.internal`, which is not a name anything was told to serve.
The fix is a judgement about what a routed name's internal form is, and 139 is the record that asks it;
this one is the evidence that the current answer publishes a third thing that is neither.
@@ -0,0 +1,49 @@
---
status: open
opened: 2026-09-30
located-in: []
fixed-by:
amended-design:
---
# 158 — The proxy re-reads and re-logs every route it serves, every two seconds
## What was observed
The route proxy on the control node logs the whole of what it serves about every two seconds —
measured 2026-09-30, **31 times in sixty seconds**, each line naming all 52 routes:
```
23:27:19 serving 52 route(s): autoconfig.novox.be, autoconfig.novox.internal, … www.praktijkdespiegel.be
23:27:21 serving 52 route(s): autoconfig.novox.be, autoconfig.novox.internal, … www.praktijkdespiegel.be
23:27:23 serving 52 route(s): autoconfig.novox.be, autoconfig.novox.internal, … www.praktijkdespiegel.be
```
Nothing is changing. The route set is identical every time.
## Why it matters
**It buries the only line that matters.** Between two of those entries sits the one error that explained
a failing name:
```
23:27:23 http: TLS handshake error from 10.10.0.3:33480:
no public route for "keycloak.novox.be.internal" in this mesh, so no certificate is asked for
```
One line of signal to roughly 5,000 characters of repetition, and the diagnosis it belonged to
([issue 157](../157-a-routed-names-internal-alias-is-served-by-nothing/00-report.md)) was found by
grepping past it. A log that says the same true thing every two seconds is a log nobody reads, which is
the same failure as one that says nothing — and it is the mesh's own accuracy rule pointed the other
way: a report that is loud about the unchanging is not reporting.
Whether the *re-read* is also wasteful is secondary and unmeasured — it may be a cheap file stat. The
logging is not in question.
## Where to look
Not localised. The proxy is `mesh-controller examples/route-proxy`; whether it re-reads on a timer or on
a file watch, and whether it logs unconditionally or only on change, is the first thing to read. Saying
what changed — or saying nothing — is the behaviour wanted, and the mesh already has the rule written
down for its own reports: a log that is quiet on success and loud on failure reads as broken when it is
working, and one that is loud always reads as nothing.