Issues 138 and 139: a seat with interchangeable holders that are not, and an internal route name that resolves to the wrong machine

This commit is contained in:
2026-09-28 21:46:20 +02:00
parent db3868e2b0
commit 5dbde0b13a
2 changed files with 107 additions and 0 deletions
@@ -0,0 +1,56 @@
---
status: open
opened: 2026-09-28
located-in: [mesh-controller internal/catalogue, mesh-catalog]
fixed-by:
amended-design:
---
# 138 — Two modules claim one seat and are not interchangeable, and nothing says so
## What was observed
Three modules claim the node-scoped uplink seat: one for each network manager a machine here might
run. [ADR 0117](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md) gives each of them the same
job — ask the manager the machine already runs to leave the resolver file alone and to leave the mesh's
interface alone — and deliberately keeps the machine's own links out of the mesh's hands.
A seat means one holder and an interchangeable holder. These are interchangeable in what they *ask*
and not in what they *do*:
- None installs, enables, starts or stops the manager. That is on purpose: stopping it takes every
link down, including the mesh's own way in.
- None carries an address, a route or a wireless credential, for the same reason.
- **Nothing checks that the module holding the seat names the manager the machine is actually
running.** Assigning the systemd-networkd holder to a machine running NetworkManager writes a file
for a daemon that is inactive and disabled, the seat reports held, and the two things the seat
exists to arrange are arranged for nobody. NetworkManager goes back to rewriting the resolver file
on every lease, which is the failure the module's own comment describes.
The machine reports which service manager and which units are active, so the fact needed to catch this
is already in the report the mesh holds.
## Why it matters beyond this instance
**A seat is the mesh's promise that a role is filled.** If the holder can be a module for software
that is not running, the seat says a role is filled while nothing fills it — and the surface that
would tell an operator says "held".
**It is the same shape as two faults found the same day.** A module named a firewall front-end the
machine does not have ([issue 136](../136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md)),
and the filter named address ranges one runtime happens to use
([issue 137](../137-converging-a-machine-cut-off-its-own-guests/00-report.md)). Each is a claim about
the machine that nothing on the machine checks.
**And it decides whether the seat is worth having.** Either the holder must match what the machine
runs, which is a condition the mesh can check from the report it already has, or the holders must be
able to switch the manager, which ADR 0117 refuses for a reason that has not changed.
## Open questions
- Should a seat's conditions of holding include a capability the machine reports, so a holder naming
absent or inactive software is refused rather than recorded?
- Is "the uplink" one seat at all, if its holders are three dialects of the same two requests? The
alternative is one module that speaks whichever dialect the machine needs, chosen from the report.
- What should happen on a machine that switches manager afterwards? The seat would then be held by the
wrong module, and the machine is the only place that knows.
@@ -0,0 +1,51 @@
---
status: open
opened: 2026-09-28
located-in: [mesh-controller internal/catalogue]
fixed-by:
amended-design:
---
# 139 — An internal route name resolves to the consumer's node, not the one that serves it
## What was observed
A module that requires a route is given two names: a public one composed under the serving node's
domain, and an internal one composed under the consumer's own machine — `<label>.<node>.internal`.
The two are published differently:
- The **public** name is written into every machine's hosts file at the address of the node whose
proxy answers it. The mesh computes that deliberately, so any container resolving a routed name
reaches the proxy.
- The **internal** name is resolved by the machine's own resolver, which answers every name under
`<node>.internal` with that node's address — the consumer's, because the name was composed from it.
Where the proxy runs beside the consumer these are the same machine, which is every case on this mesh
today, and both names work. Measured on 2026-09-28: the internal name of a service on the control node
answers with a certificate from the mesh's internal authority, and the public name with one from the
public authority.
Where the proxy is on another machine they disagree. The internal name sends the client to a machine
that runs no proxy and has nothing listening on the port, while the public name sends it to the one
that does.
## Why it matters beyond this instance
**It is latent exactly where the mesh is heading.** `route` is provided mesh-wide precisely so a
module can be routed by a proxy on another machine. The first module assigned that way gets an
internal name that does not work, and the public one that does — with no error anywhere, because both
names resolve.
**A per-machine name is what an operator will reach for.** `<service>.<machine>.internal` reads like a
promise that the service on that machine is reachable there, and the wildcard makes every such name
resolve whether or not anything answers.
## Open questions
- Should the internal name be composed under the serving node, like the public one, or should it stay
the consumer's and be published at the serving node's address like the public name is?
- Is a per-node route holder the real answer — a proxy on every machine that serves its own names —
and if so, is `route` still one mesh-wide provision or a node-scoped seat with a mesh-wide fallback?
- What certifies the name in either case? The certificate is obtained by whoever terminates TLS, and
that is the question above in another form.