Merge pull request 'Issues 138 and 139' (#168) from issue/138-the-uplink-seat-and-139-an-internal-route-name into main
This commit was merged in pull request #168.
This commit is contained in:
@@ -0,0 +1,56 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-28
|
||||
located-in: [mesh-controller internal/catalogue, mesh-catalog]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 138 — Two modules claim one seat and are not interchangeable, and nothing says so
|
||||
|
||||
## What was observed
|
||||
|
||||
Three modules claim the node-scoped uplink seat: one for each network manager a machine here might
|
||||
run. [ADR 0117](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md) gives each of them the same
|
||||
job — ask the manager the machine already runs to leave the resolver file alone and to leave the mesh's
|
||||
interface alone — and deliberately keeps the machine's own links out of the mesh's hands.
|
||||
|
||||
A seat means one holder and an interchangeable holder. These are interchangeable in what they *ask*
|
||||
and not in what they *do*:
|
||||
|
||||
- None installs, enables, starts or stops the manager. That is on purpose: stopping it takes every
|
||||
link down, including the mesh's own way in.
|
||||
- None carries an address, a route or a wireless credential, for the same reason.
|
||||
- **Nothing checks that the module holding the seat names the manager the machine is actually
|
||||
running.** Assigning the systemd-networkd holder to a machine running NetworkManager writes a file
|
||||
for a daemon that is inactive and disabled, the seat reports held, and the two things the seat
|
||||
exists to arrange are arranged for nobody. NetworkManager goes back to rewriting the resolver file
|
||||
on every lease, which is the failure the module's own comment describes.
|
||||
|
||||
The machine reports which service manager and which units are active, so the fact needed to catch this
|
||||
is already in the report the mesh holds.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
**A seat is the mesh's promise that a role is filled.** If the holder can be a module for software
|
||||
that is not running, the seat says a role is filled while nothing fills it — and the surface that
|
||||
would tell an operator says "held".
|
||||
|
||||
**It is the same shape as two faults found the same day.** A module named a firewall front-end the
|
||||
machine does not have ([issue 136](../136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md)),
|
||||
and the filter named address ranges one runtime happens to use
|
||||
([issue 137](../137-converging-a-machine-cut-off-its-own-guests/00-report.md)). Each is a claim about
|
||||
the machine that nothing on the machine checks.
|
||||
|
||||
**And it decides whether the seat is worth having.** Either the holder must match what the machine
|
||||
runs, which is a condition the mesh can check from the report it already has, or the holders must be
|
||||
able to switch the manager, which ADR 0117 refuses for a reason that has not changed.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should a seat's conditions of holding include a capability the machine reports, so a holder naming
|
||||
absent or inactive software is refused rather than recorded?
|
||||
- Is "the uplink" one seat at all, if its holders are three dialects of the same two requests? The
|
||||
alternative is one module that speaks whichever dialect the machine needs, chosen from the report.
|
||||
- What should happen on a machine that switches manager afterwards? The seat would then be held by the
|
||||
wrong module, and the machine is the only place that knows.
|
||||
@@ -0,0 +1,51 @@
|
||||
---
|
||||
status: open
|
||||
opened: 2026-09-28
|
||||
located-in: [mesh-controller internal/catalogue]
|
||||
fixed-by:
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 139 — An internal route name resolves to the consumer's node, not the one that serves it
|
||||
|
||||
## What was observed
|
||||
|
||||
A module that requires a route is given two names: a public one composed under the serving node's
|
||||
domain, and an internal one composed under the consumer's own machine — `<label>.<node>.internal`.
|
||||
|
||||
The two are published differently:
|
||||
|
||||
- The **public** name is written into every machine's hosts file at the address of the node whose
|
||||
proxy answers it. The mesh computes that deliberately, so any container resolving a routed name
|
||||
reaches the proxy.
|
||||
- The **internal** name is resolved by the machine's own resolver, which answers every name under
|
||||
`<node>.internal` with that node's address — the consumer's, because the name was composed from it.
|
||||
|
||||
Where the proxy runs beside the consumer these are the same machine, which is every case on this mesh
|
||||
today, and both names work. Measured on 2026-09-28: the internal name of a service on the control node
|
||||
answers with a certificate from the mesh's internal authority, and the public name with one from the
|
||||
public authority.
|
||||
|
||||
Where the proxy is on another machine they disagree. The internal name sends the client to a machine
|
||||
that runs no proxy and has nothing listening on the port, while the public name sends it to the one
|
||||
that does.
|
||||
|
||||
## Why it matters beyond this instance
|
||||
|
||||
**It is latent exactly where the mesh is heading.** `route` is provided mesh-wide precisely so a
|
||||
module can be routed by a proxy on another machine. The first module assigned that way gets an
|
||||
internal name that does not work, and the public one that does — with no error anywhere, because both
|
||||
names resolve.
|
||||
|
||||
**A per-machine name is what an operator will reach for.** `<service>.<machine>.internal` reads like a
|
||||
promise that the service on that machine is reachable there, and the wildcard makes every such name
|
||||
resolve whether or not anything answers.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Should the internal name be composed under the serving node, like the public one, or should it stay
|
||||
the consumer's and be published at the serving node's address like the public name is?
|
||||
- Is a per-node route holder the real answer — a proxy on every machine that serves its own names —
|
||||
and if so, is `route` still one mesh-wide provision or a node-scoped seat with a mesh-wide fallback?
|
||||
- What certifies the name in either case? The certificate is obtained by whoever terminates TLS, and
|
||||
that is the question above in another form.
|
||||
Reference in New Issue
Block a user