Compare commits

...
Author SHA1 Message Date
jschoubben dcdfcf104e Issues 140 and 141: an endpoint's reach is declared nowhere, and the forward chain follows constants instead of the modules
Found preparing the control-node's convergence. Reach is settled independently by
the filter, the proxy's names and the certificate authority, so "this must not be
public" cannot be written and a public certificate is obtained regardless. And the
forward chain allows two hardcoded ranges plus a typed list, though the mesh
already knows which networks exist because its own modules declared them — a range
wide enough to keep four of them would have forwarded two predecessor leftovers too.
2026-09-28 22:57:38 +02:00
mesh-admin d23ace1646 Merge pull request 'Issues 138 and 139' (#168) from issue/138-the-uplink-seat-and-139-an-internal-route-name into main 2026-09-28 19:46:22 +00:00
jschoubben 5dbde0b13a Issues 138 and 139: a seat with interchangeable holders that are not, and an internal route name that resolves to the wrong machine 2026-09-28 21:46:20 +02:00
mesh-admin db3868e2b0 Merge pull request 'ADR 0137: a machine says which networks it routes' (#167) from decision/0137-a-machine-says-which-networks-it-routes into main 2026-09-28 19:31:03 +00:00
4 changed files with 255 additions and 0 deletions
@@ -0,0 +1,56 @@
---
status: open
opened: 2026-09-28
located-in: [mesh-controller internal/catalogue, mesh-catalog]
fixed-by:
amended-design:
---
# 138 — Two modules claim one seat and are not interchangeable, and nothing says so
## What was observed
Three modules claim the node-scoped uplink seat: one for each network manager a machine here might
run. [ADR 0117](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md) gives each of them the same
job — ask the manager the machine already runs to leave the resolver file alone and to leave the mesh's
interface alone — and deliberately keeps the machine's own links out of the mesh's hands.
A seat means one holder and an interchangeable holder. These are interchangeable in what they *ask*
and not in what they *do*:
- None installs, enables, starts or stops the manager. That is on purpose: stopping it takes every
link down, including the mesh's own way in.
- None carries an address, a route or a wireless credential, for the same reason.
- **Nothing checks that the module holding the seat names the manager the machine is actually
running.** Assigning the systemd-networkd holder to a machine running NetworkManager writes a file
for a daemon that is inactive and disabled, the seat reports held, and the two things the seat
exists to arrange are arranged for nobody. NetworkManager goes back to rewriting the resolver file
on every lease, which is the failure the module's own comment describes.
The machine reports which service manager and which units are active, so the fact needed to catch this
is already in the report the mesh holds.
## Why it matters beyond this instance
**A seat is the mesh's promise that a role is filled.** If the holder can be a module for software
that is not running, the seat says a role is filled while nothing fills it — and the surface that
would tell an operator says "held".
**It is the same shape as two faults found the same day.** A module named a firewall front-end the
machine does not have ([issue 136](../136-a-module-may-name-a-program-the-machine-does-not-have/00-report.md)),
and the filter named address ranges one runtime happens to use
([issue 137](../137-converging-a-machine-cut-off-its-own-guests/00-report.md)). Each is a claim about
the machine that nothing on the machine checks.
**And it decides whether the seat is worth having.** Either the holder must match what the machine
runs, which is a condition the mesh can check from the report it already has, or the holders must be
able to switch the manager, which ADR 0117 refuses for a reason that has not changed.
## Open questions
- Should a seat's conditions of holding include a capability the machine reports, so a holder naming
absent or inactive software is refused rather than recorded?
- Is "the uplink" one seat at all, if its holders are three dialects of the same two requests? The
alternative is one module that speaks whichever dialect the machine needs, chosen from the report.
- What should happen on a machine that switches manager afterwards? The seat would then be held by the
wrong module, and the machine is the only place that knows.
@@ -0,0 +1,51 @@
---
status: open
opened: 2026-09-28
located-in: [mesh-controller internal/catalogue]
fixed-by:
amended-design:
---
# 139 — An internal route name resolves to the consumer's node, not the one that serves it
## What was observed
A module that requires a route is given two names: a public one composed under the serving node's
domain, and an internal one composed under the consumer's own machine — `<label>.<node>.internal`.
The two are published differently:
- The **public** name is written into every machine's hosts file at the address of the node whose
proxy answers it. The mesh computes that deliberately, so any container resolving a routed name
reaches the proxy.
- The **internal** name is resolved by the machine's own resolver, which answers every name under
`<node>.internal` with that node's address — the consumer's, because the name was composed from it.
Where the proxy runs beside the consumer these are the same machine, which is every case on this mesh
today, and both names work. Measured on 2026-09-28: the internal name of a service on the control node
answers with a certificate from the mesh's internal authority, and the public name with one from the
public authority.
Where the proxy is on another machine they disagree. The internal name sends the client to a machine
that runs no proxy and has nothing listening on the port, while the public name sends it to the one
that does.
## Why it matters beyond this instance
**It is latent exactly where the mesh is heading.** `route` is provided mesh-wide precisely so a
module can be routed by a proxy on another machine. The first module assigned that way gets an
internal name that does not work, and the public one that does — with no error anywhere, because both
names resolve.
**A per-machine name is what an operator will reach for.** `<service>.<machine>.internal` reads like a
promise that the service on that machine is reachable there, and the wildcard makes every such name
resolve whether or not anything answers.
## Open questions
- Should the internal name be composed under the serving node, like the public one, or should it stay
the consumer's and be published at the serving node's address like the public name is?
- Is a per-node route holder the real answer — a proxy on every machine that serves its own names —
and if so, is `route` still one mesh-wide provision or a node-scoped seat with a mesh-wide fallback?
- What certifies the name in either case? The certificate is obtained by whoever terminates TLS, and
that is the question above in another form.
@@ -0,0 +1,74 @@
---
status: open
opened: 2026-09-28
located-in: []
fixed-by:
amended-design:
---
# 140 — An endpoint's reach is not declared, so three mechanisms each decide it separately
## What was observed
Preparing to converge the mesh's control-node — the last machine still running the firewall it
had before the mesh — the question came up for one module: the forge serves git over ssh, and that
port must stay reachable from outside the private network. Where is that said?
The manifest declares the port with a source of `mesh`, so the derived filter would close it to
everything but the private network. Looking for the place an assignment says otherwise, there are
two per-node settings keys: one that gives a module's declared port a machine port, and one that
overrides a declared port's source. The second has exactly one caller — the function that builds
the node's filter rules. Nothing else in the control plane reads it.
A module's routed endpoint is declared somewhere else entirely: a route contribution naming a label
and a port. It says nothing about reach. The proxy composes a **public** name and an **internal**
name for every route it is given, and obtains a certificate for each from a different authority.
Measured on that machine the same day: an identity provider's public name signed by the public
authority for 90 days, its internal name signed by the mesh's own intermediate for 24 hours and
renewed daily. Both names exist, and both certificates, because the proxy makes every name it can.
No assignment asked for either.
So the forge's ssh endpoint has a firewall source and nothing else — no name, no certificate, and no
way to say it should be public other than a key the filter alone reads. And the forge's web endpoint
has two names and two certificates that nobody requested.
## Why it matters beyond this instance
**Reach is stated twice, in two vocabularies, in two places that cannot disagree out loud.** A port
may be exposed to anywhere while the module contributes no public route; a public route may be served
for a module whose own listen is private. Nothing reconciles the pair or refuses it. Each mechanism
is separately defensible and the combination is unstated.
**The vocabulary belongs to the filter, not to reachability.** *Public, internal, or both* cannot be
expressed. A source of `anywhere` is one rule on one chain; it says nothing about which names should
exist or which authority should sign them. So "this endpoint must not be public" has no way to be
written, and is therefore enforced by nothing — while a public certificate for that very name is
obtained automatically.
**An endpoint is not a thing in the model.** A module has ports, and separately it has routes.
Nothing binds a port to a name to a certificate, which is why three mechanisms each decide reach on
their own and none of them is wrong. This is
[ADR 0045](../../02-DECISIONS/0045-a-machine-firewall-is-the-sum-of-what-it-listens-on.md)'s fault
one level up: that record closed "a declaration that reads as a restriction and restricts nothing"
for the packet filter. Here the declaration is absent altogether and the mechanisms guess.
**It blocks the certificate work.** The open question recorded for certificates — a name that must
not be public needs either DNS-01 or the internal authority only — cannot be answered while no
assignment states whether a name should be public. Neither can expiry reporting, revocation, or what
happens to a name when a machine leaves: all of them need to know which names were *meant*.
## Open questions
- Should an assignment name each of a module's endpoints, bind it to a node-level port, and state
whether it is reachable publicly, internally or both — with the filter, the proxy's names and the
certificate authority all derived from that one statement?
- What is an endpoint that is neither routed nor certified? Git over ssh is public reach with no name
and no certificate; the model has to hold that without inventing one.
- Are the two existing settings keys the same statement, half-built? If so, is this a new declaration
or the completion of theirs?
- Does an internal-only endpoint get a certificate at all, and from which authority — and does that
settle [issue 129](../129-nothing-makes-a-machine-trust-the-meshs-authority/00-report.md), where nothing
installs the mesh's own root?
- Does declaring reach per assignment also settle
[issue 139](../139-an-internal-route-name-resolves-to-the-consumers-node/00-report.md), where an
internal name resolves to the consumer's machine instead of the one serving the endpoint?
@@ -0,0 +1,74 @@
---
status: open
opened: 2026-09-28
located-in: []
fixed-by:
amended-design:
---
# 141 — The forward chain does not follow the modules, though the modules declare their networks
## What was observed
[ADR 0137](../../02-DECISIONS/0137-a-machine-says-which-networks-it-routes.md), decided the same
week, gave a machine a way to say which networks it routes for its guests, because the derived
filter's forward chain had until then allowed two ranges named as constants in the control plane's
own source — the container runtime's default bridge pool, and half of the pool its compose files
are given.
Checking the last machine still to be converged, the same fault was found to be live there, and the
declaration needed to work around it turned out to be wrong in kind.
That machine hosts twenty-one container networks. Nine fall inside the runtime's bridge pool and
are forwarded. Twelve sit in the other private range, and **six of those fall below the lower bound
of the constant**, so the flip would have cut their guests off exactly as it did on the workstation
that produced 0137.
Naming a range to cover the six was the obvious move, and is what 0137 provides for. But of those
six networks, **four are networks the mesh's own modules declare** — they appear as network
resources in the node's plan, created by the host because a module asked for them — and **two are
leftovers of the predecessor**, compose networks of services the mesh does not run. A range wide
enough to keep the four would have forwarded the two as well: a firewall widened by hand to protect
networks that should not exist.
The mesh already knows which of the twenty-one are its own. It made them.
## Why it matters beyond this instance
**The node's configuration is supposed to follow the modules assigned to it.** That is the mesh's
founding shape — the machine runs modules, and its files, its filter and its accounts are composed
from what runs there. The forward chain is the one derived thing that does not: it consults two
constants and, since 0137, a list a person types. A module added tomorrow brings a network the filter
will not forward; a module deprecated leaves a range in the list that outlives it.
**A typed range cannot distinguish the mesh's networks from what was left behind.** It is stated in
addresses, and addresses are what the runtime allocates, so the only honest declaration is one wide
enough to include whatever else the runtime has handed out. The derivation is narrower than anything
a person can safely write, because it names networks rather than ranges.
**0137 rejected deriving this, and was right about what it rejected.** It considered deriving the
list from *what the machine reports* and refused, on two grounds: a test bed creates its bridge
between one declaration and the next, and a filter that follows whatever appeared on the machine is a
firewall that widens itself. Deriving from the **declaration** is neither. The set is known before
the network exists, because a module declared it; and it cannot widen itself, because only a network
some module asked for is ever forwarded. What remains genuinely for a machine to say is guests no
module declares — a test bed's pool — which is a much smaller residue than the list as it stands.
**The gap is invisible in the one place that should show it.** The converge preview lists what
*listens*, and routing is not a listener. It says in one line what the machine routes, and a reader
has to know the runtime's allocations to tell whether that line is sufficient. On the machine
measured here it read as though nothing needed saying.
## Open questions
- Should the forward chain be derived from the network resources the node's modules declare, with the
host resolving each declared network to its address the way it already resolves a container by name?
The controller cannot render the address itself: a module's network resource carries a name, and the
runtime allocates the subnet at creation.
- What remains of `node networks` once that exists — only guests no module declares, such as a test
bed's pool? And should it then be named for that, rather than for all routing?
- The runtime's own default bridge, which containers attach to when no module network is named, is
not a module's network. Is it derived from the machine, declared by the module that owns the
runtime, or left as the one constant?
- Should the preview say which of a machine's networks are the mesh's and which are not, so a range
that exists to protect a leftover is visible as such?