21 Commits
Author SHA1 Message Date
jochen 0042ca9258 0119: link ADR 0118 now that it is on main 2026-09-27 00:58:55 +02:00
jochen 63c19456b4 0119 review: rollback needs the private network unassigned first; a configuration written back is retired again with the first original kept; only the interface's own file, never a link 2026-09-27 00:58:54 +02:00
jochen 2f195d501e to-be 08: the found tunnel's configuration is retired once the take is proven (ADR 0119) 2026-09-27 00:58:54 +02:00
jochen 8a78ff4efe ADR 0119: a taken tunnel's predecessor is retired once the take is proven 2026-09-27 00:58:54 +02:00
jschoubben 762300a380 Merge pull request 'issue 130: undeclaring a service stops it, even one the mesh only reloads or keeps running' (#142) from issue/130-undeclaring-a-service-stops-it into main 2026-09-26 22:58:43 +00:00
jochen 248c99ca6c 0118/130: link ADR 0117 now that it is on main 2026-09-27 00:58:17 +02:00
jochen 13208f0f48 0118: give a unit back the state it was found in — never-stop broke the converge rollback; process removal found and fixed 2026-09-27 00:58:17 +02:00
jochen ebd19c4c6c ADR 0118: undeclaring removes what the mesh made, gives back what it changed, leaves the machine's units as they are — resolves issue 130 2026-09-27 00:58:17 +02:00
jochen dd4cbabffb 130: ADR 0117 named, not linked, until it is on main 2026-09-27 00:58:17 +02:00
jochen 5b76a09da6 issue 130: undeclaring a service stops it, even one the mesh only reloads or keeps running 2026-09-27 00:58:17 +02:00
jschoubben a363a605cb Merge pull request 'issue 129: nothing makes a machine trust the mesh's own certificate authority' (#141) from issue/129-nothing-makes-a-machine-trust-the-meshs-authority into main 2026-09-26 22:57:30 +00:00
jschoubben 24835ab710 Merge pull request 'issue 128: the machine's hosts file is written whole, and on a workstation it is shared' (#140) from issue/128-the-hosts-file-is-written-whole into main 2026-09-26 22:57:02 +00:00
jschoubben ba397d4cbe Merge pull request 'ADR 0117: a machine's uplink is a seat — the mesh configures the manager, never the link' (#139) from decision/0117-the-uplink-is-a-seat into main 2026-09-26 22:56:27 +00:00
jochen df4a3538c3 0117 review: a holder's service declares no state — the manager's lifecycle is the machine's; a start-only setting's gap on a fresh machine 2026-09-27 00:07:42 +02:00
jochen 90b8eeff6b 129 review: the example name is a routed name, not a doubled suffix 2026-09-27 00:03:36 +02:00
jochen a865fc7d79 128 review: located; the fix as built (block, at, never held, order) 2026-09-27 00:03:30 +02:00
jochen c3313f6e17 0117 review: seat table row + decision cited, networkd/dhcpcd lines match the modules, block placement, dispatcher scope, references 2026-09-27 00:03:15 +02:00
jochen 60a53f9b18 issue 129: nothing makes a machine trust the mesh's own certificate authority 2026-09-26 23:58:48 +02:00
jochen ee31f9f761 issue 128: the machine's hosts file is written whole, and on a workstation it is shared 2026-09-26 23:58:48 +02:00
jochen 0e066473b3 ADR 0117 accepted; a manager that cannot reload takes the setting at its next start (dhcpcd, measured) 2026-09-26 23:58:48 +02:00
jochen 504adef221 ADR 0117: a machine's uplink is a seat — the mesh configures the manager, never the link 2026-09-26 23:58:48 +02:00
8 changed files with 462 additions and 3 deletions
@@ -0,0 +1,135 @@
---
topic: what runs on it
status: accepted
date: 2026-09-26
deciders: jochen
reconstructed: false
extends: 0110-a-seat-is-a-module-assignment-from-a-closed-set.md
---
# 117. A machine's uplink is a seat: the mesh configures the manager, never the link
## Context
The mesh installs on top of a machine's own networking. The private network's generator says
so in as many words: a machine has an address and a route to the broker *before* the mesh
exists, the broker's address travels in the enrolment token rather than being resolved, and the
private network is something the mesh installs on top, like anything else. Nothing in the mesh
says who manages that uplink, or what the mesh needs from whoever does.
Adopting the first workstations showed that the mesh does need something from it, and gets it
by accident:
- **The resolver the mesh owns depends on a file the mesh does not.** `resolv-conf` writes
`/etc/resolv.conf` and names the mesh's resolver. On a machine running NetworkManager, the
manager rewrites that file on every connectivity change unless it is told `dns=none`; on one
running dhcpcd, every lease renewal rewrites it unless it is told `nohook resolv.conf`. On
the adopted machines both settings exist only because the predecessor wrote them. No module
declares them. Remove the predecessor's file and the mesh's resolver is silently replaced the
next time a laptop changes network, while every surface of the mesh still reads green.
- **`resolv-conf` cannot declare them itself.** Which setting is needed depends on which
manager runs, and a `service` resource for a manager that is not installed fails the
declaration. A resolver module that knew about network managers would be the wrong module
knowing the wrong thing.
- **The private network's interface is exposed to the manager.** A manager that considers
every interface its own may try to configure `mesh0`, or tear it down on a profile change.
Nothing tells it not to.
- **Two managers on one machine go unnoticed.** Among the machines adopted so far, one runs
NetworkManager *and* dhcpcd at once: two programs that each believe they own the machine's
addresses and its resolver file.
Nothing detected it, because nothing in the mesh knows the role exists.
The machines differ in a way that matters: servers are wired and never move, while
workstations join wireless networks, captive portals and phone hotspots wherever they are.
## Considered Options
**1. The mesh manages the uplink: links, addressing, wireless networks and their
credentials.** Rejected. The mesh reaches a machine only over that link. A declaration that
gets it wrong — a mistyped network, a stale credential, a manager that fails to start — takes
the machine off the network, and with it the only channel a fix could arrive on. That is the
one failure the sshd module's `listens` rule forbids the firewall to arrange; a mesh that owned
the link could arrange it with any push. And a wireless network is joined at the machine, by
the person using it, in the moment. A declaration composed elsewhere cannot answer a captive
portal.
**2. Leave the uplink unmanaged; accept the implicit dependency.** Rejected. It keeps the
resolver working only for as long as a predecessor's file survives, and it leaves two managers
on one machine undetectable.
**3. The uplink is a seat. The module holding it configures the manager's relationship to the
mesh, and never the link.** Chosen.
## Decision
**`the-uplink` is a node-scoped seat** in the closed set ([ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md)).
It delivers no provision. It is held by the module for the program that manages the machine's
own network, one per manager: `networkmanager`, `systemd-networkd`, and `dhcpcd` for a machine
with nothing more. Assigning a second is refused, naming the first.
**What a holder declares** — only what keeps the manager and the mesh from contradicting each
other:
- the manager's package, present — and its service **with no state**: the manager's lifecycle is
the machine's. The mesh never starts, stops, enables or disables it, because stopping it takes
the link down, and a holder unassigned by mistake — or the wrong holder assigned — must not be
able to do that, nor start a second manager beside the one the machine runs. The service is
declared only so a change to the holder's settings reaches a *running* manager;
- the manager's own configuration that leaves the resolver file to the mesh (`dns=none` for
NetworkManager, `nohook resolv.conf` for dhcpcd, and nothing for systemd-networkd, which
never writes the resolver file);
- the manager's own configuration that leaves the private network's interface alone
(NetworkManager's `unmanaged-devices` naming `mesh0`; dhcpcd's `denyinterfaces mesh0`; for
systemd-networkd a network file of the module's matching `mesh0` as `Unmanaged=yes`);
- each as a drop-in beside the manager's main file where the manager reads one, and written
*into* a shared file otherwise, as a marked region the host owns
([ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md)'s idea for text files),
placed where the manager reads it as global — at the start of `dhcpcd.conf`, above any
`interface` line, because every line after one belongs to that interface;
- the service **reloaded** when a drop-in changes, never restarted — a restart drops the link,
and the link is the mesh's own channel to the machine. **A manager that cannot reload is not
restarted instead:** its setting takes effect at the manager's next start. Measured on the
adopted machines: NetworkManager (1.58) and systemd-networkd (systemd 261) both report
`CanReload=yes`; dhcpcd (10.3) reports `CanReload=no`, so its module declares no trigger at
all. Whether each setting is actually *applied* by a reload is confirmed on a machine before
the module is taken there, not assumed.
**What a holder never declares:** a link, an address, a route, a connection profile, a
wireless network or its credentials. Those are the operator's, in the sense of
[ADR 0051](0051-shared-data-is-the-operators.md): the mesh does not create, change or delete
them, and the module's `access`, if it needs one, is read-only.
## Consequences
- `resolv-conf` stays generic. The condition it could not express — "only if NetworkManager
runs" — is expressed by assigning the module for the manager that does.
- The dependency on the predecessor's `dns=none` file becomes a declared resource. On an
adopted node the holder's drop-in arrives beside the predecessor's; both say the same thing,
and the predecessor's is retired by hand after the take, like any other file the mesh
replaced under another name.
- A setting a manager reads only at its start is not in force until then. On an adopted machine
the predecessor's identical line normally already is; on a machine that was not adopted,
dhcpcd's resolver hook keeps rewriting the resolver file until dhcpcd next starts, and the
operator restarts it once, in a window of their choosing.
- A machine running two managers is found at assignment: the second holder is refused, and the
operator decides which manager the machine keeps before either module is taken.
- Workstations keep joining networks the way they always have. Under NetworkManager and
systemd-networkd the host already cooperates with the manager — its dispatcher hook wakes it
on every connectivity change — and nothing here changes that. A dhcpcd-only machine has no
such hook, and nothing here adds one.
- The seat table gains one entry: `the-uplink`, node scope, delivering nothing, decided here.
- **Not decided here:** whether the mesh should ever *offer* known networks to a machine — a
sealed, add-only list the operator curates once for all workstations. That is a different
question (the mesh holding credentials for links it must never be able to break) and gets its
own record if it is wanted.
## References
- [ADR 0110](0110-a-seat-is-a-module-assignment-from-a-closed-set.md): the closed set this seat joins;
[to-be 26](../03-DESIGN/01-to-be/26-the-seats.md): the seat table
- [ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md): written into, never over
- [ADR 0051](0051-shared-data-is-the-operators.md): what is the operator's stays the operator's
- mesh-controller `internal/catalogue/seats.go` (the seat), `internal/overlay/generator.go` (the
mesh installs on top of the machine's own networking)
- mesh-catalog `modules/networkmanager`, `modules/systemd-networkd`, `modules/dhcpcd`
- mesh-host `internal/apply/block.go` (a file written into a marked region, `at` start or end)
@@ -0,0 +1,119 @@
---
topic: what runs on it
status: accepted
date: 2026-09-27
deciders: jochen
reconstructed: false
extends: 0102-the-mesh-writes-into-a-shared-file-never-over-it.md
---
# 118. Undeclaring removes what the mesh made, and gives a unit back the state it was found in
## Context
When a resource stops being declared — its module unassigned, the node sent a
deliberately-empty declaration ([issue 127](../04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)),
or a new catalogue version renaming its id — the host undoes it. The host's own code states
the rule it means to follow: **it removes what it made and leaves what it merely configured.**
For almost every resource it does exactly that:
- a container, a network, a process's unit, a directory it created: removed;
- a file it created: removed; a file it replaced: its kept original put back
([ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md));
- keys and list members it wrote into a shared file: given back as they were
([ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md));
- a package: left installed — the host cannot know it is unused;
- an operator's path it was given access to: never touched
([ADR 0051](0051-shared-data-is-the-operators.md)).
**A service is the exception.** A `service` resource never installs a unit: it puts one that
already exists — the distribution's, the operator's — into a state. Undeclared, the host stops
it. That contradicts the rule above, and in practice it is the most dangerous thing an
undeclare can do. Found reviewing the uplink modules
([issue 130](../04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md)):
- the private network declares the container runtime's unit only so a change to the registry
trust reloads it — unassigning the private network stops the runtime, and every container on
the machine, the mesh's and not;
- the sshd module declares the ssh daemon — unassigning it stops ssh, the lockout that module's
own `listens` rule forbids;
- the uplink modules would have stopped the network manager, taking the machine off the only
link the mesh reaches it by.
[ADR 0117](0117-a-machines-uplink-is-a-seat.md) answered that for its own modules with a
service declared with no `state`. Every other module that declares a unit it did not make is
exposed in the same way, and relying on each author to remember an opt-out is how the next one
is missed.
## Considered Options
**1. Undeclaring touches nothing on the machine.** Rejected. What the mesh made would outlive
the module that made it: a container nobody manages keeps serving and stops being patched; a
unit the mesh wrote keeps running a bundle nothing updates; a name collides when the module
is assigned again. An undeclare that leaves the mesh's own work behind is an orphan factory.
**2. Keep stopping services; make "leave it running" an opt-in per resource.** Rejected. It
keeps the dangerous behaviour as the default for exactly the units that matter most — the
runtime, the ssh daemon, the network — and each new module is one forgotten field away from a
machine that goes dark when it is unassigned.
**3. Never stop a unit the mesh did not create.** Rejected, found while implementing it. The
mesh's packet filter is a unit the distribution installed and the mesh started at converge;
returning a node to adopted ([ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md))
unloads it by undeclaring it. Never stopping it would leave the mesh's filter loaded beside the
predecessor's firewall re-enabled — the one rollback a converge promises, broken. Who wrote the
unit file is not the line; what the mesh *did* to the unit is.
**4. Give the unit back the state it was found in.** Chosen.
## Decision
**Undeclaring removes what the mesh made and gives back what it changed.** For a unit the mesh
did not create, what it changed is the unit's state, so that is what is given back: **the host
records the state it first found the unit in, and undeclaring returns the unit to it.**
- **Recorded once**, the first time the host applies the service — whether it was running, and,
where the declaration sets it, whether it was enabled at boot — and carried in the host's
record from then on. Later applies never overwrite it: by then the unit's state is the mesh's
doing.
- **A unit found running is left running.** The container runtime, the ssh daemon, a network
manager: running before the mesh arrived, running after it leaves.
- **A unit the mesh started is stopped again**, and one it enabled is disabled again — the packet
filter a converge loaded, which returning to adopted unloads.
- **Never started on the way out.** A unit the mesh stopped is not started again when its
declaration goes; starting something is a decision, and the operator makes it.
- **Unknown is left alone.** A record written before the host kept what it found says nothing
about the unit before the mesh; the unit is left exactly as it is. A unit left running can be
stopped by the operator; one stopped by mistake may be the link the operator needed to do it.
- A unit the mesh *did* create — a `process` resource's unit and bundle — is stopped and removed
with its declaration. That is the mesh's own code. (Before this record there was no way to
remove one at all: an undeclared process failed every apply on its node.)
- The service's settings the mesh wrote are given back by their own resources (a kept original
restored, a region or keys removed). A running service keeps running on what it read until it
next reads its configuration; the mesh does not restart it to make it notice.
- A service declared with no `state` (ADR 0117) remains the way to say the mesh must not
**start** a unit either; undeclared, it is forgotten.
## Consequences
- Unassigning the private network no longer stops the container runtime; unassigning sshd no
longer stops ssh; no uplink module can take a machine's network down on its way out.
- The host's removal report says what it gave back — "restored: stopped again, as the host
found it" — or "forgotten: it was running before the mesh; left as it is" where it used to say
"stopped". Its plan names each unit an undeclare will stop, before it does.
- On a fresh machine where the mesh installed and started a service, unassigning its module
stops it again — the mesh gave, the mesh takes back. An operator who wants it kept declares it
in a module of their own, or starts it themselves after.
- A daemon can keep running after its module is gone, on configuration that was taken back from
under it. That is a visible, running process the operator can see and stop; the alternative
was an invisible outage.
- **Not decided here:** an unassign preview that lists what an undeclare will remove and what it
will leave running. Issue 130 asks for it; it is the controller's to build.
## References
- [issue 130](../04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md): the finding
- [ADR 0117](0117-a-machines-uplink-is-a-seat.md): the uplink modules, and a service with no state
- [ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md), [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md):
what is given back, and how
- mesh-host `internal/apply/apply.go` (`remove`, the service case)
@@ -53,7 +53,7 @@ is removed from where its unit reads it.**
reporting it.
- **The mesh never brings it back.** Undeclaring the private network does not restore the found
tunnel: the mesh stopped it, and nothing is started on the way out
(ADR 0118, in review). A machine whose
([ADR 0118](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md)). A machine whose
private network is unassigned has no tunnel until it is assigned again — which is what
unassigning it means.
@@ -83,5 +83,5 @@ is removed from where its unit reads it.**
- [ADR 0105](0105-the-mesh-adopts-the-predecessors-tunnel-in-place.md): the take, and why it keeps
the found configuration during it
- [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md): kept originals
- ADR 0118 (in review): nothing is started on the way out
- [ADR 0118](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md): nothing is started on the way out
- mesh-host `internal/apply/takeover.go`
+2
View File
@@ -200,6 +200,8 @@ python3 00-META/checks/index.py fail if stale
- **0113** — [The vault makes every shared secret, a provider makes resources and data, and the mesh carries both](0113-the-vault-makes-every-secret.md) *(proposed)*
- **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md) *(proposed)*
- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md) *(proposed)*
- **0117** — [A machine's uplink is a seat: the mesh configures the manager, never the link](0117-a-machines-uplink-is-a-seat.md)
- **0118** — [Undeclaring removes what the mesh made, and gives a unit back the state it was found in](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md)
### How it is built
+3 -1
View File
@@ -8,11 +8,12 @@ code:
- mesh-controller cmd/mesh-controller/source.go
- mesh-controller internal/inventory/migrations/0032-a-source-may-live-on-a-seat.sql
- mesh-catalog modules/gitea/module.json
updated: 2026-09-26
updated: 2026-09-27
decisions:
- 02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md
- 02-DECISIONS/0111-a-build-source-is-on-the-git-seat-or-external.md
- 02-DECISIONS/0109-a-package-registry-seat-is-one-per-ecosystem.md
- 02-DECISIONS/0117-a-machines-uplink-is-a-seat.md
---
# 26 — The seats
@@ -64,6 +65,7 @@ nobody argued for is an entry nobody can explain.
| `the-private-network` | node | — | the private network the mesh runs over |
| `the-resolver-configuration` | node | — | whichever of the alternative resolver configurations is chosen |
| `the-showcase` | node | — | the showcase module |
| `the-uplink` | node | — | the program that manages the machine's own network ([ADR 0117](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md)) |
The controller holds this set in code, and a test asserts both its size and that every entry names
the record that made it a seat. **This table and [ADR 0110](../../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md)
@@ -0,0 +1,73 @@
---
status: located
opened: 2026-09-26
located-in: [mesh-controller internal/catalogue/facts.go, mesh-host internal/apply]
---
# 128 — the machine's hosts file is written whole, and on a workstation it is shared
## What was observed
The private network asks for the `node-names` fact, and the mesh delivers it as
`/etc/hosts`. `nodeNames` composes a **complete** file — its own header, `localhost`, the
machine's own name, and every name in the mesh — and the host writes it over whatever is there.
On an adopted workstation the file the mesh holds contains, besides the predecessor's block of
mesh names:
- the distribution's own lines (`localhost`, the machine's `.localdomain` name);
- two marked blocks (`# BEGIN … # END …`) maintained by a local-development tool, pointing a
dozen development hostnames at `127.0.0.1` — rewritten by that tool whenever its project
list changes;
- hand-added entries of the operator's.
Today the file is only **held** ([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)):
the private network was assigned, not yet taken, so nothing was lost. Taking it — or
converging the node, which takes everything — replaces the file. The development tool's
entries disappear, its projects stop resolving, and every later write it makes is overwritten
at the next change to the mesh's names (a machine joins, a route is contributed), silently and
without a failure anywhere: the development tool thinks it wrote its block, and the mesh thinks
it owns the file.
This is [ADR 0102](../../02-DECISIONS/0102-the-mesh-writes-into-a-shared-file-never-over-it.md)'s
failure exactly — a file the mesh shares with software it did not install, written over — in a
file 0102 did not name, because its merge verb is structured (`into: json`) and a hosts file is
not JSON.
A second, smaller finding from the same reading: the fact's contents depend on which machines
hold the private network. A machine that is enrolled but not yet assigned the private network
is in neither `node-names` nor `node-zones`; its name resolves on the others only for as long
as a predecessor's hosts block survives. Taking the hosts file before every machine is on the
private network loses that name too.
## What would have prevented it
- A **marked-region** merge in the host's vocabulary: `into: "block"` (or similar) — the host
owns only the lines between its own begin and end markers, keeps everything outside them
byte for byte, records what the region held before, and on undeclare removes the region and
nothing else. The shape local tools already use for this very file.
- The `node-names` fact written as that region — no header of its own, no `localhost`, no
machine name — so the distribution's lines and every other tool's stay where they are.
- A converge preview that names a held file the take would replace *whole*, with its line
count before and after, so a person sees "hosts: 31 lines → 12" before the flip.
## The fix, as built (in review)
- **Host:** a file resource may say `"into": "block"`. The host owns only the lines between
`# BEGIN mesh <id>` and `# END mesh <id>` and keeps everything outside them byte for byte. A
new region goes at the `end` by default, or at the `start` (`"at": "start"`) for files where a
line's meaning depends on what stands above it; a region already present is never moved.
Undeclared, what the region held before is put back, or the region is removed and nothing
else. Replacing nothing, it is written on an adopted node without being held — so a machine
gets the mesh's names before its private network is taken.
- **Controller:** `node-names` is a fact written into a shared file, emitted as that region: the
mesh's names only, no header, no `localhost`, no `127.0.1.1` line.
- **Order:** a host older than the block mode refuses the whole declaration on an unknown
`into`, so hosts are upgraded before the controller that emits it.
## Evidence to carry into diagnosis
- `internal/catalogue/facts.go`, `nodeNames`: the complete file is built here.
- The host's file resource supports `into: "json"` only; anything else is a whole write.
- `node show <node>` on the adopted workstation: `holds file /etc/hosts
mesh-wireguard.fact-node-names`, original kept.
@@ -0,0 +1,62 @@
---
status: open
opened: 2026-09-26
located-in: [mesh-controller, mesh-catalog step-ca]
---
# 129 — nothing makes a machine trust the mesh's own certificate authority
## What was observed
On an enrolled, adopted workstation — on the private network, resolving the mesh's names
through the mesh's resolver — every HTTPS name the mesh serves internally fails verification:
```
curl https://<a name the mesh routes internally>/
curl: (60) SSL certificate OpenSSL verify result: unable to get local issuer certificate (20)
```
The route proxy presents a certificate issued by the mesh's internal authority (step-ca, the
`internal-acme-ca` provision). The machine's trust store holds the **predecessor's** authority
and a developer tool's local root, and nothing of the mesh's. No module installs the mesh's
root, and no fact carries it: step-ca's only consumers are proxies, which obtain certificates
over ACME and never need the root on the machine they run on.
[Issue 048](../048-nothing-makes-a-machine-trust-the-mesh-registry/00-report.md) found the same
shape for the mesh's registry and resolved it by treating the private network as the transport
security ([ADR 0082](../../02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md)):
the runtime pulls in the clear, over the tunnel. That answer does not carry over. A browser, git
over HTTPS, a package manager and every TLS client a person or a module uses verify the
certificate chain, and there is no "insecure registries" for them — nor should there be.
Consequences today, all silent until someone tries:
- a person on a workstation cannot open any internal HTTPS name without a warning;
- git over HTTPS to the mesh's forge fails, so the working clone URL is ssh-only;
- a module on a non-hub machine that calls another module's internal HTTPS name fails
verification unless its image happens to carry the root;
- the predecessor's authority cannot be retired from any machine while anything there still
speaks TLS to a mesh name, because it is the only authority those machines trust.
## What would have prevented it
- A **mesh fact carrying the internal authority's root** (public material; the controller or
the step-ca module is its source), written onto every machine on the private network — the
same reasoning that has the private network write the registry trust and the names: being on
the network is what makes a machine one that speaks to the mesh's names.
- A resource that puts it where the machine's TLS clients look — on Arch,
`/etc/ca-certificates/trust-source/anchors/` — and **refreshes the extracted bundles**
(`update-ca-trust`). The refresh is the open design question: it is a command, and the link
may not carry an action ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)). A declared
one-shot unit, or a host primitive for "trust this anchor", are the obvious candidates.
- Removal symmetric to arrival: undeclared, the anchor goes and the bundles are refreshed again,
so a machine leaving the mesh stops trusting it.
## Evidence to carry into diagnosis
- `step-ca` module: provides `acme-ca` / `internal-acme-ca`, listens on 9000 for proxies; no
resource writes its root anywhere but its own state directory.
- The private network's generator writes `/etc/hosts` and the registry trust, and nothing
about certificates.
- On the workstation, the trust anchors present are the predecessor's authority and a local
development root; `trust list` shows no entry for the mesh.
@@ -0,0 +1,66 @@
---
status: located
opened: 2026-09-27
located-in: [mesh-host internal/apply/apply.go (remove)]
amended-design: 02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md
---
# 130 — undeclaring a service stops it, even one the mesh only reloads or only keeps running
## What was observed
Reviewing the uplink modules ([ADR 0117](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md))
found that the host's `remove` path stops every `service` resource that is no longer declared:
`SetServiceState(..., "stopped")`, reported as "stopped; the unit file is not the host's to
delete". `store.Orphans` matches by id alone. So any of these stops the unit:
- the module is unassigned — by mistake, or to switch it for another;
- the node is sent a deliberately-empty declaration ([issue 127](../127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md));
- a later catalogue version renames the resource's `id`.
That is right for a service the mesh brought into being. It is wrong for a unit the mesh
declares only to act on — and the catalogue already has two:
- **The private network declares `docker.service`** (`registry-trust-reload`, state `running`)
so that a change to the registry trust reloads the runtime ([ADR 0102](../../02-DECISIONS/0102-the-mesh-writes-into-a-shared-file-never-over-it.md)).
Unassigning the private network stops the container runtime, and every container on the
machine with it — including ones the mesh does not manage.
- **The sshd module declares `sshd.service`.** Unassigning it stops the machine's ssh daemon:
the lockout the same module's `listens` rule says a firewall must never arrange.
The uplink modules would have added a third and a fourth: unassigning the network manager's
module would have stopped the network manager, taking the machine off the only link the mesh
reaches it by.
## What would have prevented it
- A service resource that says the unit's **lifecycle is the machine's**: declared with no
`state`, the mesh never starts, stops, enables or disables it; it only reloads or restarts a
*running* unit when a trigger changes; undeclared, it is left exactly as it is. (Being built
on mesh-host `feat/a-file-written-into-a-marked-block` for the uplink modules.)
- Then: `registry-trust-reload` declared that way (the runtime is the machine's), and the sshd
module's service too — a machine's ssh daemon outlives any module that configures it.
- A plan or unassign preview that names every unit an undeclare will stop, so the consequence
is read before it happens.
## Resolution
[ADR 0118](../../02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md):
undeclaring removes what the mesh made and gives back what it changed. The host records the state
it first found a unit in, and undeclaring returns the unit to it — a unit found running (the
container runtime, sshd, a network manager) is left running; one the mesh started (the packet
filter a converge loaded) is stopped again; nothing is started on the way out; a record from
before the host kept what it found leaves the unit alone. That covers the runtime, sshd and the
uplink modules at once, without each module opting out; the private network and the sshd module
need no change.
A first draft — never stop a unit the mesh did not create — was rejected while implementing it:
returning a converged node to adopted unloads the mesh's filter by exactly this path.
Found on the way: an undeclared `process` failed every apply on its node (`remove` had no case
for it). Now removed with its unit, timer and bundle — the mesh's own code. `user` and `archive`
have the same gap and are left for their own decisions: removing a login or unpacked files is not
something to settle in passing.
The unassign preview is partly answered — the host's plan names each unit it will stop — and the
controller's side is left open.