issue 130: undeclaring a service stops it, even one the mesh only reloads or keeps running #142

Merged
jschoubben merged 5 commits from issue/130-undeclaring-a-service-stops-it into main 2026-09-26 22:58:43 +00:00
3 changed files with 186 additions and 0 deletions
@@ -0,0 +1,119 @@
---
topic: what runs on it
status: accepted
date: 2026-09-27
deciders: jochen
reconstructed: false
extends: 0102-the-mesh-writes-into-a-shared-file-never-over-it.md
---
# 118. Undeclaring removes what the mesh made, and gives a unit back the state it was found in
## Context
When a resource stops being declared — its module unassigned, the node sent a
deliberately-empty declaration ([issue 127](../04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)),
or a new catalogue version renaming its id — the host undoes it. The host's own code states
the rule it means to follow: **it removes what it made and leaves what it merely configured.**
For almost every resource it does exactly that:
- a container, a network, a process's unit, a directory it created: removed;
- a file it created: removed; a file it replaced: its kept original put back
([ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md));
- keys and list members it wrote into a shared file: given back as they were
([ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md));
- a package: left installed — the host cannot know it is unused;
- an operator's path it was given access to: never touched
([ADR 0051](0051-shared-data-is-the-operators.md)).
**A service is the exception.** A `service` resource never installs a unit: it puts one that
already exists — the distribution's, the operator's — into a state. Undeclared, the host stops
it. That contradicts the rule above, and in practice it is the most dangerous thing an
undeclare can do. Found reviewing the uplink modules
([issue 130](../04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md)):
- the private network declares the container runtime's unit only so a change to the registry
trust reloads it — unassigning the private network stops the runtime, and every container on
the machine, the mesh's and not;
- the sshd module declares the ssh daemon — unassigning it stops ssh, the lockout that module's
own `listens` rule forbids;
- the uplink modules would have stopped the network manager, taking the machine off the only
link the mesh reaches it by.
[ADR 0117](0117-a-machines-uplink-is-a-seat.md) answered that for its own modules with a
service declared with no `state`. Every other module that declares a unit it did not make is
exposed in the same way, and relying on each author to remember an opt-out is how the next one
is missed.
## Considered Options
**1. Undeclaring touches nothing on the machine.** Rejected. What the mesh made would outlive
the module that made it: a container nobody manages keeps serving and stops being patched; a
unit the mesh wrote keeps running a bundle nothing updates; a name collides when the module
is assigned again. An undeclare that leaves the mesh's own work behind is an orphan factory.
**2. Keep stopping services; make "leave it running" an opt-in per resource.** Rejected. It
keeps the dangerous behaviour as the default for exactly the units that matter most — the
runtime, the ssh daemon, the network — and each new module is one forgotten field away from a
machine that goes dark when it is unassigned.
**3. Never stop a unit the mesh did not create.** Rejected, found while implementing it. The
mesh's packet filter is a unit the distribution installed and the mesh started at converge;
returning a node to adopted ([ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md))
unloads it by undeclaring it. Never stopping it would leave the mesh's filter loaded beside the
predecessor's firewall re-enabled — the one rollback a converge promises, broken. Who wrote the
unit file is not the line; what the mesh *did* to the unit is.
**4. Give the unit back the state it was found in.** Chosen.
## Decision
**Undeclaring removes what the mesh made and gives back what it changed.** For a unit the mesh
did not create, what it changed is the unit's state, so that is what is given back: **the host
records the state it first found the unit in, and undeclaring returns the unit to it.**
- **Recorded once**, the first time the host applies the service — whether it was running, and,
where the declaration sets it, whether it was enabled at boot — and carried in the host's
record from then on. Later applies never overwrite it: by then the unit's state is the mesh's
doing.
- **A unit found running is left running.** The container runtime, the ssh daemon, a network
manager: running before the mesh arrived, running after it leaves.
- **A unit the mesh started is stopped again**, and one it enabled is disabled again — the packet
filter a converge loaded, which returning to adopted unloads.
- **Never started on the way out.** A unit the mesh stopped is not started again when its
declaration goes; starting something is a decision, and the operator makes it.
- **Unknown is left alone.** A record written before the host kept what it found says nothing
about the unit before the mesh; the unit is left exactly as it is. A unit left running can be
stopped by the operator; one stopped by mistake may be the link the operator needed to do it.
- A unit the mesh *did* create — a `process` resource's unit and bundle — is stopped and removed
with its declaration. That is the mesh's own code. (Before this record there was no way to
remove one at all: an undeclared process failed every apply on its node.)
- The service's settings the mesh wrote are given back by their own resources (a kept original
restored, a region or keys removed). A running service keeps running on what it read until it
next reads its configuration; the mesh does not restart it to make it notice.
- A service declared with no `state` (ADR 0117) remains the way to say the mesh must not
**start** a unit either; undeclared, it is forgotten.
## Consequences
- Unassigning the private network no longer stops the container runtime; unassigning sshd no
longer stops ssh; no uplink module can take a machine's network down on its way out.
- The host's removal report says what it gave back — "restored: stopped again, as the host
found it" — or "forgotten: it was running before the mesh; left as it is" where it used to say
"stopped". Its plan names each unit an undeclare will stop, before it does.
- On a fresh machine where the mesh installed and started a service, unassigning its module
stops it again — the mesh gave, the mesh takes back. An operator who wants it kept declares it
in a module of their own, or starts it themselves after.
- A daemon can keep running after its module is gone, on configuration that was taken back from
under it. That is a visible, running process the operator can see and stop; the alternative
was an invisible outage.
- **Not decided here:** an unassign preview that lists what an undeclare will remove and what it
will leave running. Issue 130 asks for it; it is the controller's to build.
## References
- [issue 130](../04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md): the finding
- [ADR 0117](0117-a-machines-uplink-is-a-seat.md): the uplink modules, and a service with no state
- [ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md), [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md):
what is given back, and how
- mesh-host `internal/apply/apply.go` (`remove`, the service case)
+1
View File
@@ -200,6 +200,7 @@ python3 00-META/checks/index.py fail if stale
- **0114** — [A credential two parties hold rotates over two credentials; one a single party holds rotates in place, staged; and retiring a credential never removes what it reached](0114-a-shared-credential-rotates-over-two-credentials.md) *(proposed)*
- **0115** — [One assignment of a module per node: the module's name is the assignment's identity](0115-one-assignment-of-a-module-per-node.md) *(proposed)*
- **0117** — [A machine's uplink is a seat: the mesh configures the manager, never the link](0117-a-machines-uplink-is-a-seat.md)
- **0118** — [Undeclaring removes what the mesh made, and gives a unit back the state it was found in](0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md)
### How it is built
@@ -0,0 +1,66 @@
---
status: located
opened: 2026-09-27
located-in: [mesh-host internal/apply/apply.go (remove)]
amended-design: 02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md
---
# 130 — undeclaring a service stops it, even one the mesh only reloads or only keeps running
## What was observed
Reviewing the uplink modules ([ADR 0117](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md))
found that the host's `remove` path stops every `service` resource that is no longer declared:
`SetServiceState(..., "stopped")`, reported as "stopped; the unit file is not the host's to
delete". `store.Orphans` matches by id alone. So any of these stops the unit:
- the module is unassigned — by mistake, or to switch it for another;
- the node is sent a deliberately-empty declaration ([issue 127](../127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md));
- a later catalogue version renames the resource's `id`.
That is right for a service the mesh brought into being. It is wrong for a unit the mesh
declares only to act on — and the catalogue already has two:
- **The private network declares `docker.service`** (`registry-trust-reload`, state `running`)
so that a change to the registry trust reloads the runtime ([ADR 0102](../../02-DECISIONS/0102-the-mesh-writes-into-a-shared-file-never-over-it.md)).
Unassigning the private network stops the container runtime, and every container on the
machine with it — including ones the mesh does not manage.
- **The sshd module declares `sshd.service`.** Unassigning it stops the machine's ssh daemon:
the lockout the same module's `listens` rule says a firewall must never arrange.
The uplink modules would have added a third and a fourth: unassigning the network manager's
module would have stopped the network manager, taking the machine off the only link the mesh
reaches it by.
## What would have prevented it
- A service resource that says the unit's **lifecycle is the machine's**: declared with no
`state`, the mesh never starts, stops, enables or disables it; it only reloads or restarts a
*running* unit when a trigger changes; undeclared, it is left exactly as it is. (Being built
on mesh-host `feat/a-file-written-into-a-marked-block` for the uplink modules.)
- Then: `registry-trust-reload` declared that way (the runtime is the machine's), and the sshd
module's service too — a machine's ssh daemon outlives any module that configures it.
- A plan or unassign preview that names every unit an undeclare will stop, so the consequence
is read before it happens.
## Resolution
[ADR 0118](../../02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md):
undeclaring removes what the mesh made and gives back what it changed. The host records the state
it first found a unit in, and undeclaring returns the unit to it — a unit found running (the
container runtime, sshd, a network manager) is left running; one the mesh started (the packet
filter a converge loaded) is stopped again; nothing is started on the way out; a record from
before the host kept what it found leaves the unit alone. That covers the runtime, sshd and the
uplink modules at once, without each module opting out; the private network and the sshd module
need no change.
A first draft — never stop a unit the mesh did not create — was rejected while implementing it:
returning a converged node to adopted unloads the mesh's filter by exactly this path.
Found on the way: an undeclared `process` failed every apply on its node (`remove` had no case
for it). Now removed with its unit, timer and bundle — the mesh's own code. `user` and `archive`
have the same gap and are left for their own decisions: removing a login or unpacked files is not
something to settle in passing.
The unassign preview is partly answered — the host's plan names each unit it will stop — and the
controller's side is left open.