Files
hq/02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md
jschoubben ce6ae943b7 Merge main: renumber this branch's records around the trunk's
Both lines of work numbered from the same point, so four decision records and one design
document existed twice with different content. The trunk keeps its numbers and this branch
yields — the only rule that scales, because the trunk's are already cited by what merged
before them.

  0117 the bus is the only broker        -> 0125
  0118 a module declares its own seats   -> 0126
  0119 amqp is a provision, not the bus  -> 0127
  0120 the mesh bus is required          -> 0128
  0123 a seat carries its role's protocol -> 0129
  0124 the predecessor is ending          -> 0130
  design 29, what a module declares       -> design 32

Applied to the code repositories too, because a stale reference is worse when numbers
collide than when they dangle: the reader lands on a real record that decided something
else.

Two reconciliations the merge forced, both real:

**0110 was marked wholly superseded and was not.** Its successor says in as many words that
everything 0110 decided about what a seat *is* stands untouched — and two records that
landed on the trunk rest on exactly that part. So it is accepted again, extended rather than
replaced, with a note saying which of its claims moved and where.

**A seat's protocol becomes columns, not fields.** The trunk moved the seat set out of
compiled code into a table the controller owns. This branch had added what a role accepts,
emits and serves to the Go slice. The decision is unaffected and the mechanism is better for
it: giving a role a protocol is now a write rather than a rebuild, which is the trunk's own
argument applied to what this branch added.

One check still fails and it fails on main too: a record resting on ADR 0112 while that is
still 'proposed'. Left alone — it is not this merge's to answer.
2026-09-27 18:23:41 +02:00

120 lines
7.2 KiB
Markdown

---
topic: what runs on it
status: accepted
date: 2026-09-27
deciders: jochen
reconstructed: false
extends: 0102-the-mesh-writes-into-a-shared-file-never-over-it.md
---
# 118. Undeclaring removes what the mesh made, and gives a unit back the state it was found in
## Context
When a resource stops being declared — its module unassigned, the node sent a
deliberately-empty declaration ([issue 127](../04-ISSUES/127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md)),
or a new catalogue version renaming its id — the host undoes it. The host's own code states
the rule it means to follow: **it removes what it made and leaves what it merely configured.**
For almost every resource it does exactly that:
- a container, a network, a process's unit, a directory it created: removed;
- a file it created: removed; a file it replaced: its kept original put back
([ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md));
- keys and list members it wrote into a shared file: given back as they were
([ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md));
- a package: left installed — the host cannot know it is unused;
- an operator's path it was given access to: never touched
([ADR 0051](0051-shared-data-is-the-operators.md)).
**A service is the exception.** A `service` resource never installs a unit: it puts one that
already exists — the distribution's, the operator's — into a state. Undeclared, the host stops
it. That contradicts the rule above, and in practice it is the most dangerous thing an
undeclare can do. Found reviewing the uplink modules
([issue 130](../04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md)):
- the private network declares the container runtime's unit only so a change to the registry
trust reloads it — unassigning the private network stops the runtime, and every container on
the machine, the mesh's and not;
- the sshd module declares the ssh daemon — unassigning it stops ssh, the lockout that module's
own `listens` rule forbids;
- the uplink modules would have stopped the network manager, taking the machine off the only
link the mesh reaches it by.
[ADR 0125](0117-a-machines-uplink-is-a-seat.md) answered that for its own modules with a
service declared with no `state`. Every other module that declares a unit it did not make is
exposed in the same way, and relying on each author to remember an opt-out is how the next one
is missed.
## Considered Options
**1. Undeclaring touches nothing on the machine.** Rejected. What the mesh made would outlive
the module that made it: a container nobody manages keeps serving and stops being patched; a
unit the mesh wrote keeps running a bundle nothing updates; a name collides when the module
is assigned again. An undeclare that leaves the mesh's own work behind is an orphan factory.
**2. Keep stopping services; make "leave it running" an opt-in per resource.** Rejected. It
keeps the dangerous behaviour as the default for exactly the units that matter most — the
runtime, the ssh daemon, the network — and each new module is one forgotten field away from a
machine that goes dark when it is unassigned.
**3. Never stop a unit the mesh did not create.** Rejected, found while implementing it. The
mesh's packet filter is a unit the distribution installed and the mesh started at converge;
returning a node to adopted ([ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md))
unloads it by undeclaring it. Never stopping it would leave the mesh's filter loaded beside the
predecessor's firewall re-enabled — the one rollback a converge promises, broken. Who wrote the
unit file is not the line; what the mesh *did* to the unit is.
**4. Give the unit back the state it was found in.** Chosen.
## Decision
**Undeclaring removes what the mesh made and gives back what it changed.** For a unit the mesh
did not create, what it changed is the unit's state, so that is what is given back: **the host
records the state it first found the unit in, and undeclaring returns the unit to it.**
- **Recorded once**, the first time the host applies the service — whether it was running, and,
where the declaration sets it, whether it was enabled at boot — and carried in the host's
record from then on. Later applies never overwrite it: by then the unit's state is the mesh's
doing.
- **A unit found running is left running.** The container runtime, the ssh daemon, a network
manager: running before the mesh arrived, running after it leaves.
- **A unit the mesh started is stopped again**, and one it enabled is disabled again — the packet
filter a converge loaded, which returning to adopted unloads.
- **Never started on the way out.** A unit the mesh stopped is not started again when its
declaration goes; starting something is a decision, and the operator makes it.
- **Unknown is left alone.** A record written before the host kept what it found says nothing
about the unit before the mesh; the unit is left exactly as it is. A unit left running can be
stopped by the operator; one stopped by mistake may be the link the operator needed to do it.
- A unit the mesh *did* create — a `process` resource's unit and bundle — is stopped and removed
with its declaration. That is the mesh's own code. (Before this record there was no way to
remove one at all: an undeclared process failed every apply on its node.)
- The service's settings the mesh wrote are given back by their own resources (a kept original
restored, a region or keys removed). A running service keeps running on what it read until it
next reads its configuration; the mesh does not restart it to make it notice.
- A service declared with no `state` (ADR 0125) remains the way to say the mesh must not
**start** a unit either; undeclared, it is forgotten.
## Consequences
- Unassigning the private network no longer stops the container runtime; unassigning sshd no
longer stops ssh; no uplink module can take a machine's network down on its way out.
- The host's removal report says what it gave back — "restored: stopped again, as the host
found it" — or "forgotten: it was running before the mesh; left as it is" where it used to say
"stopped". Its plan names each unit an undeclare will stop, before it does.
- On a fresh machine where the mesh installed and started a service, unassigning its module
stops it again — the mesh gave, the mesh takes back. An operator who wants it kept declares it
in a module of their own, or starts it themselves after.
- A daemon can keep running after its module is gone, on configuration that was taken back from
under it. That is a visible, running process the operator can see and stop; the alternative
was an invisible outage.
- **Not decided here:** an unassign preview that lists what an undeclare will remove and what it
will leave running. Issue 130 asks for it; it is the controller's to build.
## References
- [issue 130](../04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md): the finding
- [ADR 0125](0117-a-machines-uplink-is-a-seat.md): the uplink modules, and a service with no state
- [ADR 0102](0102-the-mesh-writes-into-a-shared-file-never-over-it.md), [ADR 0100](0100-a-node-in-use-is-adopted-before-it-is-converged.md):
what is given back, and how
- mesh-host `internal/apply/apply.go` (`remove`, the service case)