Files
hq/04-ISSUES/130-undeclaring-a-service-stops-it/00-report.md
T

67 lines
3.8 KiB
Markdown

---
status: located
opened: 2026-09-27
located-in: [mesh-host internal/apply/apply.go (remove)]
amended-design: 02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md
---
# 130 — undeclaring a service stops it, even one the mesh only reloads or only keeps running
## What was observed
Reviewing the uplink modules ([ADR 0117](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md))
found that the host's `remove` path stops every `service` resource that is no longer declared:
`SetServiceState(..., "stopped")`, reported as "stopped; the unit file is not the host's to
delete". `store.Orphans` matches by id alone. So any of these stops the unit:
- the module is unassigned — by mistake, or to switch it for another;
- the node is sent a deliberately-empty declaration ([issue 127](../127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md));
- a later catalogue version renames the resource's `id`.
That is right for a service the mesh brought into being. It is wrong for a unit the mesh
declares only to act on — and the catalogue already has two:
- **The private network declares `docker.service`** (`registry-trust-reload`, state `running`)
so that a change to the registry trust reloads the runtime ([ADR 0102](../../02-DECISIONS/0102-the-mesh-writes-into-a-shared-file-never-over-it.md)).
Unassigning the private network stops the container runtime, and every container on the
machine with it — including ones the mesh does not manage.
- **The sshd module declares `sshd.service`.** Unassigning it stops the machine's ssh daemon:
the lockout the same module's `listens` rule says a firewall must never arrange.
The uplink modules would have added a third and a fourth: unassigning the network manager's
module would have stopped the network manager, taking the machine off the only link the mesh
reaches it by.
## What would have prevented it
- A service resource that says the unit's **lifecycle is the machine's**: declared with no
`state`, the mesh never starts, stops, enables or disables it; it only reloads or restarts a
*running* unit when a trigger changes; undeclared, it is left exactly as it is. (Being built
on mesh-host `feat/a-file-written-into-a-marked-block` for the uplink modules.)
- Then: `registry-trust-reload` declared that way (the runtime is the machine's), and the sshd
module's service too — a machine's ssh daemon outlives any module that configures it.
- A plan or unassign preview that names every unit an undeclare will stop, so the consequence
is read before it happens.
## Resolution
[ADR 0118](../../02-DECISIONS/0118-undeclaring-gives-a-unit-back-the-state-it-was-found-in.md):
undeclaring removes what the mesh made and gives back what it changed. The host records the state
it first found a unit in, and undeclaring returns the unit to it — a unit found running (the
container runtime, sshd, a network manager) is left running; one the mesh started (the packet
filter a converge loaded) is stopped again; nothing is started on the way out; a record from
before the host kept what it found leaves the unit alone. That covers the runtime, sshd and the
uplink modules at once, without each module opting out; the private network and the sshd module
need no change.
A first draft — never stop a unit the mesh did not create — was rejected while implementing it:
returning a converged node to adopted unloads the mesh's filter by exactly this path.
Found on the way: an undeclared `process` failed every apply on its node (`remove` had no case
for it). Now removed with its unit, timer and bundle — the mesh's own code. `user` and `archive`
have the same gap and are left for their own decisions: removing a login or unpacked files is not
something to settle in passing.
The unassign preview is partly answered — the host's plan names each unit it will stop — and the
controller's side is left open.