issue 130: undeclaring a service stops it, even one the mesh only reloads or keeps running
This commit is contained in:
@@ -0,0 +1,43 @@
|
|||||||
|
---
|
||||||
|
status: located
|
||||||
|
opened: 2026-09-27
|
||||||
|
located-in: [mesh-host internal/apply/apply.go (remove), mesh-controller internal/overlay, mesh-catalog modules/sshd]
|
||||||
|
---
|
||||||
|
|
||||||
|
# 130 — undeclaring a service stops it, even one the mesh only reloads or only keeps running
|
||||||
|
|
||||||
|
## What was observed
|
||||||
|
|
||||||
|
Reviewing the uplink modules ([ADR 0117](../../02-DECISIONS/0117-a-machines-uplink-is-a-seat.md))
|
||||||
|
found that the host's `remove` path stops every `service` resource that is no longer declared:
|
||||||
|
`SetServiceState(..., "stopped")`, reported as "stopped; the unit file is not the host's to
|
||||||
|
delete". `store.Orphans` matches by id alone. So any of these stops the unit:
|
||||||
|
|
||||||
|
- the module is unassigned — by mistake, or to switch it for another;
|
||||||
|
- the node is sent a deliberately-empty declaration ([issue 127](../127-a-declaration-that-shrinks-to-empty-is-skipped-not-sent/00-report.md));
|
||||||
|
- a later catalogue version renames the resource's `id`.
|
||||||
|
|
||||||
|
That is right for a service the mesh brought into being. It is wrong for a unit the mesh
|
||||||
|
declares only to act on — and the catalogue already has two:
|
||||||
|
|
||||||
|
- **The private network declares `docker.service`** (`registry-trust-reload`, state `running`)
|
||||||
|
so that a change to the registry trust reloads the runtime ([ADR 0102](../../02-DECISIONS/0102-the-mesh-writes-into-a-shared-file-never-over-it.md)).
|
||||||
|
Unassigning the private network stops the container runtime, and every container on the
|
||||||
|
machine with it — including ones the mesh does not manage.
|
||||||
|
- **The sshd module declares `sshd.service`.** Unassigning it stops the machine's ssh daemon:
|
||||||
|
the lockout the same module's `listens` rule says a firewall must never arrange.
|
||||||
|
|
||||||
|
The uplink modules would have added a third and a fourth: unassigning the network manager's
|
||||||
|
module would have stopped the network manager, taking the machine off the only link the mesh
|
||||||
|
reaches it by.
|
||||||
|
|
||||||
|
## What would have prevented it
|
||||||
|
|
||||||
|
- A service resource that says the unit's **lifecycle is the machine's**: declared with no
|
||||||
|
`state`, the mesh never starts, stops, enables or disables it; it only reloads or restarts a
|
||||||
|
*running* unit when a trigger changes; undeclared, it is left exactly as it is. (Being built
|
||||||
|
on mesh-host `feat/a-file-written-into-a-marked-block` for the uplink modules.)
|
||||||
|
- Then: `registry-trust-reload` declared that way (the runtime is the machine's), and the sshd
|
||||||
|
module's service too — a machine's ssh daemon outlives any module that configures it.
|
||||||
|
- A plan or unassign preview that names every unit an undeclare will stop, so the consequence
|
||||||
|
is read before it happens.
|
||||||
Reference in New Issue
Block a user