Issue 315: a failed unit of a module was never raised

Liveness judged only long-running resources, so a module whose daemon
is a package's unit failed at every login while its machine read
healthy, and a degraded service manager was said by nothing.
This commit is contained in:
jochen
2026-10-08 11:30:08 +02:00
parent b11219e3de
commit 7121d0737e
2 changed files with 138 additions and 0 deletions
@@ -0,0 +1,49 @@
---
status: located
opened: 2026-10-08
located-in: [mesh-host (liveness judged only long-running resources; nothing read the service managers' failed units), mesh-controller (no finding for a degraded service manager)]
fixed-by: novox/mesh-host PR #53, novox/mesh-controller PR #137
amended-design:
---
# 315. A failed unit of a module was never raised
## Symptom
On 2026-10-08 the controller's view of the workstation said every part healthy: every long-running
resource, every part of its network. The same machine's reported capabilities said
`service-manager degraded`. Its service-manager seat listed four failed units:
| unit | manager | whose |
|---|---|---|
| `openrazer-daemon.service` | the operator account's own | the `openrazer` module's package `openrazer-daemon` ships it; the module is assigned to the machine |
| `greenclip.service` | the operator account's own | nobody's now: load `not-found`, the unit of the clipboard tool the `clipmenu` module replaced |
| two network mounts | the machine's | the operator's own mount table; no module declares them |
No condition was raised for any of them. The laptop showed the same Razer daemon failed, and nothing
raised that either.
The Razer daemon had failed at every login on both workstations since the module was assigned. That
is the failure ADR 0240 exists to make loud: a module's software not working while every check of it
passes.
## Why
- **Liveness judges long-running resources and nothing else** (ADR 0240 rule 1): a container that
stays up, a process, a service stated `running`. The `openrazer` module declares three packages
and no service. Its daemon is started by D-Bus activation through the unit its package ships, and
the module says on purpose that it does not enable or start that unit. So nothing the node-engine
judged was the daemon.
- **Nothing read the service managers' failed units.** The profile's `service-manager` capability
reads `is-system-running` and reports `degraded`. A capability is what a machine can do, not a
finding, and nothing raised it.
- **A failed unit that no module places had no place to be said at all.** ADR 0241 made an outside
writer of a mesh file a machine-level finding. A degraded service manager had no such finding.
## Asked
1. A failed unit that a mesh module places is that module's unhealthy resource, and it raises the
module's condition naming the unit. That includes a user-scope unit and a unit a module's package
ships.
2. A failed unit the mesh does not place is one finding for the machine, listing the units, so
`degraded` never hides. It clears when no such unit is failed.
@@ -0,0 +1,89 @@
# 315 — diagnosis
**2026-10-08.** Read-only, through the mesh's tools.
## Each failed unit
1. **The Razer daemon** (`openrazer-daemon.service`, the operator account's manager, `exit-code`).
The module's status tool and the daemon's last words agree: *User is not a member of the openrazer
group*. The driver is built and loaded, and a Razer mouse is bound to it, so a device is present.
The account is not in the `openrazer` group that the driver's udev rules give the devices to. The
module's README already names this as the operator's one-time step (add the account to the group,
log in again). The module cannot declare it, because the operator's account is already another
module's `user` resource and a second owner of one account is refused at composition. **So this is
the machine, not a module bug:** the documented migration step was never taken. The laptop is in the
same state. Nothing raised it, and this issue fixes that.
2. **The two mounts** are units generated from the machine's own mount table at boot. The
service-manager seat says neither is declared by the mesh, and no catalogue module or plan writes
the mount table. Both failed at the boot of 2026-10-04:
- one is a share on a game console's address on the home network. Mounting it failed in six seconds
with "could not connect" (the console was off). It had failed the same way at the previous boot,
in August;
- the other is a media library mount. It mounted at the August boot and timed out after 90 s at this
one. That is the stock mount timeout, the usual sign of a network share whose server did not
answer while the machine booted. The mesh gives no tool that reads the mount table, so whether its
source is the home server's media library is not confirmed here. The operator knows.
Both are the operator's own entries. They are not the mesh's to change.
3. **The leftover clipboard unit** (`greenclip.service`, the account's manager, load `not-found`). The
`clipmenu` module declares `rofi-greenclip` absent, so the package and its unit file were removed.
On the workstation the unit had also been enabled by hand. The enable link in the account's
`default.target.wants` stayed, as the module's README says ("it dangles once the package is gone").
The unit last ran at the boot of 2026-10-04, before X was up, and hit its start limit. The account's
manager still holds that failed state for a unit whose file is gone. Nothing the mesh did failed
here. The dangling link is the leftover, and the account's manager keeps the failure until it is
reset or the session ends.
## Where it lives
- **mesh-host `internal/liveness`**: `LongRunning` takes containers, services stated `running` and
long-running processes. A package's unit is outside it. A service whose lifecycle is the machine's
is too, and so is a unit file a module writes. Nothing in the node-engine read `list-units
--state=failed`.
- **mesh-host `internal/profile`**: `service-manager` reads `is-system-running` as a capability
detail, which is never raised.
- **mesh-controller**: `node_health` kept resources and network, and nothing about the machine's
managers. No condition kind existed for a degraded one.
## Fix
- **The node-engine** (mesh-host PR #53, new `internal/units`). On every liveness look it reads the
failed units of the machine's manager and of every account manager the declaration names. The two-look
rule is its own, as in ADR 0241 §2. A failed unit is a module's when:
1. the declaration states it as a service or process;
2. the unit's file is a file the declaration writes; or
3. the unit's file belongs to a package the declaration installs (asked of the package manager once
per file until the next apply).
A module's failed unit is said among the resources as that module's unhealthy resource of kind
`unit`, which raises the module's condition on the controller as it stands. Every other failed unit
is the machine's own, in a new `units` part of the health statement. A unit liveness already judges
is left to it, so each failure is said once. It reads and never acts.
- **The controller** (mesh-controller PR #137). It keeps `units` with the machine's statement
(migration 0080). It names the unit in a module's condition. It raises one warning for the machine's
own failed units, `machine.<m>.units`, listing them, and clears it on the first statement listing
none. Nothing a send moved is among them, so the gate never holds a send on it. `node show` lists
them.
On the workstation, once both are live, this raises `module.openrazer.<workstation>.unhealthy` naming
the daemon's unit, and the same on the laptop. It also raises one machine finding on the workstation
listing the two mounts and the clipboard leftover.
## Left to the operator
- **The Razer daemon:** the module's migration step (add the account to `openrazer`, log in again).
After that, `openrazer_check` answers ok and the condition clears.
- **The clipboard leftover:** remove the dangling enable link
`~/.config/systemd/user/default.target.wants/greenclip.service`, then reset the unit's failed state
in the account's manager. Both are the account's own files and state, and no module manages them.
The service-manager seat's `disable` (user scope) is the tool for the link, if it accepts a unit
whose file is gone. Otherwise remove the link by hand. Neither was done here.
- **The two mounts:** the operator's own mount table. Either mark them so a missing server is not a
failed boot unit, or accept the machine finding while the server is away.
## Ruled out
- *A catalogue bug in `openrazer`*: the daemon refuses on a group the module documents as the
operator's step, with a device present. Making the daemon not fail without a device would not
change anything here.
- *The mesh owning the mounts*: the seat says neither is declared, and no plan writes the mount table.