Merge pull request 'Issue 315: a failed unit of a module was never raised' (#199) from issues/315-a-failed-unit-was-never-raised into main
This commit was merged in pull request #199.
This commit is contained in:
@@ -0,0 +1,49 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-08
|
||||
located-in: [mesh-host (liveness judged only long-running resources; nothing read the service managers' failed units), mesh-controller (no finding for a degraded service manager)]
|
||||
fixed-by: novox/mesh-host PR #53, novox/mesh-controller PR #137
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 315. A failed unit of a module was never raised
|
||||
|
||||
## Symptom
|
||||
|
||||
On 2026-10-08 the controller's view of the workstation said every part healthy: every long-running
|
||||
resource, every part of its network. The same machine's reported capabilities said
|
||||
`service-manager degraded`. Its service-manager seat listed four failed units:
|
||||
|
||||
| unit | manager | whose |
|
||||
|---|---|---|
|
||||
| `openrazer-daemon.service` | the operator account's own | the `openrazer` module's package `openrazer-daemon` ships it; the module is assigned to the machine |
|
||||
| `greenclip.service` | the operator account's own | nobody's now: load `not-found`, the unit of the clipboard tool the `clipmenu` module replaced |
|
||||
| two network mounts | the machine's | the operator's own mount table; no module declares them |
|
||||
|
||||
No condition was raised for any of them. The laptop showed the same Razer daemon failed, and nothing
|
||||
raised that either.
|
||||
|
||||
The Razer daemon had failed at every login on both workstations since the module was assigned. That
|
||||
is the failure ADR 0240 exists to make loud: a module's software not working while every check of it
|
||||
passes.
|
||||
|
||||
## Why
|
||||
|
||||
- **Liveness judges long-running resources and nothing else** (ADR 0240 rule 1): a container that
|
||||
stays up, a process, a service stated `running`. The `openrazer` module declares three packages
|
||||
and no service. Its daemon is started by D-Bus activation through the unit its package ships, and
|
||||
the module says on purpose that it does not enable or start that unit. So nothing the node-engine
|
||||
judged was the daemon.
|
||||
- **Nothing read the service managers' failed units.** The profile's `service-manager` capability
|
||||
reads `is-system-running` and reports `degraded`. A capability is what a machine can do, not a
|
||||
finding, and nothing raised it.
|
||||
- **A failed unit that no module places had no place to be said at all.** ADR 0241 made an outside
|
||||
writer of a mesh file a machine-level finding. A degraded service manager had no such finding.
|
||||
|
||||
## Asked
|
||||
|
||||
1. A failed unit that a mesh module places is that module's unhealthy resource, and it raises the
|
||||
module's condition naming the unit. That includes a user-scope unit and a unit a module's package
|
||||
ships.
|
||||
2. A failed unit the mesh does not place is one finding for the machine, listing the units, so
|
||||
`degraded` never hides. It clears when no such unit is failed.
|
||||
@@ -0,0 +1,89 @@
|
||||
# 315 — diagnosis
|
||||
|
||||
**2026-10-08.** Read-only, through the mesh's tools.
|
||||
|
||||
## Each failed unit
|
||||
|
||||
1. **The Razer daemon** (`openrazer-daemon.service`, the operator account's manager, `exit-code`).
|
||||
The module's status tool and the daemon's last words agree: *User is not a member of the openrazer
|
||||
group*. The driver is built and loaded, and a Razer mouse is bound to it, so a device is present.
|
||||
The account is not in the `openrazer` group that the driver's udev rules give the devices to. The
|
||||
module's README already names this as the operator's one-time step (add the account to the group,
|
||||
log in again). The module cannot declare it, because the operator's account is already another
|
||||
module's `user` resource and a second owner of one account is refused at composition. **So this is
|
||||
the machine, not a module bug:** the documented migration step was never taken. The laptop is in the
|
||||
same state. Nothing raised it, and this issue fixes that.
|
||||
2. **The two mounts** are units generated from the machine's own mount table at boot. The
|
||||
service-manager seat says neither is declared by the mesh, and no catalogue module or plan writes
|
||||
the mount table. Both failed at the boot of 2026-10-04:
|
||||
- one is a share on a retro-gaming box on the home network. Mounting it failed in six seconds
|
||||
with "could not connect" (the box was off). It had failed the same way at the previous boot,
|
||||
in August;
|
||||
- the other is a media library mount. It mounted at the August boot and timed out after 90 s at this
|
||||
one. That is the stock mount timeout, the usual sign of a network share whose server did not
|
||||
answer while the machine booted. The mesh gives no tool that reads the mount table, so whether its
|
||||
source is the home server's media library is not confirmed here. The operator knows.
|
||||
|
||||
Both are the operator's own entries. They are not the mesh's to change.
|
||||
3. **The leftover clipboard unit** (`greenclip.service`, the account's manager, load `not-found`). The
|
||||
`clipmenu` module declares `rofi-greenclip` absent, so the package and its unit file were removed.
|
||||
On the workstation the unit had also been enabled by hand. The enable link in the account's
|
||||
`default.target.wants` stayed, as the module's README says ("it dangles once the package is gone").
|
||||
The unit last ran at the boot of 2026-10-04, before X was up, and hit its start limit. The account's
|
||||
manager still holds that failed state for a unit whose file is gone. Nothing the mesh did failed
|
||||
here. The dangling link is the leftover, and the account's manager keeps the failure until it is
|
||||
reset or the session ends.
|
||||
|
||||
## Where it lives
|
||||
|
||||
- **`mesh-host` `internal/liveness`**: `LongRunning` takes containers, services stated `running` and
|
||||
long-running processes. A package's unit is outside it. A service whose lifecycle is the machine's
|
||||
is too, and so is a unit file a module writes. Nothing in the node-engine read `list-units
|
||||
--state=failed`.
|
||||
- **`mesh-host` `internal/profile`**: `service-manager` reads `is-system-running` as a capability
|
||||
detail, which is never raised.
|
||||
- **mesh-controller**: `node_health` kept resources and network, and nothing about the machine's
|
||||
managers. No condition kind existed for a degraded one.
|
||||
|
||||
## Fix
|
||||
|
||||
- **The node-engine** (`mesh-host` PR #53, new `internal/units`). On every liveness look it reads the
|
||||
failed units of the machine's manager and of every account manager the declaration names. The two-look
|
||||
rule is its own, as in ADR 0241 §2. A failed unit is a module's when:
|
||||
1. the declaration states it as a service or process;
|
||||
2. the unit's file is a file the declaration writes; or
|
||||
3. the unit's file belongs to a package the declaration installs (asked of the package manager once
|
||||
per file until the next apply).
|
||||
|
||||
A module's failed unit is said among the resources as that module's unhealthy resource of kind
|
||||
`unit`, which raises the module's condition on the controller as it stands. Every other failed unit
|
||||
is the machine's own, in a new `units` part of the health statement. A unit liveness already judges
|
||||
is left to it, so each failure is said once. It reads and never acts.
|
||||
- **The controller** (mesh-controller PR #137). It keeps `units` with the machine's statement
|
||||
(migration 0080). It names the unit in a module's condition. It raises one warning for the machine's
|
||||
own failed units, `machine.<m>.units`, listing them, and clears it on the first statement listing
|
||||
none. Nothing a send moved is among them, so the gate never holds a send on it. `node show` lists
|
||||
them.
|
||||
|
||||
On the workstation, once both are live, this raises `module.openrazer.<workstation>.unhealthy` naming
|
||||
the daemon's unit, and the same on the laptop. It also raises one machine finding on the workstation
|
||||
listing the two mounts and the clipboard leftover.
|
||||
|
||||
## Left to the operator
|
||||
|
||||
- **The Razer daemon:** the module's migration step (add the account to `openrazer`, log in again).
|
||||
After that, `openrazer_check` answers ok and the condition clears.
|
||||
- **The clipboard leftover:** remove the dangling enable link
|
||||
`~/.config/systemd/user/default.target.wants/greenclip.service`, then reset the unit's failed state
|
||||
in the account's manager. Both are the account's own files and state, and no module manages them.
|
||||
The service-manager seat's `disable` (user scope) is the tool for the link, if it accepts a unit
|
||||
whose file is gone. Otherwise remove the link by hand. Neither was done here.
|
||||
- **The two mounts:** the operator's own mount table. Either mark them so a missing server is not a
|
||||
failed boot unit, or accept the machine finding while the server is away.
|
||||
|
||||
## Ruled out
|
||||
|
||||
- *A catalogue bug in `openrazer`*: the daemon refuses on a group the module documents as the
|
||||
operator's step, with a device present. Making the daemon not fail without a device would not
|
||||
change anything here.
|
||||
- *The mesh owning the mounts*: the seat says neither is declared, and no plan writes the mount table.
|
||||
Reference in New Issue
Block a user