From 7121d0737e994799257e182fe21f25a0b7216ec8 Mon Sep 17 00:00:00 2001 From: jochen Date: Thu, 8 Oct 2026 11:30:08 +0200 Subject: [PATCH 1/2] Issue 315: a failed unit of a module was never raised Liveness judged only long-running resources, so a module whose daemon is a package's unit failed at every login while its machine read healthy, and a degraded service manager was said by nothing. --- .../00-report.md | 49 ++++++++++ .../01-diagnosis.md | 89 +++++++++++++++++++ 2 files changed, 138 insertions(+) create mode 100644 04-ISSUES/315-a-failed-unit-of-a-module-was-never-raised/00-report.md create mode 100644 04-ISSUES/315-a-failed-unit-of-a-module-was-never-raised/01-diagnosis.md diff --git a/04-ISSUES/315-a-failed-unit-of-a-module-was-never-raised/00-report.md b/04-ISSUES/315-a-failed-unit-of-a-module-was-never-raised/00-report.md new file mode 100644 index 00000000..1f0734a4 --- /dev/null +++ b/04-ISSUES/315-a-failed-unit-of-a-module-was-never-raised/00-report.md @@ -0,0 +1,49 @@ +--- +status: located +opened: 2026-10-08 +located-in: [mesh-host (liveness judged only long-running resources; nothing read the service managers' failed units), mesh-controller (no finding for a degraded service manager)] +fixed-by: novox/mesh-host PR #53, novox/mesh-controller PR #137 +amended-design: +--- + +# 315. A failed unit of a module was never raised + +## Symptom + +On 2026-10-08 the controller's view of the workstation said every part healthy: every long-running +resource, every part of its network. The same machine's reported capabilities said +`service-manager degraded`. Its service-manager seat listed four failed units: + +| unit | manager | whose | +|---|---|---| +| `openrazer-daemon.service` | the operator account's own | the `openrazer` module's package `openrazer-daemon` ships it; the module is assigned to the machine | +| `greenclip.service` | the operator account's own | nobody's now: load `not-found`, the unit of the clipboard tool the `clipmenu` module replaced | +| two network mounts | the machine's | the operator's own mount table; no module declares them | + +No condition was raised for any of them. The laptop showed the same Razer daemon failed, and nothing +raised that either. + +The Razer daemon had failed at every login on both workstations since the module was assigned. That +is the failure ADR 0240 exists to make loud: a module's software not working while every check of it +passes. + +## Why + +- **Liveness judges long-running resources and nothing else** (ADR 0240 rule 1): a container that + stays up, a process, a service stated `running`. The `openrazer` module declares three packages + and no service. Its daemon is started by D-Bus activation through the unit its package ships, and + the module says on purpose that it does not enable or start that unit. So nothing the node-engine + judged was the daemon. +- **Nothing read the service managers' failed units.** The profile's `service-manager` capability + reads `is-system-running` and reports `degraded`. A capability is what a machine can do, not a + finding, and nothing raised it. +- **A failed unit that no module places had no place to be said at all.** ADR 0241 made an outside + writer of a mesh file a machine-level finding. A degraded service manager had no such finding. + +## Asked + +1. A failed unit that a mesh module places is that module's unhealthy resource, and it raises the + module's condition naming the unit. That includes a user-scope unit and a unit a module's package + ships. +2. A failed unit the mesh does not place is one finding for the machine, listing the units, so + `degraded` never hides. It clears when no such unit is failed. diff --git a/04-ISSUES/315-a-failed-unit-of-a-module-was-never-raised/01-diagnosis.md b/04-ISSUES/315-a-failed-unit-of-a-module-was-never-raised/01-diagnosis.md new file mode 100644 index 00000000..121b9216 --- /dev/null +++ b/04-ISSUES/315-a-failed-unit-of-a-module-was-never-raised/01-diagnosis.md @@ -0,0 +1,89 @@ +# 315 — diagnosis + +**2026-10-08.** Read-only, through the mesh's tools. + +## Each failed unit + +1. **The Razer daemon** (`openrazer-daemon.service`, the operator account's manager, `exit-code`). + The module's status tool and the daemon's last words agree: *User is not a member of the openrazer + group*. The driver is built and loaded, and a Razer mouse is bound to it, so a device is present. + The account is not in the `openrazer` group that the driver's udev rules give the devices to. The + module's README already names this as the operator's one-time step (add the account to the group, + log in again). The module cannot declare it, because the operator's account is already another + module's `user` resource and a second owner of one account is refused at composition. **So this is + the machine, not a module bug:** the documented migration step was never taken. The laptop is in the + same state. Nothing raised it, and this issue fixes that. +2. **The two mounts** are units generated from the machine's own mount table at boot. The + service-manager seat says neither is declared by the mesh, and no catalogue module or plan writes + the mount table. Both failed at the boot of 2026-10-04: + - one is a share on a game console's address on the home network. Mounting it failed in six seconds + with "could not connect" (the console was off). It had failed the same way at the previous boot, + in August; + - the other is a media library mount. It mounted at the August boot and timed out after 90 s at this + one. That is the stock mount timeout, the usual sign of a network share whose server did not + answer while the machine booted. The mesh gives no tool that reads the mount table, so whether its + source is the home server's media library is not confirmed here. The operator knows. + + Both are the operator's own entries. They are not the mesh's to change. +3. **The leftover clipboard unit** (`greenclip.service`, the account's manager, load `not-found`). The + `clipmenu` module declares `rofi-greenclip` absent, so the package and its unit file were removed. + On the workstation the unit had also been enabled by hand. The enable link in the account's + `default.target.wants` stayed, as the module's README says ("it dangles once the package is gone"). + The unit last ran at the boot of 2026-10-04, before X was up, and hit its start limit. The account's + manager still holds that failed state for a unit whose file is gone. Nothing the mesh did failed + here. The dangling link is the leftover, and the account's manager keeps the failure until it is + reset or the session ends. + +## Where it lives + +- **mesh-host `internal/liveness`**: `LongRunning` takes containers, services stated `running` and + long-running processes. A package's unit is outside it. A service whose lifecycle is the machine's + is too, and so is a unit file a module writes. Nothing in the node-engine read `list-units + --state=failed`. +- **mesh-host `internal/profile`**: `service-manager` reads `is-system-running` as a capability + detail, which is never raised. +- **mesh-controller**: `node_health` kept resources and network, and nothing about the machine's + managers. No condition kind existed for a degraded one. + +## Fix + +- **The node-engine** (mesh-host PR #53, new `internal/units`). On every liveness look it reads the + failed units of the machine's manager and of every account manager the declaration names. The two-look + rule is its own, as in ADR 0241 §2. A failed unit is a module's when: + 1. the declaration states it as a service or process; + 2. the unit's file is a file the declaration writes; or + 3. the unit's file belongs to a package the declaration installs (asked of the package manager once + per file until the next apply). + + A module's failed unit is said among the resources as that module's unhealthy resource of kind + `unit`, which raises the module's condition on the controller as it stands. Every other failed unit + is the machine's own, in a new `units` part of the health statement. A unit liveness already judges + is left to it, so each failure is said once. It reads and never acts. +- **The controller** (mesh-controller PR #137). It keeps `units` with the machine's statement + (migration 0080). It names the unit in a module's condition. It raises one warning for the machine's + own failed units, `machine..units`, listing them, and clears it on the first statement listing + none. Nothing a send moved is among them, so the gate never holds a send on it. `node show` lists + them. + +On the workstation, once both are live, this raises `module.openrazer..unhealthy` naming +the daemon's unit, and the same on the laptop. It also raises one machine finding on the workstation +listing the two mounts and the clipboard leftover. + +## Left to the operator + +- **The Razer daemon:** the module's migration step (add the account to `openrazer`, log in again). + After that, `openrazer_check` answers ok and the condition clears. +- **The clipboard leftover:** remove the dangling enable link + `~/.config/systemd/user/default.target.wants/greenclip.service`, then reset the unit's failed state + in the account's manager. Both are the account's own files and state, and no module manages them. + The service-manager seat's `disable` (user scope) is the tool for the link, if it accepts a unit + whose file is gone. Otherwise remove the link by hand. Neither was done here. +- **The two mounts:** the operator's own mount table. Either mark them so a missing server is not a + failed boot unit, or accept the machine finding while the server is away. + +## Ruled out + +- *A catalogue bug in `openrazer`*: the daemon refuses on a group the module documents as the + operator's step, with a device present. Making the daemon not fail without a device would not + change anything here. +- *The mesh owning the mounts*: the seat says neither is declared, and no plan writes the mount table. From 28aabd8621aebb95f7b292eea7f7a2f48ee166a5 Mon Sep 17 00:00:00 2001 From: jochen Date: Thu, 8 Oct 2026 11:30:30 +0200 Subject: [PATCH 2/2] Issue 315: say the node-engine's repository as an identifier --- .../01-diagnosis.md | 10 +++++----- 1 file changed, 5 insertions(+), 5 deletions(-) diff --git a/04-ISSUES/315-a-failed-unit-of-a-module-was-never-raised/01-diagnosis.md b/04-ISSUES/315-a-failed-unit-of-a-module-was-never-raised/01-diagnosis.md index 121b9216..fea0a57e 100644 --- a/04-ISSUES/315-a-failed-unit-of-a-module-was-never-raised/01-diagnosis.md +++ b/04-ISSUES/315-a-failed-unit-of-a-module-was-never-raised/01-diagnosis.md @@ -16,8 +16,8 @@ 2. **The two mounts** are units generated from the machine's own mount table at boot. The service-manager seat says neither is declared by the mesh, and no catalogue module or plan writes the mount table. Both failed at the boot of 2026-10-04: - - one is a share on a game console's address on the home network. Mounting it failed in six seconds - with "could not connect" (the console was off). It had failed the same way at the previous boot, + - one is a share on a retro-gaming box on the home network. Mounting it failed in six seconds + with "could not connect" (the box was off). It had failed the same way at the previous boot, in August; - the other is a media library mount. It mounted at the August boot and timed out after 90 s at this one. That is the stock mount timeout, the usual sign of a network share whose server did not @@ -36,18 +36,18 @@ ## Where it lives -- **mesh-host `internal/liveness`**: `LongRunning` takes containers, services stated `running` and +- **`mesh-host` `internal/liveness`**: `LongRunning` takes containers, services stated `running` and long-running processes. A package's unit is outside it. A service whose lifecycle is the machine's is too, and so is a unit file a module writes. Nothing in the node-engine read `list-units --state=failed`. -- **mesh-host `internal/profile`**: `service-manager` reads `is-system-running` as a capability +- **`mesh-host` `internal/profile`**: `service-manager` reads `is-system-running` as a capability detail, which is never raised. - **mesh-controller**: `node_health` kept resources and network, and nothing about the machine's managers. No condition kind existed for a degraded one. ## Fix -- **The node-engine** (mesh-host PR #53, new `internal/units`). On every liveness look it reads the +- **The node-engine** (`mesh-host` PR #53, new `internal/units`). On every liveness look it reads the failed units of the machine's manager and of every account manager the declaration names. The two-look rule is its own, as in ADR 0241 §2. A failed unit is a module's when: 1. the declaration states it as a service or process;