Files
mesh-catalog/modules/memory-pressure/README.md
T
jochen b098b64b22 Phase 3: asus-zephyrus-g14 and memory-pressure, with Go tools and long-running code
The laptop model's hardware module and a memory-pressure module for any
machine (hq research 027/03, 026/05, to-be 42 phase 3). The predecessor's
polling auto-profile and mem-guard user scripts become each module's own
Go code launched by the node runtime (ADR 0198): a profile switcher woken
by the kernel's power-supply uevents, and a guard that warns on RAM, swap
or PSI before systemd-oomd acts, on the desktop over the account's bus and
always as an event. supergfxctl and triggerhappy are kept as found
(research 027 Q1).
2026-10-04 15:42:56 +02:00

121 lines
7.8 KiB
Markdown

# memory-pressure
Compressed swap in RAM, systemd-oomd, and a guard that warns before the machine kills for memory. It
can be assigned on any machine (novox/hq research 027/03, 026/05, to-be 42 phase 3). Memory pressure
is not the laptop's alone.
## What it owns
| | what | notes |
|---|---|---|
| package | `zram-generator` | |
| file | `/etc/systemd/zram-generator.conf` | `zram0`: `min(ram / 2, 16384)` MiB, zstd, swap priority 100. These are the laptop's values, adopted (the path too, ADR 0182). They scale with the machine: 15.3 GiB on 30 GiB of RAM, 16 GiB on the 125 GiB desktop. Read at boot, so a change applies at the next boot. Resizing a live device would mean swapping it off, which pushes what it holds back into RAM |
| file | `/etc/sysctl.d/90-memory-pressure.conf` | `vm.page-cluster = 0` (no read-ahead on swap in RAM). Swappiness is left alone on purpose: the comment says why. `systemd-sysctl` is re-run |
| file | `/etc/systemd/oomd.conf.d/memory-pressure.conf` | `SwapUsedLimit=90%`, `DefaultMemoryPressureLimit=60%`, `DefaultMemoryPressureDurationSec=20s` |
| file | `/etc/systemd/system/-.slice.d/10-oomd.conf` | `ManagedOOMSwap=kill` |
| file | `/etc/systemd/system/user@.service.d/10-oomd.conf` | `ManagedOOMMemoryPressure=kill`, limit 80 % |
| service | `systemd-oomd` running, enabled | restarted on any of its drop-ins, after a `daemon-reload` |
The swap on disk is not this module's. Whether there is a swap file or a partition, and how large, is
the machine's swap layout (research 027 question 3, the `kernel` module). The two workstations differ:
the laptop has a 32 GiB swap file (priority 10) and a 16 GiB partition, and the desktop a 128 GiB
partition and no zram. This module adds the zram tier above whatever is there.
## The long-running code: the guard (ADR 0198)
The Go bundle serves the tools and runs the guard in the same process, launched by the node's runtime.
It replaces the predecessor's `mem-guard` user unit, which checked used RAM every 30 s and called
`notify-send`.
- **Three signals:** RAM used ≥ 95 % (the predecessor's line, which it had raised from 90),
swap used ≥ 80 % (oomd kills at 90 %), or memory pressure (PSI `some` avg10) ≥ 40 % (oomd acts at
60 %, or 80 % for the user manager, over 20 s). Whichever comes first warns, and the warning says
which.
- **Every 10 s**, not 30: oomd's window is 20 s, so a check every 30 s could warn after oomd had
already acted. Each check reads two files. PSI triggers would wake on pressure alone, but not on RAM
or swap filling, so one timer reads all three rather than run two mechanisms.
- **Once per episode.** After a warning the guard re-arms only when every value is below its clearing
line: 88 % RAM, 70 % swap, 10 % pressure.
- **The warning names what to close:** the three largest units by resident plus swapped memory, and
in each its largest process.
- **The desktop, from outside the session.** The runtime runs as the operator's account but outside
the graphical session, with no session bus address and no `XDG_RUNTIME_DIR` (measured on the
laptop's runtime unit). The guard names the account's own bus, `/run/user/<uid>/bus`, and calls
`org.freedesktop.Notifications.Notify` with `busctl --user`. That client is the service manager's
own, so nothing is installed, and a server pulls in no libnotify. Measured on 2026-10-04: from a
clean environment as the account's uid, the notification service on that bus answered (dunst 1.13).
The account is the runtime's `MESH_OPERATOR_ACCOUNT`. A runtime running as another uid could not
authenticate on that bus, and the guard says so rather than try.
- **And always an event:** `pressure.high` (reasons, percentages, the largest units) and
`pressure.cleared` (how long it lasted) are published through the runtime, whether or not anyone is
at a desktop. A machine without one has no session bus, and `memory_guard` says
*no session bus*.
Thresholds are constants until settings exist (issue 168).
**Known limit.** The runtime restarts a launched bundle that exits on its next tool call, not at once.
The guard recovers from a panic and reports it, but a crashed process waits for a call.
## Tools
| tool | what |
|---|---|
| `memory_status` | RAM, swap and each device with its priority, zram and its ratio, PSI some/full, and the guard's verdict on them |
| `memory_top` | the largest processes by resident plus swap, with their unit; `by=unit` sums per unit, which is what systemd-oomd chooses among |
| `memory_oom_history` | what was killed, newest first, from the journal: the kernel's OOM killer, systemd-oomd, and the service manager's *killed by the OOM killer*. `since` is checked against a short grammar before it reaches `journalctl` |
| `memory_zram` | each zram device (algorithm, size, stored, compressed, RAM used, ratio, same-filled and incompressible pages), the generator's configuration, swappiness and page-cluster |
| `memory_oomd` | whether oomd runs, and what it watches (`oomctl`) |
| `memory_guard` | the guard's thresholds, last verdict, whether a warning stands, and whether the desktop can be reached |
## Found on 2026-10-04 (read-only)
- **On the laptop, systemd-oomd guards almost nothing a person runs.** Every desktop application runs
in the login session's scope (`session-1.scope`), because the login manager and i3 start them there
and not in per-application scopes under `user@.service`. The pressure kill on `user@.service`
therefore watches 0.5 GiB. The swap kill on `-.slice`, when it fires, would choose the largest
cgroup, which is the whole session scope: X, i3 and every application at once. `memory_top by=unit`
shows it. **The fix is the graphical session's** (phase 2): launch applications in their own
scopes, for example `systemd-run --user --scope` from the launcher, and oomd then kills one
application. This module does not widen oomd's reach onto `user-.slice`, because there it would kill
the session.
- The desktop has no zram and oomd disabled. Assigning the module there is a change of behaviour:
16 GiB of zram at the next boot, and oomd enabled with the same caveat about the session scope.
## When assigned to the laptop: what changes
1. Written over found files (originals kept once): `zram-generator.conf` (same values, so nothing
until the next boot either), `-.slice.d/10-oomd.conf` and `user@.service.d/10-oomd.conf` (same
keys).
2. New: `sysctl.d/90-memory-pressure.conf` (the same `vm.page-cluster = 0` already in force) and
`oomd.conf.d/memory-pressure.conf` (the same values as the predecessor's file beside it).
3. `systemd-sysctl` is re-run (the values are unchanged), and `daemon-reload` and `systemd-oomd` are
restarted.
4. The node runtime restarts with the bundle, and the guard starts. **Until `mem-guard` is stopped,
two notifiers run.**
## Predecessor files this module makes redundant — the operator removes them once (ADR 0182)
On the laptop:
1. `systemctl --user disable --now mem-guard.service`, then delete
`~/.config/systemd/user/mem-guard.service` and `~/scripts/mem-guard.sh`.
2. `/etc/systemd/oomd.conf.d/g14-oomd.conf`: the same values as the module's drop-in.
3. `/etc/sysctl.d/90-g14-zram.conf`: the same key as the module's file.
The desktop had none of these.
## Tests
`go test ./...`. They read a tree standing in for `/proc` and `/sys`, using the laptop's own swaps,
zram statistics and PSI lines. They cover:
- the snapshot;
- the guard's three lines, its once-per-episode latch and its clearing, with the largest named;
- an unreachable desktop still publishing the event;
- the notification being one `busctl` call on the account's bus;
- grouping by unit;
- the journal's three kinds of kill, with non-UTF-8 messages skipped;
- `since` refused when it is an option;
- *no match* read as no kills;
- the manifest: tools listed equal tools served, events, no machine named, triggers exist.