Files
mesh-catalog/modules/memory-pressure
jochen 9c84cd9e6c Phase 3: asus-zephyrus-g14 and memory-pressure, with Go tools and long-running code
The laptop model's hardware module and a memory-pressure module for any
machine (hq research 027/03, 026/05, to-be 42 phase 3). The predecessor's
polling auto-profile and mem-guard user scripts become each module's own
Go code launched by the node runtime (ADR 0198): a profile switcher woken
by the kernel's power-supply uevents, and a guard that warns on RAM, swap
or PSI before systemd-oomd acts, on the desktop over the account's bus and
always as an event. supergfxctl and triggerhappy are kept as found
(research 027 Q1).
2026-10-04 12:56:23 +02:00
..

memory-pressure

Compressed swap in RAM, systemd-oomd, and a guard that warns before the machine kills for memory. It can be assigned on any machine (novox/hq research 027/03, 026/05, to-be 42 phase 3). Memory pressure is not the laptop's alone.

What it owns

what notes
package zram-generator
file /etc/systemd/zram-generator.conf zram0: min(ram / 2, 16384) MiB, zstd, swap priority 100. These are the laptop's values, adopted (the path too, ADR 0182). They scale with the machine: 15.3 GiB on 30 GiB of RAM, 16 GiB on the 125 GiB desktop. Read at boot, so a change applies at the next boot. Resizing a live device would mean swapping it off, which pushes what it holds back into RAM
file /etc/sysctl.d/90-memory-pressure.conf vm.page-cluster = 0 (no read-ahead on swap in RAM). Swappiness is left alone on purpose: the comment says why. systemd-sysctl is re-run
file /etc/systemd/oomd.conf.d/memory-pressure.conf SwapUsedLimit=90%, DefaultMemoryPressureLimit=60%, DefaultMemoryPressureDurationSec=20s
file /etc/systemd/system/-.slice.d/10-oomd.conf ManagedOOMSwap=kill
file /etc/systemd/system/user@.service.d/10-oomd.conf ManagedOOMMemoryPressure=kill, limit 80 %
service systemd-oomd running, enabled restarted on any of its drop-ins, after a daemon-reload

The swap on disk is not this module's. Whether there is a swap file or a partition, and how large, is the machine's swap layout (research 027 question 3, the kernel module). The two workstations differ: the laptop has a 32 GiB swap file (priority 10) and a 16 GiB partition, and the desktop a 128 GiB partition and no zram. This module adds the zram tier above whatever is there.

The long-running code: the guard (ADR 0198)

The Go bundle serves the tools and runs the guard in the same process, launched by the node's runtime. It replaces the predecessor's mem-guard user unit, which checked used RAM every 30 s and called notify-send.

  • Three signals: RAM used ≥ 95 % (the predecessor's line, which it had raised from 90), swap used ≥ 80 % (oomd kills at 90 %), or memory pressure (PSI some avg10) ≥ 40 % (oomd acts at 60 %, or 80 % for the user manager, over 20 s). Whichever comes first warns, and the warning says which.
  • Every 10 s, not 30: oomd's window is 20 s, so a check every 30 s could warn after oomd had already acted. Each check reads two files. PSI triggers would wake on pressure alone, but not on RAM or swap filling, so one timer reads all three rather than run two mechanisms.
  • Once per episode. After a warning the guard re-arms only when every value is below its clearing line: 88 % RAM, 70 % swap, 10 % pressure.
  • The warning names what to close: the three largest units by resident plus swapped memory, and in each its largest process.
  • The desktop, from outside the session. The runtime runs as the operator's account but outside the graphical session, with no session bus address and no XDG_RUNTIME_DIR (measured on the laptop's runtime unit). The guard names the account's own bus, /run/user/<uid>/bus, and calls org.freedesktop.Notifications.Notify with busctl --user. That client is the service manager's own, so nothing is installed, and a server pulls in no libnotify. Measured on 2026-10-04: from a clean environment as the account's uid, the notification service on that bus answered (dunst 1.13). The account is the runtime's MESH_OPERATOR_ACCOUNT. A runtime running as another uid could not authenticate on that bus, and the guard says so rather than try.
  • And always an event: pressure.high (reasons, percentages, the largest units) and pressure.cleared (how long it lasted) are published through the runtime, whether or not anyone is at a desktop. A machine without one has no session bus, and memory_guard says no session bus.

Thresholds are constants until settings exist (issue 168).

Known limit. The runtime restarts a launched bundle that exits on its next tool call, not at once. The guard recovers from a panic and reports it, but a crashed process waits for a call.

Tools

tool what
memory_status RAM, swap and each device with its priority, zram and its ratio, PSI some/full, and the guard's verdict on them
memory_top the largest processes by resident plus swap, with their unit; by=unit sums per unit, which is what systemd-oomd chooses among
memory_oom_history what was killed, newest first, from the journal: the kernel's OOM killer, systemd-oomd, and the service manager's killed by the OOM killer. since is checked against a short grammar before it reaches journalctl
memory_zram each zram device (algorithm, size, stored, compressed, RAM used, ratio, same-filled and incompressible pages), the generator's configuration, swappiness and page-cluster
memory_oomd whether oomd runs, and what it watches (oomctl)
memory_guard the guard's thresholds, last verdict, whether a warning stands, and whether the desktop can be reached

Found on 2026-10-04 (read-only)

  • On the laptop, systemd-oomd guards almost nothing a person runs. Every desktop application runs in the login session's scope (session-1.scope), because the login manager and i3 start them there and not in per-application scopes under user@.service. The pressure kill on user@.service therefore watches 0.5 GiB. The swap kill on -.slice, when it fires, would choose the largest cgroup, which is the whole session scope: X, i3 and every application at once. memory_top by=unit shows it. The fix is the graphical session's (phase 2): launch applications in their own scopes, for example systemd-run --user --scope from the launcher, and oomd then kills one application. This module does not widen oomd's reach onto user-.slice, because there it would kill the session.
  • The desktop has no zram and oomd disabled. Assigning the module there is a change of behaviour: 16 GiB of zram at the next boot, and oomd enabled with the same caveat about the session scope.

When assigned to the laptop: what changes

  1. Written over found files (originals kept once): zram-generator.conf (same values, so nothing until the next boot either), -.slice.d/10-oomd.conf and user@.service.d/10-oomd.conf (same keys).
  2. New: sysctl.d/90-memory-pressure.conf (the same vm.page-cluster = 0 already in force) and oomd.conf.d/memory-pressure.conf (the same values as the predecessor's file beside it).
  3. systemd-sysctl is re-run (the values are unchanged), and daemon-reload and systemd-oomd are restarted.
  4. The node runtime restarts with the bundle, and the guard starts. Until mem-guard is stopped, two notifiers run.

Predecessor files this module makes redundant — the operator removes them once (ADR 0182)

On the laptop:

  1. systemctl --user disable --now mem-guard.service, then delete ~/.config/systemd/user/mem-guard.service and ~/scripts/mem-guard.sh.
  2. /etc/systemd/oomd.conf.d/g14-oomd.conf: the same values as the module's drop-in.
  3. /etc/sysctl.d/90-g14-zram.conf: the same key as the module's file.

The desktop had none of these.

Tests

go test ./.... They read a tree standing in for /proc and /sys, using the laptop's own swaps, zram statistics and PSI lines. They cover:

  • the snapshot;
  • the guard's three lines, its once-per-episode latch and its clearing, with the largest named;
  • an unreachable desktop still publishing the event;
  • the notification being one busctl call on the account's bus;
  • grouping by unit;
  • the journal's three kinds of kill, with non-UTF-8 messages skipped;
  • since refused when it is an option;
  • no match read as no kills;
  • the manifest: tools listed equal tools served, events, no machine named, triggers exist.