Files
mesh-catalog/modules/disk-load/README.md
T
jochen 6fdf5293eb
mesh/merge-gate pass: builds new: modules/disk-load, sent nowhere; no bus step; every machine composes with the change as it did without (4 of 4 compose)
mesh/repo-check pass: its merge-check.sh passed
mesh/delivery delivered
Add disk-load: no tool said a machine's iowait, disk load or who does the I/O (hq issue 312)
2026-10-08 10:47:41 +02:00

5.9 KiB

disk-load

How loaded a machine's storage is, read from /proc and /sys and nothing else: the share of CPU time spent waiting on I/O, I/O pressure stall, each disk's busy share, throughput, IOPS, wait and queue depth, the filesystems' space, and the units, containers and processes doing the I/O (novox/hq issue 312). It changes nothing and declares no resources, so it can be assigned to every machine.

Why it exists, and why here

Investigating why the controller's checks ran five to seven times slower on the control node (hq issue 306) needed disk numbers, and no tool in the mesh reported any: not iowait, not a disk's utilisation, not throughput, latency or which process wrote. The operator's rule is that a missing tool is built in the module that owns the area, never worked round with a login to the machine.

No module owned the area. memory-pressure reads the same kind of thing for memory, but it owns compressed swap and systemd-oomd: assigning it to a machine to read the disks would change how that machine handles memory. No node seat holds the machine's storage either. Every seat is a role one module holds on each machine, with verbs that act through it, and a new seat costs a controller change first (hq ADR 0246). A report that reads kernel counters needs neither a role nor exclusivity. So this is a small module with tools only, like netcheck beside the network. If the mesh later gives a machine's resources a seat, these tools become its verbs.

Tools

All four only read.

tool what replaces
disk_load Samples a window (seconds, default 2, at most 10). It answers the CPU's iowait, user, system and idle %; I/O pressure (some and full, over the window and as avg10/60/300); per disk busy %, read and write MiB/s, IOPS, wait per request in ms, queue depth and requests in flight, with the filesystems each disk holds; the filesystems' space; and the busiest units and processes (processes, default 5, 0 for none). said puts each finding in one line, flagged past a line (below) iostat -x, vmstat, iotop, df
disk_top Over a window. by=unit (the default) ranks systemd units and containers by bytes moved, from their cgroups. by=container ranks containers by name, and by=process ranks processes iotop
disk_filesystems Each filesystem that stores data, said once however many places it is mounted: size, used and free GiB, used % as df counts it, inodes used % df -h, df -i
disk_devices Each disk without sampling: model, size, whether it spins, scheduler, queue limit, what it is built on, and since boot the GiB read and written, the busy share and the wait per read and per write lsblk, cat /proc/diskstats

How each number is read

  • iowait is from /proc/stat's cpu line: the share of every CPU's time in the window spent idle with I/O outstanding. Guest time is left out, because the kernel already counts it in user.
  • A disk is read from /proc/diskstats at both ends of the window, with iostat's arithmetic. Busy is the share of the window with any request in flight. Wait is the time spent on the requests completed, divided by their number. Queue depth is the weighted time divided by the window. Sectors there are always 512 bytes. A counter that went backwards (a device removed and added) reads as 0, never as a huge number. Loop devices, RAM disks and zram are left out unless all; zram is memory, and memory-pressure reports it.
  • Pressure is /proc/pressure/io. The share of the window comes from the total counter, in microseconds. A kernel without PSI is said as not available.
  • Units and containers come from each cgroup's io.stat, which every account can read. Only the innermost .service or .scope counts, because io.stat includes descendants. Only whole disks built on nothing count: a write through an encrypted or logical volume is charged to the volume and again to the disk under it, and a loop device's writes are charged again to the filesystem its file sits on. Container names come from docker ps; if the runtime cannot be asked, a container is given by its short id.
  • Processes come from /proc/<pid>/io (read_bytes, write_bytes, which is storage, not the page cache). Another account's process can be read only when the tool runner is root. The unreadable are counted and said, and the units above miss none of them. A process that started inside the window is left out.
  • Filesystems come from the process's own mount table, keeping only types that store data (ext4, xfs, btrfs, vfat, NFS and the like). Each one is checked with statfs, given 2 s per mount so a dead network mount costs one line.

The lines

finding past
iowait HIGH 20 % of CPU time
I/O pressure HIGH some task stalled 10 % of the window, or avg60 ≥ 10 %
disk SATURATED busy 90 % of the window
filesystem NEARLY FULL 90 % of its space or of its inodes

Busy % is how iostat counts it, so it overstates an NVMe disk, which serves many requests at once. Read it beside the wait and the queue depth.

Tests

go test ./.... The tests read a tree standing in for /proc and /sys. Its fixture is a server writing 100 MiB in 2 s to an encrypted root on an NVMe disk, with a spinning disk idle beside it. They cover:

  • iowait, each disk's numbers, and the encrypted root named, with what it sits on and what it holds;
  • loop, zram and partitions left out unless asked for;
  • pressure, and a kernel without it;
  • a counter reset, and busy held at 100;
  • units counted once, by innermost cgroup and bottom disk;
  • containers by name, or by short id if the runtime cannot be asked;
  • processes ranked, one that started inside the window left out, unreadable ones counted;
  • filesystems said once, mount escapes undone, a statfs that fails said;
  • the window's bounds;
  • the manifest: tools listed equal tools served, all read-only, no resources, no machine named.