Files
mesh-catalog/modules/disk-load/README.md
T
jochen 6fdf5293eb
mesh/merge-gate pass: builds new: modules/disk-load, sent nowhere; no bus step; every machine composes with the change as it did without (4 of 4 compose)
mesh/repo-check pass: its merge-check.sh passed
mesh/delivery delivered
Add disk-load: no tool said a machine's iowait, disk load or who does the I/O (hq issue 312)
2026-10-08 10:47:41 +02:00

88 lines
5.9 KiB
Markdown

# disk-load
How loaded a machine's storage is, read from `/proc` and `/sys` and nothing else: the share of CPU time
spent waiting on I/O, I/O pressure stall, each disk's busy share, throughput, IOPS, wait and queue
depth, the filesystems' space, and the units, containers and processes doing the I/O (novox/hq issue
312). It changes nothing and declares no resources, so it can be assigned to every machine.
## Why it exists, and why here
Investigating why the controller's checks ran five to seven times slower on the control node (hq issue
306) needed disk numbers, and no tool in the mesh reported any: not iowait, not a disk's utilisation,
not throughput, latency or which process wrote. The operator's rule is that a missing tool is built in
the module that owns the area, never worked round with a login to the machine.
No module owned the area. `memory-pressure` reads the same kind of thing for memory, but it owns
compressed swap and systemd-oomd: assigning it to a machine to read the disks would change how that
machine handles memory. No node seat holds the machine's storage either. Every seat is a role one
module holds on each machine, with verbs that act through it, and a new seat costs a controller change
first (hq ADR 0246). A report that reads kernel counters needs neither a role nor exclusivity. So this
is a small module with tools only, like `netcheck` beside the network. If the mesh later gives a
machine's resources a seat, these tools become its verbs.
## Tools
All four only read.
| tool | what | replaces |
|---|---|---|
| `disk_load` | Samples a window (`seconds`, default 2, at most 10). It answers the CPU's iowait, user, system and idle %; I/O pressure (some and full, over the window and as avg10/60/300); per disk busy %, read and write MiB/s, IOPS, wait per request in ms, queue depth and requests in flight, with the filesystems each disk holds; the filesystems' space; and the busiest units and processes (`processes`, default 5, 0 for none). `said` puts each finding in one line, flagged past a line (below) | `iostat -x`, `vmstat`, `iotop`, `df` |
| `disk_top` | Over a window. `by=unit` (the default) ranks systemd units and containers by bytes moved, from their cgroups. `by=container` ranks containers by name, and `by=process` ranks processes | `iotop` |
| `disk_filesystems` | Each filesystem that stores data, said once however many places it is mounted: size, used and free GiB, used % as `df` counts it, inodes used % | `df -h`, `df -i` |
| `disk_devices` | Each disk without sampling: model, size, whether it spins, scheduler, queue limit, what it is built on, and since boot the GiB read and written, the busy share and the wait per read and per write | `lsblk`, `cat /proc/diskstats` |
### How each number is read
- **iowait** is from `/proc/stat`'s `cpu` line: the share of every CPU's time in the window spent idle
with I/O outstanding. Guest time is left out, because the kernel already counts it in user.
- **A disk** is read from `/proc/diskstats` at both ends of the window, with iostat's arithmetic.
Busy is the share of the window with any request in flight. Wait is the time spent on the
requests completed, divided by their number. Queue depth is the weighted time divided by the window.
Sectors there are always 512 bytes. A counter that went backwards (a device removed and added)
reads as 0, never as a huge number. Loop devices, RAM disks and zram are left out unless `all`;
zram is memory, and `memory-pressure` reports it.
- **Pressure** is `/proc/pressure/io`. The share of the window comes from the `total` counter, in
microseconds. A kernel without PSI is said as not available.
- **Units and containers** come from each cgroup's `io.stat`, which every account can read. Only the
innermost `.service` or `.scope` counts, because `io.stat` includes descendants. Only whole disks
built on nothing count: a write through an encrypted or logical volume is charged to the volume and
again to the disk under it, and a loop device's writes are charged again to the filesystem its file
sits on. Container names come from `docker ps`; if the runtime cannot be asked, a container is
given by its short id.
- **Processes** come from `/proc/<pid>/io` (`read_bytes`, `write_bytes`, which is storage, not the
page cache). Another account's process can be read only when the tool runner is root. The unreadable
are counted and said, and the units above miss none of them. A process that started inside the
window is left out.
- **Filesystems** come from the process's own mount table, keeping only types that store data
(ext4, xfs, btrfs, vfat, NFS and the like). Each one is checked with statfs, given 2 s per mount so
a dead network mount costs one line.
### The lines
| finding | past |
|---|---|
| iowait HIGH | 20 % of CPU time |
| I/O pressure HIGH | some task stalled 10 % of the window, or avg60 ≥ 10 % |
| disk SATURATED | busy 90 % of the window |
| filesystem NEARLY FULL | 90 % of its space or of its inodes |
Busy % is how iostat counts it, so it overstates an NVMe disk, which serves many requests at once. Read
it beside the wait and the queue depth.
## Tests
`go test ./...`. The tests read a tree standing in for `/proc` and `/sys`. Its fixture is a server
writing 100 MiB in 2 s to an encrypted root on an NVMe disk, with a spinning disk idle beside it. They
cover:
- iowait, each disk's numbers, and the encrypted root named, with what it sits on and what it holds;
- loop, zram and partitions left out unless asked for;
- pressure, and a kernel without it;
- a counter reset, and busy held at 100;
- units counted once, by innermost cgroup and bottom disk;
- containers by name, or by short id if the runtime cannot be asked;
- processes ranked, one that started inside the window left out, unreadable ones counted;
- filesystems said once, mount escapes undone, a statfs that fails said;
- the window's bounds;
- the manifest: tools listed equal tools served, all read-only, no resources, no machine named.