disk-load
How loaded a machine's storage is, read from /proc and /sys and nothing else: the share of CPU time
spent waiting on I/O, I/O pressure stall, each disk's busy share, throughput, IOPS, wait and queue
depth, the filesystems' space, and the units, containers and processes doing the I/O (novox/hq issue
312). It changes nothing and declares no resources, so it can be assigned to every machine.
Why it exists, and why here
Investigating why the controller's checks ran five to seven times slower on the control node (hq issue 306) needed disk numbers, and no tool in the mesh reported any: not iowait, not a disk's utilisation, not throughput, latency or which process wrote. The operator's rule is that a missing tool is built in the module that owns the area, never worked round with a login to the machine.
No module owned the area. memory-pressure reads the same kind of thing for memory, but it owns
compressed swap and systemd-oomd: assigning it to a machine to read the disks would change how that
machine handles memory. No node seat holds the machine's storage either. Every seat is a role one
module holds on each machine, with verbs that act through it, and a new seat costs a controller change
first (hq ADR 0246). A report that reads kernel counters needs neither a role nor exclusivity. So this
is a small module with tools only, like netcheck beside the network. If the mesh later gives a
machine's resources a seat, these tools become its verbs.
Tools
All four only read.
| tool | what | replaces |
|---|---|---|
disk_load |
Samples a window (seconds, default 2, at most 10). It answers the CPU's iowait, user, system and idle %; I/O pressure (some and full, over the window and as avg10/60/300); per disk busy %, read and write MiB/s, IOPS, wait per request in ms, queue depth and requests in flight, with the filesystems each disk holds; the filesystems' space; and the busiest units and processes (processes, default 5, 0 for none). said puts each finding in one line, flagged past a line (below) |
iostat -x, vmstat, iotop, df |
disk_top |
Over a window. by=unit (the default) ranks systemd units and containers by bytes moved, from their cgroups. by=container ranks containers by name, and by=process ranks processes |
iotop |
disk_filesystems |
Each filesystem that stores data, said once however many places it is mounted: size, used and free GiB, used % as df counts it, inodes used % |
df -h, df -i |
disk_devices |
Each disk without sampling: model, size, whether it spins, scheduler, queue limit, what it is built on, and since boot the GiB read and written, the busy share and the wait per read and per write | lsblk, cat /proc/diskstats |
How each number is read
- iowait is from
/proc/stat'scpuline: the share of every CPU's time in the window spent idle with I/O outstanding. Guest time is left out, because the kernel already counts it in user. - A disk is read from
/proc/diskstatsat both ends of the window, with iostat's arithmetic. Busy is the share of the window with any request in flight. Wait is the time spent on the requests completed, divided by their number. Queue depth is the weighted time divided by the window. Sectors there are always 512 bytes. A counter that went backwards (a device removed and added) reads as 0, never as a huge number. Loop devices, RAM disks and zram are left out unlessall; zram is memory, andmemory-pressurereports it. - Pressure is
/proc/pressure/io. The share of the window comes from thetotalcounter, in microseconds. A kernel without PSI is said as not available. - Units and containers come from each cgroup's
io.stat, which every account can read. Only the innermost.serviceor.scopecounts, becauseio.statincludes descendants. Only whole disks built on nothing count: a write through an encrypted or logical volume is charged to the volume and again to the disk under it, and a loop device's writes are charged again to the filesystem its file sits on. Container names come fromdocker ps; if the runtime cannot be asked, a container is given by its short id. - Processes come from
/proc/<pid>/io(read_bytes,write_bytes, which is storage, not the page cache). Another account's process can be read only when the tool runner is root. The unreadable are counted and said, and the units above miss none of them. A process that started inside the window is left out. - Filesystems come from the process's own mount table, keeping only types that store data (ext4, xfs, btrfs, vfat, NFS and the like). Each one is checked with statfs, given 2 s per mount so a dead network mount costs one line.
The lines
| finding | past |
|---|---|
| iowait HIGH | 20 % of CPU time |
| I/O pressure HIGH | some task stalled 10 % of the window, or avg60 ≥ 10 % |
| disk SATURATED | busy 90 % of the window |
| filesystem NEARLY FULL | 90 % of its space or of its inodes |
Busy % is how iostat counts it, so it overstates an NVMe disk, which serves many requests at once. Read it beside the wait and the queue depth.
Tests
go test ./.... The tests read a tree standing in for /proc and /sys. Its fixture is a server
writing 100 MiB in 2 s to an encrypted root on an NVMe disk, with a spinning disk idle beside it. They
cover:
- iowait, each disk's numbers, and the encrypted root named, with what it sits on and what it holds;
- loop, zram and partitions left out unless asked for;
- pressure, and a kernel without it;
- a counter reset, and busy held at 100;
- units counted once, by innermost cgroup and bottom disk;
- containers by name, or by short id if the runtime cannot be asked;
- processes ranked, one that started inside the window left out, unreadable ones counted;
- filesystems said once, mount escapes undone, a statfs that fails said;
- the window's bounds;
- the manifest: tools listed equal tools served, all read-only, no resources, no machine named.