88 lines
5.9 KiB
Markdown
88 lines
5.9 KiB
Markdown
# disk-load
|
|
|
|
How loaded a machine's storage is, read from `/proc` and `/sys` and nothing else: the share of CPU time
|
|
spent waiting on I/O, I/O pressure stall, each disk's busy share, throughput, IOPS, wait and queue
|
|
depth, the filesystems' space, and the units, containers and processes doing the I/O (novox/hq issue
|
|
312). It changes nothing and declares no resources, so it can be assigned to every machine.
|
|
|
|
## Why it exists, and why here
|
|
|
|
Investigating why the controller's checks ran five to seven times slower on the control node (hq issue
|
|
306) needed disk numbers, and no tool in the mesh reported any: not iowait, not a disk's utilisation,
|
|
not throughput, latency or which process wrote. The operator's rule is that a missing tool is built in
|
|
the module that owns the area, never worked round with a login to the machine.
|
|
|
|
No module owned the area. `memory-pressure` reads the same kind of thing for memory, but it owns
|
|
compressed swap and systemd-oomd: assigning it to a machine to read the disks would change how that
|
|
machine handles memory. No node seat holds the machine's storage either. Every seat is a role one
|
|
module holds on each machine, with verbs that act through it, and a new seat costs a controller change
|
|
first (hq ADR 0246). A report that reads kernel counters needs neither a role nor exclusivity. So this
|
|
is a small module with tools only, like `netcheck` beside the network. If the mesh later gives a
|
|
machine's resources a seat, these tools become its verbs.
|
|
|
|
## Tools
|
|
|
|
All four only read.
|
|
|
|
| tool | what | replaces |
|
|
|---|---|---|
|
|
| `disk_load` | Samples a window (`seconds`, default 2, at most 10). It answers the CPU's iowait, user, system and idle %; I/O pressure (some and full, over the window and as avg10/60/300); per disk busy %, read and write MiB/s, IOPS, wait per request in ms, queue depth and requests in flight, with the filesystems each disk holds; the filesystems' space; and the busiest units and processes (`processes`, default 5, 0 for none). `said` puts each finding in one line, flagged past a line (below) | `iostat -x`, `vmstat`, `iotop`, `df` |
|
|
| `disk_top` | Over a window. `by=unit` (the default) ranks systemd units and containers by bytes moved, from their cgroups. `by=container` ranks containers by name, and `by=process` ranks processes | `iotop` |
|
|
| `disk_filesystems` | Each filesystem that stores data, said once however many places it is mounted: size, used and free GiB, used % as `df` counts it, inodes used % | `df -h`, `df -i` |
|
|
| `disk_devices` | Each disk without sampling: model, size, whether it spins, scheduler, queue limit, what it is built on, and since boot the GiB read and written, the busy share and the wait per read and per write | `lsblk`, `cat /proc/diskstats` |
|
|
|
|
### How each number is read
|
|
|
|
- **iowait** is from `/proc/stat`'s `cpu` line: the share of every CPU's time in the window spent idle
|
|
with I/O outstanding. Guest time is left out, because the kernel already counts it in user.
|
|
- **A disk** is read from `/proc/diskstats` at both ends of the window, with iostat's arithmetic.
|
|
Busy is the share of the window with any request in flight. Wait is the time spent on the
|
|
requests completed, divided by their number. Queue depth is the weighted time divided by the window.
|
|
Sectors there are always 512 bytes. A counter that went backwards (a device removed and added)
|
|
reads as 0, never as a huge number. Loop devices, RAM disks and zram are left out unless `all`;
|
|
zram is memory, and `memory-pressure` reports it.
|
|
- **Pressure** is `/proc/pressure/io`. The share of the window comes from the `total` counter, in
|
|
microseconds. A kernel without PSI is said as not available.
|
|
- **Units and containers** come from each cgroup's `io.stat`, which every account can read. Only the
|
|
innermost `.service` or `.scope` counts, because `io.stat` includes descendants. Only whole disks
|
|
built on nothing count: a write through an encrypted or logical volume is charged to the volume and
|
|
again to the disk under it, and a loop device's writes are charged again to the filesystem its file
|
|
sits on. Container names come from `docker ps`; if the runtime cannot be asked, a container is
|
|
given by its short id.
|
|
- **Processes** come from `/proc/<pid>/io` (`read_bytes`, `write_bytes`, which is storage, not the
|
|
page cache). Another account's process can be read only when the tool runner is root. The unreadable
|
|
are counted and said, and the units above miss none of them. A process that started inside the
|
|
window is left out.
|
|
- **Filesystems** come from the process's own mount table, keeping only types that store data
|
|
(ext4, xfs, btrfs, vfat, NFS and the like). Each one is checked with statfs, given 2 s per mount so
|
|
a dead network mount costs one line.
|
|
|
|
### The lines
|
|
|
|
| finding | past |
|
|
|---|---|
|
|
| iowait HIGH | 20 % of CPU time |
|
|
| I/O pressure HIGH | some task stalled 10 % of the window, or avg60 ≥ 10 % |
|
|
| disk SATURATED | busy 90 % of the window |
|
|
| filesystem NEARLY FULL | 90 % of its space or of its inodes |
|
|
|
|
Busy % is how iostat counts it, so it overstates an NVMe disk, which serves many requests at once. Read
|
|
it beside the wait and the queue depth.
|
|
|
|
## Tests
|
|
|
|
`go test ./...`. The tests read a tree standing in for `/proc` and `/sys`. Its fixture is a server
|
|
writing 100 MiB in 2 s to an encrypted root on an NVMe disk, with a spinning disk idle beside it. They
|
|
cover:
|
|
|
|
- iowait, each disk's numbers, and the encrypted root named, with what it sits on and what it holds;
|
|
- loop, zram and partitions left out unless asked for;
|
|
- pressure, and a kernel without it;
|
|
- a counter reset, and busy held at 100;
|
|
- units counted once, by innermost cgroup and bottom disk;
|
|
- containers by name, or by short id if the runtime cannot be asked;
|
|
- processes ranked, one that started inside the window left out, unreadable ones counted;
|
|
- filesystems said once, mount escapes undone, a statfs that fails said;
|
|
- the window's bounds;
|
|
- the manifest: tools listed equal tools served, all read-only, no resources, no machine named.
|