Merge pull request 'Issue 312: no tool said a machine's disk load' (#195) from issues/312-no-tool-said-a-machines-disk-load into main

This commit was merged in pull request #195.
This commit is contained in:
2026-10-08 08:55:40 +00:00
2 changed files with 94 additions and 0 deletions
@@ -0,0 +1,45 @@
---
status: located
opened: 2026-10-08
located-in: [mesh-catalog modules/disk-load]
fixed-by: novox/mesh-catalog PR #120
amended-design:
---
# 312. No tool said a machine's disk load
## Symptom
[Issue 306](../306-a-check-on-the-control-node-ran-past-its-timeout/00-report.md) found the controller's
checks running five to seven times slower on the control node than elsewhere, and only in the packages
that touch a store. The evidence pointed at the control node's disk, which the throwaway stores share with
the mesh's own store, the database server, the object store, the bus and the forge. The next question was
the obvious one: how busy is that disk, how long does a request wait, and who is writing?
No tool in the mesh could answer it. A search of every tool on every machine for disk load, I/O wait,
throughput or latency returned two tools, and both report how much space one package format takes. The
only per-container figure was the container runtime's cumulative block I/O. That figure gave the
16–20 GB per throwaway store quoted in issue 306, but it has no rate, no wait and no queue, and it
says nothing of anything outside a container. The machine's own counters, `/proc/stat`'s iowait,
`/proc/diskstats`, `/proc/pressure/io`, each cgroup's `io.stat`, and `/proc/<pid>/io`, were readable
on every machine, and nothing in the mesh read them.
The only way to get the numbers was a login to the machine, which the mesh's rule refuses
([ADR 0245](../../02-DECISIONS/0245-a-verb-says-what-it-replaces-and-the-agent-is-guarded-from-working-round-the-mesh.md)):
a missing tool is built in the module that owns the area, never worked round.
## Why it is an issue
ADR 0245 turns a missing tool from an inconvenience into a stop. An agent asked why a machine is slow
can read its memory (`memory-pressure`), its units, its containers' CPU and memory, and its network, but
not its storage. That is the one resource issue 306 needed. A diagnosis that has to guess which
resource is short guesses wrong first. Issue 306's first guess was plain overload, and its package
timings had to rule that out the long way.
## Located
No module owned the machine's storage load, and no node seat holds it. `01-diagnosis.md` says where the
tools belong and why. The fix is a new read-only catalogue module, `disk-load`. Its tools are
`disk_load`, `disk_top`, `disk_filesystems` and `disk_devices`, read from `/proc` and `/sys`, with
fixture tests. It is not a core component, so this issue needs no replay when it resolves. It resolves
once the module is merged and assigned, and its tools answer on the control node.
@@ -0,0 +1,49 @@
# 312 — Diagnosis
## 2026-10-08 — where the missing tools belong
**What exists.** The mesh's overview and the control node's machine view list every seat and module tool.
No seat verb or tool reads CPU iowait, `/proc/diskstats`, I/O pressure stall, a cgroup's `io.stat` or a
process's `/proc/<pid>/io`. The nearest matches are:
- `memory-pressure`, which reads the same kind of kernel counters for memory, including memory PSI;
- the container module's `docker_stats` and `docker_disk_usage`, which give each container's
cumulative block I/O and the runtime's own space;
- `logrotate_big_logs`, `snapd_disk_usage` and `flatpak_disk_usage`, each the space of one thing.
The records hold no design for a machine's resources. [ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md)
judges what a module runs. [ADR 0241](../../02-DECISIONS/0241-a-machine-says-how-its-network-is-and-an-outside-writer-of-a-mesh-file-is-a-finding.md)
adds the machine's networking to that judgement. Neither covers its processors, memory or storage
as a load.
**Ruled out:**
1. *A verb on an existing node seat.* Every node seat is a role one module holds on each machine:
the service manager, the packet filter, the uplink, the hostname, the backup, the build agent, the
login shell and intrusion prevention. None is the machine's storage. A disk report as a verb of the
service manager would put it under a seat whose holder is chosen for something else.
2. *A new seat for the machine's resources.* A new seat's verb is promised in the controller before
it is required ([ADR 0246](../../02-DECISIONS/0246-a-seats-new-verb-is-promised-before-it-is-required.md)).
A seat also exists so one module holds a role exclusively and acts through it. A read of kernel
counters needs no exclusivity and acts on nothing. Two modules reading `/proc/diskstats` do
not conflict. The cost buys nothing now.
3. *A tool in `memory-pressure`.* It is the nearest relative, but it owns compressed swap and
systemd-oomd. Assigning it to the control node to read the disk would add zram at the next boot
and enable oomd there, which is a change of behaviour that reading should not need. It is not
assigned on the control node today.
4. *A tool in the container module.* It sees containers. Issue 306's disk is shared with the forge
and the bus, which also run as containers, but the question is the machine's, and a host
service's writes would be missed.
**Chosen.** A small read-only module, `disk-load`, beside `netcheck`, which reads the network the same
way. It declares no resources, so it can be assigned to every machine without changing one. If the mesh
later makes a machine's resources a seat, its four tools are that seat's verbs as written.
**One finding while building it.** `/proc/<pid>/io` is readable only for the tool runner's own
account unless the runner is root. On the laptop it hid 502 of the machine's processes, including the
build agent's container, which was writing 83 MiB/s at the time. A cgroup's `io.stat` is readable by
every account. So the module ranks units and containers from their cgroups and names processes
only as a second view, and it says how many it could not read. Two kinds of double counting are
avoided: `io.stat` includes a cgroup's descendants, so only the innermost unit counts, and a write
through an encrypted volume is charged again to the disk under it, so only a disk built on nothing
counts.