Merge pull request 'Issue 312: no tool said a machine's disk load' (#195) from issues/312-no-tool-said-a-machines-disk-load into main
This commit was merged in pull request #195.
This commit is contained in:
@@ -0,0 +1,45 @@
|
||||
---
|
||||
status: located
|
||||
opened: 2026-10-08
|
||||
located-in: [mesh-catalog modules/disk-load]
|
||||
fixed-by: novox/mesh-catalog PR #120
|
||||
amended-design:
|
||||
---
|
||||
|
||||
# 312. No tool said a machine's disk load
|
||||
|
||||
## Symptom
|
||||
|
||||
[Issue 306](../306-a-check-on-the-control-node-ran-past-its-timeout/00-report.md) found the controller's
|
||||
checks running five to seven times slower on the control node than elsewhere, and only in the packages
|
||||
that touch a store. The evidence pointed at the control node's disk, which the throwaway stores share with
|
||||
the mesh's own store, the database server, the object store, the bus and the forge. The next question was
|
||||
the obvious one: how busy is that disk, how long does a request wait, and who is writing?
|
||||
|
||||
No tool in the mesh could answer it. A search of every tool on every machine for disk load, I/O wait,
|
||||
throughput or latency returned two tools, and both report how much space one package format takes. The
|
||||
only per-container figure was the container runtime's cumulative block I/O. That figure gave the
|
||||
16–20 GB per throwaway store quoted in issue 306, but it has no rate, no wait and no queue, and it
|
||||
says nothing of anything outside a container. The machine's own counters, `/proc/stat`'s iowait,
|
||||
`/proc/diskstats`, `/proc/pressure/io`, each cgroup's `io.stat`, and `/proc/<pid>/io`, were readable
|
||||
on every machine, and nothing in the mesh read them.
|
||||
|
||||
The only way to get the numbers was a login to the machine, which the mesh's rule refuses
|
||||
([ADR 0245](../../02-DECISIONS/0245-a-verb-says-what-it-replaces-and-the-agent-is-guarded-from-working-round-the-mesh.md)):
|
||||
a missing tool is built in the module that owns the area, never worked round.
|
||||
|
||||
## Why it is an issue
|
||||
|
||||
ADR 0245 turns a missing tool from an inconvenience into a stop. An agent asked why a machine is slow
|
||||
can read its memory (`memory-pressure`), its units, its containers' CPU and memory, and its network, but
|
||||
not its storage. That is the one resource issue 306 needed. A diagnosis that has to guess which
|
||||
resource is short guesses wrong first. Issue 306's first guess was plain overload, and its package
|
||||
timings had to rule that out the long way.
|
||||
|
||||
## Located
|
||||
|
||||
No module owned the machine's storage load, and no node seat holds it. `01-diagnosis.md` says where the
|
||||
tools belong and why. The fix is a new read-only catalogue module, `disk-load`. Its tools are
|
||||
`disk_load`, `disk_top`, `disk_filesystems` and `disk_devices`, read from `/proc` and `/sys`, with
|
||||
fixture tests. It is not a core component, so this issue needs no replay when it resolves. It resolves
|
||||
once the module is merged and assigned, and its tools answer on the control node.
|
||||
@@ -0,0 +1,49 @@
|
||||
# 312 — Diagnosis
|
||||
|
||||
## 2026-10-08 — where the missing tools belong
|
||||
|
||||
**What exists.** The mesh's overview and the control node's machine view list every seat and module tool.
|
||||
No seat verb or tool reads CPU iowait, `/proc/diskstats`, I/O pressure stall, a cgroup's `io.stat` or a
|
||||
process's `/proc/<pid>/io`. The nearest matches are:
|
||||
|
||||
- `memory-pressure`, which reads the same kind of kernel counters for memory, including memory PSI;
|
||||
- the container module's `docker_stats` and `docker_disk_usage`, which give each container's
|
||||
cumulative block I/O and the runtime's own space;
|
||||
- `logrotate_big_logs`, `snapd_disk_usage` and `flatpak_disk_usage`, each the space of one thing.
|
||||
|
||||
The records hold no design for a machine's resources. [ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md)
|
||||
judges what a module runs. [ADR 0241](../../02-DECISIONS/0241-a-machine-says-how-its-network-is-and-an-outside-writer-of-a-mesh-file-is-a-finding.md)
|
||||
adds the machine's networking to that judgement. Neither covers its processors, memory or storage
|
||||
as a load.
|
||||
|
||||
**Ruled out:**
|
||||
|
||||
1. *A verb on an existing node seat.* Every node seat is a role one module holds on each machine:
|
||||
the service manager, the packet filter, the uplink, the hostname, the backup, the build agent, the
|
||||
login shell and intrusion prevention. None is the machine's storage. A disk report as a verb of the
|
||||
service manager would put it under a seat whose holder is chosen for something else.
|
||||
2. *A new seat for the machine's resources.* A new seat's verb is promised in the controller before
|
||||
it is required ([ADR 0246](../../02-DECISIONS/0246-a-seats-new-verb-is-promised-before-it-is-required.md)).
|
||||
A seat also exists so one module holds a role exclusively and acts through it. A read of kernel
|
||||
counters needs no exclusivity and acts on nothing. Two modules reading `/proc/diskstats` do
|
||||
not conflict. The cost buys nothing now.
|
||||
3. *A tool in `memory-pressure`.* It is the nearest relative, but it owns compressed swap and
|
||||
systemd-oomd. Assigning it to the control node to read the disk would add zram at the next boot
|
||||
and enable oomd there, which is a change of behaviour that reading should not need. It is not
|
||||
assigned on the control node today.
|
||||
4. *A tool in the container module.* It sees containers. Issue 306's disk is shared with the forge
|
||||
and the bus, which also run as containers, but the question is the machine's, and a host
|
||||
service's writes would be missed.
|
||||
|
||||
**Chosen.** A small read-only module, `disk-load`, beside `netcheck`, which reads the network the same
|
||||
way. It declares no resources, so it can be assigned to every machine without changing one. If the mesh
|
||||
later makes a machine's resources a seat, its four tools are that seat's verbs as written.
|
||||
|
||||
**One finding while building it.** `/proc/<pid>/io` is readable only for the tool runner's own
|
||||
account unless the runner is root. On the laptop it hid 502 of the machine's processes, including the
|
||||
build agent's container, which was writing 83 MiB/s at the time. A cgroup's `io.stat` is readable by
|
||||
every account. So the module ranks units and containers from their cgroups and names processes
|
||||
only as a second view, and it says how many it could not read. Two kinds of double counting are
|
||||
avoided: `io.stat` includes a cgroup's descendants, so only the innermost unit counts, and a write
|
||||
through an encrypted volume is charged again to the disk under it, so only a disk built on nothing
|
||||
counts.
|
||||
Reference in New Issue
Block a user