From 7839703b608ab49d64a80c289b18e0bd16ff64c3 Mon Sep 17 00:00:00 2001 From: jochen Date: Thu, 8 Oct 2026 10:48:44 +0200 Subject: [PATCH] Issue 312: no tool said a machine's disk load, so issue 306 could not measure its disk --- .../00-report.md | 45 +++++++++++++++++ .../01-diagnosis.md | 49 +++++++++++++++++++ 2 files changed, 94 insertions(+) create mode 100644 04-ISSUES/312-no-tool-said-a-machines-disk-load/00-report.md create mode 100644 04-ISSUES/312-no-tool-said-a-machines-disk-load/01-diagnosis.md diff --git a/04-ISSUES/312-no-tool-said-a-machines-disk-load/00-report.md b/04-ISSUES/312-no-tool-said-a-machines-disk-load/00-report.md new file mode 100644 index 00000000..32968cec --- /dev/null +++ b/04-ISSUES/312-no-tool-said-a-machines-disk-load/00-report.md @@ -0,0 +1,45 @@ +--- +status: located +opened: 2026-10-08 +located-in: [mesh-catalog modules/disk-load] +fixed-by: novox/mesh-catalog PR #120 +amended-design: +--- + +# 312. No tool said a machine's disk load + +## Symptom + +[Issue 306](../306-a-check-on-the-control-node-ran-past-its-timeout/00-report.md) found the controller's +checks running five to seven times slower on the control node than elsewhere, and only in the packages +that touch a store. The evidence pointed at the control node's disk, which the throwaway stores share with +the mesh's own store, the database server, the object store, the bus and the forge. The next question was +the obvious one: how busy is that disk, how long does a request wait, and who is writing? + +No tool in the mesh could answer it. A search of every tool on every machine for disk load, I/O wait, +throughput or latency returned two tools, and both report how much space one package format takes. The +only per-container figure was the container runtime's cumulative block I/O. That figure gave the +16–20 GB per throwaway store quoted in issue 306, but it has no rate, no wait and no queue, and it +says nothing of anything outside a container. The machine's own counters, `/proc/stat`'s iowait, +`/proc/diskstats`, `/proc/pressure/io`, each cgroup's `io.stat`, and `/proc//io`, were readable +on every machine, and nothing in the mesh read them. + +The only way to get the numbers was a login to the machine, which the mesh's rule refuses +([ADR 0245](../../02-DECISIONS/0245-a-verb-says-what-it-replaces-and-the-agent-is-guarded-from-working-round-the-mesh.md)): +a missing tool is built in the module that owns the area, never worked round. + +## Why it is an issue + +ADR 0245 turns a missing tool from an inconvenience into a stop. An agent asked why a machine is slow +can read its memory (`memory-pressure`), its units, its containers' CPU and memory, and its network, but +not its storage. That is the one resource issue 306 needed. A diagnosis that has to guess which +resource is short guesses wrong first. Issue 306's first guess was plain overload, and its package +timings had to rule that out the long way. + +## Located + +No module owned the machine's storage load, and no node seat holds it. `01-diagnosis.md` says where the +tools belong and why. The fix is a new read-only catalogue module, `disk-load`. Its tools are +`disk_load`, `disk_top`, `disk_filesystems` and `disk_devices`, read from `/proc` and `/sys`, with +fixture tests. It is not a core component, so this issue needs no replay when it resolves. It resolves +once the module is merged and assigned, and its tools answer on the control node. diff --git a/04-ISSUES/312-no-tool-said-a-machines-disk-load/01-diagnosis.md b/04-ISSUES/312-no-tool-said-a-machines-disk-load/01-diagnosis.md new file mode 100644 index 00000000..f7d418a5 --- /dev/null +++ b/04-ISSUES/312-no-tool-said-a-machines-disk-load/01-diagnosis.md @@ -0,0 +1,49 @@ +# 312 — Diagnosis + +## 2026-10-08 — where the missing tools belong + +**What exists.** The mesh's overview and the control node's machine view list every seat and module tool. +No seat verb or tool reads CPU iowait, `/proc/diskstats`, I/O pressure stall, a cgroup's `io.stat` or a +process's `/proc//io`. The nearest matches are: + +- `memory-pressure`, which reads the same kind of kernel counters for memory, including memory PSI; +- the container module's `docker_stats` and `docker_disk_usage`, which give each container's + cumulative block I/O and the runtime's own space; +- `logrotate_big_logs`, `snapd_disk_usage` and `flatpak_disk_usage`, each the space of one thing. + +The records hold no design for a machine's resources. [ADR 0240](../../02-DECISIONS/0240-a-module-says-how-it-is-healthy-and-the-node-engine-judges-it.md) +judges what a module runs. [ADR 0241](../../02-DECISIONS/0241-a-machine-says-how-its-network-is-and-an-outside-writer-of-a-mesh-file-is-a-finding.md) +adds the machine's networking to that judgement. Neither covers its processors, memory or storage +as a load. + +**Ruled out:** + +1. *A verb on an existing node seat.* Every node seat is a role one module holds on each machine: + the service manager, the packet filter, the uplink, the hostname, the backup, the build agent, the + login shell and intrusion prevention. None is the machine's storage. A disk report as a verb of the + service manager would put it under a seat whose holder is chosen for something else. +2. *A new seat for the machine's resources.* A new seat's verb is promised in the controller before + it is required ([ADR 0246](../../02-DECISIONS/0246-a-seats-new-verb-is-promised-before-it-is-required.md)). + A seat also exists so one module holds a role exclusively and acts through it. A read of kernel + counters needs no exclusivity and acts on nothing. Two modules reading `/proc/diskstats` do + not conflict. The cost buys nothing now. +3. *A tool in `memory-pressure`.* It is the nearest relative, but it owns compressed swap and + systemd-oomd. Assigning it to the control node to read the disk would add zram at the next boot + and enable oomd there, which is a change of behaviour that reading should not need. It is not + assigned on the control node today. +4. *A tool in the container module.* It sees containers. Issue 306's disk is shared with the forge + and the bus, which also run as containers, but the question is the machine's, and a host + service's writes would be missed. + +**Chosen.** A small read-only module, `disk-load`, beside `netcheck`, which reads the network the same +way. It declares no resources, so it can be assigned to every machine without changing one. If the mesh +later makes a machine's resources a seat, its four tools are that seat's verbs as written. + +**One finding while building it.** `/proc//io` is readable only for the tool runner's own +account unless the runner is root. On the laptop it hid 502 of the machine's processes, including the +build agent's container, which was writing 83 MiB/s at the time. A cgroup's `io.stat` is readable by +every account. So the module ranks units and containers from their cgroups and names processes +only as a second view, and it says how many it could not read. Two kinds of double counting are +avoided: `io.stat` includes a cgroup's descendants, so only the innermost unit counts, and a write +through an encrypted volume is charged again to the disk under it, so only a disk built on nothing +counts.