Files
hq/02-DECISIONS/0009-modules-and-the-graph.md
T
jschoubben 554f6bd7a4 A capability may carry a value, and adding one is not free
Recorded while building the seat detector. A capability is a named fact about a
machine: its presence gates an assignment and its detail can carry a value, so
"can this run here" and "what should it be configured as" are the same fact
read two ways. A verdict has always had a detail beside its yes or no, so
panel: oled needs no new concept.

Two things that keep the set honest, both worth writing down before anyone adds
the fiftieth capability. It must be detected and the detector must say how it
knows -- so nobody can add one they cannot check, which is the whole of issue
007. And detectors ship inside the host, which is one static binary, so adding
a capability means shipping a new host everywhere. That argues for a small
general vocabulary rather than a specific one.
2026-08-29 21:12:21 +02:00

218 lines
11 KiB
Markdown

---
topic: what runs on it
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 9. Modules and the graph
*Consolidated 2026-08-28 from six records.*
## Everything is a module
One kind of thing, one manifest describing all of them. A database, a web application, a window
manager and a firewall rule set are all modules — not because they are alike, but because
**anything else means a second kind of thing with its own rules, and then a third.**
**A module is the unit of delivery**: assignable to a node, versionable, replaceable on its own.
## There are no domain modules
An earlier decision grouped modules by domain — four things constituting *how a node is
reachable* becoming one `networking` module. **That was wrong, and the correction is worth
keeping** because the observation behind it was right.
The measurement holds: reachability is the **only** place in the catalogue where modules
genuinely change together under one intent. What did not hold is the conclusion. Tight coupling
means they share an **authority** — one place that decides for all of them — and not that they
should be one artifact. `wireguard` and the proxy are deployed to different sets of nodes, so a
module containing both would be assigned where half of it is unwanted.
> **Coherence is a context. Delivery is a module.**
**Folders assert relationships; edges record them.** What grouping was for — finding things,
seeing what belongs together — is a tag and a query, neither of which anybody has to keep true by
hand.
## Three edges
| edge | means | declared? | satisfied |
|---|---|---|---|
| **presence** | that thing must exist and be reachable here | yes | at provisioning |
| **instantiation** | that thing makes something for me and hands back credentials — a database, a bucket, a route | yes | at provisioning, and again whenever it must be |
| **build** | I was compiled against that artifact | **no — read from imports** | **at build, once** |
**Instantiation implies presence; presence does not imply instantiation.**
**A route is an instantiation edge**, and it is worth noticing because the direction is the mirror
of a database: the consumer supplies a target and receives a *name*, rather than supplying nothing
and receiving credentials. Same edge.
**Provider stops being a category.** Any hosted thing can be a factory — an identity provider
grants clients, a mail server grants mailboxes. It is a facet, not a kind.
**A module may also declare what it claims**, because some things cannot coexist and that is a
fact about the module rather than about a particular node. What that means precisely is below.
### Why the build edge is a different kind
It is fixed inside an artifact rather than negotiated when something runs, and **its only remedy
is a rebuild** — nothing can re-provision it.
It is also **derived rather than declared**, and the asymmetry is deliberate: a runtime edge is an
*intention* somebody has about how the mesh should be wired, and only a person can state it. A
build edge is a *fact about code that already exists*, and a declared list of dependencies drifts
from the imports it describes.
**An artifact is out of date when its source moved, or when anything it was built against moved.**
So what is recorded is a commit *and the identity of every artifact it was built against*, which
is what makes the rebuild set computable and *is this current?* answerable without building.
**The graph measures design quality, not just build order.** A module with many inbound build
edges is one whose every change is expensive — and that is readable before anything is built. The
current shared library is exactly that, and nobody could see it because nothing drew the edges.
## Provisioning is declared, never configured by hand
A module declares what it **provides** and what it **requires**. The mesh satisfies it: a
provisioner belonging to the provider creates the resource and its credential, records the grant,
and the values are derived onto the consumer. **Neither the credential nor the topology is ever
written by hand.** A requirement may name a provider on another node, so cross-node wiring is the
same declaration.
## What a module claims, and why it is not a list of rivals
*Written 2026-08-29, replacing pairwise exclusion.*
**Exclusivity is not a property of a module. It is a property of a singular resource the module
takes over.** Two shells do not compete for anything and any number may be installed. Two display
servers both want the seat, and only one may have it.
> **A module declares what it *claims*. Two modules claiming the same thing cannot both be
> assigned within that claim's scope.**
**Not "xorg conflicts with wayland".** Pairwise exclusion has a property that only shows up later:
adding a third display server means **editing xorg and wayland to know about it**. Every new
module requires changing modules nobody who wrote it owns, and the edits grow as the square of
the count. With a claim, the third one says `claims: the seat` and nothing else changes anywhere.
**The new module is the only thing that has to know anything** — which is the difference between
a catalogue that grows and one that calcifies.
The pattern is common enough to be worth listing, because seeing it is most of understanding it:
| these coexist | these claim one thing |
|---|---|
| shells — bash, zsh, fish | display servers — xorg, wayland (*the seat*) |
| editors — vim, emacs, helix | init — systemd, openrc (*pid 1*) |
| language runtimes | container runtime — docker, podman |
| terminal emulators | reverse proxies — nginx, caddy, traefik (*ports 80/443*) |
| browsers | time — chrony, timesyncd, ntpd (*the clock*) |
| | resolvers — resolved, dnsmasq, unbound (*`/etc/resolv.conf`*) |
| | network management — NetworkManager, networkd, netctl |
| | mail — postfix, exim, msmtp (*port 25*) |
| | audio — pipewire, pulseaudio (*the device*) |
**A claim has a scope**, because not everything singular is singular per machine:
| scope | example |
|---|---|
| **node** | the seat, pid 1, port 443 |
| **site** | a DHCP server on a segment |
| **mesh** | the hub, the control plane |
The last is not new — the mesh already enforces exactly one hub with a unique index
([ADR 0007](0007-connectivity.md)). Scope is that idea, said once rather than hard-coded per case.
**Some conflicts need no claim at all.** Two modules declaring the same file, or binding the same
port, are visible from *what they declare* — the mesh already holds every resource of every
declaration. So a claim is only written for the abstract ones, where nothing in the declaration
reveals the clash. That keeps the manifest small, which is worth protecting.
## A requirement with several answers is refused, never guessed
A module requiring *a shell* may be satisfied by three. The mesh does not pick.
| candidates | what happens |
|---|---|
| exactly one | assigned, silently — there was no choice to make |
| none | refused, naming what is missing |
| several | **refused, naming them**, and a person chooses |
**This is what makes a solver unnecessary.** Counting candidates is a few lines and has no
surprising behaviour; a solver that picks has to be understood before its answer can be trusted,
and it is understood by whoever is debugging it at the time. Nothing here is lost by waiting —
a solver can be added later without changing a single manifest, and the reverse is not true.
**Requiring a module and requiring a capability are different fields**, because the remedies
differ and the message should say which:
- *i3 needs xorg, which is not assigned here* — assign it.
- *this machine has no seat* — wrong machine; nothing can be installed to fix it.
### A capability may carry a value, and that is not a new idea
A capability is a named fact about a machine, **detected and never assumed**. Its presence gates
an assignment; its detail can also carry a value — `seat: card1-DP-1`, `panel: oled`, an
architecture, an amount of memory. Nothing new is needed for that: a verdict has always had a
detail beside its yes or no.
So *can this run here* and *what should it be configured as* are answered by the same fact, read
two ways. A module that must not be assigned without an OLED panel and one that dims itself
differently on one are reading the same line.
**What keeps the set from sprawling is the cost of adding one.** A capability must be detected,
and the detector must say how it knows — so nobody can add one they cannot check, which is the
whole of [04-ISSUES/007](../04-ISSUES/007-an-installed-package-is-not-a-capability/00-report.md):
an installed package was treated as a capability and a node was assigned work it could not do.
**And detectors ship inside the host**, which is one statically linked binary. Adding a capability
means shipping a new host to every node that needs it. That is a real cost and it argues for
keeping the vocabulary small and general — `seat`, not `has-nvidia-with-two-outputs`.
## "Flavor" is retired
It was carrying three unrelated meanings — variants of a thing, a subset of one module a node
installs, and whatever the current system does, which earned two knowledge-base entries about
going wrong. **A word with three meanings cannot be reasoned about**, and every attempt to design
around it produced a rule that was right for one meaning and wrong for the others.
What it was reaching for is two ordinary things:
- **Different modules that provide the same thing.** `zsh` and `fish` both provide *a shell*. They
are two modules, not one module with a switch: they share a name and nothing else — different
packages, different configuration, different everything.
- **One module with a setting.** A monitoring module that is an agent here and a server there is
one module, configured. Nothing varies but a value.
If something is neither, it is probably two modules.
## The core library is the mesh's domain
One module everything may depend on. It holds **what is true of the mesh regardless of which
context you are in**: a module, a node, an assignment.
The test: *would this still mean the same thing in a context that had never heard of the one it
came from?* A node would. A pipeline stage would not — that is delivery's.
**Types ship with the module that owns them**, not here. A consumer needing `inventory`'s types
depends on `inventory` — one narrow, visible edge — rather than everything depending on a hub
where the relationship cannot be seen. **A library everything depends on is expensive to change
whether it holds types or code; the fan-in is what makes it expensive**, which is why *types, not
behaviour* was the wrong guard.
**It stays small on its own.** A domain model changes when what the mesh *is* changes, which is
rare. A drawer labelled *shared* changes whenever anybody writes something reusable, which is
constantly — and *who else might want this* always answers yes, which is how the current one grew.
## Consequences
- **Fewer things will be shared, and some code will be written twice.** That is the trade: the
current library exists because sharing felt free. Two similar functions in two modules is often
the better answer.
- **The check is a measurement rather than a prohibition.** Inbound build edges say when something
is becoming a hub, while it is happening rather than after.
- **Reading build edges needs a language-aware tool per language**, which is the real cost and the
reason declaring them looks tempting. It is still wrong.