Files
hq/01-RESEARCH/006-mesh-from-scratch/host-size.md
T
jschoubben 42bce02bba 006: answer the host-size question by measuring it
The skeleton's biggest unproven claim was that absorbing six concerns makes a
binary whose whole argument is having no dependencies carry six of them.

Measured against origin/main, and the question turns out to ask about the wrong
axis. By size the absorption is SMALLER than the machinery that already applies
state on a node — 2755 lines of adapters against 3059 lines of meshware,
env-sync and config-sync. The host is not a new large thing; it already exists,
spread across three core modules.

The real risk is direction, and it is two modules wide rather than six concerns
wide. Eight of ten adapters already receive derived state and only apply it, so
absorbing them moves code that has no dependency to move. Two — wireguard and
traefik — open a Postgres connection to the control plane and compute their own
configuration, which inside tier 0 would be an upward dependency and is exactly
what the tier rule forbids.

And the split has already been happening without being named: dnsmasq-app needs
the same node data as wireguard and does not query for it, because hand-
duplicated state went wrong and someone derived it centrally instead. Eight of
ten adapters are on the far side of that migration.

So the absorption is not a move, it is a split: deciding stays in tier 2,
applying goes to tier 0. The claim survives with its scope corrected — the host
carries ONE concern, apply declared state on this machine, of which the six are
instances.

Stated open rather than glossed: the two unsplit modules are the two hardest,
six concerns is still six vocabularies even at zero dependencies, and what the
host must carry versus find is issue 007 and unresolved.

Question B also recorded as answered by the operator — a node is a managed
machine, and a disconnected node is still a node in a different situation. The
question posed a class distinction; there is none, and what varies is state.
2026-08-25 02:20:32 +02:00

5.7 KiB

Is the host too large?

The skeleton absorbs overlay membership, packet filtering, package management, service supervision, the container runtime and filesystem management into tier 0, and 00-overview.md calls this "the skeleton's biggest unproven claim. A binary whose whole argument is that it has no dependencies now carries six concerns."

This is that claim, measured.

Method

Against origin/main of the monorepo at 2026-08-25 — read through git refs rather than a checkout, which sits on a feature branch 1031 commits behind with uncommitted work.

For each module that implements one of the six concerns: how much code it is, and what it depends on to do its job. The second question turned out to be the one that matters.

Finding 1 — by size, the concern is misplaced

Absorbed into the host Lines
dnsmasq-app 706
traefik 426
ufw 385
mesh-ca 353
wireguard 313
incus 285
package-manager 118
zfs 67
fail2ban 60
docker-app 42
total 2 755
Machinery that already applies state on a node Lines
hal/meshware 1 829
hal/env-sync 626
hal/config-sync 604
total 3 059

Everything being absorbed is smaller than the machinery that already exists to apply it. The six concerns are not six subsystems; they are ten thin adapters averaging 275 lines, most of which is rendering a config file and running a command.

The host is not a new large thing. It already exists, spread across three core modules, and the absorption adds less code than those three already contain.

Size is therefore the wrong axis, and the open question asked about the wrong thing.

Finding 2 — the real risk is direction, and it is narrow

What each adapter needs in order to decide what to write:

Module Gets its inputs from Applier only?
ufw ~/.hal/modules/*/module.yml on the local disk yes
dnsmasq-app process.env.DNSMASQ_*, derived centrally by env-sync yes
fail2ban, mesh-ca, package-manager, docker-app, incus, zfs local state and declared env yes
wireguard direct pg connection — nodes, node_accessors, node_wg_keys, module_env no
traefik direct pg connection — nodes, mesh_ca no

Eight of ten are already pure appliers. They receive derived state and put it on the machine. Absorbing those into tier 0 moves no dependency at all — it moves code that already has none.

Two reach upward. wireguard and traefik open a connection to the control plane's database and compute their own configuration from it. Absorbing them as they are would put a Postgres client and knowledge of the mesh schema inside tier 0 — an upward dependency, which is precisely what the tier rule forbids and what the entire bootstrap argument rests on.

So the danger in the absorption is real, and it is two modules wide rather than six concerns wide.

Finding 3 — the split has already been happening, unnamed

dnsmasq-app is the same kind of module as wireguard: it needs every node's addresses and names. It does not query for them. Its own comments record why:

"the mesh DB already holds [this] in node_accessors and the WireGuard address, duplicated by hand on all four nodes. Renaming the namespace then meant editing four override rows nobody [knew about]."

and

"now derived from node_accessors"

That is one module having already made the move the skeleton proposes — deciding centrally, applying locally — for the ordinary reason that hand-duplicated state went wrong. Nobody named it as an architectural direction; it was reached by fixing a bug.

Eight of ten adapters are on the far side of that migration. Two are not.

What this means for the open question

The absorption is not a move. It is a split, and it is mostly already done.

The question "does the host become too large" assumed six concerns would arrive whole. They do not. Each divides:

  • deciding — what this node's overlay, names, exposure and filtering should be, which needs every other node and therefore belongs in tier 2;
  • applying — putting that on the machine, which needs root and locality and therefore belongs in tier 0.

Tier 0 absorbs the applying. That is ~2 755 lines today, most of it already dependency-free, against 3 059 lines of apply machinery the host needs regardless.

The claim survives, with its scope corrected. The host does not carry six concerns; it carries one — apply declared state on this machine — of which the six are instances. That is the skeleton's own tier-0 test ("does it apply state on a machine?") applied to itself.

What remains open

  • The two unsplit modules are the two hardest. Overlay needs every node's key, address, site and endpoint reachability; the proxy needs certificates and every node's exposed names. These are where "derive centrally, apply locally" is most work, and neither has been done. The measurement says the design is right; it does not say the migration is cheap.
  • Six concerns is still six vocabularies. Absorbing them adds no dependencies but does add surface: the host must know what a WireGuard peer, an nftables rule, a package, a unit, a container and a dataset are. Nothing here measures that cost, and it is the residue of the original worry.
  • What the host must carry versus what it must find. The host manages wg, nft, pacman, docker; it does not contain them. Issue 007 is exactly this question and is unresolved.