Files
hq/01-RESEARCH/006-mesh-from-scratch/00-overview.md
T
jschoubben 72b22830f3 ADRs 0036, 0037, 0038 — what a node is, what the host does, how one joins
0036 (accepted): a node is a managed machine, and disconnection is a situation.
The open question posed a class distinction — full nodes and lesser presences.
There is none. Reachability is state, not kind, which promotes the host's local
store from a component to a requirement: it is what makes disconnection ordinary
rather than exceptional. The reduced contract the question reached for is real
but it is capability, and that belongs in the profile.

0037 (accepted): the host applies, it does not decide. Measured rather than
argued — the absorption is smaller than the machinery that already applies
state, and eight of ten adapters carry no dependency to move. The two that do
open a Postgres connection to the control plane, which inside tier 0 is the one
thing the tier rule exists to forbid. So each concern splits: deciding needs
every other node and stays in tier 2; applying needs root and locality and goes
to tier 0. The host carries ONE concern, of which the six are instances.

0038 (proposed): a node joins by linking first. The operator's two-modes
proposal, adopted as intent and corrected as structure. Two modes is two code
paths where the first runs once per mesh and rots — and the mesh already has
that fault in its worst form, as three hand-run shell scripts. Instead: one
behaviour, two sources of declaration. The first node is not a different kind of
node, it is a node whose mesh is not up yet, and its specialness is temporary
and self-erasing.

0038 also shrinks the migration 0037 called expensive: a joining node never
needs mesh-wide state, because the hard part of the overlay is only needed to
compute the WHOLE mesh. It needs one peer. The rest arrives.

Left open and said so: what may be pushed over the link and how a joining node
proves it is entitled to join, and whether one host can raise the substrate
alone.
2026-08-25 10:34:45 +02:00

6.2 KiB

status, initiated, touches, became
status initiated touches became
active 2026-08-23
02-DECISIONS/0015-mesh-brokers-nodes-host-agents-think.md
02-DECISIONS/0017-modules-outside-the-core-are-grouped-by-domain.md
03-DESIGN/00-as-is/00-overview.md
03-DESIGN/01-to-be/00-work-breakdown.md

006 — The mesh designed from nothing

What is being investigated

What the mesh would look like if it were laid out today, with the requirements known and none of the accumulated shape — expressed as a skeleton: repositories at the root, modules inside them, and whatever turns out to be the right leaf unit below that.

The deliverables are skeleton.md — tiers, repositories and the four design moves — and code-skeleton.md — the tier test, what a module looks like on disk, and where today's catalogue lands.

Why

Every structural decision so far has been a correction: eight contexts replacing thirty-three modules (ADR 0015), domains replacing single-function modules (ADR 0017). A correction inherits the frame of the thing it corrects, and two of the mesh's oldest problems look unsolvable from inside that frame:

  • The bootstrap circularity. The mesh needs a database, a bus, a registry and an identity provider. Those are modules the mesh installs. The mesh cannot install them before it exists. This has been worked around repeatedly and never designed away.
  • Participation requires privilege. Everything assumes root on a machine whose packages and services the mesh owns. A phone cannot participate on those terms, and neither can a machine someone else administers.

Designing from nothing is a way to find out which parts of the current shape are requirements and which are residue.

And there is more residue than expected, from a knowable source. This began as a dotfiles repository — the first two days of history adopt dotfiles, add per-node dotfile overrides, and introduce service symlinking with an ignore file. The flat one-directory-per-tool catalogue, linking rather than copying, adoption of already-configured machines, per-node overrides, and the desktop modules are all inherited from that, not chosen for a mesh. Recorded in 03-DESIGN/00-as-is/10-module-catalogue.md.

That makes this effort's question sharper than "what would we do differently": much of what looks like design is a generalisation of place files on my machines, never revisited because it was never stated as an assumption.

The requirements this is designed against

Stated by the operator, recorded here so the skeleton can be checked against them rather than against taste:

  1. The mesh manages multiple computers — full control, through modules installed to nodes.
  2. Mesh state lives in a database: which modules on which nodes, logs, configuration.
  3. Configuration has several touchpoints — tool surface, web interface, others — all hosted by the mesh itself.
  4. Connectivity is core: every node reachable from every other over a shared overlay, some nodes publicly exposed, firewalls configured.
  5. The mesh hosts applications — and requires some of them itself. This is the circularity.
  6. Arch Linux only for now; ideally any device, including phones, on lighter terms.
  7. The end goal is to operate an IT company on it — development, design, deployment, full circle, self-hosted. Personal cloud infrastructure.
  8. Agents make it self-improving and self-healing.
  9. It is end-to-end testable on one machine (ADR 0016).

Status

A first skeleton exists, with four design moves that the current shape does not have. It is active because two of them are unproven and one contradicts a record that is already accepted.

Finding worth stating up front: ADR 0015 names nine bounded contexts and none of them owns connectivity — no overlay, no resolution, no firewall, no ingress. Requirement 4 has no home in the accepted decomposition, while research 005 found reachability to be the only part of the catalogue where modules genuinely change together under one intent. The skeleton adds it.

Open questions

Question Why it is open
Does the record — the event log contexts integrate through — belong to the substrate or the control plane? It is infrastructure by shape and domain by content. Placing it wrong reintroduces a circularity.
One repository per tier, or per context? Already open from ADR 0015 as "catalogue destination — one repository or many". The skeleton assumes per tier and does not settle it.
Does an unprivileged node earn a place in the inventory, or only a presence? Answered 2026-08-25 by the operator: a node is a managed machine inside the mesh, not an unprivileged something — and a disconnected node is still a node, in a different situation. The question posed a class distinction; the answer is that there is none, and what varies is state. Recorded as ADR 0036.
Does absorbing overlay, filtering, packages, supervision and the container runtime make the host too large? Answered 2026-08-25 — host-size.md. Measured: the absorption is smaller than the machinery that already applies state, and eight of ten adapters already carry no dependency. The risk is not size but direction, and it is two modules wide. The claim survives with its scope corrected — the host carries one concern, apply declared state on this machine, of which the six are instances. Recorded as ADR 0037.
Four substrate services or five? The identity provider passes the tier test only if the control plane delegates authentication rather than doing it natively.
Does feature survive? The skeleton splits it in two and argues the conflation is what makes the delivery pipeline hard to reason about. Unproven.