--- effort: 006-mesh-from-scratch updated: 2026-08-23 --- # The skeleton Repositories at the root, modules inside them, parts at the leaf. Four tiers, and a dependency rule that only points downward. ## What a tier is A tier answers one question: **what has to exist before this can exist?** It is not importance, and it is not a layer in the networking sense. It is bootstrap order, made explicit. Walk a bare machine to a running mesh and the tiers fall out of the story: ``` bare machine │ one command lands ONE binary. Nothing else exists. TIER 0 host │ │ it reads a pinned file it already carries and raises a │ database, a bus, an object store and an image │ registry — locally, alone. TIER 1 substrate │ │ on those, the mesh's brain starts: which nodes exist, │ what runs where, what is reachable. TIER 2 control plane │ │ ways to talk to that brain. TIER 3 surfaces │ └ everything the mesh then carries. TIER 4 workloads ``` **Dependencies point only downward.** The substrate never references the control plane. That one constraint is the entire bootstrap answer, because it guarantees there is always a place to start. Why it earns its keep here: **the current mesh violates this, and that is the circularity that keeps recurring.** The mesh database is a module; modules are installed by the delivery pipeline; the pipeline needs the database. No order works, so a first-node script exists to paper over it, and every later substrate change has to pretend the problem is not there. A tier is not a repository and not a bounded context. Those are different cuts: a context says *who owns this concept*, a tier says *what must already be running*. ## The tree ``` mesh-host/ TIER 0 — the only thing ever installed by hand apply/ reconcile declared state on this machine inventory/ what this node is, has, and is capable of link/ the single outbound connection to the control plane store/ embedded local state — authoritative while disconnected profile/ capability detection: managed · user · edge substrate.lock pinned tier-1 descriptor, appliable with no mesh present mesh-substrate/ TIER 1 — declarations only, no logic of its own store/ relational state bus/ commands and events objects/ blobs and build artifacts images/ container images bundle.yml the pinned set tier 0 can raise alone mesh-control/ TIER 2 — the control plane record/ the event log every context integrates through inventory/ nodes · modules · assignments · versions config/ settings · secrets · derivation onto nodes connectivity/ overlay · resolution · exposure · filtering · certificates provisioning/ resource grants between modules delivery/ source → artifact → node observability/ health · logs · metrics · alerts identity/ agents · humans · services · authorisation work/ tasks · workflows · runs knowledge/ memory · documents · retrieval api/ the one interface every surface speaks to mesh-surfaces/ TIER 3 — thin; no logic lives here tools/ the agent-facing tool surface web/ the operator-facing interface cli/ the shell-facing interface mesh-catalog/ TIER 4 — what the mesh hosts / grouped per ADR 0022, list per research 005 mesh-lab/ the whole mesh, disposable, on one machine mesh-sdk/ contracts shared across tiers — types, not behaviour hq/ company-scoped, not a mesh repository — ADR 0019 ``` ## The dependency rule **A tier may depend only on tiers below it.** Substrate never references the control plane. The control plane never reaches into a node except through the host. A surface holds no logic a second surface would have to reimplement. This is the whole of the bootstrap answer, and per this repository's own rule it must say how it is checked: a dependency-direction lint in the build, failing on an upward import. A tier rule enforced by intention is the same as no tier rule — that is [ADR 0010](../../02-DECISIONS/0010-delivery.md) applied to architecture. ## Move 1 — the substrate is applied, not delivered **The problem.** The mesh needs a database, a bus, an object store and an image registry. Today those are modules, and modules are installed by the delivery pipeline, which needs the database and the bus. The first node is therefore raised by a special script that exists only because of the circularity, and every later change to the substrate has to pretend the circularity is not there. **The move.** The host can apply a declaration without anyone telling it to. The substrate is a **pinned bundle** the host carries: a fixed, versioned, self-contained descriptor of the four services and nothing else. Raising a first node is `host apply substrate.lock` — not a special path, just the ordinary one with no control plane on the other end. The circularity disappears rather than being worked around: **the substrate is applied by tier 0, the mesh is delivered by tier 2, and they are different mechanisms on purpose.** The price is real and should be named: the substrate is upgraded by bumping a pin and re-applying, not by the pipeline. It gets less machinery than everything else — no per-node selection, no provisioning, no fan-out — and that is the point. Five services justify a simpler mechanism than a hundred. ## Move 2 — the host is one binary with capability profiles **The problem.** Everything assumes root on a machine whose packages, services and network the mesh owns. A phone cannot offer that, and neither can a work laptop. The current answer would be a lightweight fork, which means two implementations and one of them rotting. **The move.** One host binary, and a **profile** it detects rather than is told: | Profile | Can | Typical | |---|---|---| | `managed` | packages, services, network, filesystem — the full surface | a machine the mesh owns | | `user` | user-level services and tools; no package or network management | a shared or administered machine | | `edge` | report presence, relay, expose a tool surface; hold nothing | a phone | A module declares which profiles it can land on. Assignment to an incapable node fails at declaration time, not at deploy time — a phone is not a machine that fails to install a firewall; it is a node the firewall module cannot be assigned to. This makes the phone case a **capability question rather than a platform question**, which is what keeps it from becoming a second implementation. It also removes the current unstated assumption that every node is equivalent — already false, and today handled by remembering. ## Move 3 — connectivity becomes a context [ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) names nine contexts and none of them owns the overlay, the resolver, the firewall or the ingress. `config` owns PKI, which is the closest thing, and it is not close. Meanwhile [research 005](../005-domain-grouping/analysis.md) measured the whole catalogue and found that reachability is the **only** place where modules genuinely change together under one intent — the proxy with the resolver, the firewall with the overlay, repeatedly, because *how a node is reachable* is one question asked in four places. So the evidence and the gap point the same way. `connectivity` owns: - the overlay every node joins, and the addresses on it - name resolution, internal and public - exposure — which services answer from outside, on which names - filtering — what may reach a node at all - certificates for both name spaces This is an addition to an accepted record, so it is a decision, not a drafting choice. It belongs in a new record that extends ADR 0001 the way [ADR 0009](../../02-DECISIONS/0009-modules-and-the-graph.md) does — not written here. ### Where the networking actually lives A context is an authority, not a running thing, so naming one does not say who brings the overlay up. That splits three ways, and the split is the design: | Concern | Tier | Why there | |---|---|---| | **Overlay membership** — this node joins, holds an address, keeps the tunnel up | **0, in the host** | Everything cross-node needs it *before* the substrate is reachable from elsewhere. A module cannot provide it, because installing a module is itself a cross-node operation. | | **Policy** — who holds which address, what resolves, what is exposed, what is filtered | **2, `connectivity`** | Bookkeeping and authority. It decides; it runs nothing. | | **Machinery** — resolver, reverse proxy, firewall backend, certificate issuance | **4, modules** | Swappable, and not every node needs them. A node without a reverse proxy is still a node. | There is a second circularity hiding here, and it has to be closed explicitly: **the host's link to the control plane does not run over the overlay.** If it did, the overlay would have to be up before the host could be told how to join it. The link is ordinary outbound internet to a public endpoint; the overlay carries node-to-node traffic only. Joining is therefore: host lands → links out with a join token → control plane returns an address and keys → host raises membership → the substrate on other nodes becomes reachable. This also makes the `edge` profile honest rather than special-cased. A phone can hold overlay membership in userspace without privilege, and cannot run the machinery. That is the profile distinction doing its job. ## Move 4 — `feature` splits in two The invitation was to check whether the concept survives. It does not, in one piece. Today a **feature** means both *a thing built once* and *a thing selected per node*, and the delivery pipeline is hard to reason about precisely because those have different cardinality and one word ([ADR 0010](../../02-DECISIONS/0010-delivery.md) is the pipeline half of the same confusion). Split it: | Concept | Is | Cardinality | |---|---|---| | **artifact** | something built and published — an image, a bundle, a package | once per module version | | **part** | an independently selectable piece of a module's desired state | chosen per node, per assignment | A module declares desired state in parts, and produces artifacts. An assignment names the node and the parts. Delivery builds artifacts once and applies parts per node — and the two words now carry the two cardinalities that the pipeline already has. This keeps what features are genuinely for — per-node opt-in of *some* of a module, which the work breakdown already calls for — and drops the conflation that makes the current model confusing. ## What the mesh is, versus what it hosts Tiers 0–3 are the mesh. Tier 4 is everything it carries, and the boundary is stated by requirement rather than by taste: **a module is part of the mesh if removing it stops the mesh managing nodes.** A media server does not. A relational store does — which is why it sits in the substrate and not the catalogue, despite being, in every other respect, an application like any other. An identity provider, notably, does **not**: the control plane authenticates its own callers, so identity is a hosted service like the media server. That test also settles the IT-company goal without a special category. Development, design and deployment tooling are **workloads** — tier 4, hosted, provisioned, delivered like anything else. The mesh does not grow a "company" feature; it hosts the tools a company runs on, and the fact that it runs its own development on them is dogfooding, not architecture. ## What agents are, structurally Self-improvement and self-healing are not a tier. Agents are participants ([ADR 0003](../../02-DECISIONS/0003-agents-are-persistent-employees.md)) that hold identity in tier 2, act through tier 3 like any other caller, and run as workloads in tier 4. This matters for one reason: **an agent must not have a privileged path**. Anything an agent can do to the mesh, a person can do through the same surface, and anything it cannot express through the tool surface is a gap in the surface rather than a reason for a back door. Self- healing built on a private channel is unreviewable, and would be the one part of the mesh with no human checkpoint. ## How this is tested The lab ([ADR 0016](../../02-DECISIONS/0016-the-lab.md)) raises the tree above on one machine: virtual machines as nodes, a real overlay between them, the real substrate bundle, the real control plane, the real delivery path. The tier rule is what makes that affordable. A scenario needing only tiers 0 and 1 is one virtual machine and a pinned bundle — which is also, exactly, the bootstrap path. **The hardest thing to test becomes the cheapest scenario to run**, and the first-node path stops being the one thing nobody exercises until it breaks. ## What this skeleton does not answer - Where the record lives. It is infrastructure by shape and domain by content, and putting it in the substrate risks recreating a circularity in the one place the design just removed one. - Whether tier 2's contexts are one repository or several. Open from ADR 0001 already. - Whether an `edge` node is in the inventory or merely present — which decides whether "node" is one concept or two. - The migration. Nothing here says how today's mesh becomes this, and the skeleton is worth little until that is costed. ## A naming near-miss, recorded The tier-0 binary was first called `mesh-agent`, because "node agent" is the reflex everywhere else in the industry. That is wrong here, and wrong in the specific way [`how-we-build.md`](../../00-META/how-we-build.md) §4 exists to catch: **Agent** is a first-class concept in this mesh — a participant, some of whom are human, holding identity and memory ([ADR 0003](../../02-DECISIONS/0003-agents-are-persistent-employees.md)). One document carried both meanings. It is the same failure as the anatomy naming in the current runtime: an evocative domain word pointing at infrastructure. [ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) supplies the fix in its own title — *the mesh brokers capabilities; nodes host; agents think.* Three verbs, three components: the control plane **brokers** (`mesh-control`), the tier-0 binary **hosts** (`mesh-host`), the participant **thinks** (`agents`, untouched). `mesh-node` was the alternative and was rejected: *Node* is the aggregate in the inventory — the record of a machine — while the binary is what runs on it and does the hosting.