Files
hq/02-DECISIONS/0006-the-substrate-and-the-control-plane.md
T
jschoubben 918dc04916 What this actually is, and three things that were assumed
Four things settled by talking them through, all of which had been true in
somebody's head and written nowhere.

It is not a mesh in the peer-to-peer sense and will not become one. 0001 now
says what it is instead: machines linked by a private network, one node holding
knowledge of all of them, modules as the way anything is built and delivered,
and agents hired onto nodes to do the work. The word describes what machines
can reach, not how they are governed. "Master" overstates it the other way --
nothing needs that node to keep running, only to change.

0006 gains the option that would make it a real mesh, recorded as considered
rather than rejected by silence: every node holding the whole inventory, a
replication process, an elected master with promotion on failure. What settles
it is not the complexity but that it still would not deliver the name, because
application databases are not replicated -- so a genuine peer-to-peer mesh
means becoming a replicated database system for every consumer's data too. That
is a larger product than the thing it would support.

Also in 0006: three central roles, not one. Losing the control plane costs
change, losing the broker costs being told anything, and losing the hub costs
nodes in different places reaching each other at all -- which is operation, not
administration. Whether they are one node is not decided.

And SSH access is identity's. It appeared three times as something that uses
the overlay and never as something the mesh provides, which reads as settled
when nothing decided it. Nobody else could: the mesh is the only thing that
knows which humans and agents exist and which nodes they may reach. Node to
node SSH stays out -- the host has no inbound control surface by decision, and
nodes reaching each other that way is a second control path through the back
door.

0007 gains the requirement underneath all of it. Reachability was recorded as a
fact to track and never as a thing some node must have. The broker's node and
the hub must be dialable by every node at a stable address, or nothing can join
and a disconnected node cannot return. A mesh entirely behind NAT cannot be
raised. That is a precondition and it belongs with the others.

The link staying on the underlay is also argued now rather than asserted. At
join time it is forced; afterwards it is a choice, and the reason is that a
repair channel carried over the thing being repaired is not one. Moving it onto
the overlay, with fallback, is recorded as open with what it would have to get
right -- a WireGuard interface has no link state to test, and a silent fallback
is this repository's recurring fault in a new place.

0010 says in one line what was the intention throughout: the module system is
the CI/CD. Not a pipeline beside the mesh. Build, test, publish and deploy are
one reconciliation seen at four points, which is why a thing that cannot be a
module cannot be delivered.
2026-08-29 13:01:36 +02:00

255 lines
14 KiB
Markdown

---
topic: the tiers
status: accepted
date: 2026-08-28
deciders: jochen
reconstructed: false
---
# 6. The substrate and the control plane
*Consolidated 2026-08-28 from six records. Extended 2026-08-29, by building it: the language, and
what must be running before the control plane starts — which this record had left not established
and could not have settled the way it was asking.*
## The control plane is what needs to know about more than one node
That is the whole test, and it follows from the host applying rather than deciding: **deciding
needs knowledge a single machine does not have.**
| question | whose |
|---|---|
| write this file, with this content, with this mode | the **host** |
| which nodes should run the store | the **control plane** |
| is this unit running | the **host** |
| which peers belong in this node's overlay | the **control plane** |
| has this node been unreachable for a week | the **control plane** — nobody else is watching |
**Anything a single machine could answer alone is not the control plane's.**
### Seven contexts and one interface
**inventory, config, connectivity, provisioning, delivery, observability, identity** — plus
`api`, the one interface every surface speaks to. Each earns its place by the test above rather
than by being ours.
**`work`, `knowledge` and `stream` are mesh-hosted applications, not control plane.** A task does
not need to know a node exists. *Being ours does not make something infrastructure.*
**`identity` owns SSH access.** *Written 2026-08-29, on noticing it was assumed everywhere and
stated nowhere.* SSH appears three times across this design and every time as something that
*uses* the overlay — "the way back in", "every node reaches every other: SSH, services, ordinary
traffic" — while nothing said who hands out the keys. Nobody else could: the mesh is the only
thing that knows which humans and agents exist and which nodes they may reach, which is
`identity`'s definition. The node end already works, since an `authorized_keys` file is a file.
It is three questions wearing one name, and only two of them are the mesh's:
| | |
|---|---|
| **humans** | their key, on the nodes they are allowed on |
| **agents** | the same, with a lifetime — and revocation that has to actually work |
| **node to node** | **not a mesh function.** The host has no inbound control surface by decision ([ADR 0004](0004-a-node-and-how-it-joins.md)); nodes SSHing to each other would be a second control path arriving through the back door |
**Where the record lives is deliberately open.** Contexts integrate through it, which makes it
load-bearing, and putting it in the substrate risks recreating the circularity the tiers just
removed. Listing it as an eighth context would settle by naming what has not been settled by
arguing.
### One node runs it, and nothing takes over
**Declared, never elected.** No promotion, no quorum, no fencing, no split brain — none of it
built, so none of it can be subtly wrong.
#### The option that would make it a real mesh, and why not
*Written 2026-08-29. It had been rejected by never being written down, which is the weakest way
to reject anything.*
A genuine peer-to-peer mesh means **no node is special**, and that has a concrete price:
- every node holds the **whole inventory**, so there is a replication process between them;
- replication needs a writer, so one node is elected **master**, and something promotes a new one
when it drops — Redis Sentinel and its whole family of problems;
- and it still would not deliver what the name promises, because **application databases are not
replicated.** A workload's store lives where the workload lives.
That last point is the one that settles it. To make the mesh genuinely peer-to-peer we would have
to become **a replicated database system for everything running on it** — not for our own
inventory, for every consumer's data too. That is a product, and a much larger one than the thing
it would be supporting.
**So there are three central roles, not one**, and it is worth seeing them separately because
only the third costs operation:
| | its loss costs |
|---|---|
| **the control plane** | nothing can be *changed*. Nothing stops running |
| **the broker** | nothing can be told anything, or report anything |
| **the hub** | nodes in different places **cannot reach each other** ([ADR 0007](0007-connectivity.md)) |
**Whether these are one node is not decided here.** All three must be dialable by every node, which
pushes toward one; nothing says they must be.
**That is sound rather than merely cheap**, because the design already tolerates its absence by
construction: a node reconciles from its own store and never needed to ask anybody to hold the
state it was last given. **The control plane being down is not a new failure mode — it is every
node in the ordinary disconnected situation at once.** What is lost is *change*, not *operation*.
The honest half: this node is a single point of failure, recovery is **restore rather than
failover** — which makes backup the availability mechanism rather than hygiene — and
**certificate renewal is the clock.** An outage outlasting a renewal window expires every public
name, which turns an inconvenience into an outage on a timer. Nothing measures that today.
## The authority is the control plane, not a database
**There is no single mesh database.** Each context owns its store exclusively, and *the mesh
database* names a thing that will not exist.
**No node reads any of them** — not for writes, not for reads. A node is *told* what to own, over
the link, in a bounded vocabulary; it **states** what it applied, and the owning context writes.
The difference is the security boundary: something that can write cannot be prevented from
writing anything.
**A node runs from its own store always, not as a fallback.** The current arrangement's nastiest
property is that *a node running from cache looks identical to a node running from the database*,
with no age on the cache and nothing reporting divergence. Under this there is no second mode to
be mistaken for the first.
**What survives from the original decision:** the repository defines what exists, the mesh defines
what runs where, and no node-to-module mapping is ever committed. That is what makes the
repositories node-agnostic and why anything about the mesh can be published at all.
**The error underneath was a category error**: *source of truth* named a storage location when it
meant an **authority**. Once the store is the answer, *which database* becomes the question, and
shared schemas follow.
## The substrate is what the control plane consumes and cannot grant itself
Every module needing a database asks provisioning for one. The control plane needs a database too
and cannot ask itself, because it is not running yet. **That circularity is the definition**, and
anything on the wrong side of it is raised from the bundle the host carries.
| role | product | |
|---|---|---|
| relational store | **PostgreSQL** | its own state lives there |
| message bus | **LavinMQ** | it cannot grant itself a virtual host — and **precedes it**, below |
| object store | **MinIO** | it cannot grant itself a bucket |
| image registry | **an OCI registry** | it cannot grant itself a repository |
| identity provider | — | **conditional**: substrate only if the control plane delegates authentication, which is undecided |
**The role and the product are both written.** The role is what the argument turns on; the product
is what gets installed and pinned, and a design that names only the role does not record that the
choice was made. **The dependency is on the protocol** — AMQP, S3, OCI — which is what keeps
naming them safe. The store is the exception: the provisioning model uses databases, roles and
schemas as PostgreSQL means them.
**A container runtime is detected, not chosen** — docker or podman, because a machine that
already has one keeps it. Only the version probe differs between them; the behavioural difference
(podman has no daemon, so containers do not return after a reboot unless a unit is enabled)
belongs in the declaration rather than the host.
**Being substrate and being in the bundle are different questions.** PostgreSQL and LavinMQ must
precede the control plane; the object store and the registry are substrate by role and ordinary by
delivery, provisioned once there is a control plane to do it.
### Why the broker precedes it too
Written after the fact, because this record first left it *not established* and framed it as
turning on whether the control plane's own contexts talk to each other over the bus.
**They do not** — they are one process and dispatch internally. Under that framing the broker is
provisioned like anything else and the bundle stays at one image.
**The framing cannot answer the question.** What decides it is not how the contexts reach each
other. It is how the control plane reaches a *node* — and that is settled above: never except over
the link, and the link is AMQP ([ADR 0002](0002-nodes-communicate-over-a-broker.md),
[ADR 0004](0004-a-node-and-how-it-joins.md)). So:
```
the bundle raises the control plane
the control plane provisions the broker ← by telling a host to run it
telling a host happens over the link
the link is the broker
```
**And it is not avoided by the first node being local.** `enrol` dials the broker at the address in
its token, which is the first node's own third step. A machine that raised the mesh still joins it
the ordinary way, and that was deliberate — its specialness lasts two commands. Making it join by
some other route would buy a smaller bundle by giving up the property the design was built to have.
The broker precedes the control plane for the same reason PostgreSQL does: **the control plane
cannot grant itself the thing it would need in order to grant it.**
**What it costs.** Two images rather than one, against the wish above to keep the bundle at roughly
one so a person can read it — two is still readable, four would not be. And three things that are
not images, each an action the bundle declares and the host runs, the way the database already is:
a virtual host, a credential on it, and **a certificate**. That last is the awkward one: a token
pins the fingerprint a host must expect *before it sends anything*, so the broker needs a
certificate at a moment when there is no mesh to issue one and no public name to obtain one for.
Self-signed and pinned is the shape that fits; how it is later replaced by the certificates in
[ADR 0007](0007-connectivity.md) is not decided here.
## The control plane is written in Go
The same language as the host, so tiers 0 and 2 are one language and not two.
The reason that decides it is not familiarity. **Its image is pinned by digest in the bundle**,
which means it is fetched and run on a machine where no mesh exists yet — nothing to check it
against, nothing watching, and a person expected to have read the bundle and believed it. A
statically linked binary makes that image the program and nothing else: no interpreter, no package
tree, no transitive dependency that arrived because something needed a date library. Everything
under that line is something somebody would have to audit, on the one image the whole mesh is
raised from.
A second reason, smaller and still real: the control plane runs a reconcile loop of its own
([ADR 0010](0010-delivery.md) — artifacts against source, as the host reconciles machine state
against declarations). Two loops of the same shape are cheaper to hold in one head when they are
also the same language.
**The option rejected** is TypeScript, matching the lab and the surfaces that will speak to this.
The argument for it is that tier 3 is web and CLI, so a TypeScript control plane would share types
with its callers rather than generating a contract. True, and it does not reach far enough:
`mesh-sdk` is *contracts shared across tiers* and **tier 0 is Go**, so the contracts cross a
language boundary whatever tier 2 is written in. The choice is between generating them for one
consumer or for two.
**What it costs, plainly:** the control plane can import nothing that exists today, and a person
moving between tier 2 and tier 3 changes language. Neither is recovered later — the language is the
most expensive thing in this record to reverse.
## The installer fetches what it pins
`substrate.lock` carries **references, not payload** — an image name and a **digest**, fetched at
apply time. A tag moves; a digest does not, and reproducibility comes from pinning the identity of
a thing rather than carrying its bytes.
The assumption that a machine might have no network came from the lab and was wrong: a machine
being adopted has one, and the sealed case is the lab.
**The lab places images by raising a registry inside the scenario**, which is what a real node
pulls from anyway — so it tests the real path rather than a stand-in for it. The digests that
registry serves are its own, and that satisfies this rule: what is required is a reference that
is **exact and cannot move**, and one it assigned is both. Assuming an upstream digest had to be
preserved is what made this look impossible for a while
([04-ISSUES/009](../04-ISSUES/009-a-digest-pinned-image-cannot-be-placed-in-the-lab/00-report.md)).
**Its contents are per operating system** even though its mechanism is not — package names, unit
names and service names all differ, so an Arch host embeds an Arch bundle.
## Consequences
- **The bundle stays small and reviewable.** A list of pinned references is something a person can
read; a bundle containing images is not.
- **An apply can fail because something is unreachable**, which a self-contained artifact could
not. That must fail *legibly*, naming what could not be fetched and from where.
- **Cross-context reporting is harder, and that is the point.** Anything wanting to see across
contexts consumes their events or calls their interfaces.
- **A queue with no limit grows until the broker's disk is full**, and the broker is what every
node depends on. The bound is per queue and is not decided.
- **The bundle carries two images and four actions**, and the substrate bootstrap grows a step.
- **Nothing in the first node's path is special-cased.** Enrolment is walked on node one.
- **The broker's certificate at bootstrap has no answer yet**, and is named as unfinished rather
than assumed. It is the first thing that will be wanted when the link is built.
- **The language cannot be revisited cheaply.** It is the one line here close to irreversible.