Files
hq/02-DECISIONS/0006-the-substrate-and-the-control-plane.md
T
jschoubben 918dc04916 What this actually is, and three things that were assumed
Four things settled by talking them through, all of which had been true in
somebody's head and written nowhere.

It is not a mesh in the peer-to-peer sense and will not become one. 0001 now
says what it is instead: machines linked by a private network, one node holding
knowledge of all of them, modules as the way anything is built and delivered,
and agents hired onto nodes to do the work. The word describes what machines
can reach, not how they are governed. "Master" overstates it the other way --
nothing needs that node to keep running, only to change.

0006 gains the option that would make it a real mesh, recorded as considered
rather than rejected by silence: every node holding the whole inventory, a
replication process, an elected master with promotion on failure. What settles
it is not the complexity but that it still would not deliver the name, because
application databases are not replicated -- so a genuine peer-to-peer mesh
means becoming a replicated database system for every consumer's data too. That
is a larger product than the thing it would support.

Also in 0006: three central roles, not one. Losing the control plane costs
change, losing the broker costs being told anything, and losing the hub costs
nodes in different places reaching each other at all -- which is operation, not
administration. Whether they are one node is not decided.

And SSH access is identity's. It appeared three times as something that uses
the overlay and never as something the mesh provides, which reads as settled
when nothing decided it. Nobody else could: the mesh is the only thing that
knows which humans and agents exist and which nodes they may reach. Node to
node SSH stays out -- the host has no inbound control surface by decision, and
nodes reaching each other that way is a second control path through the back
door.

0007 gains the requirement underneath all of it. Reachability was recorded as a
fact to track and never as a thing some node must have. The broker's node and
the hub must be dialable by every node at a stable address, or nothing can join
and a disconnected node cannot return. A mesh entirely behind NAT cannot be
raised. That is a precondition and it belongs with the others.

The link staying on the underlay is also argued now rather than asserted. At
join time it is forced; afterwards it is a choice, and the reason is that a
repair channel carried over the thing being repaired is not one. Moving it onto
the overlay, with fallback, is recorded as open with what it would have to get
right -- a WireGuard interface has no link state to test, and a silent fallback
is this repository's recurring fault in a new place.

0010 says in one line what was the intention throughout: the module system is
the CI/CD. Not a pipeline beside the mesh. Build, test, publish and deploy are
one reconciliation seen at four points, which is why a thing that cannot be a
module cannot be delivered.
2026-08-29 13:01:36 +02:00

14 KiB

topic, status, date, deciders, reconstructed
topic status date deciders reconstructed
the tiers accepted 2026-08-28 jochen false

6. The substrate and the control plane

Consolidated 2026-08-28 from six records. Extended 2026-08-29, by building it: the language, and what must be running before the control plane starts — which this record had left not established and could not have settled the way it was asking.

The control plane is what needs to know about more than one node

That is the whole test, and it follows from the host applying rather than deciding: deciding needs knowledge a single machine does not have.

question whose
write this file, with this content, with this mode the host
which nodes should run the store the control plane
is this unit running the host
which peers belong in this node's overlay the control plane
has this node been unreachable for a week the control plane — nobody else is watching

Anything a single machine could answer alone is not the control plane's.

Seven contexts and one interface

inventory, config, connectivity, provisioning, delivery, observability, identity — plus api, the one interface every surface speaks to. Each earns its place by the test above rather than by being ours.

work, knowledge and stream are mesh-hosted applications, not control plane. A task does not need to know a node exists. Being ours does not make something infrastructure.

identity owns SSH access. Written 2026-08-29, on noticing it was assumed everywhere and stated nowhere. SSH appears three times across this design and every time as something that uses the overlay — "the way back in", "every node reaches every other: SSH, services, ordinary traffic" — while nothing said who hands out the keys. Nobody else could: the mesh is the only thing that knows which humans and agents exist and which nodes they may reach, which is identity's definition. The node end already works, since an authorized_keys file is a file.

It is three questions wearing one name, and only two of them are the mesh's:

humans their key, on the nodes they are allowed on
agents the same, with a lifetime — and revocation that has to actually work
node to node not a mesh function. The host has no inbound control surface by decision (ADR 0004); nodes SSHing to each other would be a second control path arriving through the back door

Where the record lives is deliberately open. Contexts integrate through it, which makes it load-bearing, and putting it in the substrate risks recreating the circularity the tiers just removed. Listing it as an eighth context would settle by naming what has not been settled by arguing.

One node runs it, and nothing takes over

Declared, never elected. No promotion, no quorum, no fencing, no split brain — none of it built, so none of it can be subtly wrong.

The option that would make it a real mesh, and why not

Written 2026-08-29. It had been rejected by never being written down, which is the weakest way to reject anything.

A genuine peer-to-peer mesh means no node is special, and that has a concrete price:

  • every node holds the whole inventory, so there is a replication process between them;
  • replication needs a writer, so one node is elected master, and something promotes a new one when it drops — Redis Sentinel and its whole family of problems;
  • and it still would not deliver what the name promises, because application databases are not replicated. A workload's store lives where the workload lives.

That last point is the one that settles it. To make the mesh genuinely peer-to-peer we would have to become a replicated database system for everything running on it — not for our own inventory, for every consumer's data too. That is a product, and a much larger one than the thing it would be supporting.

So there are three central roles, not one, and it is worth seeing them separately because only the third costs operation:

its loss costs
the control plane nothing can be changed. Nothing stops running
the broker nothing can be told anything, or report anything
the hub nodes in different places cannot reach each other (ADR 0007)

Whether these are one node is not decided here. All three must be dialable by every node, which pushes toward one; nothing says they must be.

That is sound rather than merely cheap, because the design already tolerates its absence by construction: a node reconciles from its own store and never needed to ask anybody to hold the state it was last given. The control plane being down is not a new failure mode — it is every node in the ordinary disconnected situation at once. What is lost is change, not operation.

The honest half: this node is a single point of failure, recovery is restore rather than failover — which makes backup the availability mechanism rather than hygiene — and certificate renewal is the clock. An outage outlasting a renewal window expires every public name, which turns an inconvenience into an outage on a timer. Nothing measures that today.

The authority is the control plane, not a database

There is no single mesh database. Each context owns its store exclusively, and the mesh database names a thing that will not exist.

No node reads any of them — not for writes, not for reads. A node is told what to own, over the link, in a bounded vocabulary; it states what it applied, and the owning context writes. The difference is the security boundary: something that can write cannot be prevented from writing anything.

A node runs from its own store always, not as a fallback. The current arrangement's nastiest property is that a node running from cache looks identical to a node running from the database, with no age on the cache and nothing reporting divergence. Under this there is no second mode to be mistaken for the first.

What survives from the original decision: the repository defines what exists, the mesh defines what runs where, and no node-to-module mapping is ever committed. That is what makes the repositories node-agnostic and why anything about the mesh can be published at all.

The error underneath was a category error: source of truth named a storage location when it meant an authority. Once the store is the answer, which database becomes the question, and shared schemas follow.

The substrate is what the control plane consumes and cannot grant itself

Every module needing a database asks provisioning for one. The control plane needs a database too and cannot ask itself, because it is not running yet. That circularity is the definition, and anything on the wrong side of it is raised from the bundle the host carries.

role product
relational store PostgreSQL its own state lives there
message bus LavinMQ it cannot grant itself a virtual host — and precedes it, below
object store MinIO it cannot grant itself a bucket
image registry an OCI registry it cannot grant itself a repository
identity provider — conditional: substrate only if the control plane delegates authentication, which is undecided

The role and the product are both written. The role is what the argument turns on; the product is what gets installed and pinned, and a design that names only the role does not record that the choice was made. The dependency is on the protocol — AMQP, S3, OCI — which is what keeps naming them safe. The store is the exception: the provisioning model uses databases, roles and schemas as PostgreSQL means them.

A container runtime is detected, not chosen — docker or podman, because a machine that already has one keeps it. Only the version probe differs between them; the behavioural difference (podman has no daemon, so containers do not return after a reboot unless a unit is enabled) belongs in the declaration rather than the host.

Being substrate and being in the bundle are different questions. PostgreSQL and LavinMQ must precede the control plane; the object store and the registry are substrate by role and ordinary by delivery, provisioned once there is a control plane to do it.

Why the broker precedes it too

Written after the fact, because this record first left it not established and framed it as turning on whether the control plane's own contexts talk to each other over the bus.

They do not — they are one process and dispatch internally. Under that framing the broker is provisioned like anything else and the bundle stays at one image.

The framing cannot answer the question. What decides it is not how the contexts reach each other. It is how the control plane reaches a node — and that is settled above: never except over the link, and the link is AMQP (ADR 0002, ADR 0004). So:

the bundle raises the control plane
the control plane provisions the broker      ← by telling a host to run it
telling a host happens over the link
the link is the broker

And it is not avoided by the first node being local. enrol dials the broker at the address in its token, which is the first node's own third step. A machine that raised the mesh still joins it the ordinary way, and that was deliberate — its specialness lasts two commands. Making it join by some other route would buy a smaller bundle by giving up the property the design was built to have.

The broker precedes the control plane for the same reason PostgreSQL does: the control plane cannot grant itself the thing it would need in order to grant it.

What it costs. Two images rather than one, against the wish above to keep the bundle at roughly one so a person can read it — two is still readable, four would not be. And three things that are not images, each an action the bundle declares and the host runs, the way the database already is: a virtual host, a credential on it, and a certificate. That last is the awkward one: a token pins the fingerprint a host must expect before it sends anything, so the broker needs a certificate at a moment when there is no mesh to issue one and no public name to obtain one for. Self-signed and pinned is the shape that fits; how it is later replaced by the certificates in ADR 0007 is not decided here.

The control plane is written in Go

The same language as the host, so tiers 0 and 2 are one language and not two.

The reason that decides it is not familiarity. Its image is pinned by digest in the bundle, which means it is fetched and run on a machine where no mesh exists yet — nothing to check it against, nothing watching, and a person expected to have read the bundle and believed it. A statically linked binary makes that image the program and nothing else: no interpreter, no package tree, no transitive dependency that arrived because something needed a date library. Everything under that line is something somebody would have to audit, on the one image the whole mesh is raised from.

A second reason, smaller and still real: the control plane runs a reconcile loop of its own (ADR 0010 — artifacts against source, as the host reconciles machine state against declarations). Two loops of the same shape are cheaper to hold in one head when they are also the same language.

The option rejected is TypeScript, matching the lab and the surfaces that will speak to this. The argument for it is that tier 3 is web and CLI, so a TypeScript control plane would share types with its callers rather than generating a contract. True, and it does not reach far enough: mesh-sdk is contracts shared across tiers and tier 0 is Go, so the contracts cross a language boundary whatever tier 2 is written in. The choice is between generating them for one consumer or for two.

What it costs, plainly: the control plane can import nothing that exists today, and a person moving between tier 2 and tier 3 changes language. Neither is recovered later — the language is the most expensive thing in this record to reverse.

The installer fetches what it pins

substrate.lock carries references, not payload — an image name and a digest, fetched at apply time. A tag moves; a digest does not, and reproducibility comes from pinning the identity of a thing rather than carrying its bytes.

The assumption that a machine might have no network came from the lab and was wrong: a machine being adopted has one, and the sealed case is the lab.

The lab places images by raising a registry inside the scenario, which is what a real node pulls from anyway — so it tests the real path rather than a stand-in for it. The digests that registry serves are its own, and that satisfies this rule: what is required is a reference that is exact and cannot move, and one it assigned is both. Assuming an upstream digest had to be preserved is what made this look impossible for a while (04-ISSUES/009).

Its contents are per operating system even though its mechanism is not — package names, unit names and service names all differ, so an Arch host embeds an Arch bundle.

Consequences

  • The bundle stays small and reviewable. A list of pinned references is something a person can read; a bundle containing images is not.
  • An apply can fail because something is unreachable, which a self-contained artifact could not. That must fail legibly, naming what could not be fetched and from where.
  • Cross-context reporting is harder, and that is the point. Anything wanting to see across contexts consumes their events or calls their interfaces.
  • A queue with no limit grows until the broker's disk is full, and the broker is what every node depends on. The bound is per queue and is not decided.
  • The bundle carries two images and four actions, and the substrate bootstrap grows a step.
  • Nothing in the first node's path is special-cased. Enrolment is walked on node one.
  • The broker's certificate at bootstrap has no answer yet, and is named as unfinished rather than assumed. It is the first thing that will be wanted when the link is built.
  • The language cannot be revisited cheaply. It is the one line here close to irreversible.