Files
hq/02-DECISIONS/0006-the-substrate-and-the-control-plane.md
jschoubben 88ba81e9c1 Agents reaching nodes is the capability, not a hole in it
Correcting what I wrote an hour ago. I had recorded node-to-node SSH as "not a
mesh function" and "a second control path through the back door", reasoning
from ADR 0004's rule that the host has no inbound control surface. That
conflated two different things and got the product backwards.

There is no node-to-node SSH to forbid. The actor is always an agent; a node is
only where it happens to be running -- ADR 0001 already says a node is a place
where an agent can run and that is the entire relationship. An agent hired onto
one node reaching another to do work is the capability the whole arrangement
exists to provide.

The credential is the agent's, in its own credential directory, which ADR 0001
already established. So a node's authorized_keys lists agents and never nodes,
and three things follow: no node holds a key reaching another node, so 0004's
"a node holds its own identity and nothing else" stays literally true; a
compromised node costs the credentials of the agents that were on it rather
than a way into everything; and who may reach what stays a mesh-wide fact,
which is why it is identity's.

The rule I misapplied is about how a node's declared state changes -- over the
broker, never by being dialled. An agent with a shell is not the mesh
reconfiguring a machine, it is what a person with a terminal has always been,
and this design already depends on that working: the overlay is the way back in
when a declaration breaks something. What such a session leaves behind is
drift, and drift is what reconciliation is for.

0001 also stops underselling the fourth layer. It read as "the layer the other
three exist to carry", which is true and flat. The value is that an agent can
work across a set of machines as though they were one -- centrally configurable
machines are ordinary; that is not.
2026-08-29 13:06:40 +02:00

16 KiB

topic, status, date, deciders, reconstructed
topic status date deciders reconstructed
the tiers accepted 2026-08-28 jochen false

6. The substrate and the control plane

Consolidated 2026-08-28 from six records. Extended 2026-08-29, by building it: the language, and what must be running before the control plane starts — which this record had left not established and could not have settled the way it was asking.

The control plane is what needs to know about more than one node

That is the whole test, and it follows from the host applying rather than deciding: deciding needs knowledge a single machine does not have.

question whose
write this file, with this content, with this mode the host
which nodes should run the store the control plane
is this unit running the host
which peers belong in this node's overlay the control plane
has this node been unreachable for a week the control plane — nobody else is watching

Anything a single machine could answer alone is not the control plane's.

Seven contexts and one interface

inventory, config, connectivity, provisioning, delivery, observability, identity — plus api, the one interface every surface speaks to. Each earns its place by the test above rather than by being ours.

work, knowledge and stream are mesh-hosted applications, not control plane. A task does not need to know a node exists. Being ours does not make something infrastructure.

identity owns SSH access. Written 2026-08-29, on noticing it was assumed everywhere and stated nowhere. SSH appears three times across this design and every time as something that uses the overlay — "the way back in", "every node reaches every other: SSH, services, ordinary traffic" — while nothing said who hands out the keys. Nobody else could: the mesh is the only thing that knows which humans and agents exist and which nodes they may reach, which is identity's definition. The node end already works, since an authorized_keys file is a file.

It is three questions wearing one name, and only two of them are the mesh's:

humans their key, on the nodes they are allowed on
agents the same, with a lifetime — and revocation that has to actually work
an agent reaching another node this is the point of the mesh, not an exception to it — see below

There is no such thing as node-to-node SSH here, and that is a clarification rather than a restriction. The actor is always an agent; a node is only where it happens to be running — ADR 0001: a node is a place where an agent can run, that is the entire relationship. An agent hired onto one node reaching another to do work is the capability the whole arrangement exists to provide.

The credential is the agent's, never the node's. It lives in the agent's own credential directory (ADR 0001), so a node's authorized_keys lists agents and never nodes. Three things follow, and they are why this shape is better rather than merely allowed:

  • ADR 0004's a node holds its own identity and nothing else stays true — no node holds a key that reaches another node;
  • a compromised node costs the credentials of the agents that were on it, not a way into everything;
  • who may reach what is a mesh-wide fact, which is exactly why it is identity's and not something arranged locally.

And it does not conflict with the host having no inbound control surface. That rule is about how a node's declared state changes: over the broker, never by being dialled. An agent with a shell is not the mesh reconfiguring a machine — it is what a person with a terminal has always been, and this design already depends on it working (ADR 0007: the overlay is the way back in). What such a session can leave behind is drift, and drift is what reconciliation is for.

Where the record lives is deliberately open. Contexts integrate through it, which makes it load-bearing, and putting it in the substrate risks recreating the circularity the tiers just removed. Listing it as an eighth context would settle by naming what has not been settled by arguing.

One node runs it, and nothing takes over

Declared, never elected. No promotion, no quorum, no fencing, no split brain — none of it built, so none of it can be subtly wrong.

The option that would make it a real mesh, and why not

Written 2026-08-29. It had been rejected by never being written down, which is the weakest way to reject anything.

A genuine peer-to-peer mesh means no node is special, and that has a concrete price:

  • every node holds the whole inventory, so there is a replication process between them;
  • replication needs a writer, so one node is elected master, and something promotes a new one when it drops — Redis Sentinel and its whole family of problems;
  • and it still would not deliver what the name promises, because application databases are not replicated. A workload's store lives where the workload lives.

That last point is the one that settles it. To make the mesh genuinely peer-to-peer we would have to become a replicated database system for everything running on it — not for our own inventory, for every consumer's data too. That is a product, and a much larger one than the thing it would be supporting.

So there are three central roles, not one, and it is worth seeing them separately because only the third costs operation:

its loss costs
the control plane nothing can be changed. Nothing stops running
the broker nothing can be told anything, or report anything
the hub nodes in different places cannot reach each other (ADR 0007)

Whether these are one node is not decided here. All three must be dialable by every node, which pushes toward one; nothing says they must be.

That is sound rather than merely cheap, because the design already tolerates its absence by construction: a node reconciles from its own store and never needed to ask anybody to hold the state it was last given. The control plane being down is not a new failure mode — it is every node in the ordinary disconnected situation at once. What is lost is change, not operation.

The honest half: this node is a single point of failure, recovery is restore rather than failover — which makes backup the availability mechanism rather than hygiene — and certificate renewal is the clock. An outage outlasting a renewal window expires every public name, which turns an inconvenience into an outage on a timer. Nothing measures that today.

The authority is the control plane, not a database

There is no single mesh database. Each context owns its store exclusively, and the mesh database names a thing that will not exist.

No node reads any of them — not for writes, not for reads. A node is told what to own, over the link, in a bounded vocabulary; it states what it applied, and the owning context writes. The difference is the security boundary: something that can write cannot be prevented from writing anything.

A node runs from its own store always, not as a fallback. The current arrangement's nastiest property is that a node running from cache looks identical to a node running from the database, with no age on the cache and nothing reporting divergence. Under this there is no second mode to be mistaken for the first.

What survives from the original decision: the repository defines what exists, the mesh defines what runs where, and no node-to-module mapping is ever committed. That is what makes the repositories node-agnostic and why anything about the mesh can be published at all.

The error underneath was a category error: source of truth named a storage location when it meant an authority. Once the store is the answer, which database becomes the question, and shared schemas follow.

The substrate is what the control plane consumes and cannot grant itself

Every module needing a database asks provisioning for one. The control plane needs a database too and cannot ask itself, because it is not running yet. That circularity is the definition, and anything on the wrong side of it is raised from the bundle the host carries.

role product
relational store PostgreSQL its own state lives there
message bus LavinMQ it cannot grant itself a virtual host — and precedes it, below
object store MinIO it cannot grant itself a bucket
image registry an OCI registry it cannot grant itself a repository
identity provider — conditional: substrate only if the control plane delegates authentication, which is undecided

The role and the product are both written. The role is what the argument turns on; the product is what gets installed and pinned, and a design that names only the role does not record that the choice was made. The dependency is on the protocol — AMQP, S3, OCI — which is what keeps naming them safe. The store is the exception: the provisioning model uses databases, roles and schemas as PostgreSQL means them.

A container runtime is detected, not chosen — docker or podman, because a machine that already has one keeps it. Only the version probe differs between them; the behavioural difference (podman has no daemon, so containers do not return after a reboot unless a unit is enabled) belongs in the declaration rather than the host.

Being substrate and being in the bundle are different questions. PostgreSQL and LavinMQ must precede the control plane; the object store and the registry are substrate by role and ordinary by delivery, provisioned once there is a control plane to do it.

Why the broker precedes it too

Written after the fact, because this record first left it not established and framed it as turning on whether the control plane's own contexts talk to each other over the bus.

They do not — they are one process and dispatch internally. Under that framing the broker is provisioned like anything else and the bundle stays at one image.

The framing cannot answer the question. What decides it is not how the contexts reach each other. It is how the control plane reaches a node — and that is settled above: never except over the link, and the link is AMQP (ADR 0002, ADR 0004). So:

the bundle raises the control plane
the control plane provisions the broker      ← by telling a host to run it
telling a host happens over the link
the link is the broker

And it is not avoided by the first node being local. enrol dials the broker at the address in its token, which is the first node's own third step. A machine that raised the mesh still joins it the ordinary way, and that was deliberate — its specialness lasts two commands. Making it join by some other route would buy a smaller bundle by giving up the property the design was built to have.

The broker precedes the control plane for the same reason PostgreSQL does: the control plane cannot grant itself the thing it would need in order to grant it.

What it costs. Two images rather than one, against the wish above to keep the bundle at roughly one so a person can read it — two is still readable, four would not be. And three things that are not images, each an action the bundle declares and the host runs, the way the database already is: a virtual host, a credential on it, and a certificate. That last is the awkward one: a token pins the fingerprint a host must expect before it sends anything, so the broker needs a certificate at a moment when there is no mesh to issue one and no public name to obtain one for. Self-signed and pinned is the shape that fits; how it is later replaced by the certificates in ADR 0007 is not decided here.

The control plane is written in Go

The same language as the host, so tiers 0 and 2 are one language and not two.

The reason that decides it is not familiarity. Its image is pinned by digest in the bundle, which means it is fetched and run on a machine where no mesh exists yet — nothing to check it against, nothing watching, and a person expected to have read the bundle and believed it. A statically linked binary makes that image the program and nothing else: no interpreter, no package tree, no transitive dependency that arrived because something needed a date library. Everything under that line is something somebody would have to audit, on the one image the whole mesh is raised from.

A second reason, smaller and still real: the control plane runs a reconcile loop of its own (ADR 0010 — artifacts against source, as the host reconciles machine state against declarations). Two loops of the same shape are cheaper to hold in one head when they are also the same language.

The option rejected is TypeScript, matching the lab and the surfaces that will speak to this. The argument for it is that tier 3 is web and CLI, so a TypeScript control plane would share types with its callers rather than generating a contract. True, and it does not reach far enough: mesh-sdk is contracts shared across tiers and tier 0 is Go, so the contracts cross a language boundary whatever tier 2 is written in. The choice is between generating them for one consumer or for two.

What it costs, plainly: the control plane can import nothing that exists today, and a person moving between tier 2 and tier 3 changes language. Neither is recovered later — the language is the most expensive thing in this record to reverse.

The installer fetches what it pins

substrate.lock carries references, not payload — an image name and a digest, fetched at apply time. A tag moves; a digest does not, and reproducibility comes from pinning the identity of a thing rather than carrying its bytes.

The assumption that a machine might have no network came from the lab and was wrong: a machine being adopted has one, and the sealed case is the lab.

The lab places images by raising a registry inside the scenario, which is what a real node pulls from anyway — so it tests the real path rather than a stand-in for it. The digests that registry serves are its own, and that satisfies this rule: what is required is a reference that is exact and cannot move, and one it assigned is both. Assuming an upstream digest had to be preserved is what made this look impossible for a while (04-ISSUES/009).

Its contents are per operating system even though its mechanism is not — package names, unit names and service names all differ, so an Arch host embeds an Arch bundle.

Consequences

  • The bundle stays small and reviewable. A list of pinned references is something a person can read; a bundle containing images is not.
  • An apply can fail because something is unreachable, which a self-contained artifact could not. That must fail legibly, naming what could not be fetched and from where.
  • Cross-context reporting is harder, and that is the point. Anything wanting to see across contexts consumes their events or calls their interfaces.
  • A queue with no limit grows until the broker's disk is full, and the broker is what every node depends on. The bound is per queue and is not decided.
  • The bundle carries two images and four actions, and the substrate bootstrap grows a step.
  • Nothing in the first node's path is special-cased. Enrolment is walked on node one.
  • The broker's certificate at bootstrap has no answer yet, and is named as unfinished rather than assumed. It is the first thing that will be wanted when the link is built.
  • The language cannot be revisited cheaply. It is the one line here close to irreversible.