Every remaining cluster merged. Each was one design that had been split across
several records because it was worked out over days rather than at once.
the node host 8 -> 1 applies not decides, depends on nothing,
per operating system, root service, the
launcher, episodic, what a declaration is,
actions from the bundle only
a node and how it joins 4 -> 1 what a node is, joining, the link as
security boundary, the enrolment token
modules and the graph 7 -> 1 everything is a module, no domain modules,
three edges, provisioning, the core library
substrate and control 6 -> 1 the test, seven contexts, one control plane,
plane the authority is not a database, the named
products, the pinned bundle
connectivity 3 -> 1 a route is a grant, reachability declared,
filter rules
delivery 5 -> 1 reconciliation not a pipeline, artifacts,
the three silos, a failed step, the verdict
the lab 5 -> 1 (earlier)
how this repository 10 -> 1 (earlier)
works
Nothing was dropped. Each consolidated record carries the reasoning of the ones
it absorbs -- the measurements, the incidents, the alternatives rejected --
because that reasoning is the only reason to keep a record at all. What is gone
is the fragmentation: eight files to read to understand tier 0, when tier 0 is
one component.
The four superseded records went too. They existed to point at their
successors, and the successors now contain what they said.
The checker made this safe. Each merge left dangling links -- 38 files after
the host merge alone -- and it named every one. Nothing was found by reading,
and a manual pass would certainly have missed some, including references inside
AGENTS.md which every session loads.
6.8 KiB
status, date, deciders, reconstructed, consolidates
| status | date | deciders | reconstructed | consolidates | |||||
|---|---|---|---|---|---|---|---|---|---|
| accepted | 2026-08-28 | jochen | false |
|
48. The substrate and the control plane
Consolidated 2026-08-28 from six records.
The control plane is what needs to know about more than one node
That is the whole test, and it follows from the host applying rather than deciding: deciding needs knowledge a single machine does not have.
| question | whose |
|---|---|
| write this file, with this content, with this mode | the host |
| which nodes should run the store | the control plane |
| is this unit running | the host |
| which peers belong in this node's overlay | the control plane |
| has this node been unreachable for a week | the control plane — nobody else is watching |
Anything a single machine could answer alone is not the control plane's.
Seven contexts and one interface
inventory, config, connectivity, provisioning, delivery, observability, identity — plus
api, the one interface every surface speaks to. Each earns its place by the test above rather
than by being ours.
work, knowledge and stream are mesh-hosted applications, not control plane. A task does
not need to know a node exists. Being ours does not make something infrastructure.
Where the record lives is deliberately open. Contexts integrate through it, which makes it load-bearing, and putting it in the substrate risks recreating the circularity the tiers just removed. Listing it as an eighth context would settle by naming what has not been settled by arguing.
One node runs it, and nothing takes over
Declared, never elected. No promotion, no quorum, no fencing, no split brain — none of it built, so none of it can be subtly wrong.
That is sound rather than merely cheap, because the design already tolerates its absence by construction: a node reconciles from its own store and never needed to ask anybody to hold the state it was last given. The control plane being down is not a new failure mode — it is every node in the ordinary disconnected situation at once. What is lost is change, not operation.
The honest half: this node is a single point of failure, recovery is restore rather than failover — which makes backup the availability mechanism rather than hygiene — and certificate renewal is the clock. An outage outlasting a renewal window expires every public name, which turns an inconvenience into an outage on a timer. Nothing measures that today.
The authority is the control plane, not a database
There is no single mesh database. Each context owns its store exclusively, and the mesh database names a thing that will not exist.
No node reads any of them — not for writes, not for reads. A node is told what to own, over the link, in a bounded vocabulary; it states what it applied, and the owning context writes. The difference is the security boundary: something that can write cannot be prevented from writing anything.
A node runs from its own store always, not as a fallback. The current arrangement's nastiest property is that a node running from cache looks identical to a node running from the database, with no age on the cache and nothing reporting divergence. Under this there is no second mode to be mistaken for the first.
What survives from the original decision: the repository defines what exists, the mesh defines what runs where, and no node-to-module mapping is ever committed. That is what makes the repositories node-agnostic and why anything about the mesh can be published at all.
The error underneath was a category error: source of truth named a storage location when it meant an authority. Once the store is the answer, which database becomes the question, and shared schemas follow.
The substrate is what the control plane consumes and cannot grant itself
Every module needing a database asks provisioning for one. The control plane needs a database too and cannot ask itself, because it is not running yet. That circularity is the definition, and anything on the wrong side of it is raised from the bundle the host carries.
| role | product | |
|---|---|---|
| relational store | PostgreSQL | its own state lives there |
| message bus | LavinMQ | it cannot grant itself a virtual host |
| object store | MinIO | it cannot grant itself a bucket |
| image registry | an OCI registry | it cannot grant itself a repository |
| identity provider | — | conditional: substrate only if the control plane delegates authentication, which is undecided |
The role and the product are both written. The role is what the argument turns on; the product is what gets installed and pinned, and a design that names only the role does not record that the choice was made. The dependency is on the protocol — AMQP, S3, OCI — which is what keeps naming them safe. The store is the exception: the provisioning model uses databases, roles and schemas as PostgreSQL means them.
A container runtime is detected, not chosen — docker or podman, because a machine that already has one keeps it. Only the version probe differs between them; the behavioural difference (podman has no daemon, so containers do not return after a reboot unless a unit is enabled) belongs in the declaration rather than the host.
Being substrate and being in the bundle are different questions. Only PostgreSQL must precede the control plane; the rest are substrate by role and ordinary by delivery, provisioned once there is a control plane to do it.
The installer fetches what it pins
substrate.lock carries references, not payload — an image name and a digest, fetched at
apply time. A tag moves; a digest does not, and reproducibility comes from pinning the identity of
a thing rather than carrying its bytes.
The assumption that a machine might have no network came from the lab and was wrong: a machine being adopted has one, and the sealed case is the lab, which places images itself.
Its contents are per operating system even though its mechanism is not — package names, unit names and service names all differ, so an Arch host embeds an Arch bundle.
Consequences
- The bundle stays small and reviewable. A list of pinned references is something a person can read; a bundle containing images is not.
- An apply can fail because something is unreachable, which a self-contained artifact could not. That must fail legibly, naming what could not be fetched and from where.
- Cross-context reporting is harder, and that is the point. Anything wanting to see across contexts consumes their events or calls their interfaces.
- A queue with no limit grows until the broker's disk is full, and the broker is what every node depends on. The bound is per queue and is not decided.