Decided after measuring what renumbering actually costs: 96 references in code comments across two repositories, none of which would have failed to compile. They would have pointed at the wrong reasoning, which is worse than a broken link because nothing reports it. So a number identifies a record and never changes. It cannot also be a position -- a position moves when the set changes, and an identity that moves is not one. The reading order moves into an index generated from each record's `topic:`. Six topics, in the order somebody learns the system. The index is WRITTEN rather than only generated on demand, which reverses what this repository previously said. The reason it said otherwise is that a hand-written index drifts -- but a reader looking at the folder on a forge sees the folder, not a command, and the drift objection is answered by checking rather than by refusing to write one. That is §5's own rule: a rule states how it is checked. Two checks, both confirmed to bite. index.py fails when the written order no longer matches the records. records.py fails when a record has no topic or one nobody defined -- the quiet failure being a record that vanishes from the order rather than appearing in the wrong place.
6.7 KiB
topic, status, date, deciders, reconstructed
| topic | status | date | deciders | reconstructed |
|---|---|---|---|---|
| the tiers | accepted | 2026-08-28 | jochen | false |
6. The substrate and the control plane
Consolidated 2026-08-28 from six records.
The control plane is what needs to know about more than one node
That is the whole test, and it follows from the host applying rather than deciding: deciding needs knowledge a single machine does not have.
| question | whose |
|---|---|
| write this file, with this content, with this mode | the host |
| which nodes should run the store | the control plane |
| is this unit running | the host |
| which peers belong in this node's overlay | the control plane |
| has this node been unreachable for a week | the control plane — nobody else is watching |
Anything a single machine could answer alone is not the control plane's.
Seven contexts and one interface
inventory, config, connectivity, provisioning, delivery, observability, identity — plus
api, the one interface every surface speaks to. Each earns its place by the test above rather
than by being ours.
work, knowledge and stream are mesh-hosted applications, not control plane. A task does
not need to know a node exists. Being ours does not make something infrastructure.
Where the record lives is deliberately open. Contexts integrate through it, which makes it load-bearing, and putting it in the substrate risks recreating the circularity the tiers just removed. Listing it as an eighth context would settle by naming what has not been settled by arguing.
One node runs it, and nothing takes over
Declared, never elected. No promotion, no quorum, no fencing, no split brain — none of it built, so none of it can be subtly wrong.
That is sound rather than merely cheap, because the design already tolerates its absence by construction: a node reconciles from its own store and never needed to ask anybody to hold the state it was last given. The control plane being down is not a new failure mode — it is every node in the ordinary disconnected situation at once. What is lost is change, not operation.
The honest half: this node is a single point of failure, recovery is restore rather than failover — which makes backup the availability mechanism rather than hygiene — and certificate renewal is the clock. An outage outlasting a renewal window expires every public name, which turns an inconvenience into an outage on a timer. Nothing measures that today.
The authority is the control plane, not a database
There is no single mesh database. Each context owns its store exclusively, and the mesh database names a thing that will not exist.
No node reads any of them — not for writes, not for reads. A node is told what to own, over the link, in a bounded vocabulary; it states what it applied, and the owning context writes. The difference is the security boundary: something that can write cannot be prevented from writing anything.
A node runs from its own store always, not as a fallback. The current arrangement's nastiest property is that a node running from cache looks identical to a node running from the database, with no age on the cache and nothing reporting divergence. Under this there is no second mode to be mistaken for the first.
What survives from the original decision: the repository defines what exists, the mesh defines what runs where, and no node-to-module mapping is ever committed. That is what makes the repositories node-agnostic and why anything about the mesh can be published at all.
The error underneath was a category error: source of truth named a storage location when it meant an authority. Once the store is the answer, which database becomes the question, and shared schemas follow.
The substrate is what the control plane consumes and cannot grant itself
Every module needing a database asks provisioning for one. The control plane needs a database too and cannot ask itself, because it is not running yet. That circularity is the definition, and anything on the wrong side of it is raised from the bundle the host carries.
| role | product | |
|---|---|---|
| relational store | PostgreSQL | its own state lives there |
| message bus | LavinMQ | it cannot grant itself a virtual host |
| object store | MinIO | it cannot grant itself a bucket |
| image registry | an OCI registry | it cannot grant itself a repository |
| identity provider | — | conditional: substrate only if the control plane delegates authentication, which is undecided |
The role and the product are both written. The role is what the argument turns on; the product is what gets installed and pinned, and a design that names only the role does not record that the choice was made. The dependency is on the protocol — AMQP, S3, OCI — which is what keeps naming them safe. The store is the exception: the provisioning model uses databases, roles and schemas as PostgreSQL means them.
A container runtime is detected, not chosen — docker or podman, because a machine that already has one keeps it. Only the version probe differs between them; the behavioural difference (podman has no daemon, so containers do not return after a reboot unless a unit is enabled) belongs in the declaration rather than the host.
Being substrate and being in the bundle are different questions. Only PostgreSQL must precede the control plane; the rest are substrate by role and ordinary by delivery, provisioned once there is a control plane to do it.
The installer fetches what it pins
substrate.lock carries references, not payload — an image name and a digest, fetched at
apply time. A tag moves; a digest does not, and reproducibility comes from pinning the identity of
a thing rather than carrying its bytes.
The assumption that a machine might have no network came from the lab and was wrong: a machine being adopted has one, and the sealed case is the lab, which places images itself.
Its contents are per operating system even though its mechanism is not — package names, unit names and service names all differ, so an Arch host embeds an Arch bundle.
Consequences
- The bundle stays small and reviewable. A list of pinned references is something a person can read; a bundle containing images is not.
- An apply can fail because something is unreachable, which a self-contained artifact could not. That must fail legibly, naming what could not be fetched and from where.
- Cross-context reporting is harder, and that is the point. Anything wanting to see across contexts consumes their events or calls their interfaces.
- A queue with no limit grows until the broker's disk is full, and the broker is what every node depends on. The bound is per queue and is not decided.