Files
mesh-controller/README.md
T
jschoubben 306c4ca13b The control plane, as far as identity
Tier 2 exists now. It holds one context of seven, inventory, and does one
thing with it: brings its schema up to date. That is step 3 of the substrate
bootstrap -- the step the first node cannot get past.

Verified against a real PostgreSQL, with the built binary: applied 0001-nodes,
reported 'already up to date' on the second run, and the node table is there
with the index and the unique constraint the migration asks for.

Written in Go, and the image is FROM scratch holding one file. Confirmed by
unpacking it. That is the whole argument of ADR 0024: the bundle pins this
image by digest and runs it where nothing can check it, so everything in it is
something a person has to audit before trusting a first node.

Exclusive store ownership is built as a rule about credentials rather than
about intentions. There is no mesh-wide connection setting and no way to ask
for one -- a context reads MESH_STORE_<ITS OWN NAME> and holds nothing else, so
reaching another context's store needs a new variable, which is visible in the
declaration that runs it.

The migration runner is mostly refusals: an edited migration that already ran,
a migration numbered below one that has run, duplicate numbers, misnamed files,
empty files. All stop rather than warn, because at the moment any of them is
true nobody knows what the database holds.

It stops before identity, deliberately. What a node presents to prove who it is
has not been decided anywhere, and a migration is the most expensive place in
this system to guess.

Two tests did not defend what they claimed, and both are fixed rather than
removed. One asked only whether Open returned an error, which it did either way
-- a bad context name and a missing credential both fail, so deleting the name
check changed nothing. The other claimed to prove the migration runs in a
transaction, but PostgreSQL already wraps a multi-statement query in one of its
own, so it passed with the transaction taken out. What the transaction actually
buys is that the schema change and the row recording it commit together, and
there is now a test for that which fails when they are split.
2026-08-29 02:44:09 +02:00

5.9 KiB

mesh-control

Tier 2 of Novox Mesh — the control plane. Everything that needs to know about more than one node.

That is the whole test, and it draws the line the host cannot: the host applies and does not decide, because deciding needs knowledge one machine does not have. Which nodes should run the store, which peers belong in an overlay, whether a node has been unreachable for a week — nobody on a single machine can answer any of them.

The reasoning lives in novox/hq; this repository carries no argument that is not settled there.

What it is not

  • Not the thing that changes machines. It decides; the host applies. It never reaches into a node except through the host, over the link, in a bounded vocabulary.
  • Not a database. There is no mesh database. Each context owns its own store and nothing outside a context touches it — including nodes, which hold no credential to any of them.
  • Not privileged. It has no more access to a machine than a declaration can express.

What exists today

One context of seven, and one of the things it will do.

inventory the node records — the schema exists
config, connectivity, provisioning, delivery, observability, identity not built
the interface every surface speaks to not built; its shape is not decided
mesh-control migrate     bring each context's schema up to date
mesh-control version     what this binary is

migrate is step 3 of the substrate bootstrap — the step the first node cannot get past, run against a database raised moments earlier from the bundle the host carries.

Where this stops, and why there

At identity. A node's own identity is the next thing needed and its cryptographic form is not decided anywhere: whether a node holds a keypair whose public half the mesh keeps, or something else. Modelling it would have meant guessing, in a migration — which is the most expensive place in this system to guess, because a schema that ran is finished and the only way back is another migration.

So the node table holds what a node record is — a name, when it was made, what the machine last reported about itself, when it was last heard from — and stops before what a node presents.

Reaching a store

A context is granted only what it exclusively owns. No shared writes, no read-only role on another context's store, and no connection string that reaches more than one.

That is a rule about credentials, so it is built as one. There is no mesh-wide connection setting and no way to ask for one:

MESH_STORE_INVENTORY=postgres://…/inventory

A process granted inventory holds that variable and no other. Reaching another context's store is not a matter of restraint — it has no address for it and no credential to present. And it is how the rule is checked: what a context can reach is visible in the declaration that runs it, as the list of variables it was given.

Each context's database is named after the context. There is deliberately no database named for the mesh as a whole.

Migrations

Numbered, embedded in the binary, applied in order, each in a transaction with the row recording it. The applying is four lines; the rest is refusals, and the refusals are the point:

it stops when because
a migration that ran has since been edited the database holds the old version, the repository holds the new one, and nothing holds the difference
a migration is numbered below one that already ran usually two branches taking the same next number — applying it now runs the schema in an order nobody tested
two migrations share a number order is the entire guarantee, and two files with one number have none
a file in the directory is not a valid migration name a misnamed migration would otherwise never run and nothing would say so
a migration is empty it records that something happened and changes nothing, which cannot be told from a mistake

All of them stop rather than warn. At the moment any of them is true, nobody knows what the database contains, and there is no correct guess about a schema.

Running it again does nothing. Two copies running at once take an advisory lock, so a restart during a slow migration does not become two runners racing.

Building

make build      the binary
make image      the container image
make check      gofmt, vet, and every test against a real PostgreSQL

make check raises a throwaway PostgreSQL in a container and takes it down afterwards, including when the tests fail. Without one the tests that need a database skip and say so rather than passing quietly — make test is the honest subset, not the gate.

Tests run against a real database rather than a fake because what is being tested is the database's behaviour: that DDL is transactional, that an advisory lock serialises, that a checksum mismatch is caught against a record PostgreSQL actually kept. A fake would assert that the fake behaves as expected.

Every test here has been confirmed to fail when the behaviour it defends is removed. Two did not, when first written, and both are now commented with what they were missing.

The image

FROM scratch, holding one statically linked binary and nothing else — no shell, no package manager, no libc, no CA certificates.

Not a size optimisation. The bundle a host carries pins this image by digest, and it is fetched and run on a machine where no mesh exists to check anything and a person is expected to have read the bundle and believed it. Everything in the image is something that person would have to audit.

Where the reasoning lives

what the control plane is novox/hq ADR 0006
what it takes to run one, and why Go novox/hq ADR 0024
a context owns its store, exclusively novox/hq ADR 0008
schema changes are numbered migrations novox/hq ADR 0013
a test defends a decision novox/hq ADR 0017