Three additions, all written by trying to write a real database module and finding out what could not be said. A module may mirror an image it did not write. Naming an upstream reference directly needs every machine to reach a public registry and pins to a tag somebody else can move. A module may need a secret of its own — a superuser password is not FOR anybody, so the mechanism that hands credentials to consumers cannot express it. Per node, so three machines have three passwords. And the provisioner watches, which is what lets it be a module rather than a binary somebody places. It polls rather than watching the filesystem, because the host writes atomically and a watch on a replaced path silently stops working. One assignment now gets a working database provider: two directories, two pinned containers, a sealed password and the grants manifest.
11 KiB
layer, status, code, updated, decisions
| layer | status | code | updated | decisions | |||||
|---|---|---|---|---|---|---|---|---|---|
| to-be | designed |
|
2026-08-30 |
|
A module repository, and what builds it
Designed from what the mesh needs, not from what came before. The system this replaces has a concept of features — several independently-deployable units inside one module — and it is deliberately absent here.
Features are unnecessary, and that closes an open prerequisite
ADR 0001 lists named features with per-node opt-in as a prerequisite, on the grounds that without it "every independently deployable unit inside a context becomes a module again and the count returns."
The premise was right and the remedy already exists in another form. What features were for is three things the mesh now does separately:
| features did | what does it here |
|---|---|
| several deployable units in one thing | several modules, which is what they are |
| turning one on for one node | assignment, which is per node already |
| keeping related things together | requires, and a module with requirements and no files of its own |
networking is exactly that last row: it ships nothing, requires a private network and name
resolution, and assigning it brings both. So the module count does not return, because the thing
that made it return — a module is expensive, so put several things in one — is gone. A module
here is cheap: a manifest and, usually, nothing else.
One file at the root
module.json, and a convention somebody can look for beats a setting somebody has to find. It
says what the module is, what it provides and requires, what it claims, what capabilities it
needs, what it puts on a machine — and, if anything must be produced from the source, what to
build.
The manifest in the repository is not the manifest the mesh holds
A resource names an artifact:
{"id": "dotfiles", "type": "archive", "artifact": "config", "path": "…"}
and the built manifest names the thing:
{"id": "dotfiles", "type": "archive", "source": "…/blobs/sha256:…", "digest": "sha256:…"}
Two documents on purpose. A digest is not knowable until something is built, so a repository carrying one is a repository whose file is wrong the moment anybody edits anything — and the mesh would be pinning a value nobody could have checked. The built manifest is derived, and the record of which commit it was derived from is what makes "is this current?" answerable without building it again.
The word artifact never reaches a machine. The host's decoder is strict and would refuse it, at
the worst possible moment.
The builder runs on a node
Not in the control plane, and this is the same boundary as everywhere else. Building needs a container runtime and a working tree; what the control plane may send a machine is bounded by the declaration language (ADR 0005), and run this build is not in it. The alternative — the control plane holding a container socket — would make it the one component that can do anything on any machine, which is the property the whole design is arranged to avoid.
So the builder is a program a machine runs, given work over the broker like anything else, holding its own credential and nothing more.
A build is work, not state, and that is why it does not travel as a declaration. Everything else the control plane sends a node is what you should be, reconciled forever. A build happens once and is finished; as a declaration it would either rebuild on every reconcile or carry "and I already did this" — state about an event rather than about a machine.
So it has its own queue, and the answer comes back correlated. One queue, so several build machines share the work and each request is done exactly once, which a routing key per machine would not give.
A build machine has its own credential, and it is not a node's. It may read the build queue and write to the mesh exchange, and that is all — a node's queue carries that node's declarations, and a build machine has no business reading them.
The answer goes through the exchange, never the default one. Permission on the default exchange is granted per exchange, not per queue, so anything allowed to use it can publish into any node's queue. That is the privilege a build machine most obviously should not have. So an asker binds its own reply queue to the same routing key and filters by correlation; every asker sees every result, which is the price of the builder never needing that permission.
Three properties of the builder that are decisions:
- a request is acknowledged only once the answer is away. A builder that dies mid-build then leaves the work for another machine rather than losing it with nobody ever hearing why
- one build at a time. Five at once against one runtime finishes all five slower than it would have finished the first, and the queue is what shares work between machines
- a failure is a result. A build that fails silently is indistinguishable from a builder that is not running, and those want completely different responses — the same rule the host follows about a service that does not exist
What is kept
Every result, including the failures. A failed build that leaves no trace is indistinguishable from one nobody asked for, and the difference is the whole of whether somebody should be looking at something. A build that failed before it knew what it was building keeps the repository, which is what a person goes and looks at.
Recording is idempotent on the correlation, because a result arrives twice — once as the answer to whoever asked and once on the exchange, where the control plane is also listening. Two rows would show one build as two, and which is real is not answerable afterwards.
That is what a builds view reads, and until it existed there was nothing to read: a result was answered to the asker and kept nowhere.
Three properties that are decisions
- A fresh clone every time. A build reusing a working tree can succeed because of something a previous build left behind, and that is a build nobody can reproduce.
- Archives are packed deterministically — sorted, and carrying no timestamps, ownership or original paths. Two builds of one commit must produce one digest, or nothing downstream can tell this changed from this was built again, and every rebuild looks like a change to every machine holding it.
- Nothing is published until everything is built. Half a module in the store, under a digest the mesh never records, is reachable, unreferenced, and indistinguishable from something in use.
What a module may build, and what it may only borrow
| kind | is |
|---|---|
| image | built from a Dockerfile in this repository |
| archive | a directory in this repository, packed |
| upstream | an image somebody else built, mirrored into the mesh's own registry |
The third exists because a module usually runs software it did not write. A database module ships configuration and a provisioner and does not build a database. Naming the upstream reference directly would need every machine to reach a public registry, and would pin to a tag its owner can move — which is what pinning exists to prevent (ADR 0006). Mirroring is what the bootstrap already does by hand; this makes it something a module can say.
An upstream reference with no tag or digest is refused: what gets mirrored would be whatever
latest means today, and a module pinned to that is not pinned.
A module's own secret
A database has a superuser password, a broker an administrator, a registry an account. None of them is for anybody — they are not the credential a consumer is given, and the mechanism that hands those out has a consumer in the middle of it.
So a module says what it needs and where to put it, and the mesh generates one per node, seals it to that machine and reads it no more than it reads any other secret. Per node deliberately: a module running on three machines has three passwords, where one in the manifest would put the same secret on every machine that ever runs it, in a file anybody can read, for ever.
Made once and kept, or a running database would be handed a password it was not started with. Remade when the machine's sealing key changes. Declared and not made is refused, because a module whose own credential is silently absent starts, fails to authenticate, and the reason is three layers from the machine reporting it.
What one assignment gets you
A database module, written to see whether it could be:
directory /var/lib/mesh/postgres
directory /var/lib/mesh/postgres/grants
container the database pinned by digest, mirrored
container the provisioner pinned by digest, mirrored
file the superuser password sealed to this machine
file what its consumers asked for
The provisioner watches rather than being invoked. That is what lets it be a module: run once, it needs something to run it after every declaration — a timer, or a unit wired to a file. Watching, it is an ordinary long-running service the host already supervises. It polls rather than watching the filesystem, because the host writes atomically: the file is replaced, so a watch on the path stops seeing anything after the first replacement, and a watcher that silently stops working is worse than a poll.
Writing it found one thing wrong, and it was the manifest rather than the host: a container
declared restart-on, which is a service field, and the host refused it by name. It is right
to. A container whose own definition changes is recreated, and a file it mounts is read by the
process inside, which is that image's business.
Where artifacts go
The registry the bootstrap already pulls from, for both images and archives. An OCI registry is a content-addressed blob store that also understands images, and an archive is a content-addressed blob.
An object store beside it is the right answer for objects that are mutable, need per-reader access, or are not build output. None of that describes a digest-pinned archive, and running a second service for one kind of immutable blob is two things to run, two to back up, and two ways for an artifact to be missing. Overturnable without touching anything else: a manifest carries a URL and a digest, and neither says what served it.