Two additions to the module-repository design, both from building it. A build machine has its own credential and it is not a node's: read the build queue, write the mesh exchange, nothing else. A node's queue carries that node's declarations. The answer goes through the exchange and never the default one, because permission there is per exchange rather than per queue — anything allowed to use it can publish into any node's queue. The price is that every asker sees every result and filters by correlation, which is cheap against a builder never needing that permission. And every result is kept, failures included, because one that leaves no trace is indistinguishable from a build nobody asked for. That is what a builds view reads; the board page is corrected to say so.
147 lines
7.4 KiB
Markdown
147 lines
7.4 KiB
Markdown
---
|
|
layer: to-be
|
|
status: designed
|
|
code:
|
|
- mesh-control internal/builder
|
|
- mesh-control internal/catalogue/build.go
|
|
updated: 2026-08-30
|
|
decisions:
|
|
- 02-DECISIONS/0009-modules-and-the-graph.md
|
|
- 02-DECISIONS/0010-delivery.md
|
|
- 02-DECISIONS/0005-the-node-host.md
|
|
---
|
|
|
|
# A module repository, and what builds it
|
|
|
|
**Designed from what the mesh needs, not from what came before.** The system this replaces has a
|
|
concept of *features* — several independently-deployable units inside one module — and it is
|
|
deliberately absent here.
|
|
|
|
## Features are unnecessary, and that closes an open prerequisite
|
|
|
|
[ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) lists *named features
|
|
with per-node opt-in* as a prerequisite, on the grounds that without it "every independently
|
|
deployable unit inside a context becomes a module again and the count returns."
|
|
|
|
**The premise was right and the remedy already exists in another form.** What features were for is
|
|
three things the mesh now does separately:
|
|
|
|
| features did | what does it here |
|
|
|---|---|
|
|
| several deployable units in one thing | **several modules**, which is what they are |
|
|
| turning one on for one node | **assignment**, which is per node already |
|
|
| keeping related things together | **`requires`**, and a module with requirements and no files of its own |
|
|
|
|
`networking` is exactly that last row: it ships nothing, requires a private network and name
|
|
resolution, and assigning it brings both. So the module count does not return, because the thing
|
|
that made it return — *a module is expensive, so put several things in one* — is gone. A module
|
|
here is cheap: a manifest and, usually, nothing else.
|
|
|
|
## One file at the root
|
|
|
|
`module.json`, and a convention somebody can look for beats a setting somebody has to find. It
|
|
says what the module is, what it provides and requires, what it claims, what capabilities it
|
|
needs, what it puts on a machine — and, if anything must be produced from the source, what to
|
|
build.
|
|
|
|
## The manifest in the repository is not the manifest the mesh holds
|
|
|
|
A resource names an artifact:
|
|
|
|
```
|
|
{"id": "dotfiles", "type": "archive", "artifact": "config", "path": "…"}
|
|
```
|
|
|
|
and the built manifest names the thing:
|
|
|
|
```
|
|
{"id": "dotfiles", "type": "archive", "source": "…/blobs/sha256:…", "digest": "sha256:…"}
|
|
```
|
|
|
|
**Two documents on purpose.** A digest is not knowable until something is built, so a repository
|
|
carrying one is a repository whose file is wrong the moment anybody edits anything — and the mesh
|
|
would be pinning a value nobody could have checked. The built manifest is derived, and the record
|
|
of *which commit it was derived from* is what makes "is this current?" answerable without building
|
|
it again.
|
|
|
|
The word `artifact` never reaches a machine. The host's decoder is strict and would refuse it, at
|
|
the worst possible moment.
|
|
|
|
## The builder runs on a node
|
|
|
|
**Not in the control plane, and this is the same boundary as everywhere else.** Building needs a
|
|
container runtime and a working tree; what the control plane may send a machine is bounded by the
|
|
declaration language ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)), and *run this build*
|
|
is not in it. The alternative — the control plane holding a container socket — would make it the
|
|
one component that can do anything on any machine, which is the property the whole design is
|
|
arranged to avoid.
|
|
|
|
So the builder is a program a machine runs, given work over the broker like anything else, holding
|
|
its own credential and nothing more.
|
|
|
|
**A build is work, not state**, and that is why it does not travel as a declaration. Everything
|
|
else the control plane sends a node is *what you should be*, reconciled forever. A build happens
|
|
once and is finished; as a declaration it would either rebuild on every reconcile or carry "and I
|
|
already did this" — state about an event rather than about a machine.
|
|
|
|
So it has its own queue, and the answer comes back correlated. **One queue**, so several build
|
|
machines share the work and each request is done exactly once, which a routing key per machine
|
|
would not give.
|
|
|
|
**A build machine has its own credential**, and it is not a node's. It may read the build queue
|
|
and write to the mesh exchange, and that is all — a node's queue carries that node's declarations,
|
|
and a build machine has no business reading them.
|
|
|
|
**The answer goes through the exchange, never the default one.** Permission on the default
|
|
exchange is granted per *exchange*, not per queue, so anything allowed to use it can publish into
|
|
any node's queue. That is the privilege a build machine most obviously should not have. So an
|
|
asker binds its own reply queue to the same routing key and filters by correlation; every asker
|
|
sees every result, which is the price of the builder never needing that permission.
|
|
|
|
Three properties of the builder that are decisions:
|
|
|
|
- **a request is acknowledged only once the answer is away.** A builder that dies mid-build then
|
|
leaves the work for another machine rather than losing it with nobody ever hearing why
|
|
- **one build at a time.** Five at once against one runtime finishes all five slower than it would
|
|
have finished the first, and the queue is what shares work between machines
|
|
- **a failure is a result.** A build that fails silently is indistinguishable from a builder that
|
|
is not running, and those want completely different responses — the same rule the host follows
|
|
about a service that does not exist
|
|
|
|
## What is kept
|
|
|
|
**Every result, including the failures.** A failed build that leaves no trace is indistinguishable
|
|
from one nobody asked for, and the difference is the whole of whether somebody should be looking
|
|
at something. A build that failed before it knew what it was building keeps the repository, which
|
|
is what a person goes and looks at.
|
|
|
|
Recording is idempotent on the correlation, because a result arrives twice — once as the answer to
|
|
whoever asked and once on the exchange, where the control plane is also listening. Two rows would
|
|
show one build as two, and which is real is not answerable afterwards.
|
|
|
|
That is what a builds view reads, and until it existed there was nothing to read: a result was
|
|
answered to the asker and kept nowhere.
|
|
|
|
## Three properties that are decisions
|
|
|
|
- **A fresh clone every time.** A build reusing a working tree can succeed because of something a
|
|
previous build left behind, and that is a build nobody can reproduce.
|
|
- **Archives are packed deterministically** — sorted, and carrying no timestamps, ownership or
|
|
original paths. Two builds of one commit must produce one digest, or nothing downstream can tell
|
|
*this changed* from *this was built again*, and every rebuild looks like a change to every
|
|
machine holding it.
|
|
- **Nothing is published until everything is built.** Half a module in the store, under a digest
|
|
the mesh never records, is reachable, unreferenced, and indistinguishable from something in use.
|
|
|
|
## Where artifacts go
|
|
|
|
**The registry the bootstrap already pulls from**, for both images and archives. An OCI registry
|
|
is a content-addressed blob store that also understands images, and an archive is a
|
|
content-addressed blob.
|
|
|
|
An object store beside it is the right answer for objects that are *mutable*, need per-reader
|
|
access, or are not build output. None of that describes a digest-pinned archive, and running a
|
|
second service for one kind of immutable blob is two things to run, two to back up, and two ways
|
|
for an artifact to be missing. **Overturnable without touching anything else**: a manifest carries
|
|
a URL and a digest, and neither says what served it.
|