Files
hq/03-DESIGN/01-to-be/12-a-module-repository.md
T
jschoubben a036bac47b Record why the module's own secret is named for whose it is
The field was called needs, beside secrets, and both were name-to-path holding
something secret. What separates them is whose, not how secret — so that is
what the name says now.
2026-08-31 13:47:21 +02:00

348 lines
19 KiB
Markdown

---
layer: to-be
status: designed
code:
- mesh-control internal/builder
- mesh-control internal/catalogue/build.go
- mesh-control internal/inventory/secrets.go
- mesh-control cmd/mesh-builder
updated: 2026-08-31
decisions:
- 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0010-delivery.md
- 02-DECISIONS/0005-the-node-host.md
---
# A module repository, and what builds it
**Designed from what the mesh needs, not from what came before.** The system this replaces has a
concept of *features* — several independently-deployable units inside one module — and it is
deliberately absent here.
## Features are unnecessary, and that closes an open prerequisite
[ADR 0001](../../02-DECISIONS/0001-mesh-brokers-nodes-host-agents-think.md) lists *named features
with per-node opt-in* as a prerequisite, on the grounds that without it "every independently
deployable unit inside a context becomes a module again and the count returns."
**The premise was right and the remedy already exists in another form.** What features were for is
three things the mesh now does separately:
| features did | what does it here |
|---|---|
| several deployable units in one thing | **several modules**, which is what they are |
| turning one on for one node | **assignment**, which is per node already |
| keeping related things together | **`requires`**, and a module with requirements and no files of its own |
`networking` is exactly that last row: it ships nothing, requires a private network and name
resolution, and assigning it brings both. So the module count does not return, because the thing
that made it return — *a module is expensive, so put several things in one* — is gone. A module
here is cheap: a manifest and, usually, nothing else.
## One file at the root
`module.json`, and a convention somebody can look for beats a setting somebody has to find. It
says what the module is, what it provides and requires, what it claims, what capabilities it
needs, what it puts on a machine — and, if anything must be produced from the source, what to
build.
## The manifest in the repository is not the manifest the mesh holds
A resource names an artifact:
```
{"id": "dotfiles", "type": "archive", "artifact": "config", "path": "…"}
```
and the built manifest names the thing:
```
{"id": "dotfiles", "type": "archive", "source": "…/blobs/sha256:…", "digest": "sha256:…"}
```
**Two documents on purpose.** A digest is not knowable until something is built, so a repository
carrying one is a repository whose file is wrong the moment anybody edits anything — and the mesh
would be pinning a value nobody could have checked. The built manifest is derived, and the record
of *which commit it was derived from* is what makes "is this current?" answerable without building
it again.
The word `artifact` never reaches a machine. The host's decoder is strict and would refuse it, at
the worst possible moment.
## The builder runs on a node
**Not in the control plane, and this is the same boundary as everywhere else.** Building needs a
container runtime and a working tree; what the control plane may send a machine is bounded by the
declaration language ([ADR 0005](../../02-DECISIONS/0005-the-node-host.md)), and *run this build*
is not in it. The alternative — the control plane holding a container socket — would make it the
one component that can do anything on any machine, which is the property the whole design is
arranged to avoid.
So the builder is a program a machine runs, given work over the broker like anything else, holding
its own credential and nothing more.
**A build is work, not state**, and that is why it does not travel as a declaration. Everything
else the control plane sends a node is *what you should be*, reconciled forever. A build happens
once and is finished; as a declaration it would either rebuild on every reconcile or carry "and I
already did this" — state about an event rather than about a machine.
So it has its own queue, and the answer comes back correlated. **One queue**, so several build
machines share the work and each request is done exactly once, which a routing key per machine
would not give.
**A build machine has its own credential**, and it is not a node's. It may read the build queue
and write to the mesh exchange, and that is all — a node's queue carries that node's declarations,
and a build machine has no business reading them.
**The answer goes through the exchange, never the default one.** Permission on the default
exchange is granted per *exchange*, not per queue, so anything allowed to use it can publish into
any node's queue. That is the privilege a build machine most obviously should not have. So an
asker binds its own reply queue to the same routing key and filters by correlation; every asker
sees every result, which is the price of the builder never needing that permission.
Three properties of the builder that are decisions:
- **a request is acknowledged only once the answer is away.** A builder that dies mid-build then
leaves the work for another machine rather than losing it with nobody ever hearing why
- **one build at a time.** Five at once against one runtime finishes all five slower than it would
have finished the first, and the queue is what shares work between machines
- **a failure is a result.** A build that fails silently is indistinguishable from a builder that
is not running, and those want completely different responses — the same rule the host follows
about a service that does not exist
### And it is a module the mesh assigns
*2026-08-31. Written after `builder issue --node`, which is the part that makes the sentence
"holding its own credential" true rather than aspirational.*
A build machine is a machine that runs the builder, and there is exactly one honest way to say
which machines those are: **assign it**. So the builder is a module like any other — an image, a
container, a working directory, and a claim so a machine does not end up running two.
The one thing that could not be a module in the ordinary way is the credential. It is not
generated, because the broker has to have been told about it, and it is not written in a manifest,
because a manifest is public and the same file goes to every machine that ever runs it. So the
mesh **creates the account, seals the URL to the machine that will use it, and discards the
plaintext** — the "given, not generated" case above, and its first user.
Nothing is printed. A credential shown on a terminal is a credential in a scrollback buffer, and
the copy that matters would then exist in two places, one of which nobody is guarding.
**What this replaces:** a builder started by hand with whatever credential was to hand, which in
practice meant the broker's administrative account. *A program documented as holding its own
credential and given somebody else's is worse than one with no story at all* — the documentation
is what stops anybody checking.
*Checked in the lab by assigning it and then asking the mesh to build a module: the credential
file arrives readable only by that machine, names the scoped account rather than the broker's own,
and the build completes — which is the only proof the credential authenticates, because a
container that is up holding a credential it cannot use looks identical from outside.*
**And it is told what to check the broker against**, not only who to connect as. A mesh's broker
presents a certificate of the mesh's own, which is in no public trust store, so a URL alone reaches
only a broker somebody else vouches for — which is no mesh broker at all. The credential carries
the fingerprint beside the URL: **the same two facts a node's token carries**
([ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md)), for the same reason, arriving by
a path other than the thing being trusted.
**A builder that is a module cannot see the machine's filesystem.** It runs in a container, so a
local path exists for the machine and not for it. That is not a limitation to work around — it is
the arrangement working: a build machine shares the runtime it was given rather than the machine
it sits on. **A module is cloned from the forge over a URL**, and "build this directory" is a
convenience for a builder somebody started by hand.
### And the loop is closed
*2026-08-31.* [ADR 0010](../../02-DECISIONS/0010-delivery.md) replaced a pipeline with a comparison
and named the risk: **losing the question "did my change go out?"**. The mesh could already answer
which modules were behind their source — and then a person read that list and retyped each
repository, which is a person being the loop, and the loop is the thing the pipeline was doing
before it was taken away.
`build --behind` is the other half, and it is the mirror of `push --behind`: the mesh knows what is
stale, so it builds it. The two forms are deliberately not combined — naming a repository and
asking which need building are different requests, and guessing which was meant would sometimes
build something nobody named.
**One failing does not stop the others**, for the same reason one broken module no longer blocks a
machine's whole declaration: a mesh where one bad repository holds back nine good ones is a mesh
where nobody dares add the tenth.
**Each is built from its own recorded ref**, not from the commit the mesh happened to notice.
Pinning to that would quietly turn a tracked branch into a pin — a change of meaning nobody asked
for, arrived at by an implementation detail.
**Building is not delivering, and the two stay separate.** A machine keeps running what it has
until it is told otherwise; the mesh changing its mind is not a machine acting on it, and
collapsing the two is how a mesh comes to report success for something that has not happened.
*Checked end to end: a commit, a build, a catalogue entry, and a machine that ends up running what
the source says — with both halves that make the answer trustworthy. It is still running the old
one until it is pushed, and it stops being reported as behind once it has caught up, because a
status that says "behind" for ever is one nobody reads.*
## What is kept
**Every result, including the failures.** A failed build that leaves no trace is indistinguishable
from one nobody asked for, and the difference is the whole of whether somebody should be looking
at something. A build that failed before it knew what it was building keeps the repository, which
is what a person goes and looks at.
Recording is idempotent on the correlation, because a result arrives twice — once as the answer to
whoever asked and once on the exchange, where the control plane is also listening. Two rows would
show one build as two, and which is real is not answerable afterwards.
That is what a builds view reads, and until it existed there was nothing to read: a result was
answered to the asker and kept nowhere.
## Three properties that are decisions
- **A fresh clone every time.** A build reusing a working tree can succeed because of something a
previous build left behind, and that is a build nobody can reproduce.
- **Archives are packed deterministically** — sorted, and carrying no timestamps, ownership or
original paths. Two builds of one commit must produce one digest, or nothing downstream can tell
*this changed* from *this was built again*, and every rebuild looks like a change to every
machine holding it.
- **Nothing is published until everything is built.** Half a module in the store, under a digest
the mesh never records, is reachable, unreferenced, and indistinguishable from something in use.
## What a module may build, and what it may only borrow
| kind | is |
|---|---|
| **image** | built from a Dockerfile in this repository |
| **archive** | a directory in this repository, packed |
| **upstream** | an image somebody else built, mirrored into the mesh's own registry |
**The third exists because a module usually runs software it did not write.** A database module
ships configuration and a provisioner and does not build a database. Naming the upstream reference
directly would need every machine to reach a public registry, and would pin to a tag its owner can
move — which is what pinning exists to prevent
([ADR 0006](../../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)). Mirroring is what the
bootstrap already does by hand; this makes it something a module can say.
An upstream reference with **no tag or digest is refused**: what gets mirrored would be whatever
`latest` means today, and a module pinned to that is not pinned.
## A module's own secret
A database has a superuser password, a broker an administrator, a registry an account. **None of
them is *for* anybody** — they are not the credential a consumer is given, and the mechanism that
hands those out has a consumer in the middle of it.
So a module says what it needs and where to put it — `own-secrets`, keyed by a name of the
module's choosing — and the mesh generates one **per node**,
seals it to that machine and reads it no more than it reads any other secret. Per node
deliberately: a module running on three machines has three passwords, where one in the manifest
would put the same secret on every machine that ever runs it, in a file anybody can read, for ever.
Made once and kept, or a running database would be handed a password it was not started with.
Remade when the machine's sealing key changes. **Declared and not made is refused**, because a
module whose own credential is silently absent starts, fails to authenticate, and the reason is
three layers from the machine reporting it.
### It is named for whose it is, not how secret it is
*2026-08-31, from an audit asking whether the manifest format was becoming hard to hold in the
head.*
The field was called `needs`, beside `secrets` — which is where a **provision's** credential lands
on a consumer. Both were name-to-path, both held something secret, and the names distinguished
them not at all. **Reaching for the wrong one parsed cleanly and failed somewhere else entirely**,
which is the shape of fault this whole design exists to prevent, sitting in the manifest format.
The axis that separates them is not how secret they are — both are — but **whose**:
| | keyed by | belongs to |
|---|---|---|
| `secrets` | the provision it is for | the relationship with another machine |
| `own-secrets` | a name the module chose | this module, and nobody else |
**A manifest using the old name is told the new one** rather than refused with "unknown field".
Whoever wrote it knew what they meant, and the mesh knows what it is called now.
### Some of them the mesh cannot make
*2026-08-31, from making the builder a module — the first thing to hold one.*
A generated secret is the mesh's, and remaking it costs nothing: **nothing else ever knew the old
one.** That is the assumption the paragraph above rests on, and it is not true of every secret a
module needs.
A broker account's password exists because **the broker was told about it**. A licence key exists
because somebody bought it. The mesh's job with these is to carry the value to the machine that
will use it and then be unable to read it — the same sealing, from the other direction: **given,
not generated.**
Treating the two alike is wrong in exactly one place, and it is the place nobody looks. When a
machine rejoins it has a new sealing key, and everything sealed to the old one is remade. Remaking
a *given* secret puts thirty-two random bytes where a working credential was, and every visible
signal says it worked: the mesh sealed a secret, the machine applied it, the file is there with
the right permissions. What fails is a program authenticating to something else, hours later,
with an error that names neither the mesh nor the secret.
So **where the value came from is recorded, and a given secret is never regenerated.** A rejoined
machine asking for one is refused, naming the remedy — issue it again — because the remedy is a
command somebody runs and no amount of pushing will produce a password the broker has never heard
of.
*Checked by taking a given secret, changing the machine's sealing key, and asserting the mesh
refuses rather than answers; and by asserting that two ordinary pushes hand back the same value,
without which the refusal would be a secret that never survives at all.*
## What one assignment gets you
A database module, written to see whether it could be:
```
directory /var/lib/mesh/postgres
directory /var/lib/mesh/postgres/grants
container the database pinned by digest, mirrored
container the provisioner pinned by digest, mirrored
file the superuser password sealed to this machine
file what its consumers asked for
```
**The provisioner watches** rather than being invoked. That is what lets it be a module: run once,
it needs something to run it after every declaration — a timer, or a unit wired to a file.
Watching, it is an ordinary long-running service the host already supervises. It polls rather than
watching the filesystem, because the host writes atomically: the file is replaced, so a watch on
the path stops seeing anything after the first replacement, and a watcher that silently stops
working is worse than a poll.
Writing it found one thing wrong, and it was the manifest rather than the host: a container
declared `restart-on`, which is a service field, and the host refused it by name. **It is right
to.** A container whose own definition changes is recreated, and a file it mounts is read by the
process inside, which is that image's business.
## Where artifacts go
**The registry the bootstrap already pulls from**, for both images and archives. An OCI registry
is a content-addressed blob store that also understands images, and an archive is a
content-addressed blob.
An object store beside it is the right answer for objects that are *mutable*, need per-reader
access, or are not build output. None of that describes a digest-pinned archive, and running a
second service for one kind of immutable blob is two things to run, two to back up, and two ways
for an artifact to be missing. **Overturnable without touching anything else**: a manifest carries
a URL and a digest, and neither says what served it.
### And the mesh runs it
*2026-08-31.* Which registry is a **provision**, mesh-scoped: a build machine requires
`artifact-store` and is told where it is, the same way an application is told where its database
is. Nothing is configured with an address.
This closes the last thing the mesh depended on and did not run. The registry a bootstrap pulls
from belongs to whoever raised the machine; from the moment the mesh has one of its own, an
artifact's home is somewhere the mesh can move, replace and back up.
**The chicken and egg is the bootstrap's, resolved the same way.** A registry module is an
`upstream` artifact — mirrored from a registry that already exists into the one being started. The
first copy comes from outside, exactly once, and every copy after it is the mesh's.
*Checked in the lab by assigning it and then asking for `/v2/` — on the machine, and from a second
machine across the private network, because a mesh-scoped provision that only answers locally is
not one. A container that is running is not a registry that replies, and this project has paid for
that distinction once already.*