Files
hq/03-DESIGN/01-to-be/10-delivery.md
T
jschoubben 333356cff3 Order the records the way the system is learned
Jochen asked whether the order made sense. It did not -- it followed when
things happened to be decided, which after consolidation is fictional anyway
since record 5 alone folds decisions taken across a week.

Concretely wrong before: the domain statement sat at 8, after five engineering
rules; the constitution was scattered across 5, 12 and 17; the tiers landed at
15, 16, 21 and 22 with process records in between.

Now it walks: what the mesh is (1-3), its tiers from the bottom up (4-8), what
runs on them and how it gets there (9-10), how it is built (11-16), how it is
checked (17-18), how we work (19-23).

Two things made this safe rather than free. It is a permutation, not a
compaction, so the renames go through temporary names -- otherwise two files
want one slot and one is lost. And the reference rewrite is a single
simultaneous pass, because almost every number moved into a slot another number
was vacating; replacing one at a time would have cascaded and pointed things at
the wrong record while still resolving.

Verified: 284 [ADR NNNN](path) links across the repository, all with matching
text and target.

The ordering principle is now stated in 19 rather than left implicit -- the
repository already said "the numbering is the flow" about its folders, and
there was no reason for the records to be the exception.
2026-08-28 23:30:42 +02:00

181 lines
8.5 KiB
Markdown

---
layer: to-be
status: designed
code: []
updated: 2026-08-28
decisions:
- 02-DECISIONS/0010-delivery.md
- 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0010-delivery.md
- 02-DECISIONS/0010-delivery.md
- 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0009-modules-and-the-graph.md
---
# Modules and delivery
How a change somebody makes becomes a thing running on machines.
This is the whole of it, current, in one place. Where a decision record is cited it is for the
reasoning behind a choice, not because the answer is somewhere else.
## A module
The unit of delivery: assignable to a node, versionable, replaceable on its own.
**Not a grouping.** There is no `networking` module containing four things — there are four
modules, named individually, with edges between them. Folders assert relationships; edges record
them, and only edges can be queried or kept true automatically.
**When several modules always change together**, that means they share an *authority* — one place
that decides for all of them. It does not mean they should be one artifact. Connectivity is the
worked example: one context decides the overlay, names, routes, filtering and certificates, and
`wireguard`, the resolver, the proxy and the firewall remain four modules, because they are
deployed to different sets of nodes.
> **Coherence is a context. Delivery is a module.**
## The three edges
A module's relationships to other modules. Two are declared; one is read from the code.
| edge | means | declared? | satisfied |
|---|---|---|---|
| **presence** | that thing must exist and be reachable here | yes, in the manifest | at provisioning |
| **instantiation** | that thing makes something for me and hands back credentials — a database, a bucket, a route | yes, in the manifest | at provisioning, and again whenever it must be |
| **build** | I was compiled against that artifact | **no — derived from imports** | **at build, once** |
**Why the build edge is derived and the others are not.** A runtime edge is an *intention*
somebody has about how the mesh should be wired, and only a person can state it. A build edge is
a *fact about code that already exists* — and a declared list of dependencies drifts from the
imports it describes, so the imports are what is read.
**Why the build edge is a different kind rather than a variant.** It is fixed inside an artifact
rather than negotiated when something runs, and its only remedy is a rebuild. Nothing can
re-provision it.
## The core library
One module everything is allowed to depend on, holding **the mesh's own domain**: a module, a
node, an assignment. Those three are what every context talks about and none of them owns.
The test for whether something belongs: *would this still mean the same thing in a context that
had never heard of the one it came from?* A node would. A pipeline stage would not — that is
delivery's. A grant would not — that is provisioning's.
**Types ship with the module that owns them**, not here. A consumer needing `inventory`'s types
depends on `inventory` — one narrow, visible edge — rather than everything depending on a hub
where the relationship cannot be seen. A library everything depends on is expensive to change
whether it holds types or code; what makes it expensive is the fan-in.
**This stays small on its own**, which is the point of choosing a domain rather than a drawer. A
domain model changes when what the mesh *is* changes, which is rare. *Shared code* changes
whenever anybody writes something reusable, which is constantly.
## Delivery is a comparison, not a pipeline
The control plane holds two facts and builds the difference:
```
what source exists ─┐
├─► differ? ─► build ─► judge ─► declare ─► nodes converge
what has been built from it ─┘
```
**A change becomes a build because source is ahead of artifacts.** Not because a message arrived.
An event makes it fast; nothing makes it necessary — so a missed webhook costs latency and cannot
cost correctness.
That is the same shape the host uses on a machine, one layer up:
| | reconciles | against |
|---|---|---|
| the control plane | artifacts | source |
| the host | machine state | declarations |
**There is no pipeline as a state machine.** No stage list something can be omitted from, and no
run to lose.
### An artifact is current, or it is not
> An artifact is out of date when **its source moved, or anything it was built against moved**.
So what is recorded against an artifact is a commit **and the identity of every artifact it was
built against** — its input closure. That is what makes *is this current?* answerable without
building anything, and what makes the rebuild set computable: take the changed module, follow
inbound build edges transitively, and that is what is stale. In order, because the edges are
directed.
**A shared change is a cascade, and that is inherent.** One change to the core library
invalidates nearly everything. The ordering comes from the graph, not from a hand-written list of
levels.
### The verdict
An artifact may not be declared until something has judged it fit. Two tiers, because one gate
would be both slow and unreliable:
| | judged by | when |
|---|---|---|
| **the module's own tests** | the build | **always** — this is most of it |
| **the lab** | a raised scenario | when an assertion genuinely needs a mesh |
**A run that failed for environmental reasons is not a verdict.** A machine that would not boot
says nothing about the artifact, and recording it as *unfit* is the same untruth as recording a
dispatch as a deploy. *Outstanding* and *failed* are different results.
### Declaring, and converging
Deploy is **one write**: the affected nodes' declarations now name the new artifact. It is not
once per node, and nothing is pushed to a machine.
Each host applies what it is told, reads back, and reports. A node that is switched off does it
when it wakes.
**What a delivery result means:**
```
meshboard source X · built from X · fit · declared on 5 · applied on 3, 2 outstanding
```
Not *the job went green*. **Outstanding is not failure** — a node that has not applied yet is a
fact with a timestamp, and it resolves itself when the node comes back.
## What this is designed against
Every property above answers something that has actually gone wrong, recorded in
[`00-as-is/04`](../00-as-is/04-delivery.md):
| what happened | what prevents it |
|---|---|
| a merge created no pipeline, and nothing said so | a change is found by comparison, not by an event |
| a package install 404'd from every mirror while the job went green | the applier is the reporter, and it reads back |
| a verify stage was built and never scheduled | verification is not a stage that can be left off a list |
| a service was reported started when the command merely returned | *green proves transport, not effect* — so nothing reports transport |
| the build node parked forever while every other node deployed | there is no fan-out to be asymmetric about |
## What must exist before this can be built
Not aspirations — things without which the above does not work:
1. **The module graph, with build edges.** No graph, no rebuild set and no ordering.
2. **A recorded input closure per artifact**, so currency is answerable without building.
3. **Something that notices a reconciler is not converging.** Below.
## Open
- **Does a fit artifact declare itself?** Nothing above says who moves the declaration. If it is
automatic, merging to main deploys to production — which may be wanted, and is far too large a
property to acquire by omission.
- **A reconciler that cannot reach its target retries forever.** A failed job stops and names its
step; a loop is silent. Without something that notices *this has been trying for an hour*, this
design reintroduces the fault it removes. **The largest open risk here.**
- **Reproducible builds.** If rebuilding unchanged source against unchanged inputs produced the
same digest, a cascade would stop at the first module whose output did not move. Without them,
one core-library commit redeploys the fleet with no behavioural change.
- **How a module publishes its own types**, which differs per language.
- **How the control plane upgrades itself.** It declares its own new version and the host applies
it — but if the new one is broken, the thing that would fix it is the thing that is broken. The
host has a launcher for exactly this; the control plane has nothing.