Files

229 lines
11 KiB
Markdown

---
layer: to-be
status: in-progress
code:
- mesh-controller internal/builder
- mesh-controller cmd/mesh-controller (build, build --behind, push, status)
- mesh-controller internal/inventory/builds.go
updated: 2026-09-21
decisions:
- 02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md
- 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md
- 02-DECISIONS/0010-delivery.md
- 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0010-delivery.md
- 02-DECISIONS/0010-delivery.md
- 02-DECISIONS/0009-modules-and-the-graph.md
- 02-DECISIONS/0009-modules-and-the-graph.md
---
# Modules and delivery
How a change somebody makes becomes a thing running on machines.
This is the whole of it, current, in one place. Where a decision record is cited it is for the
reasoning behind a choice, not because the answer is somewhere else.
## A module
The unit of delivery: assignable to a node, versionable, replaceable on its own.
**Not a grouping.** There is no `networking` module containing four things — there are four
modules, named individually, with edges between them. Folders assert relationships; edges record
them, and only edges can be queried or kept true automatically.
**When several modules always change together**, that means they share an *authority* — one place
that decides for all of them. It does not mean they should be one artifact. Connectivity is the
worked example: one context decides the overlay, names, routes, filtering and certificates, and
`wireguard`, the resolver, the proxy and the firewall remain four modules, because they are
deployed to different sets of nodes.
> **Coherence is a context. Delivery is a module.**
## The three edges
A module's relationships to other modules. Two are declared; one is read from the code.
| edge | means | declared? | satisfied |
|---|---|---|---|
| **presence** | that thing must exist and be reachable here | yes, in the manifest | at provisioning |
| **instantiation** | that thing makes something for me and hands back credentials — a database, a bucket, a route | yes, in the manifest | at provisioning, and again whenever it must be |
| **build** | I was compiled against that artifact | **no — derived from imports** | **at build, once** |
**Why the build edge is derived and the others are not.** A runtime edge is an *intention*
somebody has about how the mesh should be wired, and only a person can state it. A build edge is
a *fact about code that already exists* — and a declared list of dependencies drifts from the
imports it describes, so the imports are what is read.
**Why the build edge is a different kind rather than a variant.** It is fixed inside an artifact
rather than negotiated when something runs, and its only remedy is a rebuild. Nothing can
re-provision it.
## The core library
One module everything is allowed to depend on, holding **the mesh's own domain**: a module, a
node, an assignment. Those three are what every context talks about and none of them owns.
The test for whether something belongs: *would this still mean the same thing in a context that
had never heard of the one it came from?* A node would. A pipeline stage would not — that is
delivery's. A grant would not — that is provisioning's.
**Types ship with the module that owns them**, not here. A consumer needing `inventory`'s types
depends on `inventory` — one narrow, visible edge — rather than everything depending on a hub
where the relationship cannot be seen. A library everything depends on is expensive to change
whether it holds types or code; what makes it expensive is the fan-in.
**This stays small on its own**, which is the point of choosing a domain rather than a drawer. A
domain model changes when what the mesh *is* changes, which is rare. *Shared code* changes
whenever anybody writes something reusable, which is constantly.
## Delivery is a comparison, not a pipeline
The controller holds two facts and builds the difference:
```
what source exists ─┐
├─► differ? ─► build ─► judge ─► declare ─► nodes converge
what has been built from it ─┘
```
**A change becomes a build because source is ahead of artifacts.** Not because a message arrived.
An event makes it fast; nothing makes it necessary — so a missed webhook costs latency and cannot
cost correctness.
That is the same shape the host uses on a machine, one layer up:
| | reconciles | against |
|---|---|---|
| the controller | artifacts | source |
| the host | machine state | declarations |
**There is no pipeline as a state machine.** No stage list something can be omitted from, and no
run to lose.
### An artifact is current, or it is not
> An artifact is out of date when **its source moved, or anything it was built against moved**.
So what is recorded against an artifact is a commit **and the identity of every artifact it was
built against** — its input closure. That is what makes *is this current?* answerable without
building anything, and what makes the rebuild set computable: take the changed module, follow
inbound build edges transitively, and that is what is stale. In order, because the edges are
directed.
**A shared change is a cascade, and that is inherent.** One change to the core library
invalidates nearly everything. The ordering comes from the graph, not from a hand-written list of
levels.
### The verdict
An artifact may not be declared until something has judged it fit. Two tiers, because one gate
would be both slow and unreliable:
| | judged by | when |
|---|---|---|
| **the module's own tests** | the build | **always** — this is most of it |
| **the lab** | a raised scenario | when an assertion genuinely needs a mesh |
**A run that failed for environmental reasons is not a verdict.** A machine that would not boot
says nothing about the artifact, and recording it as *unfit* is the same untruth as recording a
dispatch as a deploy. *Outstanding* and *failed* are different results.
### Declaring, and converging
Deploy is **one write**: the affected nodes' declarations now name the new artifact. It is not
once per node, and nothing is pushed to a machine.
Each host applies what it is told, reads back, and reports. A node that is switched off does it
when it wakes.
**What a delivery result means:**
```
meshboard source X · built from X · fit · declared on 5 · applied on 3, 2 outstanding
```
Not *the job went green*. **Outstanding is not failure** — a node that has not applied yet is a
fact with a timestamp, and it resolves itself when the node comes back.
## What this is designed against
Every property above answers something that has actually gone wrong, recorded in
[`00-as-is/04`](../00-as-is/04-delivery.md):
| what happened | what prevents it |
|---|---|
| a merge created no pipeline, and nothing said so | a change is found by comparison, not by an event |
| a package install 404'd from every mirror while the job went green | the applier is the reporter, and it reads back |
| a verify stage was built and never scheduled | verification is not a stage that can be left off a list |
| a service was reported started when the command merely returned | *green proves transport, not effect* — so nothing reports transport |
| the build node parked forever while every other node deployed | there is no fan-out to be asymmetric about |
## What must exist before this can be built
Not aspirations — things without which the above does not work:
1. **The module graph, with build edges.** No graph, no rebuild set and no ordering.
2. **A recorded input closure per artifact**, so currency is answerable without building.
3. **Something that notices a reconciler is not converging.** Below.
## Open
- **Does a fit artifact declare itself?** Nothing above says who moves the declaration. If it is
automatic, merging to main deploys to production — which may be wanted, and is far too large a
property to acquire by omission.
- **A reconciler that cannot reach its target retries forever.** A failed job stops and names its
step; a loop is silent. Without something that notices *this has been trying for an hour*, this
design reintroduces the fault it removes. **The largest open risk here.**
- **Reproducible builds.** If rebuilding unchanged source against unchanged inputs produced the
same digest, a cascade would stop at the first module whose output did not move. Without them,
one core-library commit redeploys the fleet with no behavioural change.
- **How a module publishes its own types**, which differs per language.
- **How the controller upgrades itself.** It declares its own new version and the host applies
it — but if the new one is broken, the thing that would fix it is the thing that is broken. The
host has a launcher for exactly this; the controller has nothing.
## What "behind" means, and what it used to mean
*2026-08-31.*
The risk this record names is losing **did my change go out?** — answerable today by opening a
pipeline, and something has to replace it or the comparison is worse to live with whatever its
other properties.
**It was answerable only for the machines that broke.** `push --behind` meant *failed or refused*,
so a machine that applied cleanly and whose declaration has since changed was not behind. For every
machine that worked, the answer was silence — and silence meant both *your change is running there*
and *your change has not been sent*, which is the question unanswered rather than answered.
**So the mesh records a digest of what it last sent each machine.** A digest rather than the
declaration: what a machine should be is recomputable at any moment, and a stored copy would be a
second account of it, able to disagree with the first. What cannot be recomputed is what was
*actually sent*.
**Recorded after the send.** A digest kept for something that failed to send would make the machine
look current for a declaration it never received — the failure mode this is meant to remove,
arrived at from the other side.
**Three situations, kept apart**, because they read differently to whoever is looking even where
the remedy is the same push:
| | |
|---|---|
| **out of date** | it was sent something, and the mesh would now send something else |
| **never told** | nobody has ever asked this machine to be anything |
| **not worked out** | the mesh cannot say what it should be — not reported here at all, because saying "waiting" about it would invent a comparison. `plan` is where that is answered |
**`status` says it and `push --behind` acts on it**, and both because the alternative is a flag that
knows something the person reading the status does not.
**A failure that repeats is said to be stuck**
([ADR 0090](../../02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md)). A machine
re-applies on its interval and reports each time, so a resource nothing can ever apply arrives as
the same failure over and over, at a fresh time each time. The mesh keeps, beside the last report,
when the current failure began and how many reports in a row have said it — the same resources by
id, whatever the words; three make the machine stuck, and `status` says so beside the failure. The host keeps trying — stuck is what the mesh
knows, not what the machine is told. *How it is checked:* an inventory test counts three identical
reports, a different one, and a clean apply; the status test asserts the word appears.