First pass of a design review, done by reading documents against code and against a raised mesh rather than against each other. Every error below was invisible to a proofread. **Statuses were stale, and nothing checked them.** Ten to-be documents said `designed` while naming working, lab-proven code — several with a *What was built* or *Raised, and observed* section. Added a `status-vs-code` check: naming a file is a claim that the file implements this, so a document that points at one has stopped being merely designed. It failed on all ten before it passed, per the rule this folder sets for its own checks. **The bundle carries three images, not two.** 07 reasoned about which substrate services go in and overlooked that the control plane is in there too — it is what the substrate exists to start, and there is nothing to fetch it with yet. Counted, not deduced. **The bootstrap uses four shapes, not six.** It listed `file` and `directory`, which substrate-first-node.lock never asks for. The claim that mattered — nothing is blocked on the host — was true either way, which is why the wrong count survived. **The eight capabilities were documented nowhere.** Implemented in internal/profile/detectors.go and enumerated in no document, including the one about the host that detects them. A vocabulary modules write against, readable only by reading the code. Now written down, with the seat/graphical-session distinction that is wrong in both directions if collapsed. **MinIO swept out of the to-be layer** per 0028. The gate now fails on one thing left deliberately: ADR 0024 is `proposed` while two documents rest on it and the feature it decides is built and lab-proven. Accepting a decision is not mine to do.
10 KiB
layer, status, code, updated, decisions
| layer | status | code | updated | decisions | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| to-be | in-progress |
|
2026-08-31 |
|
Modules and delivery
How a change somebody makes becomes a thing running on machines.
This is the whole of it, current, in one place. Where a decision record is cited it is for the reasoning behind a choice, not because the answer is somewhere else.
A module
The unit of delivery: assignable to a node, versionable, replaceable on its own.
Not a grouping. There is no networking module containing four things — there are four
modules, named individually, with edges between them. Folders assert relationships; edges record
them, and only edges can be queried or kept true automatically.
When several modules always change together, that means they share an authority — one place
that decides for all of them. It does not mean they should be one artifact. Connectivity is the
worked example: one context decides the overlay, names, routes, filtering and certificates, and
wireguard, the resolver, the proxy and the firewall remain four modules, because they are
deployed to different sets of nodes.
Coherence is a context. Delivery is a module.
The three edges
A module's relationships to other modules. Two are declared; one is read from the code.
| edge | means | declared? | satisfied |
|---|---|---|---|
| presence | that thing must exist and be reachable here | yes, in the manifest | at provisioning |
| instantiation | that thing makes something for me and hands back credentials — a database, a bucket, a route | yes, in the manifest | at provisioning, and again whenever it must be |
| build | I was compiled against that artifact | no — derived from imports | at build, once |
Why the build edge is derived and the others are not. A runtime edge is an intention somebody has about how the mesh should be wired, and only a person can state it. A build edge is a fact about code that already exists — and a declared list of dependencies drifts from the imports it describes, so the imports are what is read.
Why the build edge is a different kind rather than a variant. It is fixed inside an artifact rather than negotiated when something runs, and its only remedy is a rebuild. Nothing can re-provision it.
The core library
One module everything is allowed to depend on, holding the mesh's own domain: a module, a node, an assignment. Those three are what every context talks about and none of them owns.
The test for whether something belongs: would this still mean the same thing in a context that had never heard of the one it came from? A node would. A pipeline stage would not — that is delivery's. A grant would not — that is provisioning's.
Types ship with the module that owns them, not here. A consumer needing inventory's types
depends on inventory — one narrow, visible edge — rather than everything depending on a hub
where the relationship cannot be seen. A library everything depends on is expensive to change
whether it holds types or code; what makes it expensive is the fan-in.
This stays small on its own, which is the point of choosing a domain rather than a drawer. A domain model changes when what the mesh is changes, which is rare. Shared code changes whenever anybody writes something reusable, which is constantly.
Delivery is a comparison, not a pipeline
The control plane holds two facts and builds the difference:
what source exists ─┐
├─► differ? ─► build ─► judge ─► declare ─► nodes converge
what has been built from it ─┘
A change becomes a build because source is ahead of artifacts. Not because a message arrived. An event makes it fast; nothing makes it necessary — so a missed webhook costs latency and cannot cost correctness.
That is the same shape the host uses on a machine, one layer up:
| reconciles | against | |
|---|---|---|
| the control plane | artifacts | source |
| the host | machine state | declarations |
There is no pipeline as a state machine. No stage list something can be omitted from, and no run to lose.
An artifact is current, or it is not
An artifact is out of date when its source moved, or anything it was built against moved.
So what is recorded against an artifact is a commit and the identity of every artifact it was built against — its input closure. That is what makes is this current? answerable without building anything, and what makes the rebuild set computable: take the changed module, follow inbound build edges transitively, and that is what is stale. In order, because the edges are directed.
A shared change is a cascade, and that is inherent. One change to the core library invalidates nearly everything. The ordering comes from the graph, not from a hand-written list of levels.
The verdict
An artifact may not be declared until something has judged it fit. Two tiers, because one gate would be both slow and unreliable:
| judged by | when | |
|---|---|---|
| the module's own tests | the build | always — this is most of it |
| the lab | a raised scenario | when an assertion genuinely needs a mesh |
A run that failed for environmental reasons is not a verdict. A machine that would not boot says nothing about the artifact, and recording it as unfit is the same untruth as recording a dispatch as a deploy. Outstanding and failed are different results.
Declaring, and converging
Deploy is one write: the affected nodes' declarations now name the new artifact. It is not once per node, and nothing is pushed to a machine.
Each host applies what it is told, reads back, and reports. A node that is switched off does it when it wakes.
What a delivery result means:
meshboard source X · built from X · fit · declared on 5 · applied on 3, 2 outstanding
Not the job went green. Outstanding is not failure — a node that has not applied yet is a fact with a timestamp, and it resolves itself when the node comes back.
What this is designed against
Every property above answers something that has actually gone wrong, recorded in
00-as-is/04:
| what happened | what prevents it |
|---|---|
| a merge created no pipeline, and nothing said so | a change is found by comparison, not by an event |
| a package install 404'd from every mirror while the job went green | the applier is the reporter, and it reads back |
| a verify stage was built and never scheduled | verification is not a stage that can be left off a list |
| a service was reported started when the command merely returned | green proves transport, not effect — so nothing reports transport |
| the build node parked forever while every other node deployed | there is no fan-out to be asymmetric about |
What must exist before this can be built
Not aspirations — things without which the above does not work:
- The module graph, with build edges. No graph, no rebuild set and no ordering.
- A recorded input closure per artifact, so currency is answerable without building.
- Something that notices a reconciler is not converging. Below.
Open
- Does a fit artifact declare itself? Nothing above says who moves the declaration. If it is automatic, merging to main deploys to production — which may be wanted, and is far too large a property to acquire by omission.
- A reconciler that cannot reach its target retries forever. A failed job stops and names its step; a loop is silent. Without something that notices this has been trying for an hour, this design reintroduces the fault it removes. The largest open risk here.
- Reproducible builds. If rebuilding unchanged source against unchanged inputs produced the same digest, a cascade would stop at the first module whose output did not move. Without them, one core-library commit redeploys the fleet with no behavioural change.
- How a module publishes its own types, which differs per language.
- How the control plane upgrades itself. It declares its own new version and the host applies it — but if the new one is broken, the thing that would fix it is the thing that is broken. The host has a launcher for exactly this; the control plane has nothing.
What "behind" means, and what it used to mean
2026-08-31.
The risk this record names is losing did my change go out? — answerable today by opening a pipeline, and something has to replace it or the comparison is worse to live with whatever its other properties.
It was answerable only for the machines that broke. push --behind meant failed or refused,
so a machine that applied cleanly and whose declaration has since changed was not behind. For every
machine that worked, the answer was silence — and silence meant both your change is running there
and your change has not been sent, which is the question unanswered rather than answered.
So the mesh records a digest of what it last sent each machine. A digest rather than the declaration: what a machine should be is recomputable at any moment, and a stored copy would be a second account of it, able to disagree with the first. What cannot be recomputed is what was actually sent.
Recorded after the send. A digest kept for something that failed to send would make the machine look current for a declaration it never received — the failure mode this is meant to remove, arrived at from the other side.
Three situations, kept apart, because they read differently to whoever is looking even where the remedy is the same push:
| out of date | it was sent something, and the mesh would now send something else |
| never told | nobody has ever asked this machine to be anything |
| not worked out | the mesh cannot say what it should be — not reported here at all, because saying "waiting" about it would invent a comparison. plan is where that is answered |
status says it and push --behind acts on it, and both because the alternative is a flag that
knows something the person reading the status does not.