--- layer: to-be status: in-progress code: - mesh-controller internal/builder - mesh-controller cmd/mesh-controller (build, build --behind, push, status) - mesh-controller internal/inventory/builds.go updated: 2026-09-21 decisions: - 02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md - 02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md - 02-DECISIONS/0010-delivery.md - 02-DECISIONS/0009-modules-and-the-graph.md - 02-DECISIONS/0009-modules-and-the-graph.md - 02-DECISIONS/0010-delivery.md - 02-DECISIONS/0010-delivery.md - 02-DECISIONS/0009-modules-and-the-graph.md - 02-DECISIONS/0009-modules-and-the-graph.md --- # Modules and delivery How a change somebody makes becomes a thing running on machines. This is the whole of it, current, in one place. Where a decision record is cited it is for the reasoning behind a choice, not because the answer is somewhere else. ## A module The unit of delivery: assignable to a node, versionable, replaceable on its own. **Not a grouping.** There is no `networking` module containing four things — there are four modules, named individually, with edges between them. Folders assert relationships; edges record them, and only edges can be queried or kept true automatically. **When several modules always change together**, that means they share an *authority* — one place that decides for all of them. It does not mean they should be one artifact. Connectivity is the worked example: one context decides the overlay, names, routes, filtering and certificates, and `wireguard`, the resolver, the proxy and the firewall remain four modules, because they are deployed to different sets of nodes. > **Coherence is a context. Delivery is a module.** ## The three edges A module's relationships to other modules. Two are declared; one is read from the code. | edge | means | declared? | satisfied | |---|---|---|---| | **presence** | that thing must exist and be reachable here | yes, in the manifest | at provisioning | | **instantiation** | that thing makes something for me and hands back credentials — a database, a bucket, a route | yes, in the manifest | at provisioning, and again whenever it must be | | **build** | I was compiled against that artifact | **no — derived from imports** | **at build, once** | **Why the build edge is derived and the others are not.** A runtime edge is an *intention* somebody has about how the mesh should be wired, and only a person can state it. A build edge is a *fact about code that already exists* — and a declared list of dependencies drifts from the imports it describes, so the imports are what is read. **Why the build edge is a different kind rather than a variant.** It is fixed inside an artifact rather than negotiated when something runs, and its only remedy is a rebuild. Nothing can re-provision it. ## The core library One module everything is allowed to depend on, holding **the mesh's own domain**: a module, a node, an assignment. Those three are what every context talks about and none of them owns. The test for whether something belongs: *would this still mean the same thing in a context that had never heard of the one it came from?* A node would. A pipeline stage would not — that is delivery's. A grant would not — that is provisioning's. **Types ship with the module that owns them**, not here. A consumer needing `inventory`'s types depends on `inventory` — one narrow, visible edge — rather than everything depending on a hub where the relationship cannot be seen. A library everything depends on is expensive to change whether it holds types or code; what makes it expensive is the fan-in. **This stays small on its own**, which is the point of choosing a domain rather than a drawer. A domain model changes when what the mesh *is* changes, which is rare. *Shared code* changes whenever anybody writes something reusable, which is constantly. ## Delivery is a comparison, not a pipeline The controller holds two facts and builds the difference: ``` what source exists ─┐ ├─► differ? ─► build ─► judge ─► declare ─► nodes converge what has been built from it ─┘ ``` **A change becomes a build because source is ahead of artifacts.** Not because a message arrived. An event makes it fast; nothing makes it necessary — so a missed webhook costs latency and cannot cost correctness. That is the same shape the host uses on a machine, one layer up: | | reconciles | against | |---|---|---| | the controller | artifacts | source | | the host | machine state | declarations | **There is no pipeline as a state machine.** No stage list something can be omitted from, and no run to lose. ### An artifact is current, or it is not > An artifact is out of date when **its source moved, or anything it was built against moved**. So what is recorded against an artifact is a commit **and the identity of every artifact it was built against** — its input closure. That is what makes *is this current?* answerable without building anything, and what makes the rebuild set computable: take the changed module, follow inbound build edges transitively, and that is what is stale. In order, because the edges are directed. **A shared change is a cascade, and that is inherent.** One change to the core library invalidates nearly everything. The ordering comes from the graph, not from a hand-written list of levels. ### The verdict An artifact may not be declared until something has judged it fit. Two tiers, because one gate would be both slow and unreliable: | | judged by | when | |---|---|---| | **the module's own tests** | the build | **always** — this is most of it | | **the lab** | a raised scenario | when an assertion genuinely needs a mesh | **A run that failed for environmental reasons is not a verdict.** A machine that would not boot says nothing about the artifact, and recording it as *unfit* is the same untruth as recording a dispatch as a deploy. *Outstanding* and *failed* are different results. ### Declaring, and converging Deploy is **one write**: the affected nodes' declarations now name the new artifact. It is not once per node, and nothing is pushed to a machine. Each host applies what it is told, reads back, and reports. A node that is switched off does it when it wakes. **What a delivery result means:** ``` meshboard source X · built from X · fit · declared on 5 · applied on 3, 2 outstanding ``` Not *the job went green*. **Outstanding is not failure** — a node that has not applied yet is a fact with a timestamp, and it resolves itself when the node comes back. ## What this is designed against Every property above answers something that has actually gone wrong, recorded in [`00-as-is/04`](../00-as-is/04-delivery.md): | what happened | what prevents it | |---|---| | a merge created no pipeline, and nothing said so | a change is found by comparison, not by an event | | a package install 404'd from every mirror while the job went green | the applier is the reporter, and it reads back | | a verify stage was built and never scheduled | verification is not a stage that can be left off a list | | a service was reported started when the command merely returned | *green proves transport, not effect* — so nothing reports transport | | the build node parked forever while every other node deployed | there is no fan-out to be asymmetric about | ## What must exist before this can be built Not aspirations — things without which the above does not work: 1. **The module graph, with build edges.** No graph, no rebuild set and no ordering. 2. **A recorded input closure per artifact**, so currency is answerable without building. 3. **Something that notices a reconciler is not converging.** Below. ## Open - **Does a fit artifact declare itself?** Nothing above says who moves the declaration. If it is automatic, merging to main deploys to production — which may be wanted, and is far too large a property to acquire by omission. - **A reconciler that cannot reach its target retries forever.** A failed job stops and names its step; a loop is silent. Without something that notices *this has been trying for an hour*, this design reintroduces the fault it removes. **The largest open risk here.** - **Reproducible builds.** If rebuilding unchanged source against unchanged inputs produced the same digest, a cascade would stop at the first module whose output did not move. Without them, one core-library commit redeploys the fleet with no behavioural change. - **How a module publishes its own types**, which differs per language. - **How the controller upgrades itself.** It declares its own new version and the host applies it — but if the new one is broken, the thing that would fix it is the thing that is broken. The host has a launcher for exactly this; the controller has nothing. ## What "behind" means, and what it used to mean *2026-08-31.* The risk this record names is losing **did my change go out?** — answerable today by opening a pipeline, and something has to replace it or the comparison is worse to live with whatever its other properties. **It was answerable only for the machines that broke.** `push --behind` meant *failed or refused*, so a machine that applied cleanly and whose declaration has since changed was not behind. For every machine that worked, the answer was silence — and silence meant both *your change is running there* and *your change has not been sent*, which is the question unanswered rather than answered. **So the mesh records a digest of what it last sent each machine.** A digest rather than the declaration: what a machine should be is recomputable at any moment, and a stored copy would be a second account of it, able to disagree with the first. What cannot be recomputed is what was *actually sent*. **Recorded after the send.** A digest kept for something that failed to send would make the machine look current for a declaration it never received — the failure mode this is meant to remove, arrived at from the other side. **Three situations, kept apart**, because they read differently to whoever is looking even where the remedy is the same push: | | | |---|---| | **out of date** | it was sent something, and the mesh would now send something else | | **never told** | nobody has ever asked this machine to be anything | | **not worked out** | the mesh cannot say what it should be — not reported here at all, because saying "waiting" about it would invent a comparison. `plan` is where that is answered | **`status` says it and `push --behind` acts on it**, and both because the alternative is a flag that knows something the person reading the status does not. **A failure that repeats is said to be stuck** ([ADR 0090](../../02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md)). A machine re-applies on its interval and reports each time, so a resource nothing can ever apply arrives as the same failure over and over, at a fresh time each time. The mesh keeps, beside the last report, when the current failure began and how many reports in a row have said it — the same resources by id, whatever the words; three make the machine stuck, and `status` says so beside the failure. The host keeps trying — stuck is what the mesh knows, not what the machine is told. *How it is checked:* an inventory test counts three identical reports, a different one, and a clean apply; the status test asserts the word appears.