Files

11 KiB

layer, status, code, updated, decisions
layer status code updated decisions
to-be in-progress
mesh-controller internal/builder
mesh-controller cmd/mesh-controller (build, build --behind, push, status)
mesh-controller internal/inventory/builds.go
2026-09-21
02-DECISIONS/0090-a-failure-that-repeats-is-said-to-be-stuck.md
02-DECISIONS/0082-the-registry-is-reached-by-name-and-trusted-by-the-overlay.md
02-DECISIONS/0010-delivery.md
02-DECISIONS/0009-modules-and-the-graph.md
02-DECISIONS/0009-modules-and-the-graph.md
02-DECISIONS/0010-delivery.md
02-DECISIONS/0010-delivery.md
02-DECISIONS/0009-modules-and-the-graph.md
02-DECISIONS/0009-modules-and-the-graph.md

Modules and delivery

How a change somebody makes becomes a thing running on machines.

This is the whole of it, current, in one place. Where a decision record is cited it is for the reasoning behind a choice, not because the answer is somewhere else.

A module

The unit of delivery: assignable to a node, versionable, replaceable on its own.

Not a grouping. There is no networking module containing four things — there are four modules, named individually, with edges between them. Folders assert relationships; edges record them, and only edges can be queried or kept true automatically.

When several modules always change together, that means they share an authority — one place that decides for all of them. It does not mean they should be one artifact. Connectivity is the worked example: one context decides the overlay, names, routes, filtering and certificates, and wireguard, the resolver, the proxy and the firewall remain four modules, because they are deployed to different sets of nodes.

Coherence is a context. Delivery is a module.

The three edges

A module's relationships to other modules. Two are declared; one is read from the code.

edge means declared? satisfied
presence that thing must exist and be reachable here yes, in the manifest at provisioning
instantiation that thing makes something for me and hands back credentials — a database, a bucket, a route yes, in the manifest at provisioning, and again whenever it must be
build I was compiled against that artifact no — derived from imports at build, once

Why the build edge is derived and the others are not. A runtime edge is an intention somebody has about how the mesh should be wired, and only a person can state it. A build edge is a fact about code that already exists — and a declared list of dependencies drifts from the imports it describes, so the imports are what is read.

Why the build edge is a different kind rather than a variant. It is fixed inside an artifact rather than negotiated when something runs, and its only remedy is a rebuild. Nothing can re-provision it.

The core library

One module everything is allowed to depend on, holding the mesh's own domain: a module, a node, an assignment. Those three are what every context talks about and none of them owns.

The test for whether something belongs: would this still mean the same thing in a context that had never heard of the one it came from? A node would. A pipeline stage would not — that is delivery's. A grant would not — that is provisioning's.

Types ship with the module that owns them, not here. A consumer needing inventory's types depends on inventory — one narrow, visible edge — rather than everything depending on a hub where the relationship cannot be seen. A library everything depends on is expensive to change whether it holds types or code; what makes it expensive is the fan-in.

This stays small on its own, which is the point of choosing a domain rather than a drawer. A domain model changes when what the mesh is changes, which is rare. Shared code changes whenever anybody writes something reusable, which is constantly.

Delivery is a comparison, not a pipeline

The controller holds two facts and builds the difference:

what source exists          ─┐
                             ├─►  differ?  ─►  build  ─►  judge  ─►  declare  ─►  nodes converge
what has been built from it ─┘

A change becomes a build because source is ahead of artifacts. Not because a message arrived. An event makes it fast; nothing makes it necessary — so a missed webhook costs latency and cannot cost correctness.

That is the same shape the host uses on a machine, one layer up:

reconciles against
the controller artifacts source
the host machine state declarations

There is no pipeline as a state machine. No stage list something can be omitted from, and no run to lose.

An artifact is current, or it is not

An artifact is out of date when its source moved, or anything it was built against moved.

So what is recorded against an artifact is a commit and the identity of every artifact it was built against — its input closure. That is what makes is this current? answerable without building anything, and what makes the rebuild set computable: take the changed module, follow inbound build edges transitively, and that is what is stale. In order, because the edges are directed.

A shared change is a cascade, and that is inherent. One change to the core library invalidates nearly everything. The ordering comes from the graph, not from a hand-written list of levels.

The verdict

An artifact may not be declared until something has judged it fit. Two tiers, because one gate would be both slow and unreliable:

judged by when
the module's own tests the build always — this is most of it
the lab a raised scenario when an assertion genuinely needs a mesh

A run that failed for environmental reasons is not a verdict. A machine that would not boot says nothing about the artifact, and recording it as unfit is the same untruth as recording a dispatch as a deploy. Outstanding and failed are different results.

Declaring, and converging

Deploy is one write: the affected nodes' declarations now name the new artifact. It is not once per node, and nothing is pushed to a machine.

Each host applies what it is told, reads back, and reports. A node that is switched off does it when it wakes.

What a delivery result means:

meshboard   source X · built from X · fit · declared on 5 · applied on 3, 2 outstanding

Not the job went green. Outstanding is not failure — a node that has not applied yet is a fact with a timestamp, and it resolves itself when the node comes back.

What this is designed against

Every property above answers something that has actually gone wrong, recorded in 00-as-is/04:

what happened what prevents it
a merge created no pipeline, and nothing said so a change is found by comparison, not by an event
a package install 404'd from every mirror while the job went green the applier is the reporter, and it reads back
a verify stage was built and never scheduled verification is not a stage that can be left off a list
a service was reported started when the command merely returned green proves transport, not effect — so nothing reports transport
the build node parked forever while every other node deployed there is no fan-out to be asymmetric about

What must exist before this can be built

Not aspirations — things without which the above does not work:

  1. The module graph, with build edges. No graph, no rebuild set and no ordering.
  2. A recorded input closure per artifact, so currency is answerable without building.
  3. Something that notices a reconciler is not converging. Below.

Open

  • Does a fit artifact declare itself? Nothing above says who moves the declaration. If it is automatic, merging to main deploys to production — which may be wanted, and is far too large a property to acquire by omission.
  • A reconciler that cannot reach its target retries forever. A failed job stops and names its step; a loop is silent. Without something that notices this has been trying for an hour, this design reintroduces the fault it removes. The largest open risk here.
  • Reproducible builds. If rebuilding unchanged source against unchanged inputs produced the same digest, a cascade would stop at the first module whose output did not move. Without them, one core-library commit redeploys the fleet with no behavioural change.
  • How a module publishes its own types, which differs per language.
  • How the controller upgrades itself. It declares its own new version and the host applies it — but if the new one is broken, the thing that would fix it is the thing that is broken. The host has a launcher for exactly this; the controller has nothing.

What "behind" means, and what it used to mean

2026-08-31.

The risk this record names is losing did my change go out? — answerable today by opening a pipeline, and something has to replace it or the comparison is worse to live with whatever its other properties.

It was answerable only for the machines that broke. push --behind meant failed or refused, so a machine that applied cleanly and whose declaration has since changed was not behind. For every machine that worked, the answer was silence — and silence meant both your change is running there and your change has not been sent, which is the question unanswered rather than answered.

So the mesh records a digest of what it last sent each machine. A digest rather than the declaration: what a machine should be is recomputable at any moment, and a stored copy would be a second account of it, able to disagree with the first. What cannot be recomputed is what was actually sent.

Recorded after the send. A digest kept for something that failed to send would make the machine look current for a declaration it never received — the failure mode this is meant to remove, arrived at from the other side.

Three situations, kept apart, because they read differently to whoever is looking even where the remedy is the same push:

out of date it was sent something, and the mesh would now send something else
never told nobody has ever asked this machine to be anything
not worked out the mesh cannot say what it should be — not reported here at all, because saying "waiting" about it would invent a comparison. plan is where that is answered

status says it and push --behind acts on it, and both because the alternative is a flag that knows something the person reading the status does not.

A failure that repeats is said to be stuck (ADR 0090). A machine re-applies on its interval and reports each time, so a resource nothing can ever apply arrives as the same failure over and over, at a fresh time each time. The mesh keeps, beside the last report, when the current failure began and how many reports in a row have said it — the same resources by id, whatever the words; three make the machine stuck, and status says so beside the failure. The host keeps trying — stuck is what the mesh knows, not what the machine is told. How it is checked: an inventory test counts three identical reports, a different one, and a clean apply; the status test asserts the word appears.