Compare commits

..
Author SHA1 Message Date
jschoubben ab6db9369b Issue 133: the control plane's schema is migrated at birth and never again
The mesh replaced its own control plane with a build carrying a migration, applied none of it, and
then recorded no build for three quarters of an hour while saying everything was fine. ADR 0052
already prescribes the shape — a run-once step that gates the server — and the control plane was the
one module that did not use it.
2026-09-28 10:27:30 +02:00
mesh-admin 98d94ef71e Merge pull request 'Design 28: 5.5 done, the mesh has one bus; issue 131 resolved' (#156) from design/28-one-bus-issue-131-resolved into main 2026-09-28 01:59:41 +00:00
jschoubben 4e13280604 Design 28: 5.5 done, the mesh has one bus; issue 131 resolved
The AMQP transport is gone from the control plane and the hosts (mesh-controller #112,
mesh-host #39). On the way: no build had ever recorded its bases, so every bases-first order
walked an empty graph; the builder now reports what it was handed and the graph is read from
builds (mesh-controller #113/#114). Issue 131 is resolved by the forge module's merge event,
the control plane following it, and those edges.
2026-09-28 03:59:39 +02:00
3 changed files with 198 additions and 3 deletions
+22 -3
View File
@@ -5,7 +5,7 @@ code:
- mesh-catalog modules/nats
- mesh-controller internal/catalogue
- mesh-lab scenarios
updated: 2026-09-27
updated: 2026-09-28
decisions:
- 02-DECISIONS/0116-the-bus-is-built-in-five-steps.md
- 02-DECISIONS/0106-the-bus-is-nats.md
@@ -694,8 +694,27 @@ healthy while reacting to nothing.
> shutting it down ends the path that reaches this installation's machines from a workstation.
> The rollout is driven from the node, or before the broker stops — a sequencing constraint on
> 5.2, not an afterthought.
- [ ] 5.5 **the AMQP transport is deleted from the control plane and the hosts**, and the variable
that selected a transport is refused at start as unknown. One bus, nothing to select ([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)).
- [x] 5.5 **the AMQP transport is deleted from the control plane and the hosts**. One bus, nothing
to select ([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)).
**Done 2026-09-28.** The control plane's old consume loop, build request, tool call, management
API and account scoping went, and the host's old dialling and enrolment paths with them; a
membership or token naming any other bus is refused before anything is sent. Nothing selects a
transport any more: the variable that once did (`MESH_BUS_NATS`) now only names where the
control plane reads its own bus credential, the way any module reads a secret. **Checked by the
build**: neither repository's module file names the AMQP client library, so a line that still
used it would not compile. The store-window guarantee ([issue 083](../../04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/00-report.md))
is tested against a bus-less fake rather than the old transport's memory, which is what let
that memory go — the one thing it did that the stream does not (superseding a held report) is
the staleness check on the message itself (design 25 §3).
Found on the way: **no build had ever recorded what it stood on.** A recipe reads its base from
a build argument, so the digest was never in the file the builder derived edges from, and every
order that says *bases first* — `build --on`, `build --behind`, the merge follow-up of
[issue 131](../../04-ISSUES/131-nothing-tells-the-mesh-a-source-moved/00-report.md) — walked a
graph with no edges. The builder now reports the bases it was handed, the control plane records
them by artifact path, and the graph is read from the newest build of each module — a recorded
manifest carries no `build.on`, so the edge is derived from the build or it does not exist.
> **The old 5.4 note is history.** It recorded that a retirement *condition* was wrong from
> [ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md) onward, which framed the old broker
@@ -0,0 +1,107 @@
---
status: resolved
opened: 2026-09-27
located-in: [mesh-catalog modules/gitea, mesh-controller cmd/mesh-controller, mesh-controller internal/builder]
fixed-by: mesh-catalog #124 — the forge module watches for merged pull requests and emits `pull.merged` with the merge commit and the clone address; mesh-controller #110/#111 — the control plane follows that event on the bus, marks every module built from that repository and branch as moved, and builds them bases first, stopping when a base fails; mesh-controller #113/#114 — a build records the bases it was handed and the graph is read from builds, without which "bases first" had no edges to order by.
amended-design: 03-DESIGN/01-to-be/28-building-the-bus.md
---
# 131 — Nothing tells the mesh a source moved, and it reports itself current anyway
## What was observed
Six changes were merged to the trunk of six repositories in one sitting. The build machine built
nothing. Its last build, minutes before the first merge, was still the one it reported; no build was
requested, refused or failed, because none was ever asked for.
Asked afterwards what was wrong, the mesh said:
> 4 machine(s), all doing what they were told, all heard from, running what the mesh would send them,
> and every module current with its source
Every one of the six had moved. The last clause was false for all of them, and it is the clause a
person reads to decide whether there is anything to do.
## Why it matters beyond this instance
**The mesh learns a source moved by being told, and there is no longer anything to tell it.** The
command exists — a person names the module and the commit — and so does the question the overview
answers. What is missing is whatever used to connect the two. One repository still carries a forge
webhook aimed at a port; the rest carry none, and the port belongs to a different service than the
one the arrangement implies. So the state is not "the trigger is broken" but "there is no trigger,
and nothing says so".
**A wrong answer is worse here than no answer.** "Every module current with its source" is
indistinguishable, to a reader, from a mesh that has genuinely caught up. The overview is built to be
the thing you check instead of checking by hand, so a confident false negative removes the habit that
would otherwise have caught it. Nothing in the mesh is at fault for being out of date — it is at
fault for saying it is not.
**It is also why "current with its source" cannot be a stored fact.** The mesh compares what it built
against what it was last told the source was, and calls that agreement. Two facts agreeing tells you
nothing when both come from the same place.
## The intended shape, which is decided and not built
The forge emits what happened to it — a pull request merged — and the build machine reacts by
building what that commit affects. That keeps the forge ignorant of the build system and the build
machine ignorant of the forge's internals, which is the same argument
[ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) makes for addressing an event
to its emitter: the merge is a fact about the forge, and what should be rebuilt because of it is not
the forge's business to know.
The forge's module already declares the event. The build machine declares that it consumes nothing.
## What the trigger cannot be
**Not one build per changed module.** The modules form a graph: several are built from one
repository, and some are the base another is compiled on — a runtime image, a compiler base, a
repository whose context a second module builds from. Firing a build for each changed module
independently would start work that cannot succeed yet and produce a failure per dependent, for one
cause.
Observed while catching the mesh up by hand on 2026-09-27: a compiler base had to move before
anything compiled against it could build, and when it failed, the right behaviour was for its
dependents to wait rather than each fail the same way. Fifteen modules shared the cause. A trigger
that reports it fifteen times has buried it.
So whatever reacts to the forge's event resolves what changed into an order, builds the bases first,
and holds a dependent while its base is unbuilt or failed. That is a larger thing than "rebuild what
the commit touched", and knowing it now is cheaper than discovering it from fifteen identical
failures.
## Open questions
- Is "the source moved" still a thing a person can assert by hand once the event path exists, or does
the hand-operated form become the thing that made this failure possible?
- Which commit does the build machine act on — the merge, or each commit it brought — and what does
it do when several arrive for one module at once?
- How does the overview stop being able to lie? Comparing what was built against what was recorded
will always agree. Whether the trunk has moved is a question only the forge can answer, so either
the overview asks it, or it stops claiming to know.
- Does this want to be the same mechanism as the build request on the bus
([ADR 0129](../../02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md)), or does it sit in
front of it?
## What was done (2026-09-28)
The shape above was built as described: the forge's module emits the merge, the control plane
consumes it, and nothing on either side knows the other's internals. The build is asked for the
merge commit, not each commit the merge brought — the trunk moved once, to one place. Several merges
for one module arriving in a row are followed in turn, each moving the recorded source to its own
commit, so the last one to arrive is the one the mesh ends up built from.
The hand-operated form stays. `module moved` is how a source is recorded without a forge — a module
built from a repository elsewhere, or a mesh whose forge module is down — and it is the same act the
event performs, so the two cannot disagree about what "moved" means.
**Bases first needed edges, and there were none.** The order this report asked for was written and
walked a graph that no build had ever recorded: a recipe reads its base from a build argument, so the
digest was never in the file the builder read edges from. A build now reports what it was handed, the
control plane records it by artifact path, and the order is read from each module's newest build.
**What still can lie.** The overview compares what was built against where it was last told the
source is; the forge's event is now what moves that mark, so it is right for as long as the forge
module was listening. A merge made while that module was down is a merge the mesh does not know of
until the module next polls — it announces what merged since it last looked, so the gap closes when
it comes back, and not before. The overview does not say so.
@@ -0,0 +1,69 @@
---
status: resolved
opened: 2026-09-28
located-in: [mesh-controller module.json]
fixed-by: mesh-controller — the control plane's module declares a run-once `migrate` step before its server, which is the shape ADR 0052 prescribes for exactly this. A step's record of having run is the digest of its declaration and the image is part of that digest, so a new build of the control plane re-runs it; and because a run-once step gates what the declaration places after it, a migration that fails stops the new server from starting at all rather than letting it run against a schema it does not have.
amended-design:
---
# 133 — The control plane's schema is migrated at birth and never again
## What was observed
On 2026-09-28 at 08:17 the control plane was replaced, by the mesh's own upgrade path, with a build
whose code writes a column that a migration **in that same build** creates. Nothing ran the migration.
For the next three quarters of an hour the mesh built things and recorded none of them. Every build
answered:
> ERROR: column "built_contexts" of relation "build" does not exist (SQLSTATE 42703)
and that sentence went only to whoever happened to be waiting on a build's reply. The overview kept
saying the mesh was fine. The builds themselves worked — images were built and published — so the
registry filled up with artifacts the mesh has no record of, and the graph stopped learning without
anything saying so.
The schema was created once, at genesis, by an action in the foundation bundle that runs the same
binary's `migrate`. Nothing runs it again. The mesh has updated its own control plane many times since
that bundle, and every one of those updates carried whatever migrations the new build brought and
applied none of them. This is the first time a build needed one.
## Why it matters beyond this instance
**The schema and the code that needs it ship as one artifact and are applied by two mechanisms, only
one of which is automatic.** A module's version is atomic everywhere else in the mesh — the manifest,
the image and what the machine runs move together. Its schema did not, so "the mesh updates itself on
a push" was true of the code and false of what the code needs.
**The failure is quiet exactly where quiet is worst.** A build that cannot be recorded is a build that
happened and left no trace, which is the fault [issue 050](../050-the-catalogue-knows-nothing-built-before-it/00-report.md)
and [issue 131](../131-nothing-tells-the-mesh-a-source-moved/00-report.md) are both about. The mesh
has three mechanisms for noticing a module is behind its source and none for noticing that what it
recorded was refused.
**The shape was already decided, and the control plane was the one module that did not use it.**
[ADR 0052](../../02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md) says a run-once container is a
step the host runs to completion before whatever the declaration places after it, and names migrating
a schema as the case it exists for. The genesis code's own comment says a manifest may name its image
in more than one resource — "a migrate step beside the server". The control plane's manifest had no
such step; it went straight from a state directory to the server.
## What is still true
**Additive migrations are load-bearing, not a style preference.** The step runs before the *new*
server starts, which means the old binary briefly runs against the new schema. A migration that
removes or renames something would break the running control plane in the window between the two.
**The mesh now has two shapes for one problem.** The catalogue module migrates its own schema in its
own code when it starts; the control plane migrates in a step the host gates on. Both work and the
reasons differ — a module that owns its store entirely can do it at start, while a step is visible in
the declaration and refuses to let a broken upgrade serve. Which one the mesh should standardise on is
a decision, not a fix, and it is not made here.
## Open questions
- Should a module be refusable at registration when it ships migrations and declares no step and no
other way to apply them? The mesh can see both halves.
- Should a record the store refuses reach the overview? Today the only reader of that failure is
whoever asked for the thing that failed, and for an event arriving on the bus there is no such
person.