Author SHA1 Message Date
mesh-admin 98d94ef71e Merge pull request 'Design 28: 5.5 done, the mesh has one bus; issue 131 resolved' (#156) from design/28-one-bus-issue-131-resolved into main 2026-09-28 01:59:41 +00:00
jschoubben 4e13280604 Design 28: 5.5 done, the mesh has one bus; issue 131 resolved
The AMQP transport is gone from the control plane and the hosts (mesh-controller #112,
mesh-host #39). On the way: no build had ever recorded its bases, so every bases-first order
walked an empty graph; the builder now reports what it was handed and the graph is read from
builds (mesh-controller #113/#114). Issue 131 is resolved by the forge module's merge event,
the control plane following it, and those edges.
2026-09-28 03:59:39 +02:00
mesh-admin 8783a13448 Merge pull request 'Design 28: the mesh runs on the new bus' (#155) from design/28-the-mesh-runs-on-nats into main 2026-09-28 00:40:29 +00:00
jschoubben a31cfcf461 Design 28: 5.3 is built and was used for the hand-over 2026-09-28 02:28:06 +02:00
jschoubben 75b3861911 Design 28: the mesh runs on the new bus
Tasks 4.3, 5.2 and 5.4 are done as of 2026-09-28 02:25: every machine reports on
the new bus, the seat is held by the module that provides it, the old broker is
unassigned and forgotten, and every credential was minted afresh at the end.

5.2 records what it took, in the order it was found and each fixed on the trunk
before the next step, and how the bootstrap loop was broken once, by hand.
2026-09-28 02:27:49 +02:00
2 changed files with 152 additions and 33 deletions
+45 -33
View File
@@ -5,7 +5,7 @@ code:
- mesh-catalog modules/nats
- mesh-controller internal/catalogue
- mesh-lab scenarios
updated: 2026-09-27
updated: 2026-09-28
decisions:
- 02-DECISIONS/0116-the-bus-is-built-in-five-steps.md
- 02-DECISIONS/0106-the-bus-is-nats.md
@@ -535,7 +535,7 @@ it, and the beds that need a mesh living on NATS can finally run.
The outcome carries the module name, because only the manifest says what was built and one
message now has three readers. A failed build names none: it produced no module version, and
the catalogue would otherwise place something that was never made.
- [~] 4.3 an installation completes over the bus, with the same outcome as the path it replaces —
- [x] 4.3 an installation completes over the bus, with the same outcome as the path it replaces —
**the installer can raise it**: a foundation template that stands up the server, writes the
server's own settings and the mesh's first user list beside them, and starts a controller
reaching the new bus. What remains is running it, which is 4.1's bed.
@@ -655,42 +655,35 @@ healthy while reacting to nothing.
- [ ] 5.1 the cutover bed: a mesh on AMQP with a predecessor stand-in on the deprecated broker
moves its bus in one rollout, every node reporting on NATS afterwards, the stand-in's own
client still connected throughout
- [~] 5.2 the rollout: accounts composed, then the controller, every host and every runtime
- [x] 5.2 the rollout: accounts composed, then the controller, every host and every runtime
together; every node confirmed heard before AMQP stops.
**The readiness half is in and is the half worth having.** The move takes every node at once, so
there is nothing to inspect afterwards and no half to roll back — either the mesh was ready or it
was not. `rollout check` answers that from records, with one dial: is a bus answering, does a
machine hold the seat, has it been sent the composed user list, does every machine and every
module that speaks have a credential. Each missing thing names its own next step, because "not
ready" that cannot be acted on is not an answer at the point where the next step is irreversible.
**Done 2026-09-28, 02:25.** Every machine reports on the new bus, the seat is held by the
module that provides it, the old broker is unassigned and forgotten, and every credential was
minted afresh at the end because two had been printed on the way. What it took, in the order
it was found, each fixed on the trunk before the next step: the control plane's `serve` and
`push` never selected the new transport (task 4.3, open until then); a machine's user was
granted neither the asking nor the delivery of its own consumer; the account had no JetStream
of its own; the control plane's client verified the bus's certificate by name instead of
pinning it; the seat table's rows carried no protocol, so no role's work queue was raised; the
build machine decided its bus from a variable its container never received; and a rotation
put new hashes on the bus before three machines had received their new memberships — which
is why there is now `rollout hand <node>` and a host adopts a delivered membership at start.
The bootstrap loop — a bus that can only be raised by a declaration that can only arrive
over that bus — was broken once, by hand: the mesh's own composed configuration started the
server, and the controller binary was run on the node directly until the managed container
could be rebuilt over the bus it was on.
**A machine with no credential is what must stop it.** It keeps running, cannot come back, and
afterwards there is no bus to tell it anything over.
The move itself is deliberately not written yet, and the command says so rather than pretending:
it waits on the check having been run against a real mesh. Writing the irreversible half before
the question it depends on has ever been asked of something real is how the plan's own rule about
beds gets broken by another route.
> **What this costs if it goes wrong, measured rather than assumed.** Nothing in a served
> request's path goes over the mesh's own bus: modules serve from their own containers. What a
> failed move costs is the mesh's ability to *change* anything — pushes, tool calls, new
> provisioning — until it is finished or undone. That is worth knowing before rather than
> after, and it is why the operator's "as long as my services keep running" is a reasonable
> position rather than a gamble. **Measured on 2026-09-27**, when a seat emptied itself
> mid-change: 52 containers stayed up and the broker never stopped; the control plane
> crash-looped for two hours and nothing could be deployed until it was repaired by hand.
> An earlier version of this note said the old broker stays as an ordinary provider of
> `amqp` ([ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md)); that is withdrawn by [ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md) — see 5.4.
- [ ] 5.3 **the seat changes hands as one act.** A command takes a seat and the assignment taking it
- [x] 5.3 **the seat changes hands as one act.** A command takes a seat and the assignment taking it
over, and the seat is never empty in between — the emptiness is the outage of 2026-09-27, when
the control plane, which finds its own bus through this seat, lost the address and looped.
Today only `seat rename` exists. This is what 5.2 uses to move `mesh-broker` from the old
**Built 2026-09-27** (`seat_holding`, migration 0039; design 26 says how it is checked), and used
live the next night to hand `mesh-broker` from the old broker's assignment to the new one's. This
is what 5.2 uses to move `mesh-broker` from the old
broker's assignment to the new one's, and it is built first ([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)).
- [ ] 5.4 **the old broker and everything that named AMQP leave the mesh** ([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md),
- [x] 5.4 **the old broker and everything that named AMQP leave the mesh** ([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md),
superseding [ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md)): the two modules that
required `amqp` are removed, the broker's module is unassigned and removed, registration refuses
required `amqp` are removed, the broker's module is unassigned and removed (**done 2026-09-28**; the predecessor's own tooling, which rode the same adopted broker, went dark with it, as [ADR 0130](../../02-DECISIONS/0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md) accepted), registration refuses
a manifest that provides or requires `amqp`, and a whole-catalogue check asserts none does. Not
a retirement condition — a decision, taken, with the operator's "I don't care if the predecessor
breaks" on record ([ADR 0130](../../02-DECISIONS/0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md)).
@@ -701,8 +694,27 @@ healthy while reacting to nothing.
> shutting it down ends the path that reaches this installation's machines from a workstation.
> The rollout is driven from the node, or before the broker stops — a sequencing constraint on
> 5.2, not an afterthought.
- [ ] 5.5 **the AMQP transport is deleted from the control plane and the hosts**, and the variable
that selected a transport is refused at start as unknown. One bus, nothing to select ([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)).
- [x] 5.5 **the AMQP transport is deleted from the control plane and the hosts**. One bus, nothing
to select ([ADR 0131](../../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)).
**Done 2026-09-28.** The control plane's old consume loop, build request, tool call, management
API and account scoping went, and the host's old dialling and enrolment paths with them; a
membership or token naming any other bus is refused before anything is sent. Nothing selects a
transport any more: the variable that once did (`MESH_BUS_NATS`) now only names where the
control plane reads its own bus credential, the way any module reads a secret. **Checked by the
build**: neither repository's module file names the AMQP client library, so a line that still
used it would not compile. The store-window guarantee ([issue 083](../../04-ISSUES/083-other-control-messages-are-lost-while-the-store-restarts/00-report.md))
is tested against a bus-less fake rather than the old transport's memory, which is what let
that memory go — the one thing it did that the stream does not (superseding a held report) is
the staleness check on the message itself (design 25 §3).
Found on the way: **no build had ever recorded what it stood on.** A recipe reads its base from
a build argument, so the digest was never in the file the builder derived edges from, and every
order that says *bases first* — `build --on`, `build --behind`, the merge follow-up of
[issue 131](../../04-ISSUES/131-nothing-tells-the-mesh-a-source-moved/00-report.md) — walked a
graph with no edges. The builder now reports the bases it was handed, the control plane records
them by artifact path, and the graph is read from the newest build of each module — a recorded
manifest carries no `build.on`, so the edge is derived from the build or it does not exist.
> **The old 5.4 note is history.** It recorded that a retirement *condition* was wrong from
> [ADR 0127](../../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md) onward, which framed the old broker
@@ -0,0 +1,107 @@
---
status: resolved
opened: 2026-09-27
located-in: [mesh-catalog modules/gitea, mesh-controller cmd/mesh-controller, mesh-controller internal/builder]
fixed-by: mesh-catalog #124 — the forge module watches for merged pull requests and emits `pull.merged` with the merge commit and the clone address; mesh-controller #110/#111 — the control plane follows that event on the bus, marks every module built from that repository and branch as moved, and builds them bases first, stopping when a base fails; mesh-controller #113/#114 — a build records the bases it was handed and the graph is read from builds, without which "bases first" had no edges to order by.
amended-design: 03-DESIGN/01-to-be/28-building-the-bus.md
---
# 131 — Nothing tells the mesh a source moved, and it reports itself current anyway
## What was observed
Six changes were merged to the trunk of six repositories in one sitting. The build machine built
nothing. Its last build, minutes before the first merge, was still the one it reported; no build was
requested, refused or failed, because none was ever asked for.
Asked afterwards what was wrong, the mesh said:
> 4 machine(s), all doing what they were told, all heard from, running what the mesh would send them,
> and every module current with its source
Every one of the six had moved. The last clause was false for all of them, and it is the clause a
person reads to decide whether there is anything to do.
## Why it matters beyond this instance
**The mesh learns a source moved by being told, and there is no longer anything to tell it.** The
command exists — a person names the module and the commit — and so does the question the overview
answers. What is missing is whatever used to connect the two. One repository still carries a forge
webhook aimed at a port; the rest carry none, and the port belongs to a different service than the
one the arrangement implies. So the state is not "the trigger is broken" but "there is no trigger,
and nothing says so".
**A wrong answer is worse here than no answer.** "Every module current with its source" is
indistinguishable, to a reader, from a mesh that has genuinely caught up. The overview is built to be
the thing you check instead of checking by hand, so a confident false negative removes the habit that
would otherwise have caught it. Nothing in the mesh is at fault for being out of date — it is at
fault for saying it is not.
**It is also why "current with its source" cannot be a stored fact.** The mesh compares what it built
against what it was last told the source was, and calls that agreement. Two facts agreeing tells you
nothing when both come from the same place.
## The intended shape, which is decided and not built
The forge emits what happened to it — a pull request merged — and the build machine reacts by
building what that commit affects. That keeps the forge ignorant of the build system and the build
machine ignorant of the forge's internals, which is the same argument
[ADR 0126](../../02-DECISIONS/0126-a-module-declares-its-own-seats.md) makes for addressing an event
to its emitter: the merge is a fact about the forge, and what should be rebuilt because of it is not
the forge's business to know.
The forge's module already declares the event. The build machine declares that it consumes nothing.
## What the trigger cannot be
**Not one build per changed module.** The modules form a graph: several are built from one
repository, and some are the base another is compiled on — a runtime image, a compiler base, a
repository whose context a second module builds from. Firing a build for each changed module
independently would start work that cannot succeed yet and produce a failure per dependent, for one
cause.
Observed while catching the mesh up by hand on 2026-09-27: a compiler base had to move before
anything compiled against it could build, and when it failed, the right behaviour was for its
dependents to wait rather than each fail the same way. Fifteen modules shared the cause. A trigger
that reports it fifteen times has buried it.
So whatever reacts to the forge's event resolves what changed into an order, builds the bases first,
and holds a dependent while its base is unbuilt or failed. That is a larger thing than "rebuild what
the commit touched", and knowing it now is cheaper than discovering it from fifteen identical
failures.
## Open questions
- Is "the source moved" still a thing a person can assert by hand once the event path exists, or does
the hand-operated form become the thing that made this failure possible?
- Which commit does the build machine act on — the merge, or each commit it brought — and what does
it do when several arrive for one module at once?
- How does the overview stop being able to lie? Comparing what was built against what was recorded
will always agree. Whether the trunk has moved is a question only the forge can answer, so either
the overview asks it, or it stops claiming to know.
- Does this want to be the same mechanism as the build request on the bus
([ADR 0129](../../02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md)), or does it sit in
front of it?
## What was done (2026-09-28)
The shape above was built as described: the forge's module emits the merge, the control plane
consumes it, and nothing on either side knows the other's internals. The build is asked for the
merge commit, not each commit the merge brought — the trunk moved once, to one place. Several merges
for one module arriving in a row are followed in turn, each moving the recorded source to its own
commit, so the last one to arrive is the one the mesh ends up built from.
The hand-operated form stays. `module moved` is how a source is recorded without a forge — a module
built from a repository elsewhere, or a mesh whose forge module is down — and it is the same act the
event performs, so the two cannot disagree about what "moved" means.
**Bases first needed edges, and there were none.** The order this report asked for was written and
walked a graph that no build had ever recorded: a recipe reads its base from a build argument, so the
digest was never in the file the builder read edges from. A build now reports what it was handed, the
control plane records it by artifact path, and the order is read from each module's newest build.
**What still can lie.** The overview compares what was built against where it was last told the
source is; the forge's event is now what moves that mark, so it is right for as long as the forge
module was listening. A merge made while that module was down is a merge the mesh does not know of
until the module next polls — it announces what merged since it last looked, so the gap closes when
it comes back, and not before. The overview does not say so.