Both lines of work numbered from the same point, so four decision records and one design document existed twice with different content. The trunk keeps its numbers and this branch yields — the only rule that scales, because the trunk's are already cited by what merged before them. 0117 the bus is the only broker -> 0125 0118 a module declares its own seats -> 0126 0119 amqp is a provision, not the bus -> 0127 0120 the mesh bus is required -> 0128 0123 a seat carries its role's protocol -> 0129 0124 the predecessor is ending -> 0130 design 29, what a module declares -> design 32 Applied to the code repositories too, because a stale reference is worse when numbers collide than when they dangle: the reader lands on a real record that decided something else. Two reconciliations the merge forced, both real: **0110 was marked wholly superseded and was not.** Its successor says in as many words that everything 0110 decided about what a seat *is* stands untouched — and two records that landed on the trunk rest on exactly that part. So it is accepted again, extended rather than replaced, with a note saying which of its claims moved and where. **A seat's protocol becomes columns, not fields.** The trunk moved the seat set out of compiled code into a table the controller owns. This branch had added what a role accepts, emits and serves to the Go slice. The decision is unaffected and the mechanism is better for it: giving a role a protocol is now a write rather than a rebuild, which is the trunk's own argument applied to what this branch added. One check still fails and it fails on main too: a record resting on ADR 0112 while that is still 'proposed'. Left alone — it is not this merge's to answer.
8.3 KiB
layer, status, code, updated, decisions
| layer | status | code | updated | decisions | ||
|---|---|---|---|---|---|---|
| to-be | proposed | 2026-09-27 |
|
30 — The mesh updates itself on a push
Today the mesh does not update itself; a person drives the pipeline by hand, and one class of
change freezes it. A code change lands in mesh-controller or mesh-catalog, and getting it onto
the machines is a sequence somebody types. The predecessor's pipelines rebuilt and redeployed on a
push without anyone watching; the successor should too. This records the process as it is done by
hand now — so it can be read, and then coded — and the two things that make it more than "add a
webhook".
The process, as done by hand
An ordinary (non-breaking) change — new module code, a bug fix, a manifest tweak that changes no seat or schema:
module moved <module> <commit>— tell the mesh its source advanced (the controller repo has no trigger, so this is manual; the catalogue's webhook does it automatically — see below).build --behind(orbuild <repo> [--ref] [--path <subdir>]) — the build machine rebuilds and records the new image.- The mesh reconciles on its own: the module's declaration now names the new image, the next push/heartbeat sends it, and the host swaps the container. For the control plane this is a self-upgrade — the running controller composes its own new image and the host replaces it. No restart is typed.
A breaking change — a manifest schema the controller parses differently (a fact's shape, a seat's name), where the new control plane cannot read the manifests the old one stored:
- Land the code (controller + catalogue together — they are one change).
- Rebuild + deploy the new controller (steps 1–3). The moment it is live it refuses the still-old-shape stored manifests, and composition freezes for every node that runs an affected module. Running services are untouched; only new declarations stop.
- Re-register each affected manifest under the new shape, which the new controller accepts —
module add <file> -source <repo> -ref <ref> -commit <commit>. This writes the manifest to the store without a build, so it is the fast way to lift the freeze. (The controller container is distroless:docker cpthe file to the container root/x.json;/tmpdoes not exist; the root filesystem is writable. The file is lost when the container is recreated on the next image swap, so copy it after the swap.) push --behind, then verifystatusis clean andseats(or the relevant surface) shows the new shape held by the right holders.
The freeze in a breaking change has been paid three times in one session (a fact-shape change, the
/etc/hosts region, a seat rename); each time it lasted seconds and no service dropped. It is
recoverable, but it is not something a push should trigger unwatched — which is the crux of what
automating this must solve.
Why it is more than "add a webhook"
1. The trigger today is HAL's, not the mesh's
Build-on-push works for the catalogue because its repository has a Gitea webhook pointing at
http://host.docker.internal:9877/webhook/gitea — and that receiver is hal-gitea-tools.service
(~/.hal/modules/hal/gitea/tools/server.js), a predecessor component. The nox builder consumes
build work; it does not receive Git events. So the mesh's own build pipeline currently rides on a
HAL service, and:
- the
mesh-controllerrepository was never wired to it, which is why the control plane is the one thing that does not self-update — every controller deploy this session wasmodule moved+buildby hand; - when HAL is retired, build-on-push stops for the whole mesh.
The mesh needs its own forge-webhook→build trigger, a nox component (a module, and likely a
seat — mesh-forge-trigger or folded into the git seat's holder) that receives Git events and turns
them into build work over the broker, for every repository including mesh-controller. Replacing
hal-gitea-tools is the concrete first build. Its logic already exists to copy: match the pushed
repository (and changed paths, for a monorepo like the catalogue) against the build-context
repository of every registered module, and rebuild the matches.
2. The builder validates too — and a breaking change deadlocks it
The build machine embeds the same catalogue package the controller does, so it validates a manifest against its own compiled-in seat/schema set. A breaking change therefore couples four things, not two: the controller, the builder, every affected manifest, and every node's host. This session's seat rename rebuilt the controller but not the builder, and the stale builder then refused every manifest claiming a renamed seat.
Worse, one rename deadlocked the builder: the build machine's own seat was renamed
(the-build-machine → mesh-build-machine). To refresh the builder you must build it; to build it
the running (old) builder must accept the new builder's manifest — which claims the new name it
does not know. The old builder cannot build the new builder. Escapes:
- Never rename a seat whose holder validates manifests in an ordinary pass — the build machine's
seat belongs with the deferred delivering seats (ADR 0121). Reverting
mesh-build-machinetothe-build-machine(deferred) lets the old builder build the new builder, which then knows the new names. - Or bootstrap a new builder image out of band (build locally, publish to the registry, register the module at that digest), the way genesis loads the first builder — bypassing the old builder's validation once.
Either way, self-update for breaking changes needs a transition discipline so a push does not auto-freeze: the new control plane (and builder) should accept the old and new shape together for one release — deprecated aliases in the seat set, a schema that reads both — then a later release drops the old. With that, a breaking change rolls out on a push like any other: everything reads both, the manifests migrate, the compatibility is removed. Without it, self-update would simply automate the freeze.
What to build
- A nox forge-webhook trigger (replaces
hal-gitea-tools): receives Git events for every mesh repository, dispatches build work to the builder over the broker, and recordsmodule movedautomatically. Wiremesh-controllerto it so the control plane self-updates like everything else. - A transition discipline for breaking changes: the control plane and builder accept old+new for one release; the tooling that lands a schema/seat change emits the compatibility shim and the follow-up that removes it. This is what makes step 4–7 above safe to trigger unwatched.
- Config/package modules need no builder —
module addregisters their manifest directly (this is how the uplink managers and the re-registrations above were done). Only image-bearing modules need the build machine, which narrows what the deadlock above can block.
Why now, and why not yet
Why it matters: self-update is the difference between a mesh a person maintains by typing pipeline steps and one that maintains itself, and it is a stated goal (parity with the predecessor's pipelines). The HAL trigger dependency also makes it a retirement blocker: build-on-push dies with HAL.
Why not reflexively: the trigger is a new component with the broker and forge in its blast radius, and the transition discipline changes how every breaking change is written. Both should be designed, not bolted on beside a freeze. The manual process above is the interim, and it works.
References
- ADR 0121 — the seat rename whose migration and builder deadlock this record is drawn from
- ADR 0128 — the fact-shape change that first showed the breaking-change freeze
hal-gitea-tools.service(~/.hal/modules/hal/gitea/tools/server.js) — the predecessor webhook receiver on:9877the mesh currently rides on- mesh-controller
cmd/mesh-builder(the build machine),internal/catalogue(the seat/schema validation the builder shares with the controller)