Records the manual update process (module moved -> build -> reconcile; and the breaking-change freeze/re-register recovery), and the two things that make self-update more than a webhook: the build-on-push trigger is currently HAL's (hal-gitea-tools on :9877), a retirement gap the mesh must replace with its own forge-webhook trigger wired to every repo including mesh-controller; and the builder validates manifests too, so a breaking change couples controller + builder + manifests + hosts, and renaming the builder's own seat deadlocks its rebuild. Names the transition discipline (accept old+new for one release) that self-update needs so a push does not auto-freeze.
138 lines
8.3 KiB
Markdown
138 lines
8.3 KiB
Markdown
---
|
||
layer: to-be
|
||
status: proposed
|
||
code: []
|
||
updated: 2026-09-27
|
||
decisions:
|
||
- 02-DECISIONS/0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md
|
||
- 02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md
|
||
---
|
||
|
||
# 30 — The mesh updates itself on a push
|
||
|
||
**Today the mesh does not update itself; a person drives the pipeline by hand, and one class of
|
||
change freezes it.** A code change lands in `mesh-controller` or `mesh-catalog`, and getting it onto
|
||
the machines is a sequence somebody types. The predecessor's pipelines rebuilt and redeployed on a
|
||
push without anyone watching; the successor should too. This records the process as it is done by
|
||
hand now — so it can be read, and then coded — and the two things that make it more than "add a
|
||
webhook".
|
||
|
||
## The process, as done by hand
|
||
|
||
**An ordinary (non-breaking) change** — new module code, a bug fix, a manifest tweak that changes no
|
||
seat or schema:
|
||
|
||
1. `module moved <module> <commit>` — tell the mesh its source advanced (the controller repo has no
|
||
trigger, so this is manual; the catalogue's webhook does it automatically — see below).
|
||
2. `build --behind` (or `build <repo> [--ref] [--path <subdir>]`) — the build machine rebuilds and
|
||
records the new image.
|
||
3. The mesh **reconciles on its own**: the module's declaration now names the new image, the next
|
||
push/heartbeat sends it, and the host swaps the container. For the control plane this is a
|
||
self-upgrade — the running controller composes its own new image and the host replaces it. No
|
||
restart is typed.
|
||
|
||
**A breaking change** — a manifest schema the controller parses differently (a fact's shape, a
|
||
seat's name), where the new control plane cannot read the manifests the old one stored:
|
||
|
||
4. Land the code (controller + catalogue together — they are one change).
|
||
5. Rebuild + deploy the new controller (steps 1–3). **The moment it is live it refuses the
|
||
still-old-shape stored manifests, and composition freezes for every node that runs an affected
|
||
module.** Running services are untouched; only new declarations stop.
|
||
6. **Re-register each affected manifest under the new shape**, which the *new* controller accepts —
|
||
`module add <file> -source <repo> -ref <ref> -commit <commit>`. This writes the manifest to the
|
||
store without a build, so it is the fast way to lift the freeze. (The controller container is
|
||
distroless: `docker cp` the file to the container root `/x.json`; `/tmp` does not exist; the
|
||
root filesystem is writable. The file is lost when the container is recreated on the next image
|
||
swap, so copy it *after* the swap.)
|
||
7. `push --behind`, then verify `status` is clean and `seats` (or the relevant surface) shows the
|
||
new shape held by the right holders.
|
||
|
||
The freeze in a breaking change has been paid three times in one session (a fact-shape change, the
|
||
`/etc/hosts` region, a seat rename); each time it lasted seconds and no service dropped. It is
|
||
recoverable, but it is not something a push should trigger unwatched — which is the crux of what
|
||
automating this must solve.
|
||
|
||
## Why it is more than "add a webhook"
|
||
|
||
### 1. The trigger today is HAL's, not the mesh's
|
||
|
||
Build-on-push works for the catalogue because its repository has a Gitea webhook pointing at
|
||
`http://host.docker.internal:9877/webhook/gitea` — and **that receiver is `hal-gitea-tools.service`**
|
||
(`~/.hal/modules/hal/gitea/tools/server.js`), a *predecessor* component. The nox builder consumes
|
||
build work; it does not receive Git events. So the mesh's own build pipeline currently rides on a
|
||
HAL service, and:
|
||
|
||
- the `mesh-controller` repository was never wired to it, which is why the control plane is the one
|
||
thing that does **not** self-update — every controller deploy this session was `module moved` +
|
||
`build` by hand;
|
||
- when HAL is retired, build-on-push stops for the whole mesh.
|
||
|
||
**The mesh needs its own forge-webhook→build trigger**, a nox component (a module, and likely a
|
||
seat — `mesh-forge-trigger` or folded into the git seat's holder) that receives Git events and turns
|
||
them into build work over the broker, for **every** repository including `mesh-controller`. Replacing
|
||
`hal-gitea-tools` is the concrete first build. Its logic already exists to copy: match the pushed
|
||
repository (and changed paths, for a monorepo like the catalogue) against the build-context
|
||
repository of every registered module, and rebuild the matches.
|
||
|
||
### 2. The builder validates too — and a breaking change deadlocks it
|
||
|
||
The build machine embeds the same catalogue package the controller does, so **it validates a
|
||
manifest against its own compiled-in seat/schema set**. A breaking change therefore couples *four*
|
||
things, not two: the controller, the **builder**, every affected manifest, and every node's host.
|
||
This session's seat rename rebuilt the controller but not the builder, and the stale builder then
|
||
refused every manifest claiming a renamed seat.
|
||
|
||
Worse, one rename **deadlocked** the builder: the build machine's own seat was renamed
|
||
(`the-build-machine` → `mesh-build-machine`). To refresh the builder you must build it; to build it
|
||
the *running* (old) builder must accept the new builder's manifest — which claims the new name it
|
||
does not know. The old builder cannot build the new builder. Escapes:
|
||
|
||
- **Never rename a seat whose holder validates manifests** in an ordinary pass — the build machine's
|
||
seat belongs with the deferred delivering seats (ADR 0121). Reverting `mesh-build-machine` to
|
||
`the-build-machine` (deferred) lets the old builder build the new builder, which then knows the
|
||
new names.
|
||
- Or bootstrap a new builder image **out of band** (build locally, publish to the registry, register
|
||
the module at that digest), the way genesis loads the first builder — bypassing the old builder's
|
||
validation once.
|
||
|
||
Either way, self-update for breaking changes needs a **transition discipline** so a push does not
|
||
auto-freeze: the new control plane (and builder) should accept the *old and new* shape together for
|
||
one release — deprecated aliases in the seat set, a schema that reads both — then a later release
|
||
drops the old. With that, a breaking change rolls out on a push like any other: everything reads
|
||
both, the manifests migrate, the compatibility is removed. Without it, self-update would simply
|
||
automate the freeze.
|
||
|
||
## What to build
|
||
|
||
- **A nox forge-webhook trigger** (replaces `hal-gitea-tools`): receives Git events for every mesh
|
||
repository, dispatches build work to the builder over the broker, and records `module moved`
|
||
automatically. Wire `mesh-controller` to it so the control plane self-updates like everything else.
|
||
- **A transition discipline for breaking changes**: the control plane and builder accept old+new for
|
||
one release; the tooling that lands a schema/seat change emits the compatibility shim and the
|
||
follow-up that removes it. This is what makes step 4–7 above safe to trigger unwatched.
|
||
- **Config/package modules need no builder** — `module add` registers their manifest directly
|
||
(this is how the uplink managers and the re-registrations above were done). Only image-bearing
|
||
modules need the build machine, which narrows what the deadlock above can block.
|
||
|
||
## Why now, and why not yet
|
||
|
||
**Why it matters:** self-update is the difference between a mesh a person maintains by typing
|
||
pipeline steps and one that maintains itself, and it is a stated goal (parity with the predecessor's
|
||
pipelines). The HAL trigger dependency also makes it a retirement blocker: build-on-push dies with
|
||
HAL.
|
||
|
||
**Why not reflexively:** the trigger is a new component with the broker and forge in its blast
|
||
radius, and the transition discipline changes how every breaking change is written. Both should be
|
||
designed, not bolted on beside a freeze. The manual process above is the interim, and it works.
|
||
|
||
## References
|
||
|
||
- [ADR 0121](../../02-DECISIONS/0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md)
|
||
— the seat rename whose migration and builder deadlock this record is drawn from
|
||
- [ADR 0120](../../02-DECISIONS/0120-a-roster-fact-carries-its-format-as-a-template.md) — the
|
||
fact-shape change that first showed the breaking-change freeze
|
||
- `hal-gitea-tools.service` (`~/.hal/modules/hal/gitea/tools/server.js`) — the predecessor webhook
|
||
receiver on `:9877` the mesh currently rides on
|
||
- mesh-controller `cmd/mesh-builder` (the build machine), `internal/catalogue` (the seat/schema
|
||
validation the builder shares with the controller)
|