Author SHA1 Message Date
mesh-admin 017b1401e4 Merge pull request 'ADR 0135: a module version prepares its state before it runs' (#161) from decision/0133-0134-migrations-and-deploy-facts into main 2026-09-28 09:57:47 +00:00
jschoubben 476cda417d ADR 0135 supersedes 0133: a module version prepares its state before it runs
Two faults in 0133, both caught on review. It put the declaration on a container — one resource kind
the host applies — so every author would restate the machine's arrangement and a module's own
lifecycle would be tied to how its artifact happens to run. A module declares entrypoints for its
tools and its provisioner; preparing its state is the same vocabulary and nothing about a runtime.

And it derived the scope from the machine, which the facts already answer: a consumer is a module on
a machine (issue 022, migration 0015), so what the mesh provisions is per consumer. A module on three
machines has three databases, there is no shared state to race over, and the lock obligation 0133
invented was for a situation the mesh does not produce. The level question HAL answered with stages
dissolves — the scope of preparation is the scope of the state, and the mesh knows it.

0133 keeps its reasoning and gains a pointer; design 32 and issue 133 name the live record.
2026-09-28 11:57:16 +02:00
mesh-admin 74ba3ff1d4 Merge pull request 'ADR 0133 and 0134: who runs migrations, and the mesh saying what it applied' (#160) from decision/0133-0134-migrations-and-deploy-facts into main 2026-09-28 09:45:49 +00:00
jschoubben 891c7a945e ADR 0133 and 0134: who runs migrations, and the mesh saying what it applied
0133 — a module owns its migrations and the mesh owns when they run. A container declares what must
run before it; the mesh derives the gated step from the resource it precedes, so the image, the
environment and the credentials come from the one place they are described. The module owns the SQL,
the dialect and the lock; the mesh owns the moment and refuses to start a version whose step failed.
Per node, with no level: a step that ran once somewhere leaves every other machine ungated, and
'once, mesh-wide' is what holding a seat already means.

0134 — the pipeline is observable from a merge to an artifact and goes dark at the machine. What a
node now runs, and what it refused, become facts under the control plane's own seat, emitted when
what a machine runs changes rather than on every convergence pass.

Design 32's lifecycle carries both; issue 133 points at them as what ends the matter it opened.
2026-09-28 11:45:47 +02:00
mesh-admin 83cbeb8db8 Merge pull request 'Issue 133: the control plane's schema is migrated at birth and never again' (#159) from issue/133-the-control-plane-migrates-before-it-serves into main 2026-09-28 08:27:32 +00:00
jschoubben ab6db9369b Issue 133: the control plane's schema is migrated at birth and never again
The mesh replaced its own control plane with a build carrying a migration, applied none of it, and
then recorded no build for three quarters of an hour while saying everything was fine. ADR 0052
already prescribes the shape — a run-once step that gates the server — and the control plane was the
one module that did not use it.
2026-09-28 10:27:30 +02:00
mesh-admin 84571b4825 Merge pull request 'ADR 0132: a seat carries the tools its holder must serve' (#158) from decision/0132-a-seat-carries-the-tools-its-holder-must-serve into main 2026-09-28 08:17:01 +00:00
jschoubben d57196102d ADR 0132: a seat carries the tools its holder must serve
A role's tools belong to the role, not to whichever module holds it today: the seat declares them
with their schemas, serving them is a condition of occupying the seat, and what the mesh can do
becomes a read of its own records rather than a question nothing answers. A module keeps its own
tools — the same module may run without the seat, and then only its own name is true.

Design 33 follows: the three families, addressing a node-scoped seat, discovery, and what serves
this to an agent.
2026-09-28 10:16:59 +02:00
mesh-admin e6402cf777 Merge pull request 'Issue 132: a module can be recorded without the directory it lives in' (#157) from issue/132-a-module-can-be-recorded-without-its-directory into main 2026-09-28 07:20:07 +00:00
jschoubben 4f0d144833 Issue 132: a module can be recorded without the directory it lives in
Nine modules could not be rebuilt: their record named the repository and no directory, so every
build looked for a manifest at a repository root that has never had one. Resolved by mesh-controller
— `module add` takes the directory and the forge, and the rule is checked rather than described.
2026-09-28 09:20:05 +02:00
mesh-admin 98d94ef71e Merge pull request 'Design 28: 5.5 done, the mesh has one bus; issue 131 resolved' (#156) from design/28-one-bus-issue-131-resolved into main 2026-09-28 01:59:41 +00:00
9 changed files with 897 additions and 1 deletions
@@ -0,0 +1,150 @@
---
topic: the mesh
status: accepted
date: 2026-09-28
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md
---
# 132. A seat carries the tools its holder must serve
## Context
[ADR 0129](0129-a-seat-carries-the-protocol-of-its-role.md) gave a seat the protocol of its role in
three parts: the work it accepts, the events it emits, and the verbs it **serves** — request and
reply, awaited. The bus already derives authority from all three: a holder subscribes
`mesh.seat.<seat>.tool.<verb>`, and a module that uses the seat may publish it and nothing else.
**The serving third has never been used.** The mesh defines 14 seats, 8 mesh-scoped and 6
node-scoped. Exactly one carries a protocol at all — the build machine, which accepts `build` and
emits `built`. Not one seat declares a single verb it serves. The mechanism is built, enforced, and
empty.
Meanwhile every tool on the mesh is addressed to a module. Of 72 modules in the catalogue, 45 serve
tools, about 203 of them, each on `mesh.mod.<module>.tool.<name>`. So a caller binds to the module
that happens to hold a role rather than to the role, and replacing that module breaks every caller —
which is the thing seats exist to prevent everywhere else.
**Nothing can say what tools exist.** Measured on 2026-09-28, with the bus carrying the whole mesh: a
workstation client holding an operator credential connected, the bus accepted the account, and
`mesh call gitea.gitea_list_repos` answered with real repositories. The same client's `mesh tools`
found nothing, because it asks `mesh-catalog.catalog_tools` and no module serves that: the catalogue
serves `catalog_modules`, `catalog_module`, `catalog_provides`, `catalog_dependents` and
`catalog_stale`. An agent can therefore call any tool it already knows the name of and discover none.
MCP's `tools/list` is that same question, so the MCP surface is a working transport over an empty
catalogue.
**And there is nowhere for a tool's definition to live.** A manifest has a `tools` field: 0 of the 45
modules that serve tools fill it. That is not neglect, it is the arrangement failing — the field was
the bus grant's source for what a module may subscribe, and because nothing filled it every module
that served a tool was refused its own subscription on the new bus, live, until the grant was changed
to the module's own namespace. Today a tool's name, description and argument schema exist only in the
module's code.
Two facts about the machinery matter for what follows. A seat's protocol is not in the store: the seat
rows lack the ADR 0129 columns, so the protocol comes from compiled defaults and is merged in when a
row is read. And `seatSubject` is flat — `mesh.seat.<seat>.<kind>.<verb>` with no node in it — so a
node-scoped seat's tool call would reach every node's holder at once, and the holders' queue group
would hand it to whichever answered first.
## Decision
**A seat's protocol carries its tools in full**: the verb, what it does, and the schema of its
arguments and of its answer. The seat is the definition of the role's interface; the holder is an
implementation of it.
**Serving the seat's tools is a condition of holding the seat.** A module that does not serve every
verb the seat declares may not occupy it. This is checked where the other conditions of holding are
checked — registration and handover — and refused by naming the verbs that are missing.
**A role's tools are addressed to the role.** `mesh.seat.<seat>.tool.<verb>` mesh-wide. A node-scoped
seat carries the node in the address, because one subject reaching six machines' holders is not an
address, and the queue group that made it look like one would silently pick a winner.
**A module keeps its own tools, and both exist.** `gitea_list_repos` stays, because gitea can run
without holding the `git` seat — a second forge, an instance kept for one purpose. The module's name
answers *this gitea*; the seat's verb answers *whoever is the forge*. Which of the two a caller wants
is a decision in the running session, not one the mesh makes for it.
**What answers "what tools exist" follows where the definition lives.** A seat's tools are read from
the mesh's own records. A module's own tools are answered by the module, from the code that defines
them. Discovery is therefore a read for the durable half and a question to the running mesh for the
free half.
**A seat's tools are an interface, and change like one.** Additive within a version; a change that
would break a caller takes the version token the subject already has room for (design 29 §8), and the
two run side by side until nothing is bound to the old one.
**The mesh's own verbs are the `mesh-controller` seat's tools.** `status`, `push`, `build`, `assign`
and the rest are a role's interface, not a container's, and the audit point [ADR 0095](0095-the-control-plane-is-the-way-to-ask-a-module.md)
asks for is the seat's holder.
## Options considered
1. **The manifest declares each module's tools.** Rejected. The list is then written twice — in the
manifest and in the code — and a schema in a manifest goes stale silently, which is the worst kind
of wrong for something an agent reads to decide what to call. It is also the arrangement that has
already failed once: the field exists, 0 of 45 modules fill it, and the grant that depended on it
refused every tool subscription on the mesh.
2. **Every runtime answers an introspection call, and something aggregates them.** Rejected as the
shape for a role's tools, kept for a module's own. An aggregator needs permission to publish into
every module's namespace, which is a widening the mesh otherwise gives only to the control plane;
and the answer is only as available as the modules are, so a mesh whose catalogue cannot say what a
role answers while its holder is down cannot plan against it.
3. **The control plane answers everything.** Rejected. It puts a tool surface on the control plane for
tools it does not implement, and makes discovery depend on the one component that must stay
answerable while it is itself being replaced. The mesh's own verbs are its to answer, and it answers
them as the holder of a seat.
4. **Seats only; no module tools.** Rejected. Most modules hold no seat, and inventing a seat per
module to give its tools a home would dilute what a seat is: one holder of a role the mesh needs
exactly one of.
## Consequences
**One capability can have two names, deliberately.** A forge that holds the `git` seat answers both
`mesh.seat.git.tool.list_repos` and `mesh.mod.gitea.tool.gitea_list_repos`. This is the one place the
mesh accepts two names for one thing, because they are answers to different questions and the second
one survives the module not holding the seat. The glossary rule stands everywhere else.
**A seat becomes a contract to implement.** Adding a verb to a seat is a change every holder must
make, and a claim that was valid becomes invalid until it does. That is the point, and it is also the
reason a seat's tools should be few and durable while a module's own stay free.
**Three prerequisites, none of them in place.** The seat's protocol must be in the store rather than in
compiled defaults, or discovery reads a binary rather than the mesh. The protocol must become richer
than a list of verbs, because a verb without a schema is not something an agent can call. And a
node-scoped seat needs the node in its subject before any of its tools can exist.
**Discovery becomes cheap for the half that matters.** What roles the mesh has and what each answers is
a query, with no fan-out and nothing to be up. An agent's authority can then be role-shaped — *the
forge's tools* — rather than a list of module-specific names that changes when a module is replaced.
**The MCP surface belongs inside the mesh.** Once the tools are the mesh's own records, the thing that
serves them to an agent is a module the mesh assigns to the machine where the agent sits, with a
credential the mesh minted and authority derived from what it may call — not a program started by hand
with a credential printed to a terminal.
## How this is checked
- **Holding is refused without the verbs.** The condition sits with the other conditions of holding a
seat, so registration and a handover both refuse a module that does not serve what the seat declares,
and the refusal names the missing verbs. A test per condition, as the other seat conditions have.
- **The grant is derived from the seat, and already is.** A holder's subscription and a user's publish
come from the seat's protocol, so a verb nobody declared is a subject nobody may use, and a verb the
seat declares reaches exactly its holder. The golden composition of the bus's user list is the test
that keeps it honest.
- **Discovery is a read, and is tested as one.** What the mesh answers for a seat's tools equals what
the seat's records declare — no call to a module in the path, so the test needs no running module.
- **A node-scoped seat's subject carries its node**, checked by the same test that checks the subject
table: two nodes holding one node-scoped seat derive two addresses.
## References
- [ADR 0129](0129-a-seat-carries-the-protocol-of-its-role.md) — the protocol this widens
- [ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md) — the role, its work and its events
- [ADR 0095](0095-the-control-plane-is-the-way-to-ask-a-module.md) — a tool call passes one process where an audit belongs
- [ADR 0126](0126-a-module-declares-its-own-seats.md) — an event is addressed to its emitter, for the same reason a role's verb is addressed to its role
- [`03-DESIGN/01-to-be/26-the-seats.md`](../03-DESIGN/01-to-be/26-the-seats.md) — how a seat is held and handed over
- mesh-controller #116, #117, #118 — the grants as they now stand: a module serves its own namespace, the control plane may ask any tool
- Measured 2026-09-28 on the live mesh: an operator credential calling a module's tool over the bus answers; `tools/list` finds nothing
@@ -0,0 +1,159 @@
---
topic: what runs on it
status: superseded
date: 2026-09-28
deciders: jochen
reconstructed: false
extends: 02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md
superseded-by: 02-DECISIONS/0135-a-module-version-prepares-its-state-before-it-runs.md
---
# 133. A module owns its migrations, and the mesh owns when they run
## Context
On 2026-09-28 the mesh replaced its own control plane, through its own upgrade path, with a build
carrying a migration. Nothing applied it. For the next three quarters of an hour every build the mesh
made was refused by the store with one line — *column "built_contexts" does not exist* — which reached
only whoever happened to be waiting on that build's reply. The images were built and published, so the
registry filled with artifacts the mesh has no record of, and the overview went on reporting that every
module was current ([issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md)).
The schema had been created once, at genesis, by an action in the foundation bundle. Nothing ran it
again, through many updates of the control plane since.
**The mechanism to do this right already existed and one module used it wrong.**
[ADR 0052](0052-a-step-that-runs-once-before-a-container.md) makes a run-once container a step the host
runs to completion before whatever the declaration places after it, and names migrating a schema as the
case it exists for. Three facts about how it is used today:
- The control plane's manifest had no step at all. The immediate fix was to write one by hand, and that
hand-written step repeats three environment variables and three volume mounts from the server
resource it precedes — six chances to drift from the thing it prepares.
- Two other modules hand-write the same shape for the same reason: gitea's admin bootstrap and
mosquitto's dynsec seed, each repeating its sibling's image, environment and mounts. One of them
ends in `|| true`, which is a lock implemented as a shrug.
- The catalogue module takes the other road: it migrates its own schema in its own code when it starts.
That failure mode is a crash loop rather than a stop — the catalogue restarted 338 times this
morning on an unrelated start-time failure, and nothing anywhere said the mesh's graph had a gap.
**What the mesh already has, and what HAL needed stages for.** Ordering a provider before its consumer
is `providersFirst`, which topologically orders a node's modules. Ordering within a module is
declaration order, and a run-once container gates everything after it. Remembering that a step has
already run is the digest of its declaration, recorded only after it exits 0
([ADR 0018](0018-a-picture-is-read-from-what-runs.md)) — and the image is part of that digest, so a new
build re-runs it. Three of the four things a stage system provides are therefore already here. The
fourth — that a module has a schema at all — is the only thing missing.
**Nothing in the catalogue ships a migrations directory.** Of 72 modules, none has one; the modules that
migrate do it in their own code. So this is not a decision about where SQL files live. It is a decision
about who runs them and when.
**Two facts bound what is safely expressible.** A node converges toward its own declaration without
waiting on any other node. And of the five modules that run on more than one machine today — dnsmasq,
fail2ban, networking, networkmanager, sshd — not one wants a store; every module with a database is on
exactly one machine.
## Decision
**A container may declare steps to run before it.** The same container, run to completion, with
different arguments, in order, before it starts. The mesh derives the run-once resources from that
declaration, so the image, the environment, the volumes, the network and the credentials come from the
one place they are already described and cannot drift from it.
**A module's migrations are the first user of this, and the module owns them entirely.** The SQL, the
order, the idempotence, the lock, and which dialect it speaks. The mesh never learns that postgres and
mssql differ, because it runs the module's own image with the module's own arguments against the
module's own binding and requires exit 0. A module needing both stores runs one step that does both.
**The mesh owns the moment, and the gate is the guarantee.** Whether a version may serve when its
schema is not there yet is a deployment question, and the mesh is the only thing that can answer it,
because the mesh is what starts the container. A step that fails stops the container it precedes, so
a failed migration is a version that does not serve rather than a version serving against a store it
does not match.
**Per node, and there is no level.** The step runs wherever the module runs. A step that ran "once,
somewhere" would leave every other machine with no gate at all, and additive migrations protect old
code against a new schema, never new code against an old one. The cost is an obligation a migration
runner already carries: a version table and a lock.
**"Once, mesh-wide" is what holding a seat means.** A step that is not idempotent — seeding an
account, sending a notice, taking a backup — belongs to a module that holds a seat, where the mesh
already guarantees one holder, on record, handed over deliberately. That is the answer to the level
question rather than a field that has to invent an election and keep it somewhere.
**Migrations are forward-only and additive.** The step runs before the *new* container starts, so the
old one is still running against the new schema for the length of the apply.
**Declared, never inferred.** The control plane cannot see inside an image, so a module that ships
migrations and declares no step is not refusable at registration; it breaks on its first upgrade. This
record says so rather than implying a check that cannot exist.
## Options considered
1. **Each module migrates itself when it starts** — what the catalogue does today. Rejected: it turns a
schema failure into a crash loop instead of a stop, it is invisible in the declaration so nothing can
say the module even has a schema, and two machines running the module both migrate at start with
nothing sequencing them.
2. **The mesh applies migrations itself**, with a driver and a version table per store — HAL's shape.
Rejected: the mesh would have to know one store type from another, hold another module's store
credentials, and reach a machine to use them, which [ADR 0005](0005-the-node-host.md) forbids. It is
also the reason that shape needs levels: something central has to decide where the once happens.
3. **A hook lifecycle** — pre-build, post-build, pre-deploy, post-deploy. Rejected: there is no deploy
event here to hook. A declaration is a desired state applied in order and reconciled forever, so
"pre-deploy" is exactly "a step before this container", pre- and post-build are what a Dockerfile and
the artifact list already are, and "post-deploy" has no moment to name.
4. **A hook level** — once per module, or once per module-node assignment. Rejected as a field, kept as
a property: see the decision. A once-per-module step needs cross-node ordering underneath it to be
safe, and a node converging without waiting on its neighbours is worth losing on purpose rather than
by accident.
5. **Every module hand-writes its own run-once step** — the immediate fix for the control plane.
Rejected as the general answer: it duplicates the resource it precedes, in three places already, and
a hand-written step is one the next module forgets. Forgetting it is the fault this record exists
for.
6. **Record a schema level per module in the store.** Rejected: gating makes the invariant true by
construction, so a level is a second account of the same fact and the first one to go stale.
## Consequences
**Three hand-written steps collapse into one line each**, and the control plane's own migrate step stops
repeating its server's environment and mounts.
**The catalogue's self-migration becomes the exception to remove.** One shape, and the mesh's own
control plane is not an exception to it either.
**A module on two machines with one shared store must lock.** Today none is, so this is an obligation
stated before it is needed rather than discovered by two concurrent migrations.
**There is still no readiness-gated step.** Only an action carries `verify`; a container has no health
notion, so "run this once the service answers" remains unexpressible and seeding through a running
service's API has no home. That is its own decision about a container's readiness, and this record does
not make it.
**Genesis keeps its own action.** At birth there is no control plane to derive anything from, which is
what [ADR 0067](0067-genesis-is-a-pivot.md) already says about that moment.
## How this is checked
- **The composition carries the step.** A test on a node's composed declaration: every container that
declares steps before it is preceded by them, and the derived step's image, environment, volumes and
network equal the container's — so the two cannot drift, which is the failure the hand-written kind
has.
- **A failed step stops what follows.** The host already refuses to go on past a run-once step that did
not exit 0; the test for that is extended to a derived one, so the gate is checked rather than
assumed.
- **The mesh's own schema is covered by the same mechanism as everything else.** The control plane
declares its step in its own manifest, so the case that failed on 2026-09-28 is the case the test
covers.
- **A module claiming a seat for a once-only step is checked where seats are checked** — the conditions
of holding, not a new mechanism.
## References
- [ADR 0052](0052-a-step-that-runs-once-before-a-container.md) — the step this extends
- [ADR 0018](0018-a-picture-is-read-from-what-runs.md) — a digest is the record that something happened
- [ADR 0005](0005-the-node-host.md) — the control plane decides and never touches a machine
- [ADR 0067](0067-genesis-is-a-pivot.md) — why genesis does it differently, once
- [issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md) — the failure that produced this record
- [`03-DESIGN/01-to-be/32-what-a-module-declares.md`](../03-DESIGN/01-to-be/32-what-a-module-declares.md) §6 — the lifecycle this sits in
- Measured 2026-09-28: three hand-written run-once steps repeating their sibling's resource; 0 of 72 modules with a migrations directory; 5 modules on more than one machine, none of them wanting a store
@@ -0,0 +1,126 @@
---
topic: the mesh
status: accepted
date: 2026-09-28
deciders: jochen
reconstructed: false
---
# 134. The mesh says what it applied
## Context
The pipeline is observable on the bus from a merge to an artifact, and modules already plug into it:
the forge emits `pull.merged`, the build machine's seat emits `built`, the catalogue emits `registered`,
`upgraded` and `rebuild-needed`, providers emit `postgres.database.provisioned` and its siblings. Things
consume them today — the catalogue consumes `built`, each provider consumes its own provisioning events,
model-usage consumes `*.usage.*`, the audit logger consumes `**`. Nothing had to be invented for any of
that; subscribing *is* plugging in.
**It goes dark at the moment it touches a machine.** A host applies a declaration and reports to the
control plane on the control branch, which only the control plane may read — correctly, because a report
carries what a machine is and enrolment travels the same way. So nothing on the mesh says *this machine
now runs version Y of module Z*, or that it refused to, or why.
What that cost on 2026-09-28, in one morning:
- A build result the store refused was visible only to whoever was waiting on that build's reply. For
three quarters of an hour the mesh built things and recorded none of them, while the overview said
every module was current ([issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md)).
- A module crash-looping at start — 338 restarts — was found by reading a container's logs by hand.
Nothing on the bus said the mesh's graph had stopped learning.
- A run-once step that fails now stops an upgrade by design ([ADR 0133](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md)),
and the same silence would cover it: the version simply would not appear.
**And the one thing the control plane does emit is refused by its own account.** Answering a catalogue
that asks to catch up, it publishes each recorded build under `mesh.mod.control-plane.event.…` — a
module namespace for a module that does not exist. Its own permissions refuse it, so a catalogue that
restarts gets nothing and keeps its gap. The control plane has facts to state and nowhere to state
them.
## Decision
**The mesh emits the deploy half of the pipeline as facts on the bus.** What a machine now runs, and
what it refused to run and why. Both are facts about the mesh doing its work, in the same form as every
other fact on the bus, so anything that wants them subscribes the way the catalogue subscribes to
`built`.
**The control plane states them, as the holder of the `mesh-controller` seat.** Its facts live under the
seat's own namespace, which is where a role's events belong
([ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md),
[ADR 0129](0129-a-seat-carries-the-protocol-of-its-role.md)) and which survives the control plane being
replaced. That is also what gives the catch-up replay a subject it may publish instead of an invented
module namespace.
**Emitted when what a machine runs changes, not on every convergence pass.** A host reconciles
continuously and reports each time; a fact per pass would be a fact per minute per machine that says
nothing. The report carries the declaration it applied and what changed, so the control plane has what
it needs to speak only when there is something to say.
**A refusal is a fact with a subject in it** — which machine, which resource, and the reason as the host
gave it. A refusal that names only the machine is the silence this record is about, one level up.
**Reports stay where they are.** A node's report remains control traffic that only the control plane
reads. The deploy facts are derived from it, which makes them second-hand on purpose: one emitter, one
ordering, and no widening of the narrowest account in the mesh.
## Options considered
1. **Leave it as it is, and let whatever cares ask the control plane.** Rejected: asking for a fact that
already arrives is the shape the mesh removed everywhere else, and nothing can react at the moment a
machine changes — which is exactly when a graph, an audit or an operator wants to know.
2. **Each node emits its own facts.** Rejected: it widens every host's account to an event namespace, and
a host's authority is deliberately the narrowest in the mesh. Its report already reaches the one thing
that can speak for it.
3. **Widen who may read the control branch.** Rejected: that branch carries what machines say *to* the
control plane, enrolment included. Widening its readers widens that too, for an unrelated reason.
4. **A registry of deploy hooks** — something registers interest and is called. Rejected: an event is
already the mechanism; there is nothing to register, and a callback is an address the mesh spent
[issue 102](../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)
learning not to keep.
5. **Put the facts on the node's own declaration stream.** Rejected: that stream is last-per-subject by
design, so a machine away for an hour gets exactly the current declaration and nothing older. A
history of what happened cannot live in a stream built to forget.
## Consequences
**The audit logger gets the deploy half for nothing**, because it consumes everything.
**A failure becomes visible where the mesh is watched** rather than where someone happened to be
looking. That answers the open question [issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md)
left about a record the store refused.
**The catch-up replay stops being a burst of events.** With the control plane able to state its own
facts, replaying history as if it were happening now is a choice rather than the only option — and the
better shape is the question the catalogue is actually asking, answered once
([design 33](../03-DESIGN/01-to-be/33-the-tools-the-mesh-answers.md)).
**The facts are second-hand.** The control plane says what a machine reported, so a machine that cannot
reach the bus produces no fact at all. Absence is not health, and what a machine was last heard from
stays the place that says so.
**The events stream carries more.** Bounded by emitting on change rather than on every pass, and each
fact is small; the stream's own limits remain what keeps it finite.
## How this is checked
- **The control plane's grant names exactly the subjects it emits**, derived from its seat like every
other principal's, and the composed user list is compared against a golden file — so a fact it cannot
publish fails a test rather than a catalogue's replay.
- **A convergence that changed nothing emits nothing.** A test with two identical reports and one
expected fact, because the failure this guards against is a fact per minute per machine.
- **A refusal names its resource.** A test where a host reports a failed resource and the emitted fact
carries which one and why, not merely that something went wrong.
- **What the mesh emits is what something consumes.** The subject a module declares it consumes derives
to the subject the control plane publishes — the same agreement test that already keeps the
controller's own subscriptions honest.
## References
- [ADR 0041](0041-events-are-a-relationship.md) — an event is a relationship, not a call
- [ADR 0126](0126-a-module-declares-its-own-seats.md) — an event is addressed to its emitter, because the emitter's identity is the meaning
- [ADR 0121](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md), [ADR 0129](0129-a-seat-carries-the-protocol-of-its-role.md) — a role's events belong to the role
- [ADR 0083](0083-one-push-leaves-the-mesh-consistent.md) — a report is held for the store rather than lost
- [ADR 0133](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md) — the gate whose failure this makes visible
- [`03-DESIGN/01-to-be/32-what-a-module-declares.md`](../03-DESIGN/01-to-be/32-what-a-module-declares.md) §6 — the lifecycle, which ends today at a report nobody else may read
- Measured 2026-09-28: 45 minutes of builds recorded nowhere with the overview reporting health; a module at 338 restarts found by hand; the control plane's only emitted event refused by its own permissions
@@ -0,0 +1,153 @@
---
topic: what runs on it
status: accepted
date: 2026-09-28
deciders: jochen
reconstructed: false
supersedes: 0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md
---
# 135. A module version prepares its state before it runs
## Context
[ADR 0133](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md) settled who runs a
module's migrations and when, and it said so in the wrong vocabulary. It put the declaration on a
*container* — "a container may declare steps to run before it" — and derived the scope of the work from
the *machine*. Both are wrong at the level a module author works at, and the second is wrong on the
facts.
**A container is one resource kind the host applies.** A module has code, state and a version; whether
its artifact is an image, a bundle or something later is the mesh's business. The module-facing
vocabulary for a module's own code already exists and has nothing to do with a container runtime: a
module declares **entrypoints** — this file is my tools, this file is my provisioner — and the mesh runs
them. A manifest that says "run this container with these arguments, and here are the volumes and
environment again" has an author writing down the machine's business twice.
**And the scope is not the machine's to decide, because the mesh already decided what a state is.** A
consumer is a module *on a machine* (migration 0015, from
[issue 022](../04-ISSUES/022-one-credential-per-node-per-provision-not-per-module/00-report.md)):
the mesh derives a login per consumer and the provider creates a database owned by exactly that login
([ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md)). So a module on three machines is three
consumers, three credentials and three databases. There is no shared state for two machines to race over,
and ADR 0133's central caveat — that a module's migrations must take a lock because two machines might
migrate at once — describes a situation the mesh does not currently produce.
That correction makes the whole "level" question HAL answered with stages disappear: the scope of
preparation is the scope of the state, and the mesh knows it.
What the earlier record got right and this one keeps: the module owns the work, the mesh owns the moment,
the gate is the guarantee, migrations stay forward-only, and none of it can be inferred from inside an
artifact. What produced it also stands — the control plane was replaced with a build carrying a migration,
nothing applied it, and for three quarters of an hour every build was refused by the store with one line
that reached only whoever was waiting on a reply
([issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md)).
## Decision
**A module version declares an entrypoint that prepares its state.** One name in the manifest, in the
same vocabulary as the entrypoints it already declares for its tools and its provisioner. No container,
no command line, no environment, no mounts — those are how a machine runs the module's code, and the
module already said that once.
**The mesh runs it as it runs that module's own code, to completion, in the module's own context.** Every
binding, credential and setting the module's code would receive, because it *is* the module's code. How a
machine does that is the host's business and stays there: for an image artifact it is the step
[ADR 0052](0052-a-step-that-runs-once-before-a-container.md) already defines, and a later kind of artifact
changes the host, not the manifest.
**Preparation gates the version.** A version whose preparation did not succeed does not run — anywhere.
Since the rollout already sends machines one at a time and stops at the first that does not take a
version, a preparation that fails stops the rollout there, leaving every other machine on the version
that works.
**Preparation is scoped to the state, and the mesh derives that scope.** State the mesh provisions is per
consumer — a module on a machine — so preparation happens once per consumer. State the module keeps on
the machine is per machine, which is the same answer. A module that holds an exclusive seat has one of
itself, so its preparation happens once by definition. No level, no election, no cross-node ordering, and
no lock obligation invented for a race the mesh does not create.
**Once per version per state.** A version bump attempts preparation once against each state it has; the
module's own runner decides there is nothing to do, which is what a runner with a version table does
anyway. A retry after a partial failure runs it again, so the work is the module's to make safe against
that — the one obligation no design can remove.
**Forward-only and additive.** Preparation runs while the previous version is still serving, so a
migration that removes or renames what the old code reads breaks the mesh in the window between the two.
**Declared, never inferred.** The control plane cannot see inside an artifact, so a module that ships
migrations and declares no entrypoint is not refusable at registration. It breaks on its first upgrade,
and this record says so rather than implying a check that cannot exist.
## Options considered
1. **A container declares steps before it** — [ADR 0133](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md).
Superseded, not because the mechanism is wrong but because the *declaration* is in the wrong place: it
makes every module author restate the machine's arrangement, and it ties a module's own lifecycle to
one resource kind. The host-side mechanism it named is retained and is now an implementation detail.
2. **Each module prepares itself when it starts** — what the catalogue does today. Rejected: a schema
failure becomes a crash loop rather than a stop, nothing in the declaration says the module has a
state to prepare, and the version serves the moment it starts rather than after the state is right.
3. **The mesh applies migrations itself**, with a driver and a version table per store type. Rejected:
the mesh would have to know one store from another, hold another module's credentials and reach a
machine with them, which [ADR 0005](0005-the-node-host.md) forbids. It is also what forces a stage
system: something central has to decide where the work happens.
4. **A hook lifecycle** — pre-build, post-build, pre-deploy, post-deploy. Rejected: a declaration is a
desired state reconciled forever, so there is no deploy moment to hook. "Pre-deploy" is exactly this
record; pre- and post-build are what a recipe and the artifact list already are; "post-deploy" names
nothing that happens.
5. **A declared level** — once per module, or once per assignment. Rejected: the mesh already knows what a
state is, so asking an author to choose is asking them to restate a fact the mesh holds, with a chance
of contradicting it.
6. **Record a preparation level per module in the store.** Rejected for the reason ADR 0133 gave and this
record keeps: gating makes the invariant true by construction, and a level is a second account of the
same fact.
## Consequences
**An author's whole contract is one line, once.** Write the migration in the module's code, name the
entrypoint that runs it, and every later version rolls out as: build, prepare, run — with nothing
per-version to remember and nothing about the machine to restate. That is the property this exists for.
**Three hand-written steps in the catalogue collapse**, and the control plane's own migrate step stops
repeating its server's environment and mounts.
**The catalogue's self-preparation becomes the exception to remove.** One shape, and the mesh's own
control plane is not an exception either.
**A module scaled across machines with one shared state is not expressible**, and this record does not
make it so. The mesh gives each consumer its own state; a deliberately shared one is a different
provision model, and the place the "once, mesh-wide" question would genuinely return. Named here so it is
a decision when it happens rather than a surprise.
**There is still no readiness-gated step.** Only an action carries `verify`; nothing declares that a
service answers, so preparation that must happen *after* something is serving — seeding through its own
API — remains unexpressible.
**Genesis keeps its own action.** At birth there is no control plane to derive anything, which is what
[ADR 0067](0067-genesis-is-a-pivot.md) says about that moment.
## How this is checked
- **The composition carries the preparation, in the module's own context.** A test on a node's composed
declaration: a version declaring a preparation entrypoint is preceded by it, and what it is given
equals what the module's own code is given — asserted equal rather than written twice, which is the
drift the superseded shape invited.
- **A preparation that fails stops the version.** The host does not go past a step that did not complete,
and the rollout stops at the first machine that did not take a version. Both are existing behaviours
with existing tests; the test for preparation asserts the two together — the machine does not run it,
and the machines after it are left alone.
- **Once per version per state.** A test that a second convergence of the same version prepares nothing,
and that a new version prepares again.
- **The mesh's own control plane declares one.** The case that failed on 2026-09-28 is the case the tests
cover, rather than a case a comment says is covered.
## References
- [ADR 0133](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md) — what this supersedes, and why
- [ADR 0052](0052-a-step-that-runs-once-before-a-container.md) — the host-side step that implements it for an image artifact
- [ADR 0048](0048-a-provider-creates-the-credential-the-mesh-minted.md), [issue 022](../04-ISSUES/022-one-credential-per-node-per-provision-not-per-module/00-report.md) — a consumer is a module on a machine, which is what makes the scope derivable
- [ADR 0018](0018-a-picture-is-read-from-what-runs.md) — a digest is the record that something happened
- [ADR 0005](0005-the-node-host.md) — the control plane decides and never touches a machine
- [ADR 0134](0134-the-mesh-says-what-it-applied.md) — what makes a failed preparation visible
- [issue 133](../04-ISSUES/133-the-control-planes-schema-is-migrated-at-birth-and-never-again/00-report.md) — the failure that produced both records
+4
View File
@@ -140,6 +140,8 @@ python3 00-META/checks/index.py fail if stale
- **0129** — [A seat carries the protocol of its role](0129-a-seat-carries-the-protocol-of-its-role.md)
- **0130** — [The predecessor is ending, and its broker goes with it](0130-the-predecessor-is-ending-and-its-broker-goes-with-it.md)
- **0131** — [Everything on the mesh speaks to the broker seat, and AMQP is not a provision](0131-everything-on-the-mesh-speaks-to-the-broker-seat.md)
- **0132** — [A seat carries the tools its holder must serve](0132-a-seat-carries-the-tools-its-holder-must-serve.md)
- **0134** — [The mesh says what it applied](0134-the-mesh-says-what-it-applied.md)
### Its tiers, from the bottom up
@@ -212,6 +214,8 @@ python3 00-META/checks/index.py fail if stale
- **0120** — [A roster fact carries its format as a template: the mesh owns the data, the module owns the format](0120-a-roster-fact-carries-its-format-as-a-template.md)
- **0121** — [A system seat is named for its scope, and a module may define its own](0121-a-system-seat-is-named-for-its-scope-and-modules-define-their-own.md)
- **0122** — [A seat is data the controller owns, and a rename is a database update](0122-a-seat-is-data-a-rename-is-a-database-update.md)
- **0133** — [A module owns its migrations, and the mesh owns when they run](0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md) *(superseded)*
- **0135** — [A module version prepares its state before it runs](0135-a-module-version-prepares-its-state-before-it-runs.md)
### How it is built
@@ -2,7 +2,7 @@
layer: to-be
status: proposed
code: []
updated: 2026-09-27
updated: 2026-09-28
decisions:
- 02-DECISIONS/0126-a-module-declares-its-own-seats.md
- 02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md
@@ -12,6 +12,8 @@ decisions:
- 02-DECISIONS/0112-a-module-definition-names-no-node-mesh-or-path.md
- 02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md
- 02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md
- 02-DECISIONS/0135-a-module-version-prepares-its-state-before-it-runs.md
- 02-DECISIONS/0134-the-mesh-says-what-it-applied.md
---
# 32. What a module declares, and what the bus makes of it
@@ -258,10 +260,28 @@ queue.
and publishes it last-per-subject. A node that was away gets exactly the current one, never a
queue of superseded ones, and a replayed older one is refused by sequence.
**A version prepares its state before it runs.** A module version may declare an entrypoint that brings
its state to the shape that version needs — the same vocabulary as the entrypoints it declares for its
tools and its provisioner, and nothing about how a machine runs it. The mesh runs that entrypoint as it
runs the module's own code, to completion, in the module's own context, and a version whose preparation
did not succeed does not run: the rollout stops at the first machine that did not take it
([ADR 0135](../../02-DECISIONS/0135-a-module-version-prepares-its-state-before-it-runs.md), superseding
[ADR 0133](../../02-DECISIONS/0133-a-module-owns-its-migrations-and-the-mesh-owns-when-they-run.md)).
Once per state, and the mesh derives what a state is: a consumer is a module on a machine, so what the
mesh provisions is per consumer and preparation is too. No level to choose, and no race to lock against.
**Applying is reported to a role.** The host applies and reports to the `mesh-controller` seat —
not to an address it was given at genesis. Held and retried while the store restarts
([ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)).
**And the mesh says what it applied.** A report is control traffic only the control plane reads, so the
chain above went dark at the moment it touched a machine: nothing said which version a machine now runs,
or that it refused to. The control plane states those as facts under its own seat's namespace, when what
a machine runs changes rather than on every convergence pass, and anything that cares subscribes the way
the catalogue subscribes to `built` ([ADR 0134](../../02-DECISIONS/0134-the-mesh-says-what-it-applied.md)).
The facts are second-hand by design — one emitter, one ordering — and a machine that cannot reach the bus
produces none, so absence is not health.
What disappears across that chain is every address. No webhook URL, no registered callback, no
"which node is the builder on", no controller endpoint baked into a joining node. That is the
class of bug
@@ -439,6 +459,10 @@ it is the residue of a question the rest of §8 answers and the part a fingerpri
**Whether a module may declare a seat it does not itself claim** — the contract as one thing, the
implementation as another, which is how two competing implementations would ever exist.
**Whether a container should have a readiness notion.** Only an action carries `verify`, so a step that
must run once a service *answers* — seeding through its own API — cannot be declared at all. Named here
because the steps above make the gap obvious, not because they caused it.
**Whether `consumes` naming another module couples too tightly.** It is kept here deliberately —
an event's provenance is its meaning — but a consumer of `billing.order.placed` does depend on
billing existing under that name.
@@ -447,6 +471,11 @@ billing existing under that name.
- **A manifest holds no subject.** A catalogue test: no manifest contains a string matching the
subject grammar. The rule is worthless if it is followed by convention.
- **A preparation is given what the module is given.** A composition test: what the preparation
entrypoint receives equals what the module's own code receives, asserted rather than written twice —
which is the drift a hand-written step invites, three times over in the catalogue today.
- **A convergence that changed nothing says nothing.** Two identical reports, one emitted fact: what is
guarded against is a fact per minute per machine, which is a stream nobody reads.
- **Permissions are exactly the three namespaces.** A composition test per module: the derived
permission set equals what its declaration implies, and a hand-written addition to it fails.
- **A sender cannot read the queue it writes to.** A bed: a module declaring `uses` is refused
@@ -0,0 +1,137 @@
---
layer: to-be
status: designed
code: []
updated: 2026-09-28
decisions:
- 02-DECISIONS/0132-a-seat-carries-the-tools-its-holder-must-serve.md
- 02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md
- 02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md
---
# 33 — The tools the mesh answers
**An agent can call the mesh's tools and cannot find out what they are.** Both halves were measured
on the live mesh on 2026-09-28: a client holding an operator credential connected to the bus, asked a
module for its repositories and got them; the same client's request for the tool list found nothing
serving it. The transport works, the account model works, the adapter that speaks the agent protocol
works. What is missing is the mesh being able to say what it can do.
This design is the answer to that question, and it has three families in it, because a tool belongs to
whoever is accountable for answering it.
## 1. Three families, and why the split is not arbitrary
| Family | Addressed to | Where the definition lives | Example |
|---|---|---|---|
| A **role's** tools | the seat: `mesh.seat.<seat>.tool.<verb>` | the seat's protocol, in the mesh's records | ask *the forge* to list its repositories |
| A **module's** tools | the module: `mesh.mod.<module>.tool.<name>` | that module's code | ask *this gitea* for `gitea_list_repos` |
| The **mesh's** own verbs | the `mesh-controller` seat | the seat's protocol, as above | `status`, `push`, `build`, `assign` |
The split follows accountability. A role is something the mesh guarantees exactly one holder of, so
what the role answers is the mesh's to define and a holder's to implement
([ADR 0132](../../02-DECISIONS/0132-a-seat-carries-the-tools-its-holder-must-serve.md)). A module's
own tools are nobody's business but the module's, and their definitions live where they are
implemented, because a copy kept anywhere else drifts from the code that answers.
The mesh's own verbs are the third family only in where they come from, not in kind: the control plane
holds a seat like anything else, and its tools are that seat's. This is what keeps them addressable
while the control plane is being replaced, which is the moment they are most needed.
**Both names for one capability is deliberate and bounded to this.** A forge holding the `git` seat
answers the role's `list_repos` and its own `gitea_list_repos`, because the same module may run
without the seat — a second instance, kept for one purpose — and then only the second name is true.
The caller chooses which question it is asking. Nothing else in the mesh gets two names.
## 2. What a seat's tool is
A verb, what it does, and the schema of its arguments and its answer. A name alone is not callable by
something that has never seen the mesh before, which is the whole population this surface exists for.
The protocol a seat carries today is three lists of bare verbs
([ADR 0129](../../02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md)), and it must widen to
carry the rest. Two constraints on that widening:
- **It lives in the mesh's records, not in the control plane's binary.** Today a seat's protocol comes
from compiled defaults, merged in as a row is read, because the seat rows never gained the columns.
Discovery that reads a binary is discovery that disagrees with the mesh the moment the two are on
different versions.
- **The schema is stated in the form an agent protocol already uses**, so nothing translates between a
seat's idea of an argument and the caller's. A translation layer would be a second definition of
what a tool is.
## 3. Holding a seat means serving its tools
A module may not occupy a seat unless it serves every verb that seat declares. This joins the
conditions of holding that already exist — providing what the seat delivers, being assigned at the
seat's scope — and is refused the same way: at registration and at handover, naming the verbs that are
missing rather than the fact that something is.
A module knows which seats it claims, so knowing which tools it must serve is not a discovery problem
for the module: the seat says, the module implements, and anything beyond that is its own.
## 4. Addressing a node-scoped seat
A seat's subject is flat today — `mesh.seat.<seat>.<kind>.<verb>` — which is correct for a seat the
mesh has one holder of and wrong for the six node-scoped seats, where one subject would reach every
machine's holder and the holders' queue group would hand the call to whichever answered first. A
node-scoped seat's tool therefore carries the node it is asked of. Nothing about a mesh-scoped seat
changes.
## 5. Discovery
**What a role answers is a read.** The seats and their protocols are records, so the list is a query
against the mesh's own store: no call to a module in the path, nothing that has to be running, and an
answer that stays true while a holder is restarting or being replaced.
**What a module answers comes from the module.** Its definitions live in its code, so it is asked, and
the answer is as available as the module is — which is the right coupling for a tool that only exists
while that module does.
A caller therefore gets one list assembled from two sources, and the difference is visible in it: a
role's tool names a seat, a module's names a module. An agent that wants to survive a holder being
replaced binds to the first.
## 6. What serves this to an agent
A module the mesh assigns to the machine where the agent runs, holding a credential the mesh minted,
with authority derived from what it may call — not a program started by hand with a credential printed
to a terminal. The adapter itself already exists and is thin by design; what changes is that it stops
being something a person carries and becomes something the mesh runs, on a node, like everything else.
An agent's authority can then be role-shaped: *the forge's tools*, rather than a list of
module-specific names that changes the day the forge is replaced.
## 7. Versioning
A seat's tools are an interface and change like one. Additive within a version. A change that would
break a caller takes the version token the subject already has room for, and the two versions run side
by side until nothing is bound to the old one.
## How it is checked
- **A holder missing a verb cannot take the seat.** One test per condition of holding, as the existing
conditions have, and the live refusal names the verbs.
- **A verb nobody declared is a subject nobody may use.** The bus grants are derived from the seat's
protocol already, and the golden composition of the user list is what keeps that honest: a holder is
granted exactly the seat's verbs, a user of the seat exactly the publish side.
- **Discovery needs no running module.** The test for a role's tools reads records and asserts the
answer equals what the seats declare — if it needed a module up, it would not be a read.
- **Two nodes holding one node-scoped seat derive two addresses.** Checked by the same test as the rest
of the subject table.
## What this does not settle
- Which verbs each seat should serve. That is a decision per seat, and the reason to do it slowly: a
seat's tools bind every future holder.
- Whether a module's own tool definitions should also be recorded when a build resolves its manifest.
There is an argument for it — the mesh could then answer for a module that is down — and an argument
against, which is that a recorded copy of a live definition is a copy that can be wrong.
## References
- [ADR 0132](../../02-DECISIONS/0132-a-seat-carries-the-tools-its-holder-must-serve.md) — the decision this designs
- [ADR 0129](../../02-DECISIONS/0129-a-seat-carries-the-protocol-of-its-role.md) — a seat carries the protocol of its role
- [ADR 0095](../../02-DECISIONS/0095-the-control-plane-is-the-way-to-ask-a-module.md) — a tool call passes one process where an audit belongs
- [`26-the-seats.md`](26-the-seats.md) — what a seat is, how it is held and handed over
- [`25-the-bus-on-nats.md`](25-the-bus-on-nats.md) §7 — a person's account, their inbox, and the adapter
@@ -0,0 +1,60 @@
---
status: resolved
opened: 2026-09-28
located-in: [mesh-controller cmd/mesh-controller]
fixed-by: mesh-controller — `module add` takes `--path` and `--self`, so a module handed over by hand records the whole location it came from; a record naming a repository and no directory says so in the reply; the rule is one function with a test beside it. The nine records already wrong were corrected by rebuilding each with its real directory, which is the same act through the same door.
amended-design:
---
# 132 — A module can be recorded without the directory it lives in
## What was observed
Nine modules on one mesh could not be rebuilt. Each attempt failed the same way:
> has no module.json at its root, so there is nothing saying what it is
All nine were recorded as coming from a repository that holds many modules, each in its own
directory — and each record named the repository and no directory. So every build cloned the
repository and looked for a manifest where there has never been one.
The failure only surfaced when something asked for all of them at once. Before that, the overview
said every module was current with its source, because what it compares is what was built against
what the mesh was last told the source has, and neither half knows whether the source can be found
at all.
## Why it matters beyond this instance
**A module is a repository and a directory inside it** ([ADR 0069](../../02-DECISIONS/0069-a-module-is-a-repository-and-a-path.md)),
and one of the two doors into the catalogue could record only the first half. A build records the
directory it was given, so a module that arrived by being built is always whole; a module handed over
by hand had no way to say where it lived, and the flag to say it did not exist. The rule was decided
and enforced on one path out of two.
**Half a location reads exactly like a whole one.** Nothing in the record is empty in a way a person
would notice: the repository is there, the branch is there, the commit is there. The mesh only finds
out at the moment it needs the manifest, which is the moment it is trying to rebuild — and the module
stays on whatever it last built, indefinitely, with nothing saying why.
**It is the same shape as [131](../131-nothing-tells-the-mesh-a-source-moved/00-report.md).** A
comparison between two facts the mesh holds about itself will agree with itself. Whether the source
can be found is a question only an attempt to read it answers, and the answer had nowhere to go.
## What was done
`module add` takes the directory and which forge holds the repository, so a hand-registered module
records the same whole location a built one does. What a record must say to be worth anything is one
function with a test beside it, rather than a paragraph in a help string: provenance together or not
at all, a directory needs a repository to be inside, a path on the mesh's own forge is not an address.
And a record that names a repository but no directory says so when it is made — not refused, because a
module really at a repository's root is ordinary, but said, because the person adding it is the one
who knows which it is.
The nine wrong records were corrected by building each with its real directory, which re-records it.
No row was written by hand.
## What is still true
A directory that does not exist in the repository cannot be refused when the module is added: the
control plane does not clone, and inventing a check there would mean it did. The first build says so
plainly, which is one build rather than nine, and the record it leaves behind is right from then on.
@@ -0,0 +1,78 @@
---
status: resolved
opened: 2026-09-28
located-in: [mesh-controller module.json]
fixed-by: mesh-controller — the control plane's module declares a run-once `migrate` step before its server, which is the shape ADR 0052 prescribes for exactly this. A step's record of having run is the digest of its declaration and the image is part of that digest, so a new build of the control plane re-runs it; and because a run-once step gates what the declaration places after it, a migration that fails stops the new server from starting at all rather than letting it run against a schema it does not have.
amended-design: 03-DESIGN/01-to-be/32-what-a-module-declares.md
---
# 133 — The control plane's schema is migrated at birth and never again
## What was observed
On 2026-09-28 at 08:17 the control plane was replaced, by the mesh's own upgrade path, with a build
whose code writes a column that a migration **in that same build** creates. Nothing ran the migration.
For the next three quarters of an hour the mesh built things and recorded none of them. Every build
answered:
> ERROR: column "built_contexts" of relation "build" does not exist (SQLSTATE 42703)
and that sentence went only to whoever happened to be waiting on a build's reply. The overview kept
saying the mesh was fine. The builds themselves worked — images were built and published — so the
registry filled up with artifacts the mesh has no record of, and the graph stopped learning without
anything saying so.
The schema was created once, at genesis, by an action in the foundation bundle that runs the same
binary's `migrate`. Nothing runs it again. The mesh has updated its own control plane many times since
that bundle, and every one of those updates carried whatever migrations the new build brought and
applied none of them. This is the first time a build needed one.
## Why it matters beyond this instance
**The schema and the code that needs it ship as one artifact and are applied by two mechanisms, only
one of which is automatic.** A module's version is atomic everywhere else in the mesh — the manifest,
the image and what the machine runs move together. Its schema did not, so "the mesh updates itself on
a push" was true of the code and false of what the code needs.
**The failure is quiet exactly where quiet is worst.** A build that cannot be recorded is a build that
happened and left no trace, which is the fault [issue 050](../050-the-catalogue-knows-nothing-built-before-it/00-report.md)
and [issue 131](../131-nothing-tells-the-mesh-a-source-moved/00-report.md) are both about. The mesh
has three mechanisms for noticing a module is behind its source and none for noticing that what it
recorded was refused.
**The shape was already decided, and the control plane was the one module that did not use it.**
[ADR 0052](../../02-DECISIONS/0052-a-step-that-runs-once-before-a-container.md) says a run-once container is a
step the host runs to completion before whatever the declaration places after it, and names migrating
a schema as the case it exists for. The genesis code's own comment says a manifest may name its image
in more than one resource — "a migrate step beside the server". The control plane's manifest had no
such step; it went straight from a state directory to the server.
## What is still true
**Additive migrations are load-bearing, not a style preference.** The step runs before the *new*
server starts, which means the old binary briefly runs against the new schema. A migration that
removes or renames something would break the running control plane in the window between the two.
**A hand-written step is one the next module forgets**, which is why this fix is not where the matter
ends: [ADR 0135](../../02-DECISIONS/0135-a-module-version-prepares-its-state-before-it-runs.md) makes it
derived and puts it where an author works: a module version declares an entrypoint that prepares its
state, and the mesh composes the gated work from it, so the control plane stops being the only module
that had to remember. That record also settles the level question HAL answered with stages — a consumer
is a module on a machine, so the scope of preparation is the scope of the state — and
[ADR 0134](../../02-DECISIONS/0134-the-mesh-says-what-it-applied.md) answers the second open question
below: what a machine applied, and what it refused, become facts on the bus rather than a line in a log.
**The mesh now has two shapes for one problem.** The catalogue module migrates its own schema in its
own code when it starts; the control plane migrates in a step the host gates on. Both work and the
reasons differ — a module that owns its store entirely can do it at start, while a step is visible in
the declaration and refuses to let a broken upgrade serve. Which one the mesh should standardise on is
a decision, not a fix, and it is not made here.
## Open questions
- Should a module be refusable at registration when it ships migrations and declares no step and no
other way to apply them? The mesh can see both halves.
- Should a record the store refuses reach the overview? Today the only reader of that failure is
whoever asked for the thing that failed, and for an event arriving on the bus there is no such
person.