One container restarted 2286 times over five days while the mesh reported the machine as doing what it
was told. Its overlay address was five days out of date: the host compares a container by a digest of
its spec, and the mesh's names were not in it, so a container whose image and files never changed was
left alone holding a name that no longer resolved. Forty-eight others were current only because
something else had recreated them.
The same fault as issue 045, in the field that was left out. Resolved by putting the names in the
digest.
0120 was already accepted; what failed the check was that it rests on 0112, still marked proposed —
and so do four designs. The decision stands: a definition names no node, no mesh and no host path, and
everything a module needs is a requirement the mesh resolves.
Accepting it makes the gap visible rather than hiding it, so issue 134 states it. 0112 says how it is
checked — 'a catalogue test finds no domain name in any definition value' — and there is no such test.
Asked by hand: seven modules name this installation in a value the mesh acts on, and eight mention a
public name in prose nothing reads. The two are not the same fault and the fixes differ, which is why
the issue separates them rather than counting to fifteen.
Two of its statements are built — a version preparing its state, and the mesh saying what it applied —
so the document is in-progress rather than proposed, and names the code that owns them. Resting on a
live record rather than a superseded one: 0127 was replaced by 0131.
What this exposes is pre-existing: it also rests on ADR 0112, which is still proposed, and a document
that is not itself proposed may not. ADR 0120 has rested on it the same way for a while. Accepting or
superseding 0112 is a decision, not a cleanup, so it stays visible in the check rather than papered
over.
ADR 0135 made a step something the mesh derives for any module that prepares its state, which turned
ADR 0052's reach into a fault: a module whose database is briefly unreachable would stop every module
declared after it on that machine — the fault issue 011 already removed for every other shape, and
the reason the catalogue migrates itself at start rather than in a step.
A step now stops the rest of its own module and nothing else; an action still gates the machine,
because genesis is a row of them and they belong to no module. What was not attempted is reported as
skipped rather than left to be inferred from silence.
Two faults in 0133, both caught on review. It put the declaration on a container — one resource kind
the host applies — so every author would restate the machine's arrangement and a module's own
lifecycle would be tied to how its artifact happens to run. A module declares entrypoints for its
tools and its provisioner; preparing its state is the same vocabulary and nothing about a runtime.
And it derived the scope from the machine, which the facts already answer: a consumer is a module on
a machine (issue 022, migration 0015), so what the mesh provisions is per consumer. A module on three
machines has three databases, there is no shared state to race over, and the lock obligation 0133
invented was for a situation the mesh does not produce. The level question HAL answered with stages
dissolves — the scope of preparation is the scope of the state, and the mesh knows it.
0133 keeps its reasoning and gains a pointer; design 32 and issue 133 name the live record.
0133 — a module owns its migrations and the mesh owns when they run. A container declares what must
run before it; the mesh derives the gated step from the resource it precedes, so the image, the
environment and the credentials come from the one place they are described. The module owns the SQL,
the dialect and the lock; the mesh owns the moment and refuses to start a version whose step failed.
Per node, with no level: a step that ran once somewhere leaves every other machine ungated, and
'once, mesh-wide' is what holding a seat already means.
0134 — the pipeline is observable from a merge to an artifact and goes dark at the machine. What a
node now runs, and what it refused, become facts under the control plane's own seat, emitted when
what a machine runs changes rather than on every convergence pass.
Design 32's lifecycle carries both; issue 133 points at them as what ends the matter it opened.
The mesh replaced its own control plane with a build carrying a migration, applied none of it, and
then recorded no build for three quarters of an hour while saying everything was fine. ADR 0052
already prescribes the shape — a run-once step that gates the server — and the control plane was the
one module that did not use it.
A role's tools belong to the role, not to whichever module holds it today: the seat declares them
with their schemas, serving them is a condition of occupying the seat, and what the mesh can do
becomes a read of its own records rather than a question nothing answers. A module keeps its own
tools — the same module may run without the seat, and then only its own name is true.
Design 33 follows: the three families, addressing a node-scoped seat, discovery, and what serves
this to an agent.
Nine modules could not be rebuilt: their record named the repository and no directory, so every
build looked for a manifest at a repository root that has never had one. Resolved by mesh-controller
— `module add` takes the directory and the forge, and the rule is checked rather than described.
The AMQP transport is gone from the control plane and the hosts (mesh-controller #112,
mesh-host #39). On the way: no build had ever recorded its bases, so every bases-first order
walked an empty graph; the builder now reports what it was handed and the graph is read from
builds (mesh-controller #113/#114). Issue 131 is resolved by the forge module's merge event,
the control plane following it, and those edges.
Tasks 4.3, 5.2 and 5.4 are done as of 2026-09-28 02:25: every machine reports on
the new bus, the seat is held by the module that provides it, the old broker is
unassigned and forgotten, and every credential was minted afresh at the end.
5.2 records what it took, in the order it was found and each fixed on the trunk
before the next step, and how the bootstrap loop was broken once, by hand.
Until now the holder was derived — assigned and claiming — and a second eligible
assignment was refused, so a seat could not pass from one holder to the next
without a moment where nobody held it. The controller finds its own bus through
one of these seats, and that moment took the control plane down on 2026-09-27.
The holder is now a row the controller keeps, written by `seat <name> --to
<node>/<module>` in the same write that removes the previous one. No row means the
old rule, so nothing changes for a mesh that never hands a seat over; with a row,
another eligible assignment is silent rather than refused, which is what lets the
next holder run beside the current one until the switch. A holding is the
assignment's and goes when it does. Each rule names the test that checks it.
Under ADR 0131; design 28 task 5.3 is the work.
Taken during the outage of 2026-09-27, when the protocol leaked into the seat's
contract: to hold mesh-broker a module had to provide amqp, so the module that
will carry the bus could not hold the seat that names the bus, while the module
being retired could. Supersedes 0127. Modules depend on the seat and reach the
bus through the sdk; no manifest provides or requires amqp; the old broker's
module and the two modules that required it leave the catalogue; the AMQP
transport is deleted once every node reports on the new bus.
Design 28 step 5 rewritten under it: the seat handover becomes its own task and
is built first, because the seat the control plane dereferences cannot be empty
in between — that emptiness was the outage. The cost note now carries what was
measured rather than what was assumed.
0128 and 0130 extended 0127; each now rests on 0131 with a dated note and
changes nothing it decided. Every other citation of 0127 names its replacement.
records.py still fails on 0120/0112, which predates this branch.
Derived while merging and visible from no single repository, so it belongs
written down rather than re-derived later: the client library before the
catalogue, because it is where the subject is derived and a converted module
against the old one publishes the local name itself; the catalogue before the
controller, because the controller refuses an old-style name outright and would
make every unconverted module unregisterable.
Two tested properties are what make it safe and not merely ordered. An old-style
name passes through the derivation untouched, so an unconverted module keeps
working at every step. And a converted name derives to exactly the key the old bus
published, so nothing moves on the wire until 5.2 sets the variable.
The failure avoided is 127's own, which is why this is worth a table: a publisher
and a subscriber disagreeing about a subject log nothing anywhere.
Correcting design 26 to match what the merged code does, and fixing issue 112's
status, which used a word the vocabulary does not have.
A claim on a seat the manifest does not itself declare is refused at registration
rather than by the parser. A module may hold a seat another module declared —
that is why ADR 0126 has a caller name the seat and not its provider — so whether
the name exists is a fact about the whole catalogue.
And a declared seat may promise nothing. That is a marker seat, and most
node-scoped seats are markers: which module is this machine's packet filter. ADR
0126's "a declared seat carries a protocol" says what a holder must satisfy, not
that every seat offers something.
`records.py` still fails on ADR 0120 resting on a proposed ADR 0112, which is not
this branch's and not mine to decide.
Both lines of work numbered from the same point, so four decision records and one design
document existed twice with different content. The trunk keeps its numbers and this branch
yields — the only rule that scales, because the trunk's are already cited by what merged
before them.
0117 the bus is the only broker -> 0125
0118 a module declares its own seats -> 0126
0119 amqp is a provision, not the bus -> 0127
0120 the mesh bus is required -> 0128
0123 a seat carries its role's protocol -> 0129
0124 the predecessor is ending -> 0130
design 29, what a module declares -> design 32
Applied to the code repositories too, because a stale reference is worse when numbers
collide than when they dangle: the reader lands on a real record that decided something
else.
Two reconciliations the merge forced, both real:
**0110 was marked wholly superseded and was not.** Its successor says in as many words that
everything 0110 decided about what a seat *is* stands untouched — and two records that
landed on the trunk rest on exactly that part. So it is accepted again, extended rather than
replaced, with a note saying which of its claims moved and where.
**A seat's protocol becomes columns, not fields.** The trunk moved the seat set out of
compiled code into a table the controller owns. This branch had added what a role accepts,
emits and serves to the Go slice. The decision is unaffected and the mechanism is better for
it: giving a role a protocol is now a write rather than a rebuild, which is the trunk's own
argument applied to what this branch added.
One check still fails and it fails on main too: a record resting on ADR 0112 while that is
still 'proposed'. Left alone — it is not this merge's to answer.
ADR 0119 rejected giving the old broker a retirement condition and said why: "its
clients are not only the predecessor's, so the retirement condition describes a day
that will not come". The operator has said that day is coming — the predecessor is
deprecated, some of it still running, none of it being migrated, left to stop rather
than moved.
Recorded because three documents reason from the premise it overturns. Design 25 §5's
"no day anything is waiting for", §9's "the predecessor's clients never notice", and
design 28's closing note that the predecessor's world does not need to move.
**And it needs no new machinery, which is 0119 being paid off rather than revised.**
Because that record made the broker an ordinary provider rather than a compatibility
module, ending it is unassigning a provider whose provision nothing requires — something
the module system has done since it existed. So step 5.3 finishes instead of trailing
off, and the transitional double announcement of a build outcome has a date.
The consequence worth planning around: the predecessor's own mesh talks over that
broker, so shutting it down ends the tooling that reaches this installation's machines
from a workstation. The rollout is driven from the node, or driven before the broker
stops. That is a sequencing constraint on 5.2, not an afterthought.
What survives is `amqp` as a provision: a module that genuinely needs an AMQP broker can
still be given one. What retires is this broker's role as the predecessor's.
`rollout check` answers from records whether this mesh could move its bus, and names the
next step for each thing missing. The move itself waits on that check having been run
against a real mesh — writing the irreversible half before its question has ever been
asked of something real breaks the plan's own rule about beds by another route.
And the cost of being wrong is written down rather than assumed: the old broker stays for
its other clients, nothing in a served request's path goes over the mesh's own bus, and
what a failed move costs is the mesh's ability to change things rather than the services
its modules serve.
Three ways to do the replay were weighed and the right answer was that the bus being
moved to already does it. A queue on the old bus receives only what is published after
it is bound, so everything built before the catalogue existed was announced to nobody. A
stream is a log and a consumer is a position in it: a consumer created later starts at
the beginning and the builds are simply there. Checked against a running server, because
the decision rested on it.
So "who replays" has no answer because nothing replays. The mechanism was never about
builds — it was about a queue that could not remember, and carrying it across would have
carried a workaround for a limitation that no longer exists, with nothing looking wrong.
Retiring it belongs to step 5, with the rest of what only the old bus needs.
The account existed as a permission model and as nothing a person could be given; there
is a record and three commands now. The client is two surfaces over one thing, a command
line and an MCP server, both using the client a module's runtime uses — so what a person
may do is answered by the same permission list that answers it for a module.
Design 25 §7 says nothing of this is built before its bed passes, and this was built
before. Noted in the task rather than quietly ignored.
A node's intrusion filter should be composed from its assigned modules, like
its firewall (the Filtering mechanism): a service module (postgres, mssql,
mailu) declares its jail in its manifest (filter + stanza, no node/path per ADR
0112), and the mesh writes the jails of a node's modules into the fail2ban
holder's jail.d. The base (sshd, recidive, ignoreip=mesh-range) stays the
fail2ban module's. Records the model after novox's HAL per-module jails were
lost as dangling symlinks; the ignoreip is now safe on disk, the service jails
need this to be restored.
A foundation template that stands the server up, writes its settings and the mesh's
first user list beside them, and starts a controller on the new bus. The first user list
is the installer's because at genesis there is no mesh to compose one — a bootstrap
credential, rotated like the store's.
The carried list is checked against what the controller derives, since a mesh cannot be
raised twice to discover they disagreed. That check immediately found the composer
granting a role's whole event branch as well as the one event it follows.
What is left of 4.3 is running it, which is 4.1's bed.
Both sides behind a seam, one implementation per bus, and the outcome is the role's
own event so one publish reaches the asker, the controller and the catalogue. Checked
against a running server, including the part the decision rests on: a third party
hears the same outcome the asker does.
Reviews 0110/0121 after a session where renaming seats cost three freezes, a
builder deadlock, and hand-resolved manifests. The seat rules were right; the
set being a compiled Go slice referenced by name-string everywhere was the
mistake. Seats become a table keyed by a stable id; claims/held/production code
reference the id; a rename is one UPDATE, no rebuild, no re-registration, no
freeze. The build machine reads the set from the mesh instead of embedding it,
removing the controller/builder seat coupling. Closed set and scope naming
unchanged; only storage and reference change. Outstanding renames (registry
seats, private-network scope) wait for this — as data each is a write.
The mesh's own seats said who does a job and nothing about what may be said to
them or by them, and that gap showed up three times in one day looking like three
different problems: a build machine with three audiences for one outcome and no way
to derive a grant for any of them; an event genuinely about a role with nowhere to
live but the namespace of whichever module holds that role today; and a catalogue
catching up on builds, where every option needed a grant the design refuses.
One cause — the mesh has roles it cannot describe. So the `mesh-*` seats take the
same three fields a module's seat has, and the machinery that already derives
authority, queues and consumers from a declared seat does it for these too.
Builds become work submitted to a role, and `mesh.build.request`,
`mesh.control.built` and the BUILDS stream retire. A work queue shared by several
build machines is exactly what a seat's `accepts` is, so a second mechanism for it
was two places a permission could be wrong. The outcome is the seat's own event,
which means one publish still reaches whoever asked, the controller that records it
and the catalogue that places it — the fan-out a shared exchange gave for free,
written as a subject the mesh derived rather than a topology somebody configured.
That also avoids the grant that ruled out the alternatives: no holder needs
permission to publish into an asker's inbox.
The blocking gap is now named rather than incidental: the shared library has no way
for a module to publish on a seat. The build machine is Go and reaches the bus
directly, so it is unaffected; the artifact-store event waits.
Records the manual update process (module moved -> build -> reconcile; and the
breaking-change freeze/re-register recovery), and the two things that make
self-update more than a webhook: the build-on-push trigger is currently HAL's
(hal-gitea-tools on :9877), a retirement gap the mesh must replace with its own
forge-webhook trigger wired to every repo including mesh-controller; and the
builder validates manifests too, so a breaking change couples controller +
builder + manifests + hosts, and renaming the builder's own seat deadlocks its
rebuild. Names the transition discipline (accept old+new for one release) that
self-update needs so a push does not auto-freeze.
Every module named its events the way the old bus spelled a routing key, so on the
new bus every cross-module subscription pointed at a namespace nobody publishes to.
Nothing failed — the services started and none of them reacted. Converted, and the
rule now has checks at both scales: at registration for one manifest, and as a test
across the whole catalogue where a consumed event's emitter is present.
It was larger than the report said, in two directions nobody had looked. Forty-three
files of module code pass the event name at runtime, so the code mattered as much as
the manifests. And both clients had to learn the mapping — without that, converting
the modules would have broken the mesh that is actually running, which is the
opposite of what fixing this was for.
Design 29 gained three things it did not say: what a wildcard is (`*` for one name,
`**` for the rest, spelled the mesh's way and derived to each bus's own), that an
event about a role belongs on the seat and why that is not yet possible, and how the
rule is checked — because "a subscription that matches nothing is silence" is exactly
why nobody noticed thirty-seven manifests being wrong the same way.
4.2 and 4.3 are unblocked. The catch-up half of 4.5 is not: it is a decision, and it
narrowed rather than closed. It cannot be a reply to a module's inbox, because that
needs the blanket grant design 25 §4 refuses.
Records the reversal: distribution stays as the mesh's OCI registry (it serves
every artifact-store:// image); only verdaccio, a redundant second npm registry,
is removed. The 'consolidate onto gitea / retire distribution' direction was
dropped. Also records that the node-* rename was executed as one controlled
migration with a brief compose freeze, and why the delivering registry seats
are deferred rather than folded in.
The control plane's seats grew a second naming style (the-*) beside mesh-*,
and the closed set was the only place any seat could be defined. This settles
both: system seats are mesh-* (one, mesh-wide) or node-* (one per node), named
for scope; a module may define its own seat outside the closed set. Folds in
the seat review: mesh-build-machine (scope fix), mesh-private-network (one
server + client modules, dropping per-node VPN choice), showcase becomes the
first module-defined seat, node-uplink, and the node-* renames — plus the
registry consolidation onto gitea, which reshapes the registry seats and gates
retiring distribution/verdaccio. Records why the renames are a coordinated
migration and why distribution cannot be removed until gitea serves images.
Minting on both halves, the file delivered per push, and the bus's objects
asserted on every start — verified against a real server that asserting twice
changes nothing, that a machine joining an already-raised bus is accepted, that
each node's consumer is bound to its own declaration subject, and that CONTROL
does not dead-letter before the controller gives up.
Which bus the mesh is on is one fact, and being told about both is refused at
start rather than warned about: a mesh half on each is one where a declaration
goes out on one bus and the report comes back on the other while every component
logs success — ADR 0074's failure arriving through configuration rather than code.
So 4.1 no longer waits on code. Both links speak NATS, the composition happens,
and every claim behind them has a unit test or a check against a running server.
What none of those can stand in for is a mesh raising itself, which is what the bed
is — this is where the code stops and the lab starts.
Two corrections of fact, both found by building the module's image and connecting
to it as a host would.
The first composed configuration said `verify: true`, which makes the server
demand a *client* certificate — and nothing in the mesh presents one. A host pins
this server's exact certificate and authenticates with the password the mesh
minted, and so does a module's runtime. Every connection in the mesh would have
died at the TLS handshake before any password was looked at, with an error that
reads as a fault in the client. TLS is still required; verify only decides whether
client certificates are checked. Mutual TLS is a later question and would need
machinery the mesh does not have — a certificate per module per node.
And §4 read as though the controller wrote the whole file. It writes the user list
and nothing else: ports, TLS paths and a store directory belong to the container
the module raises. The two files share one directory of necessity, because an
absolute include path is resolved relative to the including file's own directory.
The decision stands in both cases — accounts are composed, not called for, and
passwords are minted and sealed. What changed is what the file says and who writes
which half.
What is in: a bus user's hash is recorded and its plaintext returned once, and
the user list is derived from the machines, what each runs, every manifest and
which machines hold a live token. Permissions stay derived rather than stored,
because a stored copy could disagree with the records it came from while both
looked internally consistent.
What is out, with what each needs, so the next person does not rediscover it:
delivery, which has one open question about what a module declares in order to
receive the file — design 29's ground, not this document's; minting, which is
transport-coupled because an enrolment reply carries one password and a node on
the old bus must not be handed a credential for the new one; and calling the
assertions from a start path.
A roster fact may be shared — written into a marked region of the machine's
file (into: block, hq 128) rather than as the whole file. The template
renders the content; shared decides how the host lays it down. Composes with
hq 128: the region mechanism is the host's, the format is the module's.
Correcting a tick and a claim I made one commit ago. 4.1 does not wait on an
enrolment user per live token; it waits on the whole composition, of which that
user is one input.
Tasks 1.3 and 1.4 are honest about what they built — the composer, the
derivation, the permission model, the stream and consumer definitions, the
asserter, all pure and held by unit tests and a golden composition. Nobody wrote
the caller. Measured: outside the package that defines them there is not one use
of the composer, the permission derivation, the stream set, the stream asserter or
the principal type. Step 1's "done when" claims every account and permission
composed from the manifests, and a mesh raised today would stand up a server with
no user list at all.
It also needs state the mesh does not keep. Design 25 §4 says the file holds
bcrypt hashes and that passwords are minted and sealed exactly as today — but
today the mesh mints one, hands it to the broker through a management call, seals
the plaintext to the holder and keeps nothing. With no management call the hash
has to survive every later recomposition, because the first thing a new module or
a person's access change touches is a file that must still hold every other
user's password. No bcrypt hash is stored anywhere in the controller.
Named as its own task rather than folded into 1.3, so the gap between "the parts
of step 1 exist" and "the mesh does any of it" is visible.
The facts mechanism formatted the roster in Go in the control plane — one
formatter per fact, in the consumer's own configuration language. ADR 0120
makes a fact a path and a template: the mesh owns the data, the module owns
the format, and the control plane holds no format at all.
to-be 29 (operator accounts + what lives under a home) is rewritten to ride
it: the ssh files become roster templates, the whole ~/.ssh is owned with a
found/owned boundary that cannot lock the operator out, keys are mesh-owned
through an SSH CA (existing keys adopted not regenerated, the operator's
personal key signed not minted), and the ssh-agent is a user-scoped service.
All three halves of the host's link are through seams, and the reply address in
the payload is now proved from both ends rather than one — the test asserts the
transport's own field held the consumer's ack address by the time the request
arrived, so a server that stopped claiming it fails a test instead of letting the
reason become folklore.
Two things had to be built for the host to hear anything at all: a node's
declaration consumer, which only the controller may create, and the enrolment
user's inbox, which design 25 §6 names and the composer granted none of. Both
were silent gaps — a node with either missing looks correct and hears nothing.
What remains is a single piece: something that composes an enrolment user per
live token. On the old bus that account is made imperatively through the broker's
management API; here there is no management API, so issuing a token has to
recompose the server's configuration. It is the only thing between the two links
and a mesh raised on NATS from nothing, so 4.1 now says so.
The controller's inbound is through a seam with both transports behind it, and
the store window is now the server's rather than the controller's memory. Seven
claims about that were asked of a running server rather than reasoned.
Wiring the controller's own subscription is what found issue 127: every event
name in the catalogue is still written the way a routing key is, so design 29's
derivation turns a consumer's declaration into a subject no emitter publishes.
Thirty-seven manifests, one that cannot be composed at all. It fails on the first
mesh raised on the new bus and not before, which is why nothing had caught it —
the conformance fixtures pin one emitter against one subject, and both halves of
that pair are correct.
The node-facing flows are unaffected: those subjects are the mesh's own and
derive from nothing a module declares.
It waited on a retirement condition 0119 abolished when it made the
deprecated broker an ordinary provider. A step waiting for a condition
nobody set would sit open forever.
A claim the whole enrolment handshake rests on, now measured against a
running server rather than reasoned from documentation — and held by a test
so it cannot become folklore if a server version changes.
The guarantee is the same and the mechanism is simpler — a nak with a
delay, no parked list, nothing lost when the controller restarts. It costs
one thing: a naked message comes back whatever happened meanwhile, so an
older report is redelivered after a newer was applied. A report already
carries the digest of the declaration it answers, so supersession becomes a
check rather than memory — ordering settled by what a message says, not by
when it arrived.
Two things. A paragraph from the superseded 0117 survived beside the 0119
correction that reversed it, so §5 said both that the amqp interface
retires and that it does not. The stale one is gone.
And the framing. §1 opened with "the bus carries five kinds of traffic
today, and this design keeps the five", with a column mapping each to the
queue it used to be — which describes the mesh's nervous system as a port
of something that did a fraction of this. It now says what the bus is: a
role addressable without knowing its holder, the mesh's own state, work
that queues until somebody can do it, and permissions derived from what a
module declared. Conditions, observation and a person's client land there
too as they are built.
Glossary gains `bus` and `the deprecated broker`, with a note on why not to
say "compatibility broker" or name it after a protocol — the second invites
exactly the backwards framing this commit removes.
The mesh models machines but not the people on them — a node record
holds no username, and no module places anything under a home. So who
you are on each node (jochens/ace/jochen) is unknown to the mesh, and
nothing owns ~/.ssh, dotfiles or ~/.config. HAL knew it; the nox mesh
dropped it. Proposes the account as a node fact and a home-scoped
resource class (the ~/ mirror of ADR 0112's /var/lib placement), with
the login key staying the operator's (ADR 0051). Not urgent — HAL's
generators still run — load-bearing at node-by-node retirement. Found
generating ~/.ssh/config from HAL's registry, which nox has no
equivalent for.
Task 3.2. ADR 0074's model is untouched — floor plus capabilities, partial
implementations legitimate, identity from the credential, dedup on
x-event-id, conformance as executable fixtures. The transport beneath it is
rewritten: exchanges and queues become subjects and streams.
Statements marked *verified* were checked against a running server while
the runtime's client was written, not reasoned from documentation. Three
of them are things the specification would otherwise have got wrong:
- the payload is the body alone, with metadata in NATS headers; an
implementation that nested the whole envelope would agree with nobody
- a durable name may not contain a dot, while the ack subject joins two
names with one — conflating them looks right in a permission list and is
refused as a consumer name
- a certificate must carry a name the bus is dialled by, because the NATS
client has no hook to replace hostname verification the way pinning did
on AMQP
And one limitation lifts: a module may now call another's tool. Issue 049
recorded that a scoped account could not declare the reply queue a caller
needs, and ADR 0095 routed every ask through the control plane because of
it. Per-account inbox prefixes plus allow_responses replace that. ADR 0095
is not reversed — the control plane is still how a person asks — but
module-to-module calling stops being a question about capability and
becomes one about policy, which `uses` already answers.
The NATS client has no checkServerIdentity hook, so pinning no longer makes
the name check redundant — the bus's certificate must carry a SAN matching
the address nodes dial.
A holding is derived at resolution from manifests, never stored, so there
are no recorded old names to rewrite. The work is an edit plus a kept rename
table — kept because a module lives in its own repository and may be
registered long after the catalogue stopped using an old name.
The manifest already uses serves for a provision's facts, so a module's
tools take their own key. Declaring them is itself new — until now a
module's tools existed only in a runtime environment variable.
Storage is not a property of a subject — a stream is a separate object that
covers one — so the question is always how many streams, not which topics
are durable.
Three facts decide it, two of them verified rather than assumed: NATS
refuses overlapping streams instead of merging them, so a shared stream
plus a per-module one is not available at all; a filter cannot express an
exception; and a stream per module turns one cross-module consumer into one
per module. So one stream, with per-subject caps for the fairness that
matters. Per-module age is genuinely unavailable, and a module that needs
it declares a seat.
2.3 was already true and is now proved — the seat refusal is generic, and
three tests pin what matters: a second bus is refused by name, a different
bus implementation is refused too (which is what makes the bus replaceable),
and the AMQP broker no longer contends so both run on one mesh.
2.1/2.2 turned out not to be a no-op. The host keeps a container only when
its spec matches exactly; genesis raises the upstream image and the module
declares the mesh-built one carrying the entrypoint, so assigning it
recreates the container. That is ADR 0067's pivot and it is safe only
because the bus carries nothing yet — which is why step 2 comes before
anything speaks NATS. After it, never again: the config is a directory
mount, so accounts change without touching the container's spec.
Design 29 said no module requires the bus. The catalogue disagrees: 49 of
72 modules take a broker credential and 23 do not, so an ambient connection
mints an account for a third of the catalogue that never speaks — and the
49 each hand-write the path it lands at, which is provisioning done badly
by hand.
The bootstrap argument that made it ambient was narrower than it looked.
"A provisioner needs an account before it can run" is true of a provisioner
process and says nothing about a provision the controller answers, and the
controller is not waiting on a bus account to compose one.
So: the mesh-broker seat delivers mesh-bus; a module requires it and gets an
address, a sealed credential and the trust to verify the server; a module
that requires nothing has no account at all. The requirement delivers the
connection, the declarations shape the authority, and declaring a subject
without requiring the bus is refused as incoherent.
mesh-bus and nats are deliberately two names: a module may run its own NATS
as a backing service exactly as one provides amqp, and a manifest saying
"nats" would otherwise mean either the mesh's nervous system or a private
queue.
The seat's Delivers was wrong twice today — amqp, then empty — and the
comment says so rather than reading as though it were always right.
0117 went a step further than it had grounds for. It was right that the bus
is the only bus, and wrong that the amqp interface must therefore retire —
because it conflated two reasons to want a broker. Using one to reach
another module is a second bus and stays refused. Needing an AMQP broker as
a backing service, the way something needs a database, is ordinary, and
forbidding it would make the mesh unable to run normal software while
calling that architecture.
So the broker becomes a plain provider module: no seat, not foundation,
never raised at genesis, no retirement condition. lavinmq now claims nothing
and provides amqp; nats claims mesh-broker and provides nothing.
The rule that survives is about direction, not software: inter-module
communication goes over the bus. A module may hold a broker for itself; it
may not use one as a channel to another module. That is a review judgement
where 0117 could have used a parser, which is the honest cost.
0106's progressive insight was itself wrong and is corrected by a second one
there — nothing moves off the old broker, so its "one purpose" sentence does
not become true, it is just not what that server is.
The insight check needed two fixes it found itself: a date may carry
trailing words, and a bold run with a link is discussing an insight rather
than marking one. All four bad shapes still fire.
1.1 to 1.4 built and tested. 1.5 turned out to need no controller change:
it already resolves the broker by seat and names no broker module in its
source, which is what ADR 0079 was for. The genesis module set naming is
scenario and installer config, carried with the bed.
Recorded what must NOT change yet: the amqps:// credential shape and the
5671 default are correct until the rollout, because steps 1-4 leave every
node on AMQP.
29 put a module's events and tools in one namespace, 25 kept mesh.events.*
and mesh.tools.*. One namespace is right — a module's authority over its own
name becomes a single pattern the server enforces — but it needs a kind
token, because a stream is a subject filter and mesh.mod.*.> would persist
every tool call in the mesh. Tools stay on core NATS for the reason 25
already gives.
So: mesh.mod.<module>.event.<name>, .tool.<name>, and seats the same shape.
Found composing the first real configuration. Scoping every inbox to its
owner is right and leaves a responder unable to reply, because the answer
goes to the caller's inbox. The fix is not a wider grant but the server's
own allow_responses: one reply to the subject of a message the user actually
received. Without it every tool call times out while the permission list
looks correct.
Versioning: additive is free; a breaking change is refused while callers are
bound, and the refusal names them, because the mesh already holds the uses
graph; a real break versions the subject, not the seat name, so the role
does not fork; binding is a recorded pin, not a drift to whatever is newest.
Semantic change stays open — no fingerprint sees it, and saying so beats
implying the check is complete.
Provisioning: a provisioner's create/remove/holds IS a serves protocol, so a
provision interface is a seat that also delivers a credential — which is why
design 26 already allowed that. The per-consumer resource is what stops the
two collapsing into one.
Secrets: sealed, so the bus is never trusted with plaintext — but sealed is
not enough, because a stream persists and a durable ciphertext is an archive
the day a key leaks. So a secret never enters a stream: core request/reply
only, and a declaration names a secret rather than carrying one, which is
0098's fetch-don't-store applied where carrying is worst. The vault's own
credential and the bus's own accounts are the two bootstrap exceptions,
resolved the way 0067 resolves the control plane.
Also rewrote the addresses paragraph, which was too compressed to follow:
on-bus addresses disappear because nothing stores them, off-bus ones are
untouched and still 0098's problem, and the bus's own address is the one
that cannot be a subject.
The architecture 0117 opened needs a module to offer a service as a role on
the bus — one holder, addressed by what it does. A closed table in the
controller cannot express that: a capability a module contributes would
require changing the mesh itself.
But 0110 closed the set for a good reason — nothing could say what seats a
mesh had, and the hand count came out at eleven of thirteen. That argues for
enumerable, not hardcoded, and 0110 weighed free-form against a fixed table
without considering a third option: closed at any moment and derived from
the catalogue. A derived list cannot drift, which is how the count broke.
So: the mesh's seats stay the mesh's, reserved by the mesh- prefix so the
prefix is the rule and there is no list to maintain; ten seats are renamed
to restore 0079's convention; everything 0110 decided about what a seat IS
survives untouched.
Design 29 carries the declaration model: three namespaces, subjects derived
from local names so a manifest survives the wire changing, queues never
declared, five relationships (the job and state shapes 0041 had no room
for), and the build-publish-deploy lifecycle with hard, soft and build-time
dependencies distinguished.
0041 gets a progressive insight: "no per-consumer setup, only a
subscription" was a fact about a topic exchange, and a JetStream durable
consumer is a real object someone creates.
WBS 1.3/1.4 were wrong and say so: streams come at registration and
consumers at assignment, so only the foundation set belongs at genesis.
Issue 115 is resolved and converted four modules away from named volumes;
the bus's own data is not the place to bring one back. Also: NATS carries
TLS on the client port rather than beside a plaintext one, so there is no
5671/5672 pair to mirror.
A module does declare requirements the provisioner fulfils — but the broker
it gets that way is a private vhost, the analog of a database, not the
mesh's bus. Two modules of the new mesh depend on it, so the compatibility
broker was never single-purpose and its retirement would have stranded them.
NATS is the heart: one bus, a module's messaging is subjects on it scoped by
what it declares, and no module is handed a server of its own. The seat
delivers nothing; the interface retires with the broker. Also closes the
EVENTS question — one stream, on the bootstrap argument, not preference.
Designs 25 and 28 go in-progress: step 1 is starting.
The insight check caught a false positive on its own first real use — its
bold-run pattern crossed newlines and joined an unrelated `**` to the
marker. Constrained to one line, still catching all four bad shapes.
Both repository checks fail on main. 0115 is cited by no design, and it is
marked accepted while resting on 0112, which is proposed.
Design 27 already states the rule the record decides — "a module is assigned
at most once to a node, and that pair is the assignment's identity" — so it
is the home, and now says so. And 0112, 0113, 0114 and design 27 are all
proposed: the batch is under review, so the record is too. Promoting 0112
instead would be marking a record accepted to satisfy a check, which
check_rests_on names as a failure this repository already made once.
PR #133 landed a different 0115 while this branch was open. The bus record
is now 0116, with every citation in designs 19, 25, 28 and the index
following it.
Note: cycle.py and records.py both fail on main as merged, on that record —
nothing cites it, and it rests on 0112, which is still proposed. Both
pre-date this branch and are left for their own change.
A record can assert a fact that goes stale while the decision it supports
stays right. Superseding for that buries a sound record under a second one
and makes every reader work out which is live. So a correction of fact is
now made in place, marked and dated, with the old wording quoted — bounded
by three conditions and checked by records.py, which fires on an unmarked,
undated or back-dated note. Judgements still supersede.
Applied to 0115: no conformance suite exists to recapture, and the full
genesis bed cannot run until the links exist. Designs 25 and 28 follow.
Counting the surface first changed the plan twice: the genesis bed cannot run
until the links exist, so it belongs to step 4, and there is no conformance
suite to recapture — step 3 builds one against the current bus before moving
it. Both corrections are recorded in the breakdown rather than edited into
ADR 0115. Also indexes design 25, which was never listed.
The NATS change was recorded as one undivided item, which hid three gaps:
a mesh already running had no adoption path, the protocol specification did
not know its transport was being replaced, and nothing was runnable until
everything was. Dividing it is what surfaced them.
125: a hold is not a line in the apply report — sixteen resources held
for an untaken module while four surfaces reported success, and the
operator stopped the edge's predecessor on their word (the route-proxy
flip outage). 126: a changed volume path neither recreates a running
container nor warns, and a roll-out upgrade policy makes a build a
deployment — together they turned a data-path migration into a forge
outage (the /var/lib move). Filed as 119/121 in the novox session
before syncing; renumbered past the other session's 119-124.
Not every host path names this machine. A system file the mesh owns is at that path on every machine
of the kind — the path is the fact. The operator's shared data is already answered as an access. A
path inside a container is the software's contract. What is left, and what a node's default layout
would replace, is 514: a module's own data, and what the mesh writes for that module.
Documents the reservation model as the records already have it — a root per node, one directory per
assignment, a placement for adopted data — and the three things nothing states: where the root comes
from, what sits beneath it, and the order the 514 are retired in.
Stops there deliberately. Changing where a definition looks without moving the data does not fail: the
mesh creates the directory, the container starts, the service comes up empty. Retiring these is a data
migration with a verification step, module by module, and belongs with whoever can see the machine.
A path inside a container is not a fact about the machine — /run/secrets and the directory a server
keeps its data in are the software's own contract, true in any mesh that runs it. Only the host side
of a mount names where it landed.
The first sweep matched path-shaped strings, so it counted both halves of every mount and every
in-container location a value mentioned: 798. Counted by role — directory and file resources, the
host side of mounts, accesses, and the targets of binds, grants, receives and secrets — it is 698
across 70 definitions.
The five modules this report first named were what a first look found. A sweep of all 71: 49 public
domains across 26, the node's own name 58 times across 19, a routable IP 12 times in one, 798
absolute paths across 70 (issue 119's number, grown), and no email addresses at all.
The sharpest case is not a domain: a mail module states the node's public IPv4 as the address it
trusts a real-IP header from, so a node that moves or gains a second address stops attributing mail
correctly, silently. Two upstream resolvers are excluded deliberately — naming a public DNS service
is a policy default, true of any mesh, not a fact about this one.
The object store derives each consumer's bucket from the login the mesh minted, and never reads the
one a definition named. The consumer still has to tell its own software which bucket to use, and has
no way to be told: bound values come from the provider's serves, which is a literal block identical
for every consumer, and a provisioner returns nothing. So all three consumers wrote the answer down
by hand and one of them wrote the predecessor's bucket — a key scoped to one bucket and software
asking for another, which reads like a credential fault and is not one.
Records the general shape: any interface where the provider names the resource forces the consumer to
reproduce the provider's rule, kept in agreement by hand and checked by nothing.
Three wordings disagree, and the confusion is the damage: the glossary defines artifact as an OCI
image while the build vocabulary already names four kinds in use, two of which are not images; the
image registry's seat is named after its job while ADR 0079 names foundation seats after their servers
and ADR 0109 names package seats after their ecosystem; and prose that says 'the module's image' reads
as though a module were an image.
Records the question the naming hides: ADR 0075 keeps two provisions because packages and images are
two protocols, and already allows the forge to provide the artifact store. The second implementation
rests on a bootstrap argument, and the forge has the same upstream-server shape the store and broker
have, which ADR 0078 raises as plumbing and adopts in place.
Step 4 claimed the forge cannot exist as a container before the builder has built its image. The
forge's server is an upstream public image pinned by digest; only its runtime sidecar is built. The
store and the broker have the same shape, and ADR 0078 raises both at genesis as plumbing and adopts
them in place — so the forge can be raised the same way and serve git, packages and OCI before
anything is built.
What survives: a grant is minted by the provider's runtime sidecar, which is built, so the question
is whether raise-service, grant, build-sidecar simply works. Sequencing inside the mesh, not images.
The wrong version came from taking a record's bootstrap argument at face value instead of comparing it
to how the store and broker are raised — one command away in the manifests.
Two things were treated as one. The mesh's own artifact store holds the store seat and is internal by
design — reached by name over the overlay, no accounts, ADR 0082. Serving a registry publicly is a
service the mesh can host: a module with its own name, accounts and storage, like anything else it
runs for somebody. The conversion this report was written beside gave the seat holder a second public
door over the same filesystem, which is neither.
Keeps the original issue whole — no garbage collection, and the settings a routine needs are not
enabled — drops the two-door complication, and sharpens one thing: deletion on the only door is
deletion on a door with no accounts, which the predecessor kept behind its authenticated one.
Asked whether merging the three reviewed changes would set a precedent. It would not: five modules
already carry a name belonging to this one mesh — a workflow module stating its host, protocol and
absolute webhook URL, and two carrying a full clone URL for a repository on the mesh's own forge.
That changes what the issue is for. There is no version of this catalogue today that does not name
the mesh it was written in, so refusing three changes buys nothing and a mechanism is the only thing
that removes any of them. The three were merged on that reading, each PR saying so.
Three open module changes independently wrote one mesh's names into the catalogue — two literal
public URLs, because the software generates absolute URLs behind a proxy, and one bucket renamed to
match what exists here. None was careless: the mesh composes <label>.<public-domain> for the proxy
and never hands it back to the module that asked for the route, and no interpolation yields a public
name, so writing the answer down is the only expressible option.
Files it rather than blocking the three, because the fix is a mechanism and the instances are live
needs. The cost is stated: a second mesh installing the identity provider gets the first mesh's
hostname, and nothing distinguishes a literal domain from a version number.
Both halves are on their main branches, so the seats stop being an intention. Writes the as-is
document from the controller's code and the catalogue's manifests: the closed set of fourteen, the
three refusals a claim meets, the holder being an assignment and nothing else, and the one place a
seat changes resolution — which of several providers answers, never whether a requirement may go
unanswered.
Two things the as-is layer exists for are stated rather than smoothed over: a seat cannot answer
before it is held, which is the standing condition issue 121 records; and capacity is not
implemented at all, so the design's bench has no counterpart in the code.
Asked first whether the seats change fixed this in passing, since it landed the same day and
touches both manifests the report names. It did not: the seat's holder is consulted only where
several nodes provide the thing, and with none providing it resolution refuses outright. Genesis
has no exemption — the unchecked first pass exists to learn what each node offers, and a
declaration is never built from it.
Records the part that did change: the three tests left failing on purpose were deleted by the
controller's seats PR and replaced with passing seat-based ones, so the gap is invisible again.
Adds the resolver to located-in, since that is where the refusal is.
The seats half of to-be 27's review is settled, so the two records it rests on are accepted and
the vocabulary catches up: the glossary's *seat* becomes a named role from a closed set, held by
an assignment and possibly delivering a provision, and 23 — Choosing a provider gains the seat
step in resolution, with ambiguity still refused rather than guessed. Both were held back when
0110 was proposed, because a document may not rest on a record that is not accepted.
26 — The seats moves to in-progress rather than designed: it names the files that implement it,
and naming a file claims implementation, which is only defensible once those files are on the
owning repositories' main branches. It becomes implemented when mesh-controller #63 and
mesh-catalog #69 land.
0112, 0113 and 0114 stay proposed; to-be 27 stays proposed with them.
Renumbered from 117, which is taken on main by 'a module's own code is a container in one record
and a process in another' — two reports claimed the same number and git would not have said so.
Scrubbed the node's name and a real registry path; this repository is public.
Keeps 118: the other claimant to this number is on main as issue 119, where ADR 0112 points.
Scrubbed the service's public name — this repository is public — and completed the report's
frontmatter with the fixed-by and amended-design keys every other report carries.
The operator's wish written as intended behaviour for the NATS bus:
every loop compares against what is, repairs by the ordinary path, never
destroys, and raises a condition for what it cannot fix. What is done
before NATS is limited to what survives the move.
The harness compares against its own memory, so a backend that loses
what was provisioned (the cache's ACL users on a server restart) is
never provisioned again, silently.
The fact-check found mailu, whose user is its mailbox, so 0114 rotates
over two credentials rather than two logins, the adapter choosing what a
credential is. Also: minio keeps non-empty buckets; five backends take
their admin credential only at first init, so single-party rotation is
staged; postgres ownership moves to a non-login role; the harness keys by
consumer; rotation state lives with the vault. Consistency fixes across
0110-0113, 26 and 27; issue 103 resolved by mesh-host PR #22.
Graduates research 016. Retiring a login is separated from removing a
consumer, which closes a data-loss path in five providers; single-party
secrets rotate in place; the number of parties decides, not the provider.
Overlap as drafted in 0113 would have deleted consumer data: seven of
eight providers name the resource after the login and five drop it on
remove. Rotation is now undecided in 0113 and to-be 27, pending the
survey. Also: a requirement naming a seat resolves to its holder, a
person chooses among remaining candidates at assignment, the controller's
secrets are requirements of its definition, genesis seals to the
control-node key, and moving the vault or broker is break-glass.
Decided with the author. A credential is never changed in place: each consumer has two logins, both
derived by the mesh, and uses one at a time. An applier adds the new login beside the old through the
adapter's existing create, and confirms both work; only then are readers released to the new one and
restarted by derivation; only when every reader has confirmed is the old login retired through the
existing remove.
It closes the three cases review found in applier-first rotation: an offline reader keeps working on
the old login until it returns; a bus account's owner keeps its bus until it has moved; a provisioner
restarted mid-rotation is still delivered both values. Nobody is ever without a credential that works,
which replaces to-be 13's all-or-nothing rule with a stronger one.
No consumer module changes. The alternation is the provider loop's. A provider's adapter gains one duty,
giving both logins the same rights over the consumer's data — in postgres, membership of one role that
owns it. The mesh derives two logins per consumer, both within ADR 0049's limit, which 0113 now names
among what it amends. Every rule has a check: overlap, offline reader, bus account, restarted
provisioner, equal rights, login length, and confirmation only once the old login is gone.
Decided with the author:
- A seat is held by one assignment, not claimed by a definition. A definition says which seats a module
can hold; an assignment says which it does. The store module can run on every node and one assignment
holds mesh-store; moving a role changes an assignment, never a definition. The foundation's seats name
what the mesh itself uses and route no consumer — database and amqp consumers use co-location, the
holder included. This replaces the wrong rationale that the foundation's store is "provider to nobody",
which contradicted ADR 0078 and to-be 21. 0079's one-postgres rule becomes one mesh-store holder.
- A module is assigned at most once to a node. The instance identity in 0112 and 27 is withdrawn, and the
login-length problem with it.
Review fixes to 0113:
- The bottom of the stack: the vault is installed as soon as the shared runtime base exists, and genesis
generates everything needed until then — including the permanent controller's, the control-node
agent's, the builder's and the broker provisioner's bus accounts, and the controller's store login.
Genesis creates those accounts until the broker's provisioner runs and adopts them.
- Genesis's values are delivered recorded as the mesh's own, so 0092's never-replace rule for operator
values does not make them unrotatable.
- Backend-issued secrets (a forge's once-only API token) enter through the vault. Non-module parties
(the controller's logins, node agents' accounts) are answered the same way, the controller asking on
their behalf; an enrolment token reaches the controller only as what verifies it.
- A secret with no provisioner to apply it is marked not rotatable by the mesh and refused, instead of
a restart reported as done. Unused password generators in six provider clients are removed, and a
catalogue scan checks no module mints.
- Rotation's lock-out cases (offline reader, bus account owner, restarted provisioner) are recorded as
open, with overlap and re-confirm-with-safeguards as the two answers, to be chosen before acceptance.
0110, 0111 and 26 are marked proposed: they changed in meaning and are under review, and an accepted
record must not rest on proposed ones. To-be 23 and the glossary are restored to main; they change when
these records are accepted.
Two decisions taken with the author:
- Genesis delivers and the vault adopts. The vault cannot run first — it is built on the runtime base
the installation makes after the store, broker and controller, and it learns its work over the bus.
Genesis generates the foundation's first shared secrets, seals them to the operator key, and
delivers them to the vault through the path an operator's value takes; from then on the vault holds
and rotates them. This answers ADR 0085's own reason for rejecting vault-only minting, which 0113
now names instead of stepping around.
- Rotation re-confirms on every pass. An applier repeats its confirmation until acknowledged, so a lost
message costs one pass; an applier that stops after applying locks readers out until its supervised
restart, and that window is stated and shown, not claimed away.
Fixes:
- Scope: a shared secret is made by the vault; a private key (node sealing keys, the operator's key,
the certificate authority) is made where it is used. The inventory adds the makers the first version
missed: node and builder broker passwords, and enrolment tokens.
- Broker accounts are created by the broker's provisioner, not the controller, so the controller never
holds their plaintext; mesh-broker delivers amqp again — one broker per mesh — and only mesh-store
delivers nothing.
- secret is a reserved provision: only the mesh-vault holder may provide it, and no pin routes around it.
- A secret's contract says whether a recipient applies it or reads it at start; appliers are never
restarted for it, init-only secrets are applied, and confirmation is to-be 13's standard.
- Operator secrets are one rule everywhere: a secret requirement answered by the vault (0112 no longer
says otherwise). A data provider's adapter may return fields; the data-return check names a lab consumer.
- 'Holder' now means a seat's holder only; a secret has recipients.
A secret comes into being seven ways today: provider credentials, own secrets (54 modules), broker
accounts through a command that is easy to forget, a vault that only records what the controller
mints (6 modules), operator values, licences, and root secrets. The vault was built to end own secrets
and did not; the old path was never retired.
0113 is rewritten as a waterfall. The vault makes every secret and nothing else does. A provider that
needs a secret for a consumer requires it from the vault, declared once in its provision's contract
and expanded per consumer by resolution; the vault delivers it to both holders, each sealed to its own
node, so a provider's code is unchanged. Own secrets, broker passwords, operator values and licence
credentials take the same path. Genesis is not an exception: it raises the vault first and asks it,
so there is one way a secret is made from the first one on. The vault can sit at the bottom because it
requires nothing but a broker account.
One shared mint function in the SDK was considered and rejected: generation becomes uniform but custody
stays spread over every provider's machine, and each SDK language needs its own implementation.
Rotation is asked of the vault and is provider-first: the value goes to the holder that accepts it,
which confirms, before the holder that presents it gets it, so the lockout window shrinks to the
consumer's own restart, and an unconfirmed provider holds the rotation rather than half-doing it. The
host derives which processes to restart or recreate from the requirement a definition reads, so no
definition declares restart-on for a secret. A rotation shows unconfirmed until each consumer restarted
and passed its health check. Issue 103 becomes a prerequisite.
The file is renamed to match what it now decides. 0112 follows.
0113 — the plaintext claim was false under its own mechanism: handing a provider's answer to the
controller puts every secret on the broker and in the controller in the clear. The provider now seals
each secret field itself, to the consumer node's public key the mesh hands it, and the controller
carries sealed fields it cannot open. That is stricter than today, where the controller holds every
minted credential in the clear. Option 3 (plaintext to the controller) is recorded and rejected. The
foundation exception now covers root-secret rotation (0085) and forms like the broker admin's hash, so
no phase claims to remove the broker's bootstrap step. To-be 24 and 13 are named among what it amends.
27 — resolution is consistent with 0110: co-location and the only provider apply only where no seat
delivers the provision, so an unheld seat is refused even with one provider. The secret-field rule now
matches 0086 exactly (a declared env-file, never a container environment value). The seat placeholder
is the controller's, and the one module reading it moves to a host port. Contracts are held by the
controller and written down in phase 1, so they can be checked; every rule has a check. An operator's
secret is still the operator's, with the vault as custodian. Which seats a module holds is listed as
not settled.
0110 — the unheld-seat-with-one-provider case and the one-answer-for-everyone rule have checks; the
claim about moved manifests is corrected. 26 — the table governs and the code catches up, not the
reverse; scope and capacity agree with the glossary; moving a seat is described as it really is today.
0112 — aligned with 27, and lists 0049 and 26 among what it changes.
Issue 118 is renumbered 119: another branch took 118 first. 'Control-plane' is gone from 0110 and 0111.
The design pass. Everything a module needs is a requirement: a name, a contract, and one of four kinds
of provider — a module, the node's host, the mesh, the operator. Installing a module resolves every
requirement or refuses, naming everything missing at once. It retires six mechanisms that grew
separately: provisions through bindings, settings, assigned ports, machine facts, minted secrets and
literals in the definition.
ADR 0113, proposed: a provider makes what it provides, and the mesh carries it back sealed to the
consumer's node. It is the return path ADR 0048 left "to a separate decision", now needed three ways:
data provisions with nothing to answer with, contracts needing a value the controller cannot make, and
a vault that generates nothing. Who a consumer is stays the mesh's (ADR 0049). Genesis is the one
exception. On acceptance it supersedes 0048 and amends 0085.
ADR 0112 is revised from three sources to that single concept.
ADR 0110 is amended for two points raised in review. The vault gets the mesh-vault seat (issue 106).
A seat's holder outranks co-location for a provision it delivers. Writing that down exposed an
inconsistency: mesh-store delivering postgres-database would have sent every database consumer to the
control-node, against to-be 23's node-local stores. So a seat delivers a provision only where the mesh
has one answer for everyone — artifact store, npm registry, git, vault — and mesh-store and mesh-broker
deliver nothing. 23 and 26 follow.
'Control plane' becomes 'controller' in the records written today.
- Secrets follow ADR 0085 as amended: a module's own secret is a provision the controller mints and
the vault records. The previous commit had that backwards. Whether the vault should generate
instead is recorded as an open question, not decided.
- A directory's contract is owner and mode only. The persistence flag was the keep flag ADR 0030
refused; a directory is kept while it holds anything, and disposable data is a named volume (0107).
- An operator's shared data stays an access (ADR 0051), which rejected an operator-owned directory.
Only where its path is written moves to the assignment.
- The records it changes on acceptance are named: 0051, 0091, 0046 (settings keyed by instance),
0084 (a provider is a node and an instance), and the glossary, which gains its new words only
when the record is accepted.
- How it is checked covers every stated rule. Container-side paths are no longer flagged by the
host-path rule, and code fallbacks are covered.
- Provisions are what other modules provide. A seat's occupant is not listed as one, and the vault
is not described as selectable per assignment.
- 'Control plane' becomes 'controller'. The provider count is ten of eleven, not eleven of twelve.
The first draft listed minted secrets under what the mesh generates. ADR 0085 made a module's own
secret — a password, an internal token, an external key it was handed — a secret provision answered
by the vault, like a database by the store. What the mesh still mints is the delivery credential for
each provision a module takes (ADR 0048), the vault's own included.
Issue 118 records what a review of where module code reads its files found: 789 host-path strings
in 70 of the catalogue's 71 definitions, every one a decision the definition makes about a machine.
Mounts are checked (ADR 0091); the same paths retyped as values are not. It records what that has
already allowed — a DNS provider that would provision nobody silently, a contributions file that
names credentials by host path and so forces every provider to mount at the identical path, an SDK
loop that treats an unwritten contributions file as empty without a word, defaults in code that
disagree with their own manifests — and that no module can be assigned to one node twice, because
every identity is keyed by the module's name.
ADR 0112, proposed for review, answers it the way ADR 0038 answered ports: a definition names
variables, and installing it resolves every one or refuses, from three sources — the assignment's
own configuration, provisions the mesh resolves against a contract, and what the mesh generates or
knows. A directory becomes a provision: the module requires one by name with its owner, mode and
persistence, and where it lands is the assignment's. The mesh's own files stop carrying host paths.
An assignment gets an identity of its own, so a module may run twice on one node.
Checking copies for agreement was rejected as checking something that should not exist; rewriting
paths per assignment was rejected as inferring which strings are paths by their shape. Syntax, a
node's default layout, and when a second instance becomes possible are left to the design.
The enumeration behind the first set read manifests in two repositories and missed a claim made
in the control plane's own code: the private-network module it ships claims the-private-network at
node scope. A closed set without it would refuse the control plane's own module. Thirteen claims in
use, naming twelve seats.
Seats have been doing two jobs and neither is written down. The mechanism ADR 0009 introduced is
enforced — a second holder is refused — but any well-formed name becomes a seat by being claimed,
and nothing can say which seats a mesh has or who holds them: holdings are assembled while planning
and discarded. The enumeration done while preparing this missed the control plane's own manifest,
because core modules' manifests live in its repository rather than the catalogue.
0110 closes the set. Each seat has a name, a scope, what occupying it delivers, and the record that
made it one; a claim outside the set is refused. A seat is held by a module assignment, and what the
mesh knows about the holder is what it knows about that assignment — nothing is stored beside it. A
seat may deliver a provision, and then its holder answers for it among several providers: pin, then
the holder, then the only provider, then refused. That keeps 0009's "refused, never guessed": the
seat is the choice made once, mesh-wide, instead of a pin per consumer node. The first set is the
eleven seats already claimed plus 0109's npm-package-registry, so nothing in use is refused.
Two concepts — seats for exclusion, a new word for consumable singulars — was rejected: both mean
"this mesh's one X", and the overview a person wants is one list.
0111 gives the mesh a git seat and makes a build source one of two explicit forms: a repository on
the seat's holder, recorded by its path and cloned from wherever the holder runs at build time; or
an external URL, recorded and cloned exactly as given. Recognising self-hosted sources by matching
URLs against the forge's address was rejected — it fails in the one case it exists for, after the
forge moves. Credentials for private repositories are left undecided and said so.
Design: new to-be 26 (the seats); 23 gains the seat step in resolution; 18's source entry names
the two forms; the glossary's seat and provision entries say where they meet. 0109 is carried from
its own branch so every link here resolves.
Extends ADR 0075. Surfaced fixing builder's hand-faked package-registry
binding tonight: gitea's manifest declares the provision once with a single
npm-path, conflating what should be independently assignable per ecosystem
(npm/cargo/docker/...) the same way artifact-store and package-registry
were themselves split. Cited in 22-the-work-ahead.md's Phase 2, where the
target state this decision points at was already described a week ago.
Numbered 109, not 108: route-proxy's policy feature (mesh-controller PR
still-unwritten decision record — reserved but never committed. Renumbered
around it rather than colliding.
Fixing builder's hand-faked package-registry binding tonight (requires:
package-registry, a real mesh grant instead of a hardcoded JSON fragment)
broke three tests describing a deliberate carried-binding fallback for
exactly this: gitea's own image is built by builder, so builder cannot
yet hold a real grant from gitea the first time either has to exist.
Invisible on novox (already bootstrapped, gitea already live) — real on
any genesis from scratch. Fix left in place, tests left failing rather
than reverted or hacked, so the gap stays visible.
Filed 2026-09-24 on a branch of its own and never merged, numbered 113, which is taken. 114 is
free because a sibling branch folded it, so it takes that number and keeps its commit.
Kept separate from issue 117 rather than folded into it. 117 asks the same question of every
module and locates the missing decision; this asks it of the controller, where `network: host`
means container network isolation — the property that resource type usually buys — is not in use.
That observation is this report's own and is nowhere in 117, and folding would lose it.
Its first open question is answered by 117's diagnosis and now says so: the host's `process` shape
is built, applied and tested, restart and run-to-completion semantics included, so deciding this
does not wait on host-side work.
Filed after a session where every mesh-controller interaction went through
docker exec — its manifest runs it as a container with network: host, using
none of the isolation that resource type usually buys, while ADR 0006 makes
it the mesh's single point of coordination. Open question, not a claimed
defect: does type: container get the controller anything type: process
(supervised the way the host supervises its own unit, per ADR 0005) would not.
Asked what the "sidecar" is and whether a supervised process would do instead. The repository
answers both ways. ADR 0047 (accepted, unsuperseded) says a module with tools or events runs a
container carrying its compiled code. To-be 18 and 20 (both proposed) define a `process` resource
type — the module's own code, a unit the machine's supervisor keeps up — and the worked guide says
plainly "it is why these are `process` rather than four containers." Neither design doc names 0047,
and no decision record mentions a `process` shape at all.
Diagnosed rather than left open, because the ground truth settles what the report could not.
The shape is real: mesh-host defines TypeProcess, applies it, and tests it, and the host's
vocabulary is twelve shapes rather than the nine ADR 0029 counted. So the alternative the report
offered — that two proposed documents describe a type that does not exist — is disproven.
ADR 0029's mechanism is intact and was not enough. The vocabulary-count test names the decision
behind each addition: network 0029, access 0051, opening 0100. The eleventh names a *proposed
design document*, and TypeProcess is the only shape in the vocabulary whose doc comment cites no
ADR. Requiring every addition to name something does not require it to name a decision.
The argument this issue asked for already exists — as a Go test comment. "It is a full-host shape
rather than a portable one: it needs a process supervisor to install into. It does NOT need a
container runtime, which is the point — only software that genuinely needs isolation asks for a
container." That is a decision's context and consequences, in another repository.
What the catalogue does is a third thing: 115 container declarations against 3 process, all three
in showcase — the module to-be 20 documents. There the tools resource is a container running
`sleep infinity` on a bare upstream base with the broker credential mounted, and the tools and
provisioner entrypoints are run by nothing. That is the condition 0047 was written to end, back
in a new shape.
Where the isolation argument leaks is narrower than expected and worth having precisely: the
serving key and the credential shape both conform. But serveTools serves every registered module
over one broker connection, the runtime takes its modules from a comma-separated list, and
x-source is stamped from the single credential — so two modules in one runtime means the second's
events are attributed to the first. Nothing refuses it and no test asserts against it.
Located on hq rather than on a code repository: the implementation and the design layer agree,
and the missing thing is the record. Which shape is right is left open, deliberately — this
establishes that the question was answered in practice and never written down, not which answer
is correct.
One correction kept in the trail: the first search here was for len(Vocabulary()), found nothing,
and was two steps from being written up as "the mechanism ADR 0029 relied on is gone." The test
binds the slice to a local first. A negative search result read as a fact about the world is the
same error issue 113 recorded.
The gap is closed in the proxy: policy applies, the four capabilities exist, the table is keyed
by host and path with a total ordering, and the two failure modes that rot quietly are held by
tests — a declaration carrying a credential refused rather than served, an unreadable secret
failing closed.
Resolved rather than left open because the issue reports a gap in the proxy and that gap is
gone. But the record says plainly what it does not yet allow: an operator still cannot move the
affected routes, because that needs the mesh side — a manifest able to declare these values and
the controller minting the secret auth names. Until both exist the capability is reachable only
by writing the routes file by hand. That is the ordinary build-out of a contract this issue's
decision created, and it belongs to to-be 08 rather than here.
The open questions are marked answered and kept rather than deleted, pointing at ADR 0108 —
what was rejected and why is the half worth having, and a section still saying "the fix should
not be written before these are answered" after the fix was written reads as though nobody
looked.
One finding kept in the record: priority was read with the reader for ports, which caps at
65535, and the one real rule this reproduces is declared at 100000. It parsed to zero, so
refusal and path scoping would have shipped looking complete and doing nothing on the only case
that motivated them. A validator borrowed from a neighbouring field is a silent default.
Issue 116 found the mesh's proxy applies nothing to a request — host lookup, forward. Against
what the replaced ingress actually relies on, four capabilities are missing: authentication
(three dependents, each gating an admin surface with no login of its own), refusal scoped to a
path (one, a live incident mitigation), path-scoped routing with priority, and redirect.
Policy goes on the route rather than beside it. A proxy-side settings layer keyed by route name
would keep the grant literally clean, but then "what protects this route" is answered from two
files nothing keeps in step — and a route's protection is part of what a route is.
The set is closed at those four, so a fifth is an amendment and each addition is earned by a
dependent that exists. An open middleware surface was rejected: it recreates what is being
replaced, and narrowing one later is far harder than widening a closed one.
Where policy needs a credential the declaration names a secret and never carries the value,
which keeps the existing secret machinery the only thing holding credentials. Inlining a hash
was rejected as the first credential in a declaration — a precedent easier to set than withdraw.
This re-keys the routing table by host and path with priority, which follows from the decision
rather than being a separate one: two of the four need one host routed more than one way. Equal
priorities must resolve identically every time or the proxy stops being reproducible.
The record says how it is checked, including the negative case that rots quietly — a
declaration carrying a credential value rather than a reference must be refused, so the
rejected option cannot return by accident.
08-connectivity §3 names the record and gains the subsection; issue 116 gains amended-design.
Two things the report got wrong, and one it could not have found the way it looked.
Disclosure first: it carried a real hostname and an absolute node path, in a public
repository. Both are gone; the ingress, the modules and the routes are named by role, as the
rest of 04-ISSUES does.
The count was low. Basic authentication has three dependents in the catalogue, not one — the
key-value store's browser UI, a database web UI, and the ingress's own dashboard. All three
are credential-less admin surfaces whose only gate is a middleware the mesh's proxy lacks.
The earlier version read only the node's dynamic configuration directory, which cannot see
what modules declare as container labels; counting needs both sources, and the report now
says so.
Two gaps were missing entirely. Redirect rules: two live routes canonicalise a www name onto
its apex, they exist only on the node and not in the catalogue, and they fail silently rather
than erroring. And path-scoped routing with priority, which is the one that reorders the
issue: the table maps host to exactly one target, so a host cannot be routed two ways, and
the refusal rule matches a path on a host already routed elsewhere. Authentication and a
source filter would not make it expressible. Path scoping is a prerequisite, not a sibling.
Also corrected: the refusal rule was described as an address-scoped deny. It is an allow-list
holding a single documentation-range address — deny-everyone — so reading it as address-scoped
points at the wrong fix. And its severity was understated: its own header records it as
incident response closing an abused write primitive, which is not "a real exposure" but a live
mitigation.
The open questions now say plainly that they are design questions and the fix should not be
written before they are answered, and one is added: whether a declaration may carry a
credential at all.
Comparing route-proxy against what HAL's actual Traefik config does today,
not Traefik's general feature set, per the standing rule that the nox mesh
must do at minimum what the HAL mesh it replaces already does. Everything
else checked out even or better; these two are real, confirmed gaps —
RedisInsight has no login of its own and depends entirely on Traefik's
basicauth middleware, and the gitea-internal route depends on an IP-scoped
deny rule. Neither has any equivalent in route-proxy's single-lookup
request path.
The object-store module was repinned to a maintained fork of the withdrawn server image, its
runtime sidecar built rather than pulled, and its data moved off the predecessor's live
directory. The instance is closed; the three general points the report makes are not, and What
was done says so rather than letting a resolved status imply otherwise.
Folds in the one thing a duplicate report of this symptom had that this one did not: issue 064
asked whether the build environment can reach a declared vendor image and assumed that, once
declared, it stays fetchable. Withdrawal is the case that assumption does not cover. The
duplicate is not merged — it carried the reading this report's diagnosis retracts.
The previous version ranked candidates on whether they preserved single sign-on to the
object store's console. That is not a requirement: a "user" of the store is normally an
application, so the requirement is per-application keys scoped to buckets — which the mesh
already mints. And the console login it ranked on never worked; the module's own hook comment
records "policy claim missing", a failing login written up as progress.
It also omitted the incumbent's own maintained fork, which changes the question from "which
product replaces it" into two decisions: repoint, or migrate — and if migrating, to which.
Repointing costs an image reference; migrating costs a data copy, two handler rewrites and a
maintenance window. Repointing does not foreclose migrating, which is the argument for taking
it first.
On the corrected requirement Garage ranks first — its per-key-per-bucket model is the
requirement verbatim, its admin API matches how the mesh provisions, and the highest-risk
consumer is first-party documented against it. Its remaining gap (no versioning, no
server-side encryption, partial lifecycle) is unmeasured against the buckets and is the one
thing that could still disqualify it.
Measured and folded in: 230 GiB logical, 82,496 objects, 468 GiB raw at 2.03x, eight drive
directories on one filesystem on one machine. That last fact decides more than any feature —
the erasure coding is not buying independent-drive redundancy, so the redundancy model is
close to irrelevant and only storage overhead remains, which at this volume is a rounding
error against the headroom.
Both errors are recorded at the end of 01 rather than quietly fixed. A configured feature is
not an observed one; and when a dependency dies, "who took it over" precedes "what replaces
it" — searching for alternatives by construction returns things that are not the incumbent.
The diagnosis carried a table headed "claims that could not be substantiated", denying a
module.json, a digest pin, and an all-zeros runtime digest. All three exist. The table is
withdrawn in full and replaced with what is actually true, plus the two claims that remain
genuinely unverified rather than disproven.
The cause: one repository was searched and absence in it was written up as absence. The
catalogue of the mesh being built is a separate repository, not checked out where the search
ran, and all four claims were about that repository. Compounding it, the predecessor's
object-store module and the one being cut over to were treated as one thing — they are
different files in different repositories, one pinning a tag with no sidecar, the other a
digest with two container resources.
Also corrected in the report: located-in named the wrong repository; the "pins a tag" passage
described the predecessor; the open question about pinning by digest is struck, because this
module already does and it made no difference — a deleted digest resolves to nothing either
way. The section on why nothing broke is now scoped explicitly to the predecessor's
machinery.
The lesson kept in the record: "zero occurrences anywhere in the tree" is only as strong as
the tree searched, and a diagnosis must say which tree. A confident rebuttal of a correct
report is worse than no diagnosis — it sends the next person to the wrong place with a
written record behind them.
A thought raised mid-task needed remembering but not working on, and there was no
mechanism for that — so it was recorded by hand. This is that, made repeatable.
Records to Claude's persistent memory rather than the repository, deliberately. A parked
thought has no number, owner or status: giving it one asserts triage that deferring says
has not happened. A shared "deferred" document would be a central status file, which
AGENTS.md forbids. And a repository write means a branch, a commit and an MR — the drift
the skill exists to prevent.
Wraps no playbook, because deferring precedes the development cycle rather than being part
of it. It does say which playbook a thought would need if it graduates, and that recording
"undetermined" is the honest answer when the evidence does not say.
The stop condition is the substance: at most two lines, then return to what was in
progress. No plan, no triage question, nothing opened.
Records the rule the operator gave directly, mid-session, after checking
that HAL's own postgres and lavinmq both used a directory bind and the
mesh's adoption of them three weeks ago switched to a named volume without
a reason recorded anywhere.
Already built and rolled out on novox (mesh-catalog PR #54) before this
record -- urgent enough to fix first and write down after. Includes the
incident: the new host directories needed the container's own UID, which a
named volume gets for free and a directory bind does not; mesh-store
crash-looped on Permission denied until ownership was matched to what the
original volume already had.
Closes issue 115. Checks pass.
SeaweedFS was scoped as primary because it looked like the only candidate preserving
OIDC console login. Measured: its admin UI is Apache-2.0 but its identity-provider
integration is not — console SSO sits behind the per-TB commercial licence, alongside
point-in-time recovery and automatic EC repair. The free build gives OIDC on the S3 API
via STS and a console authenticated by local username and password.
So the answer to the gating question is that no candidate preserves the current feature
set for free, which this effort had written down as a possible outcome. Reopened across
three candidates with the requirement-by-requirement evidence in 01.
Two corrections to what the overview recorded. RustFS is not a binary-level drop-in
retaining existing data: API and on-disk compatibility are separate paths and the on-disk
one is preview-scoped. And it carries an open defect in the credential path the bucket
provision depends on, which gates it specifically.
Nothing graduates before two measurements named in 01: whether an authenticating proxy
is an acceptable answer to console SSO, and which S3 endpoints consumers actually call —
the latter because Garage does not implement the full span and cannot be ranked until
that is counted.
The decision (the mesh assigns a container's machine-side port; a module
says only what it needs) was proposed 2026-09-01, and the machinery
already implements it in full -- internal/inventory/ports.go's PortFor,
declaration.go's publishedOn. What was missing was the catalogue actually
complying: 14 of 46 modules baked a machine-side number into their own
manifest anyway. mesh-catalog PR fixes 11 of them (the two defensible
kinds -- foundation, protocol-fixed -- are left alone, per the issue's own
categories). Accepting the decision now that it's actually enforced, and
closing the issue it was blocking.
Checks pass.
Review of my own text found an unattributed claim — "the mesh's claim that a node
can be rebuilt from its declarations" — which is not a stated principle anywhere.
Replaced with the design position that genuinely covers it: to-be 07 chooses
references over payload because "reproducibility comes from pinning the identity of
a thing rather than carrying its bytes". This incident is that choice's failure mode
when the identity stops resolving, which is a sharper point than the one I made.
Scope stated honestly: the passage is about the foundation bundle and this module is
not in it, but pin-identity-fetch-bytes is how every module gets third-party images.
Also names the tension the mirroring question actually carries — mirroring is a move
away from references-over-payload, so it is a decision, not a fix.
The symptom arrived diagnosed as "the registry disabled anonymous pulls for the
whole vendor namespace". It did not hold: sibling repositories in that namespace
pull normally, the "$disabled" token field appears on every repository including
working ones and describes signing rather than access, and "actions": [] with a
401 is byte-identical to what an invented repository name returns. Both registries'
own APIs establish deletion instead.
Recorded because the correction is the expensive part to rediscover, and because
the instance was harmless while the standing condition is not: no node that does
not already hold the images can ever provision the module again, and nothing
detects that until one tries.
Research 015 scopes the replacement. It is not a redesign — the foundation design
already commits to S3 the protocol rather than the product, and the object store
is an ordinary module, so this instantiates an existing principle. The live OIDC
wiring is the requirement that gates the choice, and it is checked first.
Playbook 03 step 2: move status to diagnosing, then located once the
owner is known. located-in is filled with four confirmed packages —
the owner is known.
The carried-peer record (CarriedPeer/TunnelPeer) lives in mesh-controller
internal/inventory, not internal/catalogue. internal/catalogue is the
right package for the zone-generation side of the fix (facts.go's
nodeZones), but a different concern from where the name field itself
would go. Split the two so a decision record doesn't get pointed at the
wrong package.
Fixes, each named where it was wrong:
- reload-on is a service field; a container only has restart-on, which
recreates. Cited precedent (registry-trust-reload) is a service resource,
not a container. Fix: nats-server's own SIGHUP reload, triggered by an
in-image entrypoint watching a directory-mounted config file (issue 103's
recreate-on-change applies to a directly-mounted file, not a directory's
contents) -- asks nothing new of the host.
- A JetStream-delivered message's Reply field is already claimed by the
consumer's own ack address, so a responder using it answers nobody. Fix:
every CONTROL message needing a reply carries its reply subject in its own
payload; the controller publishes there explicitly, never via Respond().
Enrolment is the case this design actually depends on, so it's fixed there
too, not just noted.
- The listed permissions never granted publish on a durable consumer's own
ack-reply subject -- a module could receive but never ack, so every
message redelivers forever. Fixed with a scoped grant per module's own
consumer.
- One account (a deliberate choice, kept) means inbox privacy is the
permission list or nothing. The design granted 'its reply inbox' without
scoping it, which read as any user reaching any inbox. Fixed: each user's
inbox prefix is derived from its own identity and its permissions name
only that prefix.
New open question from this revision, not closed: whether the in-image
watch-and-SIGHUP shape belongs in mesh-sdk if a second module ever needs it.
Checks pass (records.py, cycle.py, index.py).
Checked why ADR 0104's forward-to-predecessor shape doesn't transfer to the
resolver the way it did the proxy: DNS is one process on one port, and
assigning the mesh's dnsmasq module replaces it in place, so there is no
predecessor process left standing to forward to.
But /etc/dnsmasq.d/hal-dns.conf's static address= lines for ace/shanks/g14
match the mesh's own carried-peer addresses from overlay show exactly. The
name a carried peer needs isn't a guess the operator has to make under
pressure — it's a transcription of a record the predecessor already has and
has been correctly serving for six days. Located in mesh-controller (no way
to attach a name to a carried peer today) and the dnsmasq module (doesn't
emit a wildcard for a named-but-uncarried peer). Not implemented.
hq cannot hold the migration's operational record — it names machines, addresses and
paths, and this repository is public — but it can say where that record is, which is what
this map is for. Asked for by the operator, who had to be told.
111, resolved: the map the control plane hands a resolution holds the machines and the
names the mesh merely serves, and the resolver's zones were given both — inventing names
under a suffix it answers authoritatively for. 112, open: adopting a tunnel gives the mesh
the peers' addresses and none of their names, so taking the resolver before they enrol
stops three machines resolving at all.
Both found by reading the plan before pushing it.
Found reviewing the resolver's conversion. Nothing fails while the node is adopted; it
fails at the flip, and it is the same split that decided which container survived the
hub's address change.
Found when adopting the tunnel moved the hub's address: the declaration followed, the
running container did not, and the forge lost its database. Issue 102's rule broken one
level down, and issue 103's fix stopping one input short.
Implemented in mesh-controller #49 and mesh-host #24. One proposal was rejected on the
record's own terms: converging the hub is not made to wait on other machines' migrations.
The addresses follow: the control plane holds the ports the node gave, a recorded build
holds no address at all, and the two forwarders that had been holding the control plane
together are removed. 097's stranded container turned out to be listening on every
interface and connected to the mesh's store — by its own old database, which is the only
reason nothing was at risk.
Both from the review of the issue-104 fix: a hand-applied file on an enrolled node is
recorded as carried and would remove the foundation, and nothing on the wire orders one
declaration against another.
Measured: AMQP is spoken in three places of the mesh's own code and in none of the
sdk or the modules; the predecessor's world is AMQP and retiring. NATS answers every
guarantee the bus relies on, durability via JetStream. Recommended: not underneath
the migration, not after it either — in parallel, one rehearsed rollout.
Two birth-address outages and a registry that would have been the third; a
container that keeps a stale environment after its file changes; a host command
that applied a converged declaration to an adopted node; the hub and the vault
without seats. And the decision the operator made under it all: the hub adopts
the predecessor's tunnel in place, key and peers and range and port.
Found checking the third cutover rather than running it. The first two were safe by
accident — both are reached through a host port, which survives a change of owner.
This is the first constraint found that decides the order of the migration.
Found at the second cutover. Carrying a value in works only for a module's own
secrets; a secret answered by the provision is minted, and six catalogue modules
take one that way.
Three modules in a row on one machine; the first was found by taking it and cost a
three-minute outage. The runbook's answer is a rule a person must remember, which is
the shape this repository says not to settle for.
The first pass answered under both ends, which review showed is the same fault seen
from the other side where two mappings share a number. Verified on the machine: the
forge is back on the port its own configuration has always advertised.
Found reading the second module's cutover rather than running it: the catalogue's
config drops a rule the machine's has, and no step puts the two side by side.
094's cause is one blind spot read from two ends, written up in its diagnosis; the fix
answers the first open question and not the other two, which become 096. 097 was found
looking at what the forge's cutover left running.
The controller kept one report per machine, replaced, so a resource nothing can
ever apply looked like a failure that had just happened, every few minutes, for
ever. It now counts identical reports and status says stuck after three.
The playbook, the README and the status skill knew five statuses; the cycle check
knew a sixth, 'fixed', and not 'wontfix'. Eleven issues sat in the sixth for weeks
with their fixes shipped, one step short of closed. They are resolved; the check
refuses the word from now on and accepts the one the playbook allows.
ADR 0069 had already placed the controller's manifest in its own repository, and the
raising design called the catalogue copy a thing to remove. The installer's step 3 was
handed that manifest by the build and step 9 read a second copy anyway.
The end-to-end design held 'the run rebuilds what it tests' for binaries and images
and not for manifests. The beds' inline copies fell into three kinds; only the first
is a stale copy. The other two are named: a mesh test wearing a catalogue module's
name (074) and a stocked runtime image the run never rebuilds (075).
A seeded file is created once (0087); the foundation filters before anything
listens (0088); a machine becomes the last thing it was told (design 05).
Each says how it is checked.
Closes issue 041 by decision and by code on the same branch: the catalogue
engine refuses a secret in a container's env, and a secret-carrying env-file
unless the container declares its reason; the controller reads all six of
its credentials from files; design 13 states the rule and how it is checked.
An audit of the six code repositories found eleven open issues fixed on main
with commits and beds to show (025, 027, 033, 036, 037, 040, 045, 047, 050,
052, 053), three partly (007, 026, 035), nine not (020, 031, 041, 046, 049,
054, 064, 065, 066) and one whose fix would live outside those repos (006).
Resolved ones name their evidence; partly ones say what remains; 041 records
that the exposure has widened since it was reported.
The installer half of the amended ADR 0085 is built and proven by the
genesis bed's root-secrets step; the foundation design closes its open item
and the installation design says what the installer does and what it still
cannot check.
A bed run from .work/<slug>/mesh-lab derives mesh-tools and mesh-sdk by
sibling path and fails at once when the directory holds only the touched
repos. Detached worktrees on main, never symlinks.
Recorded on the record, dated, before anything shipped against the sentences
that change. The vault is installed at genesis like the store and broker, one
per mesh, and holds every secret a module has for itself sealed a second time
to an operator key whose private half never enters the mesh — the break-glass
path the first version left open, without a key one place holds.
Design 24 says how; 07 and 21 say what genesis does not yet do; issue 071
names the fixed credentials the foundation is raised with today.
Design 24 flips to in-progress with mesh-catalog as its owner (playbook 04).
Starting the build surfaced two gaps the decision did not settle: a module
requiring `secret` receives exactly one value (069), and no command can
accept an operator's value into a consumer↔vault pair (070). Both opened as
issues rather than improvised around.
Also fills fixed-by on 067 and 068, which the cycle check refused as resolved
with no reference.
ADR 0084 (extends 0027) — a provision is served by a node-scoped provider the
consumer selects, defaulting to co-location; a module may instead carry a private
embedded instance that is not a provision. Design: 01-to-be/23-choosing-a-provider.
ADR 0085 (extends 0031) — a secret is a provision and the vault is the module that
provides it; a module's own local secret becomes an ordinary pair credential that
rotates through the existing machinery, while the controller's provisioning-credential
mint (0048) is unchanged. Design: 01-to-be/24-the-secrets-vault; doc 13 amended to
cross-link the non-pair secret.
Issues 067/068 marked resolved with amended-design set. ADR index regenerated;
records and index checks pass.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Store and broker seats were made ordinary modules; secret-minting is still a
privileged property of the controller that no module owns. Propose the vault
become a module that provides a secret provision (generate/hold/rotate/backup/
audit), node-scoped like every other provider (issue 067), subsuming the three
secret paths and giving local-secret rotation a home.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Naming which provider is only half of how a module gets a database. The other
half: a module may carry its own version/fork-pinned instance, module-network
only, no published port, not a provision — and the model has no word for it.
Add it as a second axis with its own open questions.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The mesh models provisions as mesh-scoped (one provider of a kind, a single
mesh-store). But node-specific services delivered to the mesh was the plan from
the start: both nodes already run their own postgres, SQL server, redis and
object store, and identity — currently single — already serves apps on a second
node. The model cannot express which provider serves a consumer, so it collapses
a deliberately per-node fleet to one. Provider scoping is a whole-mesh decision
across postgres/s3-bucket/oidc, not an SSO patch.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Surfaced while reviewing the host apply loop for the controller/host
split discussion. 065: a permanently-failing resource retries for ever
with no escalation — reported, but never raises its hand as stuck.
066: a partly-applied declaration leaves a mixed state with no rollback,
which is harmless for independent resources and unexamined for pairs
that are only correct together. Both are questions HQ must answer, not
incidents — status open, no fix proposed.
The P1-sweep PR left both at 'located'; the fixes have since merged
(mesh-controller #31, mesh-tools #10) and the bed proves them. Flip to
resolved with their fixed-by.
060: 44 of 70 modules now mesh-buildable (was 8) — the mechanical
majority, proven by direct builds and the bed. The structural
remainders: external-dependency fetch (new issue 064) and route-proxy's
cross-repo build context. 064 records the build environment's isolation
from public npm and Docker Hub.
The broker's port was accepted on input but not forwarded, so a joined
node reached a DNAT'd broker only until it restarted. Fixed in
mesh-controller (25e42b3), proven by the built-store-cross-node bed
adopting the foundation's broker (which restarts it) and the joined
node still receiving declarations.
One push leaves the mesh consistent (the 057 decision, proposed for
acceptance); the shared runtime waits for its broker (058). Fixes on
mesh-control fix/one-push-is-enough and mesh-tools
fix/the-runtime-waits-for-its-broker; the built-store-cross-node bed
enforces both.
Fixed by mesh-controller PR 30 (79a9e17), rebased onto the merged
registry-reach train and proven by a green fresh run of the no-fake
two-node bed built from that commit.
One green fresh run of the no-fake two-node bed is the proof: node2's
consumers open the store and broker the mesh built and adopted, over
the overlay. Fixed by mesh-controller PR 29, mesh-catalog PR 26, gated
by mesh-lab PR 34.
061: the broker module's provisioner never ran; its runtime container
named no command and the image default is the tool host. 062: a failed
artifact-store lookup composed the network without the registry trust,
turning a transient error into permanent silent state. Both located,
fixes on the 042/048 train branches.
Only 8 of 70 modules carry a build section; the other 62 exist through an out-of-band
build script plus placeholder rewriting, so the mesh's own build-and-deliver path has
never run for them. Found by the no-fake bed; blocks it and the migration's delivery
assumption.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Decisions were the one link the cycle checks skipped, and measuring found 19 of 70 records
orphaned — the credential flow and the module-runtime cluster among them, which is how a
stale premise about a settled decision survived in working memory. cycle.py now refuses an
accepted record nothing cites; the 19 got true homes (design frontmatter, the playbook that
implements 0021, META for the process records). The overview names the practice: spec-driven
development with provenance.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Findings of an adversarial review of the 055 fix: hub!=broker conflated, membership tested
as has-address, silent staleness, portless silent fallback, two hubs unrefused. One review
claim recorded as disputed against live wire measurements.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The flow the process overview draws — idea/symptom -> decision -> to-be design -> code ->
as-is — was enforced by nothing. cycle.py now refuses a to-be design naming no decision, an
in-progress/implemented design naming no owning code, a located/fixed issue with no owner,
a fixed/resolved issue with no fix, and a graduated research overview that does not say what
it became. AGENTS.md carries the cycle and a where-to-look table so a fresh session (or a
cleared context) finds the chain in frontmatter instead of assuming it. Grounding the check
surfaced two real gaps, fixed here: the work-ahead design named no owning code, and research
003 listed one became target twice.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The glossary's authority page still named the controller's seat the-controller in two
entries, contradicting its own seat section after ADR 0079; issue 058's heading kept the
pre-renumber 059; 055's fixed-by named branches that stop existing after merge (now merge
commits/PRs) and its located-in listed file paths where the convention wants repos; 056's
located-in named mesh-host, which received no fix, instead of mesh-catalog; and the design
layer never said the one-store/one-broker property is enforced — 07-the-foundation and the
installation table now state the seats.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Observed in the two-node bed: a cross-node runtime exits on 'timed out fetching the
broker's certificate' until the tunnel forms, then settles. The retry belongs in-process.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The store and broker were "one per mesh" by convention only. Each foundation module now
claims a mesh-scoped seat named after the server it guards — postgres/mesh-store,
lavinmq/mesh-broker — and the controller's seat is renamed the-controller -> mesh-controller
so all three follow one rule. The resolver refuses a second holder, closing 056. Glossary,
the foundation and installation docs, and the decisions index follow the new name.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
055 is proven end-to-end by the two-node bed: a consumer on a joined node reaches the
adopted broker over the overlay, its binding names the control-node, and its vhost is
minted. Three fixes — the broker's amqps port in the firewall, the broker credential
naming the overlay not the public address, and pushing the provider node after the remote
consumer arrives. The last is an operator ordering, not a code fix, and its silent-failure
edge (a cross-node consumer that never provisions and only crash-loops) is opened as 057.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A two-node bed pins it: the firewall gap (broker's 5671 not in listens) is fixed
in mesh-catalog; the real bug is modules.go:294/build.go:231 building the broker
URL with the genesis public MESH_BROKER_ADDRESS, which a joined node cannot reach.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
"One postgres, one lavinmq" (ADR 0078) holds only by genesis assigning the
adopted modules to the control-node alone; nothing enforces it. A second assign
raises a divergent second server, silently. Open questions: a mesh-scoped
exclusive seat, or explicit adoption the controller refuses to place elsewhere.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Settles the design repository now that the self-upgrade build is on main:
- Records the two decisions that shipped without a record — ADR 0077 (the
controller/foundation/node vocabulary) and ADR 0078 (the store and broker are
ordinary modules); accepts ADR 0075 and 0076, which shipped work rests on.
- Fills issue 051's amended-design and wires ADR 0078 into 07-the-foundation.
- Sweeps the repo rename (mesh-control -> mesh-controller) into the mutable docs
now that the forge repo is renamed; updates the glossary note and repos.md.
- Fixes the six broken links from the design-doc renames, indexes the glossary,
regenerates the decisions reading order.
Both checks (records.py, index.py) are green. Statuses stay honest: the build is
on main and lab-proven but not deployed as the production mesh, so the to-be docs
remain in-progress and the as-is layer (the hal mesh) is unchanged — graduation
to implemented + as-is belongs to deployment, not merge.
https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Marks WBS 3.2/3.3/3.4 done and resolves issue 051: the foundation's store and
broker are adopted in place as the postgres and lavinmq modules, upgradeable
through their stated windows, source-tracked by status. A bare-metal mesh runs
one postgres and one lavinmq, proven 22/22 in the one-node lab. The two follow-up
gaps are tracked as issues 054 and 055.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
054: the adopted servers bind 0.0.0.0 from genesis but the packet filter is
installed later, so there is a window where they are open with only bootstrap
credentials. 055: the servers bind on the control-node and the one-node bed
cannot prove a consumer on another machine can reach them over the overlay.
Both are follow-ups to issue 051's adoption (WBS Phase 3), tracked rather than
rushed.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
"control plane" -> controller and "substrate" -> foundation throughout
03-DESIGN, 00-META and the README, with 06-the-control-plane.md and
07-the-substrate.md renamed to 06-the-controller.md and 07-the-foundation.md.
The immutable 02-DECISIONS records keep their original wording (and links to
them are unchanged) — a term retired here may still appear there, which the
glossary explains how to read.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Locks the vocabulary that kept drifting in conversation — controller (not
"control plane"), foundation (not "substrate"), node and control-node, seat /
bench / claim, package vs artifact. AGENTS.md points at it as the authority.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Records the decision the package-registry work turns on — the SDK is built on a
public base and published before the toolchain that consumes it, so nothing is
circular; mesh-tools stays the thin toolchain base but resolves the SDK by
version. Reconciles docs 12/17/22 and indexes the record.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The protocol spec claimed the envelope and grant drifted across implementations.
Inspection showed the wire agrees — envelope required headers match, the two
optional ones are legitimately optional, and the grant wire (the contributions
file) is identical on both sides. The disagreement was in dead types, now removed.
So phase 1's 'make them agree' work is done by deletion and correction. What
remains is a conformance fixture as prevention — pinning the envelope and the
contributions file so a future change that breaks agreement fails a test — and a
full per-capability suite is deferred until a third language actually needs it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The record claimed the two implementations already disagreed. Inspection showed
the live wire agrees: the disagreeing grant types were dead (removed), and the
envelope's two extra headers are optional and set when relevant, not missing.
The danger was dead types contradicting the wire, not live disagreement — which
is a sharper reason for specifying the wire and checking against it, not a weaker
one. The model stands; the conformance suite's job is prevention rather than
repairing a present break.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The correction the operator pushed: stop running a 20-minute lab against a mesh
mid-transformation, debugging paths the next phase deletes. The base-build hang
is almost certainly the SDK resolving from a git URL inside a docker build (issue
053), which Phase 2 removes — so debugging it on the current shape is debugging
deprecated code.
Phase 0 folds in: the installer's own regressions are fixed and committed;
whether it runs green is the final acceptance test, after the phases that change
its build path are in. Faults that can be reasoned out of the code path are, by
reading rather than running.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Everything decided this cycle and not yet built. Phase 0 gets the installer
green, because nothing else is testable end to end without it. Phase 1 makes the
protocol one thing and fixes the Go/TS drift the installer's own provisioning
exercises. Phase 2 stands up the private package registry ADR 0014 assumes and
publishes the SDK into it. Phase 3 adopts the substrate so one postgres and one
lavinmq serve everything, which is the hardest and needs all three above.
Order is dependency, not preference. Each phase ends at a run rather than a
paragraph, because a phase that ends at a claim is how things went missing this
cycle without anything complaining.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The lavinmq module raises its own server container and provisions vhosts on it,
separate from the substrate's mesh-broker. A mesh with the module assigned runs
two LavinMQ servers where one belongs — the exact AMQP twin of the two postgres
containers.
The module's own provisioner already assumes one server: it creates a vhost per
consumer, named for the login, isolated by the vhost boundary — the analog of
postgres's database-per-login. So the mesh's own control traffic is the / vhost
and every consumer's broker is a vhost beside it, all on one server.
Adoption therefore means the module does not run a server of its own: its server
is the substrate broker, adopted, and the module contributes the provisioner,
tools, events and bootstrap against it. The isolation model is already built;
what remains to design is only raising the one broker at genesis and then holding
it as a module, over the broker it is.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The substrate's store and the postgres module are the same thing. They were two
rows only while the substrate was a different KIND of thing — a store raised from
a bundle cannot provide postgres-database, so anything wanting a database needed
a second server. It is visible on any mesh built today: mesh-store and postgres,
two containers, the same image.
Which name survives is settled by the naming rule. Where a consumer speaks a
protocol the interface is the protocol, and "database" is not a capability. The
control plane's own queries use distinct on and on conflict, so the coupling is
to postgres and a "store" module would advertise a swap that fails the first time
anyone tries it.
The broker collapses the same way with a different outcome: amqp IS a protocol
that several implementations speak, so amqp is a legitimate provision and lavinmq
is one provider of it.
So adopting the substrate is not only an upgrade path — it is two rows of a
mesh's module list becoming one, twice. And it leaves nothing that is a specialty
after the pivot, which is the claim the whole design rests on and is not true
today for exactly those two.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Reported as 'the host does not survive a reboot', which reads as a mesh that
cannot come back. mesh-host/packaging/ ships nox-mesh-host.service and two
companions. The installer declines to place them because a unit file is a
packaging decision, and the lab starts the host with --host-in-background, which
says in its own help that it does not survive a reboot.
So the lab run failing this was the lab being honest, and the gap is the step
that puts a shipped unit on a machine — narrower and more fixable than what I
wrote.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The first draft implied gitea and the registry could not share a machine, because
registry claims the-artifact-store at node scope and I carried that across to
gitea without asking what the claim is for.
A machine running gitea for git and packages alongside a registry serving
artifacts is an ordinary arrangement. They are different ports doing different
jobs, and nothing about one being the mesh's artifact store requires the other
not to exist.
The exclusivity that matters is mesh-wide and already expressed: provides at mesh
scope means two providers are two answers, and the resolver refuses until one is
assigned. Forbidding co-residence adds nothing and forbids something reasonable.
Whether registry should still hold that claim is left open rather than answered
from outside its manifest — it may be protecting something about its port or its
data directory that nobody wrote down.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Two questions circling turned out to be one asked twice: should gitea be the
mesh's registry, and where does the SDK come from. The framing that dissolves
both is that artifact-store is already a provision and the registry already
provides it — so this was never about replacing a component. It is a second
provider of an existing provision, which this mesh has a mechanism for and uses
for certificate authorities already.
So: two provisions, because they are two jobs. artifact-store is content
addressed, pinned by digest, no versions and no ranges — what the mesh delivers
to machines. package-registry is an ecosystem's own, addressed by name and
version — what code resolves when compiled. Conflating them is how a mesh that
pins everything ends up rebuilding one commit into two different things.
The small registry stays the provider genesis installs, not because it is better
but because of what it is: a directory and one container, installable where there
is no database and no control plane. Gitea needs both, and the pivot needs
somewhere to publish before either exists.
Gitea also provides artifact-store, so a mesh may choose it — and choosing it
answers issues 042 and 048 by adopting something that already has accounts and
TLS, rather than reimplementing them in a registry that has neither.
Moving between providers is a designed act with a verification step that is easy
to skip and is the only thing between it and a mesh that cannot restart its own
control plane.
And the bootstrap still has no package registry when the first build needs one.
Named rather than solved, so the next person does not discover it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Every step from a bare machine to a mesh that maintains itself, in three phases,
with each step named as the installer prints it.
The point of writing it out is the shape it exposes. The installer owns twelve
steps and ends at a mesh that RUNS. Seven more turn that into a mesh that WORKS —
the shared base, a store that is a provider rather than the control plane's own
memory, the catalogue, the replay of what was built before the catalogue existed,
the control plane rebuilt through the module path, the private network with the
node actually placed on it, and the packet filter. None of those seven is the
installer's. They are things somebody types, which is why a test had to be
written to discover they were missing.
Machines arrive last, in phase three, because a machine joining a mesh that
cannot build anything proves enrolment works and nothing else.
And five things that are not yet true are named rather than implied: phase two is
manual, a second machine cannot pull what the mesh built, ADR 0014 assumes a
private package registry that genesis has not installed when the first build
needs it, the host agent does not survive a reboot, and nothing can contradict a
claim that a machine was installed this way.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
ADR 0014 is accepted and unambiguous: each module consumes its dependencies from
the private registry, the mesh's own shared library included, and a cross-package
change is publish then consume.
So the git dependency at a pinned commit is not a mechanism under consideration.
It is the shared library being consumed a way the record rules out, and the lock
naming a sibling directory is what that looks like when nobody publishes. Issue
053 is reclassified from a question about mechanism to a violation with a
direction.
What stays open is narrower and real: which software serves the private registry,
and that a fresh mesh has none when the first SDK is built — ADR 0014 assumes one
exists, and at genesis nothing has installed it.
Recorded after arguing at length for a bespoke content-addressed alternative,
against a decision that was already made and that I had not read.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
What the SDK contains, answered by exclusion as much as by inclusion. It is the
protocol and nothing else — no configuration loader, because configuration
arrives as files the mesh wrote; no API clients, because a Plex client changes
when Plex changes and that has nothing to do with any other module; no storage,
HTTP or logging, because the language has those. The test for anything proposed
is ADR 0039's: does editing it recompile unrelated modules, and does it change
often. Both, and it stays out.
Then the worked module: events in TypeScript, tools in Go, a provisioner in Rust,
a scheduled job in Python. Four artifacts, four toolchains, four processes, one
module — and each part is an ordinary project in its language depending on the
mesh SDK the ordinary way, so a laptop resolves what a build resolves.
And publishing a package as a module capability, which makes the SDK unspecial:
it is simply the first module that published a library. A Plex client belongs to
the Plex module because that is the only thing that knows when Plex changed.
Three things left open rather than papered over: which registry (the catalogue
holds verdaccio and a forge usually serves one too, and nothing says which is
ours), who may publish (a credential that does not exist), and what a range means
in a mesh where everything else is pinned by digest — a mesh that can rebuild a
commit and get a different library is a real change, and should be decided rather
than arrived at.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A floor every implementation needs and three capabilities independent of each
other, so an SDK can implement the floor and events and be a real thing rather
than an unfinished one.
Written as a specification, which means it says what is required rather than how
anything is arranged — and says plainly where it describes behaviour that is not
yet true. Three places it does:
x-causation-id and x-schema are specified and emitted by nothing; the Go side
writes four headers and the TypeScript side declares six. A module may serve
tools and may not call them, because a caller needs a reply queue its account may
not declare. And the two implementations disagree about what a grant carries — in
TypeScript consumer is the module, in Go it is the node and the module is From.
One word, two meanings, in two halves of one mesh.
Naming those in the specification rather than leaving them for conformance to
discover, because a specification that only described what already works would
have nothing to say about the things most likely to break.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The tool runtime's manifest names the SDK as a git dependency at a pinned commit;
its lock file names a sibling directory that exists on one workstation. It builds
only because the recipe runs npm install, which tolerates a lock disagreeing with
its manifest and re-resolves from the manifest — the one command that hides this.
It matters because a lock exists to make a build reproducible and this one
describes one machine, and because it is the first thing a fresh mesh builds: the
toolchain carries the SDK and everything with code of its own compiles inside it,
so a dependency resolved differently on the build machine than on a workstation
is a difference in every module the mesh will ever build.
And it is about to be copied. Each language's toolchain will carry that
language's SDK the same way, so the shape is worth settling before there are four
of them.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Two corrections, both from the operator and both better than what was written.
An SDK is an implementation of the mesh's module protocol in one language, and
nothing more. The first draft defined it by the test it passes, which describes
how you check one rather than what one is — and leaves it sounding like a library
that helpers could accumulate in.
And the protocol is split per capability, which was missing entirely. A module
that only consumes events uses the event capability; one that serves tools uses
the tool capability; a provider uses provisioning. Nothing about consuming an
event requires knowing how a grant is answered, so an SDK need not implement all
of it to be real.
That has a precedent here: a host declares which resource kinds it can apply, and
a partial host is a real thing rather than a broken one (ADR 0005). An SDK
implementing the floor and events is exactly as legitimate, and a module written
against it is a module that does events.
Which changes what adding a language costs. A Rust SDK doing connection and
events is useful the day it exists, with tools and provisioning following when
something needs them — rather than a language being unsupported until it is
entirely supported, which is what makes adding one a project instead of a
contribution.
Conformance is therefore per capability too: a monolithic pass or fail would make
a partial implementation indistinguishable from a broken one, which is the
distinction the whole thing rests on.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
ADR 0039 settles what belongs in an SDK. It does not say what happens when there
is more than one, and there already is: the contracts are expressed as Go types
in the control plane and host and as TypeScript types in the SDK, and nobody has
felt it because both live in one repository.
They already disagree. The provision's field is "resource" in one and "Provision"
in the other; "consumer" means the module in one and the node in the other; the
envelope declares six headers on one side and emits four on both — the missing
two being x-causation-id and x-schema, the second of which is exactly what a body
needs in order to change shape without silent misreads.
That class of failure does not announce itself. Two implementations disagreeing
about an envelope do not fail to compile — they ignore each other's messages, and
a mesh where a module stops reacting looks like a mesh where nothing happened.
So the decision is to specify the wire rather than share the types, because the
shapes are the easy half. What two implementations actually disagree about is
behaviour: queue naming and durability, which headers are required and what an
unknown one means, taking identity from the sealed credential rather than the
environment, dedup on an id only the emitter can make, pinning a fingerprint
rather than trusting an authority.
And the suite is executable rather than prose, because a specification nobody can
run is a document two implementations drift from while both believe they conform.
The two existing implementations are the first made to pass it — a suite only new
SDKs must satisfy would certify every future language against a disagreement that
is already here.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A reference table goes stale the day somebody adds a field, so this one points at
mesh-catalog/modules/showcase — a module that uses all of it — and a test that
fails when it stops doing so. Read the module when the table disagrees with it.
Two rows in the coverage survey were stale because of this week's work: systemd
units were a file plus a service, which made every author write unit syntax and
is why "process" exists; and building from source was images only, where a
bundle now names a language and lets the mesh choose the toolchain.
And the two rows at the bottom of the resource table are the interesting ones.
"action" is refused to modules outright — the link may not carry a command, so a
module needing something done ships a program that reconciles. "service"
installs no unit by design, right for software shipping one and wrong for code
the mesh built, which has none until the mesh writes it.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The first draft said the builder refused archives. It does not. An archive is
packed deterministically, hashed, published by digest, fetched by the machine and
unpacked — the whole path exists. Only the local builder used at genesis refuses
one, and deliberately: an archive is bytes that mean nothing until something
serves them, and at genesis nothing does.
What an archive cannot do is compile. Its source is a directory packed as it
stands, so shipping compiled output means compiling somewhere first, which means
a Dockerfile — the burden this document is about. The gap is not the artifact
kind. It is that no recipe both builds and packs.
Found by reading the builder rather than the manifest schema, which is where the
first draft's claim came from.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
12-a-module-repository says what a module may build and where it goes. Nothing
said how a build is MODELLED, and the model is the problem: a recipe is implicit,
singular and always a Dockerfile; a toolchain is not modelled at all, arriving as
two build arguments the module hand-writes; a language is not a concept; and an
archive is declared in the manifest and refused by the builder.
The cost is measurable rather than theoretical. Adding a module with its own code
means repeating an incantation - two ARG bases, a specific working directory so
the SDK resolves upward, the compiler invoked by absolute path because the usual
symlink is resolved away when the base is assembled, a second stage, an env var
naming the entrypoints. Most of the catalogue is unconverted, and two conversions
done in one session were each wrong twice with a working example open.
So: recipe becomes explicit with three kinds, and toolchain becomes derived from
a declared language rather than written by every author. A Dockerfile stays, and
stops being compulsory - it is right for software needing a particular base and
wrong for "compile my module's code", which is the same operation every time.
The cost is stated before it is chosen: every language is permanent, and the
contracts are already expressed twice - Go structs and TypeScript types kept in
step by hand. A second language makes that drift. So language-neutral contracts
come first, or the drift gets worse while hiding.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The packet filter generates its rules from what modules declare they listen on.
The broker is not a module, so it declares nothing, so its port is not opened.
Every machine dials that port to enrol and to receive every declaration it is
ever sent.
Invisible where it is assembled and fatal on the next machine: a mesh of one
never dials its own broker across the network, so the ruleset looks right. The
first machine to join is refused at the packet filter during enrolment, several
steps from anything that reports it — and assigning the firewall before joining
machines is both the natural order and the one that breaks.
This is 051 in a second place. That issue says the substrate cannot be updated
because the mesh holds no record of it; the same absence means the firewall
cannot know it exists. Ssh is already a floor for the same reason — a machine
nobody can reach is a machine nobody can repair — and the broker may be the same
class of fact.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The store and broker come from a bundle the installer writes once, with images
pinned in it, and nothing can change them afterwards: no build, no version to be
behind, no roll-out, and no way to report being out of date, because the mesh
holds no record of them as modules at all.
That is backwards. They are what everything else depends on, so their updates
matter most, and they are the only things with no mechanism to deliver one. A
mesh with a year-old broker reports itself entirely current.
The fix probably already exists: the control plane is carried, raised and then
adopted as an ordinary module pinned to what is running. Nothing in that pattern
is specific to the control plane. It would also remove a duplication visible on
any one-node mesh — the same postgres image running twice, because a store that
cannot be a module cannot provide a database to anything.
Recorded with the two hard parts stated rather than waved at: upgrading a store
the control plane is reading from, and upgrading a broker over the broker.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A mesh raised from bare metal built six modules and its catalogue reported three:
exactly those built after it began running. Missing were the shared base, the
store, and the catalogue itself.
The hole is never random. On a fresh mesh the modules built before the catalogue
are by necessity the ones it needed in order to exist, so the foundation is
always what is absent, on every mesh, at the moment the graph is first populated.
It breaks the question the catalogue is for: build edges hang off the base, so a
catalogue with no record of it answers 'what must be rebuilt' confidently and
wrongly. And nothing reports the gap, because a catalogue cannot know what it was
never told — this was found by comparing its answer against what the mesh had
just been watched doing.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A module's broker account is scoped to what it declares it emits and consumes. A
tool call needs a reply queue, which that scope does not cover and should not. So
the account is right, the request is reasonable, and no account exists that can
make it — asking a module its own question, from its own container, with its own
credential, is refused.
It matters because a module's tools are its operator-facing surface: the
catalogue serves the five questions it exists to answer and nothing can reach
them. It is also why a running module keeps being mistaken for a working one — a
test that cannot ask anything checks a container is up, and that substitution has
hidden two faults this week.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The registry says every machine pulls from it and opens its port to the mesh for
that reason. A machine that tries is refused by its own container runtime: the
registry serves plain HTTP and anything but loopback is treated as HTTPS.
It has never failed, and that is the finding. Every proof that a machine can
fetch a mesh-built artifact was a proof about the machine that built it, where
the reference was loopback. The bed that uses a routable address gets away with
it because the harness writes the runtime's configuration before the mesh exists.
Kept separate from 042, which they are easy to confuse: that one is about not
being allowed to pull, this one about not being able to whatever the credential
says. Fixing 042 alone leaves a machine with a valid account for a registry its
runtime will not talk to.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The report covered ports the firewall cannot close. The matching fault is ports
it does close and should not: rules come from what modules declare, a migrating
mesh knows about almost nothing, and anything listening on the host is dropped.
No module declares an ssh port and the generator has no allowance for one. The
session that loads the rules survives on conntrack until it drops, and then the
machine is reached from a rescue console.
Rehearsed on three lab machines before doing it on an anchor: loading the mesh's
rules refuses an ordinary host port and leaves a published container port
reachable, same prober, same second.
It is deliberate — no forward chain, because dropping there would stop every
container — but traffic to a published port never reaches the input chain, so
the firewall is silent about it. The anchor publishes 38 such ports today, all
filtered by the system being replaced, so the cutover would open every one of
them while reporting a firewall that is up.
Found while migrating the first module. Five approaches were tried and observed
to fail, including resolving the index and naming the platform at both ends; a
speculative fix was written and reverted rather than shipped, because it did not
make the mirror work.
One module uses this today, which is why it went unnoticed — and it is the shape
most of a migration wants, because the services being moved are third-party.
The design said how the builder arrives was unsettled and that nothing
installed it — the one gap stopping a fresh mesh from producing anything. Both
are now false. What is still true is narrower and worth keeping separate:
nothing asks a raised mesh for the rest of the catalogue, and no bed asserts
that it could.
The mesh could already say which modules a base change invalidated, and could
not do anything about it: each recipe named one particular copy of the base by
fingerprint, and rebuilding produced the old one. Worse, the copy each named
existed only inside a throwaway lab, so those three modules could not be built
anywhere at all — and the line each replaced had the same fault.
The builder's arrival was the one rule the document said nothing checked. It is
checked now, by both genesis beds — and so is the thing that distinguishes a
built control plane from a carried one, which every earlier assertion accepted
either way.
The section saying the change was decided and had not happened now contradicted
the section below it. It also records the argument that failed, because a reader
will otherwise ask the same question and reach the same wrong answer.
Written an hour ago claiming a produced image must be published before anything
can fetch it, so the registry had to precede the control plane. The premise is
false: the machine that builds the image is the machine that runs it, and the
temporary control plane names a built image exactly as it names a carried one —
by the digest of its own configuration, which requires nothing to have served
it. Building changes where the bytes came from, not where they are.
Rewritten rather than superseded because nothing has been built on it and
nobody has read it: a record that contradicts itself is a draft, not a decision.
The argument is kept, because it was asked for and a negative answer is the
result.
The two questions the design record named as the one gap stopping a fresh mesh
from producing anything. They cannot be answered apart: a builder with nowhere
to publish has made a file on a disk.
The registry's role did not change — the answer to 'must it precede the control
plane' did, because the control plane's image is now produced rather than
carried, and a produced image must be put somewhere before it can be fetched.
Cloned alone, built by the mesh, every module moved onto it, and the base then
changed for real: all three went stale naming what moved. A comment-only change
correctly makes nothing stale, which found a second bug — staleness compared
commits where it should compare artifacts.
The runtime every module compiles against cannot be built by the mesh, so the
one rule that would catch it moving can never fire. And a container keeps the
values it was created with, so two good applies can leave it running on neither.
Corrects ADR 0070, written an hour earlier, which had the control plane consuming
the catalogue in order to compose a declaration. That was written before the two
graphs had been told apart and creates a dependency that need not exist: a
catalogue that is down would leave the control plane unable to compose the thing
that would repair it.
The catalogue links module-versions to each other and does not know nodes exist.
The control plane links module-versions to nodes, and holds capabilities and
claims. They meet only when something is installed, and everything between them
travels as events over the broker, on durable queues, so nothing is lost when a
receiver is away.
Build order is not computed anywhere. The builder never consults the graph and
builds what it is asked for; the catalogue asks for the next build after the
previous registration, so ordering holds by construction. Its rule is a condition
rather than a schedule — rebuild once everything a module was built against is
current — which covers a chain and a diamond alike.
Left open: whether an upgrade is applied or merely noticed, and whether a module
on several machines upgrades on all at once.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
ADR 0070 has the init builder clone the source and does not say from where, and
ADR 0067 had rejected building at genesis partly because the forge runs on the
mesh being rebuilt. That objection binds only when those are the same mesh, which
is true exactly once.
So genesis clones from a mesh by name, and if that mesh is lost the name moves to
another that holds a copy — recovery is a name pointing elsewhere rather than a
backup being restored, and every installation adds somewhere it could point.
It names a commit and checks what it got, because the forge a mesh installs from
is the trust anchor for everything that mesh will run. On 2026-09-11 that forge
was running a cryptominer and tampering with git operations in flight; nothing was
altered, but a mesh installing during those hours could not have known that.
What relationship a mesh keeps afterwards is left open and named, so that whoever
writes the init builder does not settle it by accident: a snapshot and then
independence, or a continuing upstream for core modules.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The graph had no owner: what modules are, what they require, what they claim and
what they are built against all sat in the control plane because that is where it
was written first. The control plane's own test says otherwise — anything a single
machine could answer alone is not its work, and what a module needs requires no
knowledge of any node.
So the catalogue becomes a core module beside the control plane and the builder,
owning the graph and serving tools over it. The control plane consumes it, which
is the opposite of what the tiers suggest and is therefore stated rather than
inferred.
That makes the catalogue a fourth thing that cannot arrive through the ordinary
path, so the claim written this morning that the list was closed at three is
corrected. All four are answered by one mechanism instead: the installer carries
an init builder and the core modules are built on the machine, so what is carried
is a builder rather than a result and nothing is left without a route.
Two things are open and named rather than assumed: where the init builder clones
from, given the forge normally runs on the mesh it would be rebuilding, and what
it publishes into, given the registry is installed later in the order today.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
A build machine was refused the build queue, and this was raised as a gap in what
a manifest can express. It is not: `builder issue` creates exactly that account,
three lines from the code being read at the time.
Kept rather than deleted, for the one real thing in it — the wrong verb succeeds
and reports success, producing an account that authenticates and can do nothing,
so the failure surfaces a layer away as a permissions error that reads like a
missing feature.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Raising a scenario occupies the machine and the person who started it, and
running in the background against a working copy is worse than waiting: a run
reads that copy as it goes, so editing while it runs yields a result describing a
state that never existed.
Proposed rather than accepted. The load-bearing part is the restriction — a
request names a bed and a commit and nothing else — because a request that could
say what to install and where would make the lab a second way of installing a
mesh, which is the arrangement that just cost a year of late-found faults.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The design said a module's manifest sits at a repository's root, full stop, which
means one repository per module. Nothing that exists is shaped that way: the
catalogue holds sixty-seven modules one to a directory, no code repository has a
manifest at its root, and the system being replaced has always built a module
from a repository and a path.
So the builder could be asked to build nothing that exists — pointed at the
catalogue it finds no manifest, pointed at a module's source it finds none
either. Recorded as a decision because it moves the core modules' manifests
beside their source, and corrects the design that said otherwise.
Also corrects, in the same document, how the three things the build loop cannot
produce actually arrive. They were written as though all three were carried in.
Only the control plane is: the registry is pulled from the public internet, and
the builder has no route at all — which is now stated as the open one rather than
implied to be solved.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
The one complete account of standing a mesh up was an integration test, and a
fixture is free to invent what it needs — which is how a registry that exists in
no production hid two faults for as long as the lab existed.
Written from what the installer does, not what it should do: genesis and joining
are separate moments, the lab runs the installer rather than describing
installing, and three things that are not true yet are named rather than glossed,
including one rule nothing checks.
Claude-Session: https://claude.ai/code/session_01D6qtiYU3P9jk3pnAXyAFyx
Found by deleting the lab's registry and giving the machines a real path out:
public images fetched, the operator's own could not be fetched at all. There is
no provision for a registry credential, no manifest field, and no step in
enrolment that establishes one.
It applies to the mesh's own store too, which today asks for nothing — a
decision that has never been written down as one, and so reads as an absence.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Both carried 'model access', which is not one of the six the reading order
knows, so neither had a place in it. 0050 — the record they extend, on the same
subject — is 'what runs on it'.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Nine references now name the digest their tag resolved to. What closed is the
immediate fault; the open questions stand, because a digest in a repository is
wrong the moment anybody rebuilds — which is the reason the design wants the
repository to name artifacts and the mesh to hold digests.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
0066 was proven on a four-node bed before it was ratified: one node setting
moved an entire domain, a routed name resolved inside the mesh and was issued a
certificate by the internal authority, and TLS verified against that authority
with no override. 08-connectivity rests on it and could not while it was
proposed.
0067 records what deleting the lab's registry exposed — that pinning quietly
required a registry before the thing that lets a mesh have a registry could
start — and the pivot that resolves it.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Both carried a topic outside the six the index knows, so neither had a place to
be read in — 0067 had none at all. Both are 'the tiers', beside 0036 (bootstrap
ends at a usable mesh) and 0007 (connectivity), which is what they extend.
0067 cited 0041 for tier 0's property; on this trunk 0041 is events, and the
record it meant is 0005. A citation that resolves to the wrong record reads as
corroboration, which is worse than a dead link.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The doc's own rule is that the lab verifies capability by outcome, never by
reading a setting. A path out is exactly that kind of claim — a route and a
policy can both read correctly while nothing gets through — so it is asserted by
fetching something, in the table with the rest.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A container runtime on the same workstation sets the kernel's forwarding policy
to drop, and the virtualisation daemon's own accept rules do not override it —
both are consulted and a drop anywhere is the answer. The machines then get
addresses and resolve names, because the daemon's resolver is on the bridge, and
discard every packet to anything real.
The sharpest 'available is not adequate' yet: nothing is misconfigured, nothing
logs, and it presents as every image pull hanging. A workstation that runs
containers is the ordinary case, so it is a prerequisite — verified by reaching
something, never by reading a setting.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
029 says a module providing the artifact store may not build artifacts. The
pivot shows that is the narrow case: it may not require anything the store is
needed to deliver. A route-label migration gave the registry a public name and
a route requirement, and at genesis nothing provides a route — nor can anything,
since the routing stack needs images and images need the store.
The same cycle through a door the existing wording did not cover, so the rule is
widened where the bootstrap decision states it, with a check that would catch the
next one where it is written.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
There is no installer. The complete account of standing up a mesh is an
integration test in the lab, and a fixture may invent what it needs — this one
raised a registry no production has and rewrote every image reference through
it, concealing both 039 and the fact that a first node outside the lab had no
bootstrap path at all. The bed being green said nothing about whether a mesh
could be installed.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Nine images across seven modules name a tag, not a digest. ADR 0006 forbids it
and the host refuses it by name — and the refusal has never fired in a bed,
because the lab pushed every image into its own registry and rewrote every
reference to the digest it had just assigned. The harness was supplying the
property under test. Found by deleting the harness.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The control plane's image is built from source and pushed nowhere, so it has no
manifest digest; a registry assigns those. Pinning therefore required a registry
before the thing that lets a mesh have a registry could start — a dependency the
rule created by accident. The lab hid it by raising a disposable registry no
production has.
Proposed, not accepted.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
0056 is 'the authority is the control plane, not a database', drafted on the
in-progress record chain this branch was cut from before those three records
landed. Two files would have collided at merge, which is the kind of thing that
is cheap now and confusing later. The code written against it still says 0056
and is corrected separately.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The real defect is same-node consumers announced the declared port instead of
the assigned/published one; the loopback observation was a stale pre-0038 build.
Fixed in mesh-control fix/same-node-provider-announced-port (c147a26) with a
regression test.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A from:mesh provider is announced (per 018's fix) at the node's private-network
name, but its port is published bound to loopback, so consumers dialing the
announced <node>.internal:port reach nothing. Diagnosed from the lab: DB
consumers that require the database at startup crash-loop; the provider is
healthy on loopback only.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A route contribution carries a label; the node carries its public domain; the
mesh composes <label>.<public-domain> and holds no name map. A granted route is
published into internal resolution so anything in-mesh (notably an internal ACME
authority) can resolve and reach it. That authority certifies routed names by the
same path a public one would, differing only in issuer and trusted root.
Records the decision as proposed and amends connectivity SS2/SS3/SS5 plus its Open
list with the lab findings behind it.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Extends ADR 0050: model-access, a provision, may be answered by a NODE that
hosts a model (Ollama/vLLM serving an OpenAI-compatible endpoint) as well as by
a licence record. A node-answer delivers an endpoint (base URL + model, and a
key only if the server wants one), rides the ordinary serves/binds path, and
uses no adapter; the resolver already prefers a local model over a licence. One
provision, two answers — the mesh's own model sits behind the same interface as
a vendor's. Names the mesh-scope-alongside-licence limitation and its fix.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Closes the usage-tracking half ADR 0050 left open. One row shape (0050's), recorded at two
consumer grains: the holding module (licence-level, e.g. Anthropic utilization%) and the agent
session (per-session tokens/cost, since a session IS a consumer per ADR 0026). Produced by the
vendor adapter; read on a schedule (0053); recorded as events the audit-logger keeps (0041/0042)
plus a queryable usage context store (0008); usage is not a credential and is recorded in the
clear. Reuses the session, schedule, event, and store the mesh already has. Accepted per the
user's choice to build the full feature incl. per-session token/cost.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The recurring twin of run-once (0052): a schedule modifier on the container shape, reusing
its security bound (no new host action/shape, strictly less than an action) and reversing its
gating rule — a scheduled step runs after convergence, does not gate the apply, and a failed
run is logged, not fatal. Unblocks kometa's sync and pollers. Accepted per direction to build
the primitive now.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The run-once lifecycle primitive: run-once:true on the existing container shape,
gated by declaration order + exit 0, idempotent by digest. No new host shape, no
arbitrary command — strictly less powerful than an action. Resolves issue 037.
Verified sound and faithful (ADR 0005/0047); status proposed->accepted.
A module can declare state but not a step that runs. mosquitto must seed its
dynsec admin into dynamic-security.json before the broker starts, or the plugin
aborts; the database providers need the same for first-boot migrations and
health gates (04-ISSUES/037). The old event-hook engine that did this was
powerful and flaky; this is the narrowest sound mechanism instead.
A run-once step is an ordinary container marked `run-once: true`: the host runs
it to completion, requires exit 0, and gates the apply on it — so what the
declaration places after it (the broker) starts only once it has finished.
Gating is by declaration order, not a resolved dependency (ADR 0005); the
completion marker is the recorded declaration digest (ADR 0018), so a re-apply
does not re-run it unless the declaration changed. No new host shape and no
arbitrary host command: strictly less powerful than an `action`.
Points 04-ISSUES/037 fixed-by/amended-design at the record; index regenerated;
records.py and index.py pass.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Your decision, ratified: the media library (and shared/pre-existing data) is
operator-owned and external; a module declares access, not ownership; the host
mounts but owns nothing (no create/chown/reconcile/remove); an absent accessed
path is refused clearly; several accessors co-resolve. Status proposed->accepted;
index regenerated (records + index checks pass). Implementation lives on the code
branches (mesh-control/catalog/host), held for merge after the convergence fix.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Resolves 04-ISSUES/036: eight media modules each declared the shared
library and download directories as their own resources, and the
resolver's duplicate-owner refusal — right in general — would refuse the
stack's only sensible assignment the first time two landed on one node.
The decision, from the operator: shared, pre-existing data is
operator-owned and external. The mesh does not create, chown, reconcile
or remove it. A module declares it needs access to such a path (read or
read-write); the host mounts it and owns nothing. Several modules
accessing one path is normal — the duplicate-path refusal is about
ownership, not use. An accessed path absent at apply is refused clearly,
not created. Extends ADR 0030: the third case the host had no word for,
what it neither made nor configured and must not touch.
Point 036's fixed-by/amended-design at the record; mark it located in
mesh-control, mesh-catalog and mesh-host. Regenerate the decision index.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Turn the completed vendor-agnostic analysis into HQ design. The model-access
provision stays one vendor-blind interface (extends 0024/0027); the
vendor-specific lifecycle moves into a per-vendor adapter keyed by the licence's
`vendor` field, mirroring registrar-scoped public-dns providers (0044), named at
the consumer's real coupling per 0040.
The crux is the sealing-vs-central-rotation carve-out: for refreshable-grant
vendors only, the manager node holds the refresh token encrypted at rest (a
bounded, declared exception), access tokens sealed per holder, refresh stripped
on delivery. Static-key vendors keep full sealing.
Amend 03-DESIGN/01-to-be/14-model-access.md with the adapter generalisation as a
proposed section (prose + diagram, no code); regenerate the decision index.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Init's 003 was already resolved with its own consolidated attribution
(03-DESIGN/01-to-be/08-connectivity.md); the re-homing overwrote it with the
session's ADR-0045 firewall attribution. Init is canonical and the firewall
decision is recorded in the ported ADR 0045 regardless, so 003 is restored
untouched. No initialization record is modified by this reconciliation.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The real work lived on initialization (consolidated decisions 0001-0038, the
fuller issue set 001-031, the control-plane/substrate/node-lifecycle/delivery
design, research 011/012, the checks tooling). main had diverged onto a stale
base and only carried this session's genuinely-new work. This merge makes
initialization's tree canonical on main; this session's 11 new ADRs and 6 new
issues are re-homed on top in the following commits. initialization is recorded
as a parent so its history is preserved in main's ancestry.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Surfaced converting the catalogue: a module can declare things that exist
(dir/file/network/container) but not a step that runs at a point in its
lifecycle. mosquitto's dynsec admin client must be seeded before first start;
the DB providers have nowhere for a migration or health-gate; it is the timing
face of issue 011. Framed as a missing module capability, not a defect. Records
the prior-art event hooks and their real warning — powerful but complex and
flaky — so the resolution avoids rebuilding that. Ends in open questions
(run-once resource vs general lifecycle hook, where the code runs, idempotency).
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Numbers 008 and 009 were taken on main after this branch was opened
(008-provider-runtime-has-no-seal-key, 009-runtime-config-change-does-not-restart,
both merged). Renumber the seed-file-wipe and shared-directory issues to the next
free numbers so merging records two more issues rather than duplicating two.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Records the branching-and-merging workflow for code changes across the mesh
repos, written against a failure it names: branches and MRs opened per unit of
thought, treated as done when opened not merged, and named differently per repo,
so they pile up unmerged — one session left sixteen to consolidate by hand. The
rule is one feat/<slug> shared across every repo a feature touches, isolated in
.work/<slug>/<repo> worktrees off main, pushed and opened as one MR per repo only
when the whole feature is done, then merged promptly. Adds the ground-rule
pointer in AGENTS.md and the row in the process overview.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Brings the independent ADR branches (0044-0052) onto one branch so hq lands as a
single MR, and ratifies the five that were still proposed — 0017, and 0049-0052,
which are implemented and green in the lab. With 0053/0054 already accepted here, the
whole ADR chain 0044-0054 is accepted on this branch.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
ADR 0054 accepted with option E (a declared slug). Issue 010 resolved: the login fits
via the slug, and the minted secret shrinks to 40 chars for S3's secret-key limit —
both halves of an S3 credential now fit the tightest backend.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Implementing A (bound the identity at 20) revealed the readable budget is node+module
<= 14 chars — so tight that the catalogue's own test names (workstation+keycloak, 25)
compact to an opaque hash. B's fallback would fire for the common case, not the rare
overflow, inverting A+B into mostly-opaque identities. Option E — an optional short
slug a module/node declares, preferred over the cleaned name — is the escape hatch B
wanted to be without the opacity: legible because a person chose it, and it makes an
early refusal palatable (refuse on the slug field, not the machine's name). B dropped;
E recommended over a bound of 20, composing with C later if needed.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Sketches the options for issue 010: the mesh's identityLimit (63, postgres's) is not
the shortest among the backends the derived name reaches — S3's is 20 — so CheckIdentity
lets an over-long access key through and minio fails at provision time. Options: bound
by the true minimum and refuse at assignment (recommended, with a compact fallback held
in reserve), per-interface bounds, or a provider-generated identity (rejected — breaks
"the mesh says the identity once"). Links issue 010 to it.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Found doing the per-backend provider e2e: redis and postgres accept the mesh's `as`
(mesh_<node>_<module>) verbatim, but minio's S3 access key is capped at 20 chars and
`as` is 22, so the provisioner cannot create the service account. `as` is doing two
jobs — a stable identity the two ends agree on, and a literal identifier a backend
must accept — and those are not always the same string.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
ADR 0053 accepted; adds the scope boundary the umami rework surfaced (credential
provisions vs data provisions — analytics' generated siteId return is left to a
separate decision) and records the lab proof. Issue 008 marked resolved: the sdk
harness and the four adapters are reworked, the symmetric seal removed, and
provider-uses-mesh-credential is green.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Issue 008's trace confirmed the premise in control-plane code: the mesh already
mints one password per consumer/provider pair and delivers the provider its copy
(SecretFor/SecretsFrom/grantsFor -> Grant.Sealed; the receives contribution carries
As + Secret). The provisioner's symmetric seal is an orphaned, contradictory second
model. ADR 0053 corrects the provider contract in one place (the sdk harness):
providers create the resource with the mesh-supplied login and password and drop
seal/key/return entirely. Reframe 008 as contract-first (every provider, not four).
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A cross-repo trace showed nothing writes the provisioner's grant-request files,
nothing reads its sealed credentials, and no consumer unseals — while the mesh
already mints and delivers provider/consumer credentials asymmetrically with no
shared key. The fix is to drop the symmetric seal and have providers consume the
mesh-minted password, a breaking provider-contract change that wants an ADR.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Found rolling the runtime out to the catalogue: a module's runtime reads its
settings-merged config file once at start, but a container is only recreated on a
spec change, and file content is not part of the spec. So updating settings
re-renders the file and nothing re-reads it — ADR 0051's "on the fly" holds only
for config set before first start. Services have restart-on; containers do not.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Found building the module-runtime vertical slice: a provider's provisioner
requires a seal key it has no way to receive, and the consumer no way to obtain
the matching one. The runtime cannot come up as delivered.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The runtime-model gap the review found. A module with tools or events runs one
container — the tool runtime carrying its code — holding the one scoped account
ADR 0048 gave it. A node-wide runtime can't: it would hold the union of every
module's permissions, the isolation 0048 draws. So per-module: one module, one
process, one account. Tools served per key (serve.<tool>) so a caller names a
tool and only its module answers (superseding a shared tools.invoke); events in
the same process under the same account; the runtime image is the tool runtime
plus the module's code (the audit-logger's shape, made the rule). A plain
service module runs no such process. A provider's provisioner is a runtime too —
which is why a provisioner that emits must carry a broker credential or not emit.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A module is assigned to a node (there is no mesh assignment; 'mesh' is a scope).
The manifest is what the module IS, plus defaults; the configurable values are
settings, carried by the assignment — per-node or mesh-wide, applied at
resolution, changeable live (what a meshboard edits). Extends settings from a
config file's content to the manifest fields marked settable: foremost
listens.from (postgres from:mesh by default, from:anywhere per node — the
firewall follows), and a provider's own config (a registrar's zone/domain/
ingress). Static config in a manifest is config in the wrong place: it cannot
vary per node and cannot change without a rebuild.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The manifest refuses unknown keys (DisallowUnknownFields), 'from' is the field
that scopes a port and it is rendered to nftables (AsNftables), and the firewall
module applies the rule set. The chain from a declared scope to a dropped packet
is closed. Amended-design: ADR 0050.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
0049: a public name is provisioned like any capability — a module requires
public-dns and contributes its host; a neutral interface answered by
registrar-scoped providers (cloudflare-dns, route53-dns) that create/remove
the record pointing the name at the mesh's public ingress. Pairs with route
(the proxy) and a public cert (the proxy's ACME).
0050: answers the firewall question. The firewall is NOT a provider like the
proxy — it is a machine's own filter, derived by the host as the sum of what
its modules declare they listen on, with 'from' the whole of public-vs-internal.
Enforced both ways, unknown keys refused — closing 04-ISSUES/003. A public
service is exposed through the proxy (listens from:mesh + requires route), not
by opening its own port.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Events (0046) and their wire (0047) left open how a module reaches the
broker. The code has no generic module broker-account: only node and
builder scopes exist, so emits/consumes are enforced by nothing — a
manifest declaring a scope the broker does not draw (04-ISSUES/003).
Decides: on assign, a module gets a broker account whose permissions ARE
the manifest — read on mesh.events + its own queue bound to consumes;
write to mesh.events under module.<self>.* only; nothing else. Consuming
'#' is a deliberate, auditable grant. The account is what makes the
declaration a rule the broker enforces, not a comment.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
The wire contract ADR 0046 left open: two topic exchanges (mesh.events,
mesh.rpc, kept apart so # is a clean audit); the routing key as the event
type namespaced by origin (module.*, mesh.*, node.*); metadata in AMQP
headers (required x-event-id/x-source/x-node/x-time/content-type; optional
x-causation-id/x-schema; unknown x- headers ignored) with the body only the
payload; persistent messages; per-consumer durable dead-lettered queues
with prefetch; at-least-once with idempotent consumers (no false exactly-
once). The precedent is ADR 0043 for declarations.
Supersedes the sdk's first cut (metadata in body -> headers); that and the
queue config are code to align in mesh-sdk and mesh-tools.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A module emits and consumes events, both declared (emits/consumes),
parallel to provides/requires. Events are 1:many, broadcast, credential-
free — no provisioner, just the broker's topic routing — so most inter-
module reaction should be an event, not a provision. Every event carries
source/node/time so it is auditable; the audit logger is just a module
consuming '#', no privilege. A consumes for an event nothing emits is a
dangling edge and refused, like requires. One per-node runtime serves
tools, provisioning and events alike.
Extends ADR 0045; builds on ADR 0001 (the broker) and 0044 (emit/on are
stable sdk surface; the binding and runtime are not).
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
A module is one self-contained piece of software the mesh installs and
manages; the software is its identity, and capabilities/seats/provisions
are the relationships between modules, not what a module is. Records the
three relationships (shared seat, exclusive seat, provide/require), that
interfaces are mesh-owned and providers adapt to them, and the naming
rule: draw the interface at the consumer's real coupling — neutral where
the coupling is thin (analytics), protocol-scoped where the consumer
speaks a protocol (postgres/mssql/mongodb), never false genericity.
Supersedes 0017 (domain grouping — wrong axis), refines 0002, generalises
0027's protocol-not-product rule.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Supersedes ADR 0030's 'types, not behaviour' line for mesh-sdk. The
boundary is change-frequency, not kind: the SDK holds the stable spine
(tool-serving harness, messaging/event framework, contracts, core
primitives) and refuses per-module clients, per-module tool code, and
anything volatile — because those are what turned hal/sdk into constant
maintenance and made every edit rebuild every module.
States the rule (frequent AND cascading is the disease), why the root
cause was intra-module feature-sharing leaking into inter-module
coupling, and where per-module shared code lives instead (in the
module — a shared file, or a module-local sdk for the few large ones).
Updates repos.md's canonical mesh-sdk description to match; leaves 0030
untouched (immutable).
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Both from the 2026-09-02 review of the catalogue examples, and both
design gaps rather than defects in a file: a seed file the host
reconciles back to empty over the grants that grew in it, and a module
stack refused co-assignment because six manifests each own the
directories they exist to share. Fixing either in place would have
been picking an answer the records do not yet hold.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Declarations queue and the host applies all of them, oldest first, so a
machine pushed five things in a minute spends five applies becoming the
last one. Correct every step — each declaration is the whole machine —
and wasted in all but the final step.
Only visible since a report names its declaration: the reports arriving
were about ever-older ones, while timestamp comparisons used to happen
to pass. The fix is consumption order, not the queue; the open question
is what a superseded declaration's report should say, because silence
reads as disobedience and "applied" would be a lie.
The five declarations cover relations; a module is more than its
relations. Add the facet-by-facet coverage table so the effort cannot
conclude while tools, verification and contributions are unplaced —
and weigh each candidate gap rather than adopting it: contributions
probably dissolve into declared resources, mandatory verifiers risk
trivial ones, and the tool surface is the one facet with no home.
Claude-Session: https://claude.ai/code/session_01LrgweAeERJYBg88c5cKDzF
Composing a declaration signed the machine's certificate anew each
time, and a signing carries a fresh random serial — so what the mesh
would send differed from what it had sent by one byte, for ever, and
every machine carrying a certificate stood eternally waiting.
Third find of the same rule: issued once and kept. The port had it, the
secret had it, the certificate composed fresh on every asking — and the
keeping column had existed since the serving-key migration, written by
nothing, the same shape ReleasePorts was found in.
Found by keeping the scenario standing and diffing two plans seconds
apart: one line, where four theories had none.
The registry module names its image by digest, the way the bundle names
the three a first node starts from, and goes in as a manifest. The lab
now walks that path and passes.
A manifest that provides the store and also builds artifacts is refused
at parse. The lab had been passing only because its Docker Hub stand-in
quietly received the push — a prop covering for the thing under test.
Installing the module that provides `artifact-store` requires something
that provides `artifact-store`: its image is mirrored in, mirroring
publishes to the store, and the builder refuses to run without one.
Never seen, because the lab always has a registry standing before the
mesh asks for one, and so does any mesh built on a machine that already
had one.
What it blocks is larger than a registry. The store holds images and
packed archives both — it is the module catalogue in artefact form — so
until it exists a mesh can run only what its bundle already carries.
The substrate record already answers it. It asks of each candidate
whether it can grant itself the thing it provides: the store cannot
create its own database, the broker cannot create its own virtual host,
and the registry cannot grant itself a repository. So the registry
module names its image and is never built.
With a limit worth saying out loud rather than discovering: a module
providing the store may not build artifacts of its own, a UI or a tool
server included, because there is nowhere to put them until it runs.
Such a registry is two modules.
Enforcing "declare what you mount" refuses the builder, which mounts the
container runtime's socket. That socket is not the builder's data: it
exists already, the machine owns it, and declaring it as one of the
module's directories would be a lie the host would act on.
Two kinds of mount are spelled identically today — the directory my data
lives in, and a machine facility I was granted. Until a manifest can say
which, the fourteen declared mounts are right by coincidence, which is
what this issue was opened about.
Recorded rather than decided: separating them is new vocabulary, and
inventing it to turn a check green is how a mechanism nobody chose ends
up load-bearing.
The fourteen undeclared mounts are declared. More to the point, a
manifest that does not declare one is now refused: they were right by
coincidence, and a checklist nothing enforces is a checklist that is
true until the next commit.
Refused in the control plane, because the machine cannot tell the
difference — asked to mount a path that does not exist, it makes the
directory, which is a thing it is perfectly able to do.
A module now says its port once, in `listens`, and the container's
mapping, the rule set and what a consumer is told are all derived from
one assignment. The three hand-written copies that agreed only because
one person wrote them are gone.
The half that made this an issue rather than an inconvenience was that
the substrate is not a module: nothing in the mesh had heard of its own
store, so it handed a database module the port the store already had.
The machine now says what it carries, and the mesh assigns around it.
Ports the protocol fixes became claims, which needed no new mechanism —
the mesh already had one for what is singular on a machine.
Left open, and unchanged by any of this: whether a module should publish
to the machine at all. Assignment makes publishing safe without making
it necessary.
Jochen's call, and the right one: a module cannot choose a port well,
because it is written once and assigned anywhere. Any number it picks is
a guess about a machine it has never seen, and two modules guessing the
same number is not a mistake either of them made.
Writing it up turned up something the issue had missed. The same number
appears three times in every module — the rule set, what a consumer is
told, and what the runtime publishes — and nothing checks that they
agree. They agree today because one person wrote all three. A module
whose `serves` said one thing and whose container published another
would resolve, compose, apply, and hand every consumer a port that
answers nothing.
So the decision is one source with the other two derived, and an
assignment made once and kept, as a credential is.
The part that needed thought is ports that cannot move — mail on 25,
submission on 587. Those become claims, which is what the mesh already
has for what is singular on a machine. Two modules wanting 25 is the
same shape as two wanting the seat, and gets refused by name at
assignment rather than by a container runtime at apply. That makes this
mostly a matter of pointing an existing mechanism at ports.
Left open: whether a module should publish to the machine at all.
Assignment makes publishing safe without making it necessary.
Found by fixing 027 and pushing again. The declaration is now accepted
and the database container still cannot start: the mesh's own store
holds 5432 on that machine, and the module publishes 5432.
Nothing catches it because the substrate is not a module. It arrives
from the bundle before there is a mesh to ask, so the control plane has
never heard of the store and does not know it holds a port. Resolution
can compare modules with each other and cannot compare one against what
the mesh is built on.
Nor does it compare modules with each other. A port is exclusive on a
machine in exactly the way a claim is, and the mesh has a mechanism for
that which ports do not use.
It has been met before: the end-to-end test that exercises a real
database publishes 5433 rather than 5432, inline, with nothing saying
why. That is how a constraint becomes folklore.
The open question is bigger than the bug. Whether a module should
publish to the machine at all decides how a consumer reaches it, and
changes what `serves` means.
Found by the forge failing to start. I had put `restart-on` on nine
containers so they would pick up a rotated credential; it belongs to a
service, and the host refused the whole declaration.
Removing it fixes the modules and leaves the reason I reached for it.
The mechanism is written against exactly this, in the host's own words:
a running service does not re-read its configuration, so replace the
file, find it already running, do nothing, and the machine keeps
behaving as before while every check passes. Every word of that applies
to a container, and nearly everything the mesh runs is one.
The cost is concrete. Rotation replaces the file and tells the provider
to accept the new credential. A provider reconciles, so it takes it. A
consumer is usually a container, so it does not — and the two ends hold
different passwords, which is the fault ADR 0001 records costing two
days. The test that proves rotation works uses a consumer that reads the
file on each attempt, so it does not meet this.
Two things a fix has to keep: it stays declared state rather than a
command, because the link may not carry an action; and where an env-file
changed, the honest verb is recreate rather than restart, because a
container's environment is fixed at creation.
Three corrections, two of them to things I wrote today.
026 is the serious one. Four modules mounted fourteen host paths nothing
declared — the mail spool, the databases, the object store's data. The
runtime creates those as root, so owner and mode go unapplied, and the
rule that keeps a directory holding data the mesh did not put there is
written in terms of declared directories. It reached the configuration
and missed the data. The cause was carrying compose files across: a
container shape that can express one gets filled in like one.
025 claimed nothing turns a tag into a digest. That is false, and the
answer was designed and built before I wrote it. A module names an
artifact, not an image, and `kind: upstream` mirrors somebody else's
image into the mesh's own registry, pinned by the digest it lands with.
The two-document split the issue described as the shape of a fix is the
design. Pinning twelve images by hand was treating the symptom, and left
them pointing at a public registry rather than the mesh's.
And the image store was written up as something the mesh does. It is an
ordinary module — considered for the substrate and removed, because the
test is whether the control plane needs it before its first instruction,
not whether it can grant itself one. So somebody's own registry is the
same module as the mesh's.
Recorded against phase 3, because the phase note said running them
needed images stocked and provisioners built — and missed that not one
of them named an image that exists. Sixty-four zeros where a digest
belongs, eighteen times, parsing and resolving perfectly.
The forge now runs: on a database another module provides, with a
password it did not choose and a connection string it could not have
written itself. First of these descriptions to be started rather than
planned, and it exercises the whole of the credential work.
What remains is the mechanism rather than the data. Nothing turns a tag
into a digest as part of the mesh's own work, so it was done by hand —
which is what the issue says a person should not be asked to do. Asking
a registry takes a second and pulls nothing, so the main argument for
leaving it undone is gone.
The refusal landed, and the examples pin images that exist. What is
still missing is the part that makes it unnecessary: nothing in the mesh
turns a tag into a digest, so it was done by hand — which is precisely
what the issue says a person should not be asked to do.
Recorded because the resolution mechanism turned out to be trivial:
asking a registry what a tag points at takes about a second and pulls
nothing. That removes the main argument for leaving this open.
Also records the two faults that fell out of pinning for real. The mail
system named seven repositories that do not exist, because it publishes
to a different registry than the manifest assumed, and one of the seven
had been renamed upstream. Nothing checking only the shape of a
reference could have found either.
The code and its tests went in hours before the record was touched, so
an issue that read `located` had been fixed all along. That is the exact
failure the frontmatter exists to prevent: status is meant to be
answerable from the record rather than by reading the code.
Closed with the commit that did it, and cross-referenced to 022 and 023,
which came out of the same mistaken instinct — treating the machine as
a boundary, then as an identity, then finding a consumer had a password
and no name to present with it.
The module descriptions sit in `examples/` inside the control plane, and
that name has been doing harm: everything there reads as a sketch, and
one shipped naming a container image nothing builds. A directory called
the catalogue would have made "does this work" the obvious question.
The shape of the answer turns on one measurement. Of the 126 modules in
the system being replaced, 47 are software in their own right — the
largest is 182 source files, and a speech-capture module carries a whole
daemon. Another 44 ship helper scripts. Only 35 are a description and
nothing else.
So a catalogue cannot be a folder of manifests, because two thirds of
modules are programs. That splits them four ways, and only two of the
four belong in a catalogue: things the world made that we describe, and
packages with some files. What the mesh is made of stays in the
repositories that build it. What we wrote keeps its description beside
its code, in the same commit, because nothing else can stop the two
drifting.
The mesh's list of modules is a table, not a repository, and it already
records where each module came from and at which commit. Nothing needs
inventing for modules from anywhere; a repository of ours is just the
source we curate.
The check that a description is valid should move to a command on the
control plane's binary. Today a test reaches into the control plane's
internals to parse manifests, and another reads its build file to check
images exist — two jobs tangled. A command would also give the same
check to somebody describing their own application, which is the case
that matters most and has none.
Left open: how a provisioner's image gets published and pinned, and
whether thirty-five install-a-package modules deserve to be modules at
all.
The cause was one line. Machines get a systemd-networkd unit with a
static address, so networkd finishes and reports the link configured.
The registry ran `ip addr add` inline, which leaves networkd waiting to
configure something it was never told about — and
systemd-networkd-wait-online has an infinite timeout.
So network-online.target was never reached and everything ordered after
it never started. On these machines that is Docker, so `docker load`
blocked on a socket whose daemon was queued behind a target that would
never come, and three bounded timeouts stacked to thirty-five minutes.
These machines have no DHCP by design, so that wait was never going to
end.
The hypothesis in this record was wrong and the record now says so.
Stocking had just been changed, so stocking looked guilty; stocking
takes 34 seconds and always did, timed directly before changing
anything.
Fixed with two things that made it cost hours instead of minutes: an
image placement now waits for the runtime and refuses after 120s naming
what systemd is waiting on, and the end-to-end test passes onProgress —
the raise reported every step and the test discarded it, which is why
thirty-five minutes and four minutes of silence looked the same.
The suite then ran to completion, 23 of 24, the one failure a check of
its own flagging a path as a credential because `/` is in the base64
alphabet.
Also recorded: a redirected log lags, because Node block-buffers stdout
to a file. Read as a stall twice, the second time right after the real
fix — where a buffering artifact argues the fix did not work.
Seen twice today. Once mid-run: thirteen passes, then the process ended
with no summary, no failure and no receipt. Once from the start: the
first test ran 35 minutes against a measured 4.5 and was still running
when it was stopped.
Ruled out rather than assumed: not memory (84 GiB free, no OOM), not the
daemon (the stalled machine answered `incus exec` immediately), and not
the changes under test — the anchor VM had no host log and no
containers, so the run never reached placing the host.
What changed just before is that the rebuild went from two artifacts to
six, and every one of them is pushed into the scenario's registry, which
is the step the second stall sat in. Recorded as what changed, not as
the diagnosis.
The reason this is an issue and not a slow test: the suite prints
nothing between starting a scenario and finishing its first test, so
four minutes and thirty-five look identical from outside, and the only
recourse is to guess. That is how a workstation was left unbootable in
August. And a run that ends silently after thirteen passes is a run
somebody may believe.
023 is fixed, so the design faults are gone and one concrete thing is
left: the realm provisioner does not exist. Its manifest named an image
nothing builds and no program backs, which has been removed — a manifest
describing a program nobody wrote is the same mistake as the credential
files that could never be read.
Keycloak's manifest now says what is true today, and the gap is loud: it
no longer claims to provide oidc-client, so a consumer asking for one is
refused by name at plan time instead of resolving cleanly and waiting
for a client nothing will create.
The provisioner should be written against a real Keycloak in the lab
rather than from the API documentation. The object store's took three
corrections that only a running server produced.
Tool servers — 56 modules, over half — were written up as the biggest
missing thing. They are expressible with what exists, and the first
framing was wrong in a way worth keeping: a module provides `tools` and
the session requires them does not work, because a requirement has one
answer and 56 modules offering tools would be 56 answers.
Turned around it fits exactly. The session provides `tool-host`; every
module offering tools requires it and contributes where its tools are.
Many-to-one is what `contributes` has always been, and the session
receives all of them in one file. Verified by resolving it rather than
by reading the code.
It only became possible today: until 022, several modules on one node
requiring the same thing was refused outright. Worth noting because it
means the credential fix bought more than credentials.
What remains is a decision about what a tool server is, which is work
rather than a missing shape.
The entry stays in the list rather than being deleted — a checklist that
quietly loses its biggest item reads as though nobody looked.
Both halves had one cause: the mesh knew something and did not say it.
Who a consumer is now comes from one derivation, sent to the provider in
its grant and to the consumer in its binding, so the two agree by
construction. The provisioners use the name they are given and refuse to
invent one, because a name of their own would create a login the
consumer could never guess while everything reported success.
Bound values reach the file that needs them through the symmetric twin
of the sealed placeholder — simpler, because they are not secret, so the
control plane fills them in and the host gains nothing.
The lab run meant to prove this failed in a way that looked like the fix
being wrong: rotation could not authenticate against a real database.
The cause was the suite rebuilding the control plane's image and not the
provisioner's, so an image built that minute ran against a provisioner
built the day before. That is 005's family and is recorded with the
issue, because the misleading part is worth more than the fix.
All three modules have manifests, all three parse, resolve and plan, and
none of them can start. Worth writing down before it reads as progress
or as failure, because it is neither.
The vocabulary held. Nothing in 3.1–3.3 needed a new shape — including
the mail system's several containers on a private network, which was the
one expected to break it. That was the question this phase was designed
to answer.
What did not hold was underneath: 022, now fixed, and 023, open. Both
are about credentials rather than about what a module can say.
The third fault was in the manifests, not the design: a secret declared
at a path named .env and read as one, when a sealed file holds a
password and nothing else. That is what a manifest checked only by a
parser buys, and it is why there are now two tests reading the manifests
on disk.
023 is the whole of what remains before the identity provider runs.
I wrote that a module cannot declare an action. It could — the parser
accepted one, and the refusal only came on the machine. The claim was
wrong in the direction that matters: it read as "the design prevents
this", when what prevented it was a check at the far end that nobody
would connect back to the manifest.
Health checks are still the gap most worth closing, but the shape of the
answer is different from what I wrote. An action is not available to a
module at all, so a health check needs a way to say ask this and expect
that without saying run this — closer to a listens entry than to an
action.
Also records the finding itself, because it is a recurring shape here
and not a one-off: a rule enforced only at the far end is enforced and
unusable.
The playbook offered `env-file` and `${secret:name}` as alternatives,
and that reading is what produced the bug every example module shipped
with: own-secrets pointing at a path named `.env`, mounted as env-file,
holding a bare password. The container starts with no password set —
which is a service running on the wrong credential, not a failure.
They are not alternatives. A sealed file holds a password and nothing
else, so env-file points at a file the module declares whose content
leaves a hole, and the host fills it on the machine. A provisioner is
the exception, because it reads a password file.
Written out as the three lines a module needs, with the failure it
prevents named, since the abstract version was already there and was
read the other way.
Amends the credentials page, which said "every pair has its own
credential" and meant two machines. Built that way, it was wrong in a
way that only shows on a real node: a machine running several services
against one database server had one credential between them, so the
provider refused to plan at all and the consuming node quietly gave the
first module a credential and the rest nothing.
The page already argues the case against itself — one credential with
many holders is the first of the three faults it was written to remove.
It just drew the boundary at the machine.
Two modules on one node are as separate as two on different nodes, and
one login opening both is what this page exists to prevent. It is also
what makes withdrawal possible: one role per machine cannot say that
this module has lost its login and the others still have theirs.
022 turned out to have a silent half worth recording: the provider
refuses loudly and names the modules, which reads as a decision, while
the consuming node does not refuse at all. Three modules wanting one
database produce one need, so two of them get no credential file and
each starts and fails to authenticate with nothing saying why.
023 is what remained after fixing it. A consumer now gets its own
password, in whatever shape its configuration wants, and still cannot
connect: the user name is invented by the provisioner and recorded
nowhere in the mesh, and the host and port sit in a JSON binding that an
application reading KEY=value cannot use.
The asymmetry is backwards and the coverage document now says so. The
secret is the hard case, because the mesh must not be able to read it,
and the secret is the part that arrives. The host and port are ordinary
facts the mesh holds in the clear, and they are the ones stuck.
Keycloak, Gitea, Mailu and MinIO all parse and resolve and none of them
can start. This is what stands between the module set and a running one.
Found while checking whether the module vocabulary covers real use
cases. A node running three modules that all want a database cannot be
planned at all:
anchor has 3 modules asking for "postgres-database" and they would
share one credential: gitea, keycloak, umami
The refusal is right about what it says and wrong about what it implies.
They would share one credential, and sharing is worse than refusing —
but the arrangement being refused is the ordinary one, and the node this
mesh exists to take over runs eight modules against one database server.
The cause is the key: a credential is keyed by provision, consumer node
and provider node, so `consumer` is a machine. The provisioner inherits
it and names the role `mesh_<node>`. The refusal is not a check that
caught something; it is the only honest thing that function can do with
a key that cannot tell two consumers apart.
It is the same mistake as 021 with a different face. There the machine
was treated as a trust boundary; here it is treated as an identity, as
though "who is asking" is answered by naming a host. Two modules on one
node are as separate as two on different nodes.
Worth stating plainly: without the refusal, gitea's login would have
opened keycloak's database, and nothing would have said so — from the
provisioner's side it created exactly what it was asked to create.
Not a local fix. It crosses the control plane, the grant file naming and
every provisioner that names something after a consumer.
Every manifest in the system being replaced was read and every key
counted, then set against what the new one can express. Three findings
worth more than the table.
**The most-used key was already covered and I expected a gap.**
Depending on another module — 65 manifests, the commonest thing any of
them says — is a requirement naming a module, which already means that
module rather than anything providing the name.
**The largest real gap is tool servers: 56 modules, over half.** A
module can already run one; what is missing is anything saying it offers
tools. That is plausibly a provision rather than new vocabulary, which
would need nothing added — not yet decided, and recorded as undecided.
**The gap most worth closing is health, at seven modules.** The mesh
knows a container is running, which is not whether it answers, and this
project has paid for that distinction twice. An action with a verify is
exactly the right shape and may not arrive over the link, so a module
cannot declare one.
Two things are missing deliberately and say so: stage hooks, because the
link may not carry an action and a module needing setup ships a program;
and flavours, retired in favour of claims.
Config merging is missing and should stay missing. A mechanism that
understands TOML gets asked for YAML, then INI, which is how the thing
being replaced became unholdable.
Also records what the survey found that is not about coverage: manifests
that had stopped matching what was actually brokered, one fact derived
in two places giving two answers, and a live listing returning
credentials in plaintext.
Written after porting the first real workload end to end. Every step
exists because skipping it cost something, and the ratio is recorded
because it is the lesson: six attempts, one real bug, and the mesh was
right every time.
The rule worth carrying out of it: read the host's log before
theorising. A declaration that was sent and not applied says so there
and nowhere else — it took an hour to look, and the answer was one line.
The conversion's detail is operational and names machines, so it lives
in the mesh's knowledge base rather than in this repository:
`migration/where-service-data-lives` for where every service's data
actually sits, and `troubleshooting/db-password-frozen-at-first-init`
for the lockout. This document says the rule; those say the specifics.
The lockout is the finding worth carrying here, because it is worse than
the one this plan was already guarding against and it is likelier. A
database image consumes its password variable only when its data
directory is empty. Everything keeps data on a persistent directory, so
the role holds whatever password it was created with for ever;
regenerate the variable and the application moves on while the database
does not, permanently, because nothing reconciles it.
Eight modules are in that state today and work only because nobody has
regenerated their credential since their data directory was created.
It was already documented in the knowledge base and my survey had missed
it — found by searching, which is the argument for the knowledge base
existing.
Pinned 2.5.0 rather than latest, on the suspicion that its draft
profiles extension was involved. Identical failure, so that is ruled out
and recorded — two of the three guesses in this issue have now been
tested and both were wrong, which is the useful half.
The scenario keeps the pin regardless; it should have had one from the
start.
Against a real ACME server the proxy orders, the challenge is answered
at the name on port 80 through the proxy itself, the authorisation goes
valid, finalisation is accepted, and the authority issues a certificate.
The client then posts to an empty URL to collect it, and never does.
Read from the authority's own log rather than inferred. Across one run
it issued two certificates and accepted finalise three times: the client
reaches issuance every attempt and fails at the same step after it.
Ruled out and recorded, so nobody repeats it: the directory is complete;
the authority's API certificate covers the address; the challenge path
works. A hand-written server config was suspected and was wrong —
replacing it with the server's own default, changing only the challenge
port, gives the identical error.
Filed rather than pursued because what remains is interop between two
libraries against a server that exists to be a test server, and may say
nothing about a real authority. What the mesh needed to show, it showed:
a routed name gets a certificate ordered from a configured authority,
and an unrouted one gets nothing — that second assertion passes.
Phase 1 closes with this one item partly open. Two of its four tasks
needed no code at all, the network shape was built, and the next thing
to learn comes from moving a module rather than a fourth lab run.
An earlier paragraph implied a secret becomes unrecoverable once
accepted. It does not. It is sealed to the node, which holds the private
half and writes the plaintext into the module's own file at 0600 — the
value is there, on the machine, as an ordinary file.
What does not exist is a way to ask the mesh what a secret is. That is
the property worth having and it is narrower than what was written.
The reason to capture the old system's environment first is simply that
adoption means supplying those values, not that they become
unrecoverable.
Nothing is rotated during the conversion. A service keeps the password
it is already using, because minting a new one is how a running service
stops being able to reach its own database mid-migration.
The mesh has both paths already: generate-and-seal for a new module,
accept-and-seal for an adopted one. Adoption needs the second, and it is
built.
Rotation becomes a separate act afterwards, once everything works — the
machinery is proven, and it is a thing to do deliberately rather than as
a side effect of moving a service between systems.
Records the step that has to come first and is easy to miss: read the
current environment out of the old system while it can still be read.
Once accepted, the mesh cannot show a secret back, and once the old
system is gone neither can that. A password nobody wrote down is a
service nobody can adopt.
Assumed throughout and stated nowhere — the wrong way round for the most
consequential fact about this component.
A board reachable only over the private network would sit inside the
boundary 0004 already calls the security boundary, and a login there
would guard a room whose door is inside the building. This one faces the
internet, so its login is a perimeter rather than defence in depth.
Which makes the identity provider the mesh's outermost gate. The board
presents the control plane, and the control plane's networked surfaces
can change the mesh (0035) — so whoever that provider admits can assign
modules, from anywhere. Written flatly because it is easy to arrive at
one reasonable step at a time and then be surprised by.
What follows is not the board's own design: who may log in is a decision
about the mesh rather than about an application; a public name needs a
certificate from an authority the world trusts, which is why that work
exists; and the provider going wrong in the permissive direction is a
mesh-wide exposure with no local symptom.
The command line is unaffected and is why this is tolerable — it
authenticates through nothing and answers to the machine's own login, so
the mesh stays operable by somebody standing at it whatever happens to
the gate. That is the property to protect if the rest is ever traded
away.
Bootstrap stopped when the control plane started — a mesh that runs and
cannot be used by anybody not standing at the machine, since the
networked surfaces need an identity provider and no module has been
assigned yet. It now runs through the provider and the first login.
The obstacle was not incidental. The mesh has never held a readable
secret: Make generates and seals, keeping no readable copy. An initial
administrator's credential is the first value a person must read.
Generating it and printing it once was the convenient option and is
refused. It would give the control plane a plaintext secret for the
first time — briefly, and to one terminal, but the capability would then
exist, and an exception made for one case does not stay one. The next
awkward credential gets printed too, and "a copy of the database is a
copy of nothing" stops being checkable by reading the code.
So the operator supplies it, on standard input, not echoed — the path
that already exists for a model-access key. What is created is an
account in the identity provider, not a user of the mesh; there is still
no user model.
Unattended bootstrap remains possible and the value still comes from
outside: automation supplying it is the operator supplying it. What is
refused is the mesh inventing one, so an unattended bootstrap with
nothing provided yields a mesh with no administrator — correct rather
than broken.
The mesh is operated from a command line and must be operable from a
browser and from a model's tools, without becoming three systems. The
pattern is already in the code and was unnamed: `board` serves HTTP by
calling the same functions the CLI calls, holding nothing.
Takes the decision 0034 said had to be taken deliberately rather than
arrive with a feature: the HTTP surface is not read-only, so a browser
login now carries authority over the mesh.
Names the dependency by protocol — an OAuth2 identity provider — as the
mesh does for AMQP, S3 and OCI. Keycloak is what fills the role; what
the control plane knows is that it validates a token, and replacing the
provider is a migration rather than a redesign.
Says what this must not become, because it is the failure the project
was started over: a kernel every module imports, 155 files of code from
every context. Shared surfaces are not a shared library. Three adapters
calling the same functions is not the same as logic leaving the context
that owns it.
And records the loop it creates. The networked surfaces depend on a
module the control plane assigns, so when identity is down nobody can
authenticate — including whoever is trying to fix it. The way out is the
command line, which authenticates through nothing and is available to
the account that owns the machine. Hence the rule: no capability exists
only behind an authenticated surface, because that is a capability which
disappears exactly when identity does.
Supersedes 0032, which decided the right thing and described it wrongly.
The decision is unchanged: the account that installed the host owns the
mesh, and there is no user model.
What was wrong was inventing "a surface that delegates authentication"
for the board. It is a web application with a login, in the way every
web application has a login. That is a fact about an application, not a
property of the mesh.
The cost was not cosmetic. It made the identity module look like part of
the mesh's authority — something the mesh depends on to know who anybody
is — when the mesh knows nothing about people at all and one of the
applications running on it happens to have a login.
Keeps the line that is worth writing down, and states it more plainly:
signing in to an application must not become authority over the mesh.
Today it cannot, because the board reads and does not act. The moment it
can assign a module, whoever it lets in has mesh authority — and it
would arrive as a feature rather than as a decision. So a surface that
can change the mesh is a change to who owns the mesh, and is taken as
one. Not forbidden; just not something that turns up in a pull request
titled "add assign button".
Third correction to one table today, found the same way as the other
two: by asking whether both halves of the test were answered, or only
the easy one.
0006 admits the registry because "it cannot grant itself a repository" —
true, and the second half. Nothing established that the control plane
needs one in order to run. Counted rather than argued: the bundle raises
twelve resources and no registry is among them. The registry arrives
afterwards as an ordinary module, which is exactly what the lab asserts.
0006 half-said this already, calling it "substrate by role and ordinary
by delivery, provisioned once there is a control plane to do it". A
member provisioned by the thing it supposedly precedes is not a member;
that phrase was carrying a contradiction rather than resolving one.
The registry is a closer call than the object store and the difference
is worth keeping: the control plane never touches an object store at
all, but it genuinely uses the registry. So the registry is a real
dependency of the mesh operating and not of the control plane starting —
and it is the second that the word means.
The substrate is now exactly what the bundle raises, which is the
strongest form the list can take: checkable by counting rather than by
reading an argument, and the two cannot drift.
The finding is not about substrates. A test with two conditions is a
test only when both are asked.
Answers what 0031 left open, and a question it did not ask — who owns
the mesh at all. There was no answer, and the absence was invisible
because every operation so far has been run by the person sitting at the
machine, so nothing had to say whether that was the design or the
circumstance.
The account that installed the host owns the mesh on that node. No user
model, no roles, nothing to administer. It follows from 0004 rather than
adding to it: there is no authorisation between nodes because every node
is the operator's own, so a user model inside that boundary would guard
nothing — anyone it could stop could read the node's key off the disk.
The board is different, and the difference is the network. A surface
reachable by a browser has to know who is asking, because those people
are not by construction people with a shell on the machine. So it
delegates to an OAuth provider, which is a module.
That does not make identity substrate. A surface delegating
authentication is not the control plane delegating it: the control plane
runs, applies declarations and reaches nodes with no identity provider
in existence. Only the board needs one.
Records the cost plainly: anybody with a shell on a node has full
authority there, and there is no way to give somebody authority over one
node without giving them a login on it.
Closes the last open question about what the substrate contains. 0006
left an identity provider conditional — substrate only if the control
plane delegated authentication — and said the decision had not been
taken. It is now: it delegates to nothing.
The conditional was never about machines. A node proves itself with a
keypair it generated over a broker account issued at enrolment, and
declarations are verified by signature; none of that involves an
identity provider. It was only ever about whether a person signing in to
a mesh surface would be authenticated by something else.
So the substrate is three — a relational store, a message bus, an image
registry — and with 0028 having removed the object store, no member is
conditional and every one is there for the same reason.
It does not settle how a person signs in to a surface, deliberately.
What is settled is that whatever answers that is not something which
must exist before the mesh does, so it can be decided late or replaced —
which being substrate would have prevented.
The conversion method, recorded because it decides everything else and
was not written down.
The old control plane is stopped — provisioning, coordinator, syncs, the
pipeline, anything that decides or writes. The workloads it was managing
keep running, because nothing is managing them. The new mesh then takes
ownership one module at a time.
Nothing is ever unassigned in the old system. Unassigning is how it
removes things and removing is how data is lost; it is asked to stop
having opinions, never to take anything away.
Disabled rather than merely stopped, which is the part easy to get
wrong: those units are enabled, so a stop lasts until the next reboot. A
reboot mid-conversion would bring the old control plane back to
regenerate managed files underneath the new one — the one situation
where two systems really would fight over a machine.
A service left running with nothing managing it is the safe state: it
has its data, its configuration is on disk, and nothing will change
either. The risk in a conversion is in the managing, not the running.
Also records why taking ownership piecemeal is safe: the new host's
orphan removal is per-origin, so it only removes what it recorded
itself. Services it was never told about are not orphans to it.
0030, found by asking what the conversion actually needs rather than by
reviewing anything. The host deleted a directory and everything under it
when it stopped being declared — which happens when a module is
unassigned, or when a manifest is edited to move a data folder, which is
the exact operation this plan needs. A database's files, a mail spool.
The report said "removed".
A directory still holding something is now kept and said so. No flag and
nothing to remember: emptiness is the test, and it works because the
removal order was already right — the mesh's own contents are gone by
the time the directory is reached, so what remains is by definition
something nobody declared.
The plan now says data outranks its own ordering: copy, read back
through the service that owns it, and only then point anything at the
new location. Never move and then check.
And it records where this starts — the node holding all the production
data — with what that costs stated rather than argued with. Everything
proven so far was proven on machines that could be destroyed and raised
again. A scenario proves the mechanism, not the state on that machine.
Recorded because it is load-bearing and was not written down: moving
from the current system to this one is a person at a command line, not a
migration program.
What that removes is larger than what it adds. Nothing in this plan
needs an importer, a translation layer, a compatibility shim, or a way
of keeping two systems agreeing while both are live — each of which
somebody would otherwise reasonably build, use once, and maintain for a
year.
It also settles what "safe" means for the system being retired: a fix to
it must be safe on its own, because there is no careful rollout to
sequence it into. A change needing three steps in the right order is a
change that will be half-applied. That reversed a certificate default I
had chosen this morning.
Ordering needed no change for the third time running — resources apply
in the order declared and nothing sorts them — and is now asserted,
because sorting them for any sensible reason would have passed every
other test.
Separates ordering from readiness, which the task had run together: a
container started is not a container ready. Nothing waits, and what
needs something usable retries. That is deliberate and more robust than
start ordering, since a dependency can restart long after apply.
The network was the first thing in Phase 1 that genuinely needed
building, and the first that needed a decision: 0029 records why a shape
rather than an action, and the vocabulary is nine.
A session as a licence consumer needed no change either: the two
sessions are two modules, so the existing (node, module) binding already
names them apart. 14-model-access.md's "a step toward it and not it" is
true of a worker and not of a session, and the difference is that there
is one session per node rather than many per machine.
Records what stays open: the worker half of that gap is real and
unaffected, and belongs with 0003, which is unbuilt.
Two tasks in a row that were already possible. Both were written from
the design rather than from the code — the review's own finding arriving
in the plan it produced. The remaining Phase 1 items should be checked
against the code before being started rather than after.
An object-store provision, proven against a real store with seven
assertions.
The finding is worth more than the task: the control plane
special-cases nothing. provides, requires, contributes and grants are
name-agnostic, so asking for a bucket needed no change to the mesh at
all. What was missing was a provider and the last step on the machine —
"add an object-store provision" was never mesh work, and the breakdown
now says so rather than leaving the next person to rediscover it.
Named s3-bucket by 0027: the coupling is to the API, not the product,
because swapping one store for another does not break a consumer. A
database is the other case and names its engine.
Records the assertion a database does not need, because it is the one
that will be forgotten when somebody writes the next provider: one store
holds every bucket behind one endpoint, so isolation is a policy rather
than a property, and a policy granting everything passes every test that
only checks a consumer can reach its own bucket.
**0024 accepted.** Model access was decided, built, and proven in the
lab, and two design documents rest on it; only the status had never
moved. The gate is green again.
**The work breakdown rewritten.** It planned a decomposition of the
existing system in place — extract contexts, declared features, shrink
the shared library. That is not the work. A replacement is being built
beside it, and only the old Phase 0 survived contact with reality, so
the one document meant to say what happens next was describing a system
being retired.
Now ordered by what "modules move across one at a time until the old
registry is off" actually requires:
- Phase 0 is marked done against the twenty-two lab assertions, **and
carries its own limitation**: every module exercised was written to
test the mechanism, so the vocabulary was shaped by its own fixtures.
- Phase 1 is the vocabulary gaps found by asking what real modules
need — an object-store provision, a session as a licence consumer, a
network shape with ordering, public certificate issuance.
- Phase 2 is one module, then a week of running it, because the point of
going first is to find what Phase 1 missed.
- Phase 3 picks modules that each prove something the first did not; the
mail system is last because it is the one that may send work back into
the declaration language.
- Phase 4 is switching the registry off, named as a phase so it is not
mistaken for the goal.
Keeps the rules of engagement unchanged — they were about how work is
done, not what it is — with one addition: stop and ask before anything
that touches a machine outside the lab.
Adds a section on keeping the list true, since the document it replaces
was wrong for weeks and nothing said so. A claim here is counted, not
reasoned, and a phase is done when the lab says so.
First pass of a design review, done by reading documents against code
and against a raised mesh rather than against each other. Every error
below was invisible to a proofread.
**Statuses were stale, and nothing checked them.** Ten to-be documents
said `designed` while naming working, lab-proven code — several with a
*What was built* or *Raised, and observed* section. Added a
`status-vs-code` check: naming a file is a claim that the file
implements this, so a document that points at one has stopped being
merely designed. It failed on all ten before it passed, per the rule
this folder sets for its own checks.
**The bundle carries three images, not two.** 07 reasoned about which
substrate services go in and overlooked that the control plane is in
there too — it is what the substrate exists to start, and there is
nothing to fetch it with yet. Counted, not deduced.
**The bootstrap uses four shapes, not six.** It listed `file` and
`directory`, which substrate-first-node.lock never asks for. The claim
that mattered — nothing is blocked on the host — was true either way,
which is why the wrong count survived.
**The eight capabilities were documented nowhere.** Implemented in
internal/profile/detectors.go and enumerated in no document, including
the one about the host that detects them. A vocabulary modules write
against, readable only by reading the code. Now written down, with the
seat/graphical-session distinction that is wrong in both directions if
collapsed.
**MinIO swept out of the to-be layer** per 0028.
The gate now fails on one thing left deliberately: ADR 0024 is
`proposed` while two documents rest on it and the feature it decides is
built and lab-proven. Accepting a decision is not mine to do.
**0027 — provisions.** A module written against PostgreSQL could be
matched to a provider of SQL Server, resolve as satisfied, and fail on
its first query. The name said the role, so nothing distinguished
engines. Refusing on ambiguity could not help: with one provider of
each name nothing is ambiguous. Enforced at parse rather than
documented, because the old naming was the documentation.
**0028 — the substrate.** 0006 admits an object store on the grounds
that it cannot grant itself a bucket. That answers the second half of
the test and assumes the first: the control plane does not need one.
Verified — no S3 client in mesh-control, and internal/builder/registry.go
records the deliberate choice to put artifacts in the OCI registry as
content-addressed blobs. The row was inherited from the system being
replaced, where an object store distributed module tarballs, and was
never re-tested against the definition above it.
So an object store is an ordinary module, and a mesh with nothing
needing one runs none. Migrating it is module work, not substrate work.
0028 also states what 0006 left unsaid: a substrate service and a
module of the same product are different instances. The substrate is
raised from the bundle before any mesh exists, so it is not in the
module graph — a workload depending on it would depend on something the
graph cannot see, cannot rotate a credential for, and cannot move, and
would put workload data in the store the control plane keeps its own
state in.
Both records were found by reading code against design rather than
design against itself, which is the review that should have happened
sooner.
Answers the question 15 raised: a board showing many sessions leaves
one-per-node untouched, because each is still one conversation. Only
concurrent conversations with the same session would touch 0004.
Records soulstream and herdr as the prior art to draw from, and marks
it explicitly off the provisioning path so it stays a note rather than
becoming the work.
Settles the question 15 left open: the mesh session holds its own
memory in the mesh root, rather than assembling a view over the node
sessions. Memory follows the rule the rest of the design already uses —
the context root is the whole of what makes one session a different
agent, and memory is part of what makes it that agent.
The control-plane node is what makes this load-bearing rather than
tidy. Two sessions share that machine; if memory belonged to the
machine instead of the root they would share it too, and the mesh's
recollection would be indistinguishable from that node's own — the
collision 0026 exists to avoid, arriving through the back door.
Also corrects an error made writing it up: memory is NOT declared
state. The engram and tools are — the mesh says what they are and the
host writes them (0011). Memory is written by the session itself and
declared by nobody, so a mechanism that regenerates the root wholesale
would erase it on the next heartbeat, silently, while reporting
success. The root is not uniformly managed and which parts are has to
be explicit.
A session for the mesh itself, addressed as the mesh, differing from a
node's in exactly three things: the context it starts in, its engram,
and its licence binding. Not a new kind of agent — the same mechanism
pointed at a different root. Two implementations of one mechanism drift,
and the vocabulary collision 0001 exists to undo began exactly that way.
It runs on the control-plane node, and the reasoning is easy to get
backwards: not "the important agent on the important machine", but that
this node is already the one place excepted from "compromise of a node
is compromise of that node". Placed anywhere else it would create a
second such place.
It is an addition to per-node messaging and never a replacement. 0001
holds that losing the control plane costs change, not operation — and a
mesh whose only conversational surface lived there would lose the
ability to ask anything while every machine kept running perfectly.
Writing it up exposed that the node session's setup was never designed
at all. 0004 gives behaviour and stops: nothing said how a session
starts, where its context lives, or how a broker message becomes a
prompt. That gap was invisible until something had to be built *like* a
node session. 15-the-agent-session.md covers both as one mechanism.
It also makes "a consumer that is not a machine" undeferrable. The
control-plane node now hosts two sessions that must hold different
licences, and a per-machine binding cannot express that at all. Noted in
14-model-access.md against the gap it was already recorded as.
Also completes the to-be index, which stopped at 10 and omitted four
documents. Pre-existing broken ADR references in the older rows are left
alone rather than guessed at.
Decides the question 006 narrowed to. An agent reads this repository
directly and the search consults it, so these documents surface beside
ordinary results instead of only when somebody already suspects they
exist.
A scheduled sync into the mesh's memory was the option that works with
what exists today, and lost on the ground this repository can least
afford: it makes a second copy, and the copy that is searched quietly
stops matching the copy that is edited. A design record that has
silently diverged from the reasoning it claims to carry is worse than
one that cannot be found — the first misleads, the second merely fails.
Amends what 0019 promised rather than satisfying it: these documents
will not be indexed, they will be read. The commitment that survives is
the one that mattered — that a searcher finds them without already
suspecting they exist.
Gated on an agent that does not exist yet, so 006 stays open on the
build with a decided shape. What closes it is a check that fails today
by design: search the mesh's memory for a phrase that appears only in a
design document here, and require it back.
**004 — certificate issuance.** The resolver declared no authority at
all, so the client fell to its built-in production default: there was no
setting set wrongly, there was no setting. It is now a node property
defaulting to staging, which answers the first open question. Staging by
default rather than production-with-an-override, because the alternative
leaves the safe path depending on remembering to opt out of it — 005's
lesson, in a second place. The rollout is ordered and the order is the
dangerous part; recorded, not performed.
**008 — node rescue.** Read back from running nodes as the report asked,
and one of its own claims was wrong in a way that matters: the health
timer does exist and does fire. It simply never calls the rescue script.
A trigger that exists and does not do what the script claims survives a
halfway check, which makes it worse than the absence the report
described. Resolved by making the documentation true, not by
implementing rescue — the replacement host already supervises recovery,
and wiring unattended restart into the fleet being retired is a
deliberate decision rather than a tidy-up. Two "self-healing" claims
narrowed to what they actually do.
**006 — deliberately not closed.** Re-checked today: the indexing still
does not exist. What is gone is the reason it was an issue — the claim
is no longer load-bearing, because the README names the gap and the
decision's reasoning never invoked indexing. A signpost now points here
from the knowledge base, and was measured rather than assumed: it is
reachable, it is not surfacing. Closing it while the indexing does not
exist would be this repository's own named failure, one folder from
where it names it.
Retired in favour of the lab rather than repaired — that answers the
first open question. The second finding is the one that generalises:
"nothing runs it, and nothing reports that nothing runs it" is not a
fact about that harness, it is a fact about any suite too expensive to
run on every push. The replacement inherited the fault it was replacing.
Records the three rules that now hold, and what the fix taught twice:
the remedy rebuilt the symptom inside itself, and the code that counts
results passed every test while reading nothing.
A service is reached at <service>.<node>.internal, so what resolves is anything
under a node's name. The mesh writes the data and runs no daemon; two roles,
two claims, because systemd-resolved cannot serve a wildcard at all.
Both prohibitions were found by a machine rather than by reasoning: an address
systemd already held, and reading resolv.conf for upstreams that now point at
itself.
Twice in one file, a statement about a machine that read as reasoned and was
wrong — and the module's unit tests all passed while the daemon could not
start. That is what a unit test is: it confirms the assertion was made, never
that it is true of any machine.
003 in prose rather than in a manifest key.
The field was called needs, beside secrets, and both were name-to-path holding
something secret. What separates them is whose, not how secret — so that is
what the name says now.
012 named its own closing condition — a scenario with four images coming up —
and the scenario now stocks seven and has raised cleanly many times at the
memory the wrong diagnosis had raised.
001 is answered by the host reading the package database back after installing.
002 was NOT answered and was present here too, so it is a fix rather than a
note: a stale index is now named instead of reported as a failed install.
The connectivity design still said a hub cannot be filtered — a gap recorded in
the morning and closed in the afternoon, left standing as though it were
current. Worse than a stale date: it would send somebody away from something
that works.
`restart-on` was described nowhere, including the part added today that lets a
service reflect a file another module put on the machine. A rule the host
enforces and no document mentions is a rule nobody can rely on.
And nine of fifteen design documents claimed an `updated:` older than their last
change, some by a week. That field is what cross-cutting views are generated
from, so it is not decoration.
A resolver takes over /etc/resolv.conf, which is a singular resource — ADR 0009
lists it in the table beside the seat and pid 1. So choosing between resolved,
dnsmasq and unbound is assigning a module, per machine, and the mesh refuses
two rather than letting them fight over the file.
Recorded because it was treated as an open question two days after being
decided, which is the argument for that table being a table.
Found by a container failing to resolve a name every machine could: a container
gets its own hosts file holding only its own hostname, and on the machine it
always worked, which is what made it easy to miss.
Declared containers are given the names. A container somebody starts by hand is
not the mesh's to configure — which is a second, different reason to want a
resolver, recorded beside the first rather than folded into it.
Asked whether a machine that drops off needs re-adopting: it does not, nothing
expires, and the only thing that forces re-enrolment is losing its own key.
The gap was the twenty or thirty seconds after a resume in which a node
believes it is in a mesh it has left — recovering on its own, which made it a
quality gap rather than a fault, and still a machine waiting to be told
something it already knew.
It meant failed-or-refused, so the question this record says must not be lost
was answerable only for the machines that broke. Out of date, never told, and
not worked out are kept apart: the remedy is the same push and they read
differently to whoever is looking.
Their subject matter has been built and proven for days and their frontmatter
still said code: [] — which is what the cross-cutting view is generated from,
so it was claiming nothing existed for the substrate, the node lifecycle and
delivery.
One reading answered three ways, holding nothing and touching no context's
store — which is the constraint the whole document is about, and the thing the
board being replaced gets wrong.
Rotation and the provisioner contract; model access as a provision answered by
a record, with ADR 0024's other two gaps left as gaps; exposure, which closes
the open question about revoking a route; and the delivery loop, which closes
the gap ADR 0010 left when it replaced a pipeline with a comparison.
The broker's fingerprint travels with its credential, and the machine's
filesystem does not travel at all — it runs in a container, which is the
arrangement working rather than a limitation to route around.
Found by the firewall: every packet filtered as declared, and the machine
reported as not doing what it was told, because the unit that loaded the rules
had finished. Stated as a gap rather than worked around silently.
Otherwise it succeeds into a state its verify rejects, and the host's report is
accurate and names nothing. Recorded where the vocabulary is described, because
it is a rule about writing an action rather than about one action.
A hub needs its overlay port open and a node that is not a hub does not, and
they are the same module — so listens, a static manifest field, cannot express
it while the overlay module's resources are computed per node. Written down
rather than left as an oversight for whoever first puts a firewall on a hub.
Issue 014: the node's serving key was stored in the host's own encoding, so
every check that reads the file passed and no server could start. Same shape as
013 — two halves of one mechanism designed separately, each correct about its
own half. Where a file exists so a third party can read it, the format is the
interface.
Issue 003 is answered in both halves: manifests are parsed strictly, and a
module says what it listens on and from where rather than carrying a key
nothing reads. The design records what was built and how each part is checked.
Issue 013 is new, found by reading while writing the first module that has
both a computed file and a service that needs it. The file arrived second.
It failed, then the next reconcile fixed it, which is why nothing caught it.
2026-08-31 00:37:34 +02:00
345 changed files with 30492 additions and 551 deletions
description:Use when a thought is raised that should be remembered but NOT worked on now — an aside during other work, a "we should look at X someday", a known gap nobody is assigning yet. Triggers on "defer this", "park this", "register this thought", "note this for later", "don't work on it, just remember it". Records and returns to whatever was already in progress.
---
# hq-defer
Parks a thought so it is not lost, **without moving the work off course.** The defer is the
point: the thought is recorded and the previous task resumes.
This skill wraps no playbook, because deferring is not part of the development cycle — it is
what happens *before* something enters it. A parked thought has no number, no owner and no
status, and that is correct.
## Where it goes, and why not the repository
Record it as **one memory file** in Claude's persistent memory directory for this project (the
path is given in the session's memory instructions), with `metadata.type: project`, plus a
one-line pointer in `MEMORY.md`.
**Not** in `04-ISSUES`, `01-RESEARCH` or anywhere else in the repository:
- A parked thought is not an issue or a research effort. Giving it a number asserts it has been
triaged, which is exactly what deferring says has not happened.
- A shared "deferred" or "someday" document is a **central status file**, which
[`AGENTS.md`](../../../AGENTS.md) forbids. Status lives in frontmatter on real records, and a
parked thought has no real record yet.
- A repository write means a branch, a commit and a pull request — drift, which is the one thing
this skill exists to avoid.
## Steps
1. Write the memory file. Slug is kebab-case and descriptive of the thought, not of the act of
deferring.
```markdown
---
name: <kebab-slug>
description: Deferred note — <one line>
metadata:
type: project
---
Raised and deliberately deferred on YYYY-MM-DD: **<the thought, in the user's own terms>**
**Why:** what was being worked on when it came up, and that deferring was intentional so
that work was not pulled off course.
**How to apply:** treat as an open thread, not an assignment. Do not start on it
unprompted. If it graduates it needs an HQ home first — an issue under playbook
[03](../../../00-META/process/03-issues.md) if a stated behaviour does not happen, or
research under playbook [01](../../../00-META/process/01-research.md) if it is still an
idea. Say which is undetermined, if it is.
```
2. Append one line to `MEMORY.md`: `- [<Title>](<kebab-slug>.md) — deferred YYYY-MM-DD; parked, no HQ
record, do not start unprompted`.
3. Convert relative dates to absolute before writing. "Last week" is worthless in six months.
4. Check for an existing memory covering the same thought and update it instead of adding a
duplicate.
## Then stop
Reply in **at most two lines** — what was recorded, and that it is parked — and **return to
whatever was in progress before.** Do not summarise the parked thought back at length, do not
propose a plan for it, do not ask which playbook it belongs to, and do not open anything.
If nothing was in progress, say only that it is recorded.
## Do not
- Do not create an issue, a research effort, a decision record or a design document.
- Do not create a branch, commit or pull request.
- Do not start investigating the thought, however cheap the first check looks.
- Do not name nodes, domains, addresses, absolute paths or usernames in the memory file — the
thought may later be quoted into this repository, which is public.
- Do not decide whether it is an issue or research when the evidence does not say. Recording
"undetermined" is the honest outcome and costs nothing later.
@@ -9,6 +9,7 @@ works — plus the engineering practice that holds across everything Novox build
| [`context.md`](context.md) | The environment — conditions, not aspirations |
| [`effect.md`](effect.md) | What is different when the work is done |
| [`how-we-build.md`](how-we-build.md) | The rules that hold across the mesh, each one earned. **The source of the mesh constitution** — the governed page the mesh injects into design sessions is derived from it. |
| [`glossary.md`](glossary.md) | One name per thing — the authority on vocabulary, and the words that were retired |
| [`repos.md`](repos.md) | Where implementation lives, and what each repository owns |
| [`process/`](process/) | The playbooks — how work moves through this repository, for engineers and agents alike |
@@ -26,6 +26,7 @@ indistinguishable from one that cannot.
| `numbering` | the number in the filename is the number in the heading | — |
| `topics` | every record names a topic the index knows | — |
| *(index.py)* | the written reading order matches what the records say | — |
| `status-vs-code` | a to-be document naming specific code is not still `designed` | **ten documents**, several with a *What was built* section, describing lab-proven code |
## What is deliberately not checked
@@ -47,3 +48,10 @@ what to do.
State what incident it would have caught, and make it fail before you make it pass. A check
whose failure has never been observed is a guess about its own correctness.
## cycle.py
The development cycle, checked ([ADR 0080](../../02-DECISIONS/0080-the-development-cycle-is-checked.md)):
a to-be design names a decision, an in-progress/implemented design names its owning code, a
located/fixed issue names its owner, a fixed/resolved issue says what fixed it, a graduated
research overview says what it became. `python3 00-META/checks/cycle.py`
# Glossary — the words this repository uses, and the ones it stopped using
One name per thing. This page is the authority; where an older record says something else, that
record is being superseded, not this page. It exists because the terms kept drifting in
conversation — control plane / controller / master / hub for one thing, substrate / foundation for
another — and a mesh you cannot name precisely is a mesh two people describe differently.
## The mesh and its machines
- **node** — a machine in the mesh. There are 0..n of them, and each runs the host agent. A node is
just a machine that has joined; being one implies nothing about what it runs.
- **control-node** — the one node that also holds the `mesh-controller` seat. There is exactly one
per mesh. "control-node" is not a separate kind of machine — it is a node that additionally runs
the controller (and, today, the foundation). Lose it and the other nodes keep running what they
were last told; they simply cannot be told anything new.
- ~~master / slave~~, ~~hub / peer~~ — not used. The relationship is *controller and nodes*, and no
node is subordinate: a node applies declarations on its own and survives the control-node dying.
## What runs the mesh
- **controller** — the component that decides what each node should be, holds the mesh's records,
and tells nodes over the broker. Replaces **"control plane"** (borrowed from networking's
control-plane/data-plane, and opaque here).
- **mesh-controller** — the module that runs the controller. It **claims** the `mesh-controller`
seat at mesh scope, which is what makes it singular. Replaces the module name **`mesh-control`**.
(The git repository has been renamed `mesh-control` -> `mesh-controller` on the forge; the module,
container and image it produces are `mesh-controller`.)
- **foundation** — the store and the broker, raised at genesis before any module system exists.
Replaces **"substrate"** (a biology metaphor that landed for no one). The foundation is not a
third thing beside the store and broker — it *is* those two, named together.
- **store** — the one postgres server. It holds the controller's own context databases
(`inventory`, `identity`, `licences` — a context owns its store, [ADR 0008](../02-DECISIONS/0008-a-context-owns-its-store.md))
and every module's own database. One server, many databases — never one shared "mesh database".
- **bus** — the mesh's own nervous system: NATS, one per mesh, carrying every link the mesh has —
control, declarations, builds, events, tool calls
([ADR 0106](../02-DECISIONS/0106-the-bus-is-nats.md)). A module reaches it by requiring
`mesh-bus` ([ADR 0128](../02-DECISIONS/0128-the-mesh-bus-is-required-not-ambient.md)); one that
does not require it has no account on it. Held by the `mesh-broker` seat, which is named after
the *role* rather than the server, so the server can change without the seat doing so.
- **the deprecated broker** — the lavinmq module. It was the mesh's bus and is not any more. It
keeps running as an **ordinary provider** of the `amqp` provision, for modules that need a
message broker of their own the way something needs a database
([ADR 0127](../02-DECISIONS/0127-amqp-is-a-provision-not-the-bus.md) (superseded by [ADR 0131](../02-DECISIONS/0131-everything-on-the-mesh-speaks-to-the-broker-seat.md))) — no seat, not foundation,
never raised at genesis, and a mesh that never installs it is complete.
Say *the deprecated broker*, not "the compatibility broker" (it serves the mesh's own modules,
not only the predecessor's) and not "the AMQP broker" (naming it after a protocol invites
describing the bus by contrast with it, which is backwards: the bus is the mesh's nervous
system and this is a module).
## What the mesh stores and serves
- **package** — what code resolves when it is **compiled**: an npm/cargo/pypi dependency, by
**version**. Served by the **package-registry** (gitea). Only a builder talks to it.
- **artifact** — what the mesh delivers to a machine to **install and run**: an OCI image, by
**digest**. Served by the **artifact-store** (distribution). Every node pulls from it.
- These are two protocols, not one store being weak — see [ADR 0075](../02-DECISIONS/0075-two-stores-and-which-provides-what.md).
## How modules relate to the mesh
- **seat** — a named role at a scope (node / site / mesh), held by a module assignment, from a
**closed set** the mesh defines: a claim naming a seat outside the set is refused. A seat may
**deliver a provision**, and its holder is then the mesh's answer for it when several modules
provide it ([ADR 0126](../02-DECISIONS/0126-a-module-declares-its-own-seats.md) (superseding [ADR 0110](../02-DECISIONS/0110-a-seat-is-a-module-assignment-from-a-closed-set.md))).
The set, with who holds each seat, is the overview of what a mesh has
([26 — The seats](../03-DESIGN/01-to-be/26-the-seats.md)). A seat has a **capacity**: a
capacity-1 seat is exclusive (one holder); a higher-capacity seat is a **bench** (several holders
coexist).
- **claim** — a module taking a spot on a seat. `claims: [{name, scope}]` in a manifest. A
mesh-scoped exclusive claim is how the mesh says "there is one of me". A foundation seat is
named after the server it guards: the `mesh-controller`, `postgres` and `lavinmq` modules claim
the `mesh-controller`, `mesh-store` and `mesh-broker` seats ([ADR 0079](../02-DECISIONS/0079-the-foundation-seats-are-named-after-their-servers.md)).
- **provision** — a service one module `provides` and others `require`; the mesh resolves a provider
and wires the two with an endpoint and a credential. A provision is a service you offer, a seat
is a role you occupy, and the two meet where a seat delivers a provision: occupying the seat is
what makes a module *the* provider of it.
## How this page is kept
A new name for an existing thing lands here first, in the same change that introduces it in code. A
record under `02-DECISIONS/` keeps whatever word it was written with — those are immutable — so a
term retired here may still appear there, and the mapping above is how to read it.
**Trigger.** Something that runs today must run on the mesh, or a new capability must be
declarable.
**Who runs it.** Whoever is porting or writing it.
*Written 2026-09-01 from doing this for the first time end to end. Every step below exists
because skipping it cost something.*
## Before anything: read what runs
**A module is written from the thing, not from memory of the thing.** For a port, that means its
current compose file, its environment, and where its data actually sits. Assumptions about any of
the three have been wrong every time they were not checked.
Three questions, answered from the machine:
| | why it decides something |
|---|---|
| **what containers, and how do they find each other?** | more than one means a `network`; names between them must match what the software is configured to dial |
| **where is its data?** | a bind mount moves with a path; a named volume does not; an anonymous volume is already losing data on every redeploy |
| **which values are secret, and which are merely settings?** | a secret goes in `own-secrets` or a grant; a setting goes in the manifest and may be overridden per node |
## The steps
1.**Name what it provides and requires**, if anything. A name is what a consumer is coupled to,
not the role it plays ([ADR 0027](../../02-DECISIONS/0027-a-provision-names-what-the-consumer-is-coupled-to.md)):
`postgres-database`, not `database`. Most modules provide nothing and require nothing — an
application is usually a leaf.
2.**Declare capabilities, not dependencies, for facts about the machine.**`container-runtime`,
`package-manager`, `seat`. A capability is detected and refused against; it is not something a
module can install.
3.**Write the resources in the order they must happen.** They are applied in the order written
and orphans are removed in reverse, so a `network` is written before the containers that join
it and removed after them.
4.**Put every secret in a file, never in `env`.** A declaration travels over the broker in plain
text: a password in `env` is a password the broker sees. **The mesh delivers parts; a module
that needs them combined combines them.**
**A sealed file holds the password and nothing else** — no key, no `=`, no newline that means
anything. So `env-file` must never point at one. It points at a file the module *declares*,
whose content leaves a hole:
```
own-secrets superuser → /var/lib/postgres/superuser.secret the password, alone
a file /var/lib/postgres/superuser.env, mode 0600,
content: POSTGRES_PASSWORD=${secret:superuser}
the container env-file: [/var/lib/postgres/superuser.env]
```
The host fills the hole on the machine, which is the only place both halves exist — the mesh
discarded the value
([credentials and their rotation](../../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md)).
A **provisioner** is the exception: it reads a password file, so it mounts the `.secret`
directly.
Every example module in `mesh-controller` had this wrong and shipped: `own-secrets` pointing at a
path *named* `.env`, mounted as `env-file`, holding a bare password. Docker reads that as a
malformed line and the container starts **with no password set at all** — not a failure to
start, a service running on the wrong credential. They parsed and they resolved. Two tests in
`examples/modules` now refuse both halves of it.
Add `restart-on` naming the env file, or the container keeps the credential it started with
through every rotation.
5. **Pin every image by digest.** A tag moves. The manifest in a repository names artifacts; the
manifest the mesh holds names digests, and they are not the same document.
6. **Decide generate or accept.** A new module's credential is generated. **An adopted one keeps
the credential it already has** — `secret accept` — because minting a new password for a
database that already exists locks the application out of its own data.
7. **Add a provisioner only if the software cannot read a file.** A proxy that watches a
directory needs nothing. PostgreSQL needs `CREATE ROLE`, an object store needs a bucket and a
policy, an identity provider needs a realm and a client — those need a small program beside
them. It reads what the mesh granted and reconciles; it does not decide anything.
8.**Prove it in the lab, against the real software.** Not that a container started — that the
thing works: the credential authenticates, a wrong one is refused, the containers reach each
other, the data survives a restart.
## What the first port actually cost
Six attempts, one real bug. Recorded because the ratio is the lesson: **the mesh was right every
time and the scaffolding was not.**
- A shape existed in the language and no host implemented it, so every declaration carrying one
was refused whole — correctly, and the host said exactly that. **Nobody was reading the host's
log.** Read it first; it is the only place that says why a machine did nothing.
- A blind find-and-replace renamed a provision in quotes and missed the same word bare.
- A command was tested only for the invocations that should fail, so it rejected every real one
and the suite stayed green.
- A test asserted on a helper rather than on the code that calls it, three separate times. **A
test that cannot fail when the behaviour is deleted is not defending the behaviour.**
## Rules
- **Read the host's log before theorising.** A declaration that was sent and not applied says so
there and nowhere else.
- **A failing test is kept, not skipped.** It is the reproduction.
- **Never rotate during an adoption.** Rotation is a separate act, afterwards, deliberately.
- **A data directory is never removed by the mesh** ([ADR 0030](../../02-DECISIONS/0030-data-outlives-the-mesh-that-declared-it.md)),
and that protects against the mesh only — not against a disk or a mistaken command.
@@ -17,21 +17,24 @@ and a forge address is an operational detail (see [`README`](../README.md)).
| `hal` | The monorepo — the node runtime, the module catalogue, the delivery machinery, and the bootstrap scripts. Every core module lives here. |
| `hq` | This repository, under the company organisation — mission, research, design, decisions, issue diagnosis. Company-scoped ([ADR 0019](../02-DECISIONS/0019-how-this-repository-works.md)); the mesh is its first product. The source of truth for *why*. Carries no implementation. |
| *(one per application)* | Every standalone application, site or side-project gets its own repository, with `module.yml` at the root. Registered with the mesh as a build source; built and deployed by the same pipeline as anything in the monorepo. |
| `migration` | **Private.** The record of one installation replacing the predecessor mesh with this one: the runbook, a dated log of every step and what it cost, the per-service data procedures, the readiness checks, and the scripts. Private because it is the opposite of this repository in every way that matters — it names machines, addresses, ports and paths, because a procedure that cannot be followed is not one. Where hq asks *what did we decide and why*, that repository answers *what happened on the machines, in what order, and what to do next*. Its `HANDOFF.md` is where somebody picking the work up starts. |
## What the mesh becomes
[ADR 0019](../02-DECISIONS/0019-how-this-repository-works.md) records the repositories the
monorepo decomposes into. **`mesh-lab`, `mesh-host` and `mesh-control` exist so far** — the lab is built first
([ADR 0016](../02-DECISIONS/0016-the-lab.md)); the rest are the
so far** — the lab was built first ([ADR 0016](../02-DECISIONS/0016-the-lab.md)). The tiered
decomposition below is the planned shape; the repositories built to date do not map onto it
one-for-one — `mesh-catalog`, `mesh-sdk` and `mesh-tools` exist where the table names
`mesh-foundation` and `mesh-surfaces`, and reconciling the two is itself still ahead.
| Repository | Tier | Holds |
|---|---|---|
| `mesh-host` | 0 | **exists.** The node host — one statically linked binary, requiring nothing present ([ADR 0005](../02-DECISIONS/0005-the-node-host.md)) |
| `mesh-substrate` | 1 | the four pinned services, as declarations |
| `mesh-control` | 2 | **exists.** The control plane and its contexts — one of seven built ([ADR 0006](../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)) |
| `mesh-foundation` | 1 | the four pinned services, as declarations |
| `mesh-controller` | 2 | **exists.** The controller and its contexts — one of seven built ([ADR 0006](../02-DECISIONS/0006-the-substrate-and-the-control-plane.md)) |
| `mesh-surfaces` | 3 | tools, web, cli |
| `mesh-sdk` | — | contracts shared across tiers |
| `mesh-sdk` | — | the stable spine modules build against — the tool-serving harness, the messaging/event framework, the contracts and core primitives. Holds nothing per-module and nothing volatile ([ADR 0039](../02-DECISIONS/0039-what-the-sdk-holds-and-refuses.md)). |
| `mesh-lab` | — | **exists.** The lab — scenario lifecycle, networking, placement. Ships to nobody; runs on a workstation. |
Tier 4's shape is open, and deliberately so: see ADR 0019 and
@@ -90,5 +90,5 @@ the catalogue where modules genuinely change together under one intent. The skel
| One repository per tier, or per context? | Already open from ADR 0001 as "catalogue destination — one repository or many". The skeleton assumes per tier and does not settle it. |
| ~~Does an unprivileged node earn a place in the inventory, or only a presence?~~ | **Answered 2026-08-25** by the operator: a node is a *managed machine inside the mesh*, not an unprivileged something — and a disconnected node is still a node, in a different situation. The question posed a class distinction; the answer is that there is none, and what varies is **state**. Recorded as [ADR 0004](../../02-DECISIONS/0004-a-node-and-how-it-joins.md). |
| ~~Does absorbing overlay, filtering, packages, supervision and the container runtime make the host too large?~~ | **Answered 2026-08-25** — [`host-size.md`](host-size.md). Measured: the absorption is smaller than the machinery that already applies state, and eight of ten adapters already carry no dependency. The risk is not size but direction, and it is two modules wide. The claim survives with its scope corrected — the host carries one concern, *apply declared state on this machine*, of which the six are instances. Recorded as [ADR 0005](../../02-DECISIONS/0005-the-node-host.md), designed in [`05-the-node-host.md`](../../03-DESIGN/01-to-be/05-the-node-host.md). |
| ~~Four substrate services or five?~~ | **Answered conditionally**, which is the honest form — [`07-the-substrate.md`](../../03-DESIGN/01-to-be/07-the-substrate.md). The substrate is *what the control plane consumes and cannot grant itself*. The identity provider qualifies only if the control plane delegates authentication; if it authenticates natively it is an ordinary hosted service. The count follows from a decision not yet taken, and asserting four was asserting that decision. |
| ~~Four substrate services or five?~~ | **Answered conditionally**, which is the honest form — [`07-the-foundation.md`](../../03-DESIGN/01-to-be/07-the-foundation.md). The substrate is *what the control plane consumes and cannot grant itself*. The identity provider qualifies only if the control plane delegates authentication; if it authenticates natively it is an ordinary hosted service. The count follows from a decision not yet taken, and asserting four was asserting that decision. |
| Does `feature` survive? | The skeleton splits it in two and argues the conflation is what makes the delivery pipeline hard to reason about. Unproven. |
@@ -191,6 +204,7 @@ reopened by adding severity, and is not decided here.
| Can a module say which of its settings are load-bearing? | The question that dissolves the conflict rule rather than choosing a side. A setting the module *requires* cannot be kept from the machine without producing something installed and broken; a setting it merely *prefers* should always yield. Until a module can say which is which, adoption is defaulting in the dark. Belongs with the graph. |
| Does a `failed` line still let adoption complete? | *Flags inform, they do not block* was decided about conflicts, where the mesh chose and the machine works. A failure is *we could not*, which is different in kind — and treating them alike hides the worse one behind the commoner one. |
| How is a flagged conflict reconciled, and by whom? | The briefing hands it to a session. What that session is empowered to change, and whether the resolution is recorded so the next adoption does not re-raise it, is undecided. |
| ~~How long is a machine adopted?~~ | **Decided** ([ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)) — for as long as it is being migrated: a mode per node, recorded, ended by an explicit and previewed flip. |
| Where does the kept original live, and for how long? | Whether it is recorded in the node's state so adoption is visibly reversible, and whether it is returned when the mesh stops managing the thing. |
| What shape is a briefing? | Structured enough to be acted on, prose enough to be read. It is the first thing a session on a new node sees, which makes it an interface rather than a log. |
| Does owning a package mean owning its version? | Owning configuration and owning the package are different scopes. The second means the mesh decides which version is installed, and that decision then has to survive the machine's own package manager updating it. |
# Migrating a node that is in use — adoption as a mode, not a moment
*2026-09-22. Measured on the machine that will be the control-node, which is running the
predecessor mesh today. Nothing below names it; the counts are its own.*
## The question
The mesh replaces a predecessor mesh that is running, on the same machines, with the services
people use. The control-node is decided: it is the machine that already carries the predecessor's
broker and build pipeline. So the question is not *where* the mesh starts but **how a machine
running the predecessor becomes a node of the mesh without its services noticing** — and then how
each of the other machines follows.
## The operator's proposal
Proposed by the operator, 2026-09-22, and the shape this document tests:
1.**Stop the predecessor's control on a machine** — its daemons that write configuration: the
network and firewall configuration above all, which decide what is reachable and what is
blocked. Its services keep running; only the control over their configuration stops.
2.**Bring the mesh up on that machine in adoption mode.** It takes custody of those files and
keeps what it finds in force.
3.**Migrate the modules one at a time**, data preserved, per the cutover procedure.
4.**Move to the next machine and repeat** — adopted first, keeping its local configuration, then
migrated.
5.**When every machine is migrated, flip adoption mode**, and the mesh takes full control of the
configuration it has been holding.
This is the research above made concrete. It keeps the conflict rule already decided here — *on
conflict, what is on the machine stays* — and gives the middle state, *adopted*, a length: not a
one-time import before generating starts, but a mode that lasts for as long as the machine is
being migrated, ended by an explicit act.
## What the machine actually looks like
Measured, read-only:
| | |
|---|---|
| Containers running | 60, all the predecessor's services and their stores |
| The predecessor's control | user-level daemons, separate from the services; none of the 20 running system services is the predecessor's control |
| Firewall | the predecessor's, active: 54 incoming rules and 52 forwarding rules, each served port allowed explicitly — how a default-deny firewall reads |
| Files the predecessor's configuration sync writes | 12, of which 3 are system files (an ssh server drop-in, the package manager's configuration, one service's configuration); the rest are the operator's shell and agent files |
| Per-service configuration | environment and composition files per service, written by the predecessor's service tooling |
**Stopping the predecessor's control stops nothing that serves.** The services are containers and
system units that run without it; what stops is the rewriting of their configuration. Nothing
changes on the machine until something else writes.
## What collides, measured rather than assumed
An earlier note assumed the mesh's foundation could not stand beside the predecessor because both
want the store's and the broker's standard ports. **The measurement says otherwise.** The
predecessor publishes its own store on a non-standard port and its broker on another; the standard
ports the foundation binds for its store and its bus are free.
What does collide:
| The mesh wants | Held by | When |
|---|---|---|
| the registry's port | the predecessor's registry | at genesis — the foundation raises a registry |
| the broker's management port, on loopback | the predecessor's broker | at genesis |
| the private network's port | the predecessor's own tunnel | at genesis on a control-node that is the private network's hub; otherwise when the node is placed on it |
| the web ports | the predecessor's reverse proxy | when the route proxy is assigned |
| the resolver's port | a resolver the predecessor runs | when a resolver module is assigned |
On the control-node, which is the private network's hub, three of these fall at genesis and cannot
be deferred; two only when a particular module is taken, the moment its predecessor stops anyway.
Today the foundation's ports are **fixed**: written in the installer's bundle and in the
catalogue's manifests, so a collision is found when a container fails to bind, not before — and a
port changed at genesis would be changed back when the foundation is adopted as modules, since the
applier recreates a container whose declared spec differs. The private network's address range
must also stay clear of the range the predecessor's tunnel uses; on the machines measured they are
distinct.
## What would break if the mesh came up as it is today
**The firewall.** Genesis loads a base ruleset — drop anything undeclared — in its own table
([ADR 0088](../../02-DECISIONS/0088-the-foundation-filters-before-anything-listens.md)). The
predecessor's firewall is a different table. The kernel runs every base chain registered at the
same hook, in priority order: an accept ends only its own chain and the packet goes on to the next,
and a drop in any is final — whether the other firewall's chains are nftables or legacy iptables.
So the
mesh's base ruleset would drop everything the predecessor's firewall allows and the mesh has not
declared — every web, mail and database port in the table above — the moment genesis ran.
**The files.** The host writes a declared file whatever it finds at the path, reporting it as
updated. A file the predecessor left — the ssh server drop-in, the resolver's configuration — is
replaced the first time a module declaring that path is assigned, before that module's service
has moved.
**The container names.** The applier keys a container on its name. A catalogue module whose
container carries the same name as the predecessor's service it replaces takes that container over
the moment it is assigned: today, *assigning a module is migrating it*, never a preparation.
**The published ports.** The foundation publishes its ports on every interface, and a published
container port reaches the container through the forwarded path, not the incoming one. A firewall
that filters only incoming traffic never sees it. The base ruleset is what keeps the store
unreachable from outside today — and it is the thing that cannot be loaded on this machine.
## The pipeline freezes while the control-node migrates
The predecessor's build pipeline and its coordinator run on the control-node. Stopping its control
there stops the predecessor's updates for **every** machine it manages, until the migration is
done. Their services keep running; they receive nothing new. That is the price of the proposal and
it is worth stating, not a reason against it: the migration is the period in which the predecessor
is being replaced, and it does not need to keep changing.
## What adoption mode has to mean
For the proposal to hold, *adopted* must be a **state the mesh records per node**, not an
intention, and each thing the mesh would otherwise take must say what it does in that state:
- **What is found is kept until its module is taken.** *Found* is precise: present at a declared
path or name with no record in the host's store. A found file or container is held, its original
recorded, until the operator **takes** the module on that node — the cutover, done when the
module's data has moved. Assigning prepares; taking migrates. Without the distinction the rule
never fires: the host only ever sees what assigned modules declare.
- **The firewall found on the machine stays in force.** The mesh loads no table on an adopted node
that drops by default or holds an accept. What it needs open it declares as openings the host converges
*through the found firewall*, on the incoming and the forwarded path, marked as the mesh's and
re-checked on every reconcile so a reload or reboot does not lose them. An accept in a table of
its own would not help: the found firewall's drop would still be final. What a table of its own
*can* do is refuse, and a refusal is final too — so the mesh guards the store and the broker's
management port from everyone but the private network and the machine itself, in a table that
only refuses, ahead of the container runtime's redirect — which the found firewall does not do
and cannot undo. The bus, the registry and the hub's port stay open to
anywhere: a node enrols before it has a private-network address.
- **The foundation's ports are the node's to give** — set at genesis, checked free, and kept as
that node's settings, read everywhere they are used, so adopting the foundation as modules does
not move them back.
- **The flip is per node and previewed** from what is actually reachable — listening sockets and
published ports, not the found firewall's allow list, which does not see what a container runtime
forwards.
## Options weighed
| Option | Verdict |
|---|---|
| Cut the machine over in one go: stop the predecessor's store, broker, registry and proxy, raise the foundation in their place | Rejected. Every predecessor service goes down until it has migrated, and the predecessor's other machines lose their broker. The rollback is restarting the predecessor, which is a recovery, not a step. |
| A separate machine as the control-node | Rejected. Contradicts the decision that the control-node is the machine that already carries the predecessor's broker and pipeline. |
| Converge on joining, as the mesh does today | Rejected. The base ruleset closes every predecessor port at genesis, and found files are replaced before their services move. |
| Make only the foundation's ports configurable and otherwise converge | Rejected as insufficient. It solves the bind collisions and none of the firewall or file ones. |
| Treat assigning a module as migrating it | Rejected on review. The rule that keeps found files would never fire — the host only sees what assigned modules declare — and an assignment on an adopted node would be an outage rather than a preparation. |
| **Adoption as a mode, per node, ended by an explicit flip** | The proposal. It is the conflict rule already decided here, given a duration. |
## What this leaves open
Answered by the decision record this feeds — [ADR 0100](../../02-DECISIONS/0100-a-node-in-use-is-adopted-before-it-is-converged.md)
— only for the migration's needs. The general questions above stay open: whether a module can say
which settings are load-bearing, what shape a briefing takes, how a flagged conflict is reconciled.
One is newly sharp: the mesh opening ports through a firewall it did not install needs to speak that
firewall. There is one kind on the machines measured; a machine with another is not covered until
**The predecessor's world**, now on the same broker: ~50 AMQP connections per machine from four
machines, 595 queues, four vhosts, six users. All of it AMQP, all of it retiring module by module.
NATS speaks no AMQP: that world cannot move; it does not need to.
## How NATS answers each of those
| The mesh relies on | NATS | Note |
|---|---|---|
| topic routing keys | subjects with wildcards | same shape (`mesh.events.>`) |
| RPC and tool invocation over an exchange + reply queue | request/reply, native | simpler than today |
| competing consumers | queue groups | same |
| durable control queue, hold-unacked-and-retry, catch-up | JetStream: streams, durable consumers, ack/nak with delay, replay | the 0083 guarantee moves to JetStream; core NATS alone is at-most-once and would not do |
| dead-letter | max-deliver + advisories, or a stream fed from them | different mechanism, same effect |
| TLS bus | TLS | same |
| accounts scoped by emits/consumes, vhosts | accounts (isolation) with users and per-subject publish/subscribe permissions | stronger than today; the controller writes an auth config the host declares and the server reloads, instead of calling a management API |
| management API | `nats` CLI and an HTTP monitoring endpoint | no vhost concept — accounts instead |
| MQTT | built in | same |
| a module seat `mesh-broker` | unchanged — the seat is the server, the module changes | ADR 0079 |
| multi-node | clusters and leaf nodes | not needed now; a leaf per node is a later question |
Nothing the mesh needs is missing. The differences are in the shape of durability (JetStream must
be declared, streams and consumers are objects) and of accounts (configuration, not API calls).
## The cost, honestly
The bus is the mesh's nervous system. Moving it means: a new `nats` module in the catalogue taking
the `mesh-broker` seat; the controller's and the host's link packages rewritten; the tool runtime's
client swapped behind the unchanged contract; enrolment, reports, builds and the guard's ports
re-derived; the lab beds that prove the bus (store window, enrolment, upgrades) re-run on the new
one; and eight decisions amended or superseded. Weeks, not days — and none of it can be done
halfway on a live mesh: the controller, every host and every tool runtime move together.
## The sequencing question, answered
1.**NATS first, directly.** Stalls the migration for the length of the build; the mesh's bus
changes under a half-migrated node; the predecessor's clients still need AMQP, so a second
broker runs anyway, plus a bridge for whatever crosses. Rejected.
2.**Finish on AMQP, NATS after.** Wastes nothing — no module written from here on speaks AMQP —
but leaves the decision unmade while modules are written, and the runtimes idle. Adequate.
3. **Decide NATS now; build it in the lab in parallel; cut the mesh's bus over in one rehearsed
rollout after the core is migrated.** The predecessor's clients never notice: their broker is
the one the mesh adopted, kept as a compatibility module with an end date — the day the last
AMQP client is gone. **Recommended.**
## What this leaves open
- The **decision itself**, as a record: the bus is NATS; the AMQP broker becomes the predecessor's
compatibility broker and retires with the last AMQP client. Written when the operator says so.
- **JetStream's shape for the control plane**: one stream per concern (control, nodes, builds,
events) or one with subjects; retention; what the store window guarantee looks like as ack-wait
and nak-delay. Measured in the lab, not designed on paper.
- **Accounts as configuration**: the controller writes users and permissions into a file the host
declares, reloaded on change — which is the [ADR 0102](../../04-ISSUES/102-an-address-recorded-at-genesis-or-build-does-not-follow-the-nodes-ports/00-report.md)
discipline applied from the start — or the JWT/operator model. The first is simpler and matches
how the mesh already writes everything.
- **The MCP bridge** the operator asked for the same day is written against the sdk's contract, so
it moves with the bus and is not written twice.
- Whether a **leaf node per machine** replaces the hub-and-spoke bus later — out of scope here.
Which S3-compatible implementation replaces it, and what is the migration track for the data and
the provisioning model that sit on top of it?
**Why now, and why not sooner.** Nothing is on fire: nodes that already hold the images keep
running, and issue 113 establishes that the deploy path tolerates an unfetchable-but-present
image by design. The forcing function is not an outage but a one-way door — **no node that does
not already hold the images can ever provision the module again**, so the mesh's ability to stand
a node up from its declarations is already broken for this module, and silently.
**The direction is not a departure from the design; it is the design.** The foundation document
already states the commitment:
> The dependency is on the **protocol**, not the product: AMQP for the bus, S3 for the object
> store, the OCI protocol for the registry. That is what keeps the naming safe rather than a
> commitment that cannot be revisited.
The object store is also **not** a foundation service — ADR 0028 removed it, and it is an
ordinary module required through the module graph by whatever wants one. (The "exception that is
not a swap" in that passage is the relational store, whose provisioning model borrows PostgreSQL's
own meaning of databases, roles and schemas. The object store carries no such coupling: a bucket
is a bucket.) So this effort is an instantiation of an existing principle, not a redesign — which
is the cheapest kind of decision to make and the strongest kind to cite.
## What the replacement has to carry, measured
Taken from the module's manifest, its composition, its tool surface, and a search for its
consumers across the catalogue — not from assumption.
| Requirement | Evidence in the module today |
|---|---|
| S3 API | The protocol every consumer speaks; already the design's stated dependency. |
| ~~OIDC login against the mesh's identity provider~~ | **Struck 2026-09-24. Not a requirement, and it never worked.** Six variables are wired and an entrypoint blocks on the provider, which reads as a live feature. The module's own hook comment records the end state as *"policy claim missing"* — a failing login. See [01](01-candidate-comparison.md). |
| **Per-application access keys, each scoped to a bucket** | The real requirement. A "user" of the store is normally an application; the mesh already mints a credential per provisioned bucket. |
| **One live consumer using it as opaque primary storage** | A file-sync application, since early 2023: objects named by internal id, metadata in its own database. Highest-risk consumer — a live copy drifts, and its bucket name must be preserved. |
| Erasure-coded multi-node topology | Four server nodes with two data directories each, behind a load balancer. |
| A single-node form | Declared as a flavour, for development and small nodes. |
| Buckets as a typed provision | The module declares a provision type of `bucket` on a named network; the mesh mints the credential and the provider creates it (ADRs 0048, 0084). |
| A tool surface | Bucket create/list/delete, object list/info/delete, presigned URL, and provisioning. |
| A console | Published on its own subdomain through the reverse proxy, with an unlimited request-body middleware for uploads. |
**Consumers, counted:** one application module, one capture module that takes a private bucket per
node, one workflow module's tools, and the delivery/rescue internals of the shared library. The
surface is small — the cost is concentrated in the provisioning handler, the tool handlers and the
OIDC story, not spread across the catalogue.
## Candidates
**Four candidates, not three.** The comparison was briefly narrowed to SeaweedFS on the strength of
console single sign-on; that axis turned out not to be a requirement, and the incumbent's own
maintained fork had been omitted altogether. Both errors, and why they happened, are recorded in
[01 — the candidates measured](01-candidate-comparison.md), which carries the evidence and the
requirement-by-requirement detail.
In short, and only in short:
- **The maintained fork of the incumbent** — the community edition was archived and its images
deleted, but a fork publishes, tracks CVEs, and preserves the on-disk format, S3 API and
environment surface. Costs **an image reference** where every other option costs a data
migration, two rewrites and a maintenance window. Does not end the dependence on an abandoned
codebase; buys time to choose deliberately.
- **Garage** — its permission model *is* the requirement (per access key, per bucket), its admin
API is the closest match to how the mesh provisions, and the highest-risk consumer is
first-party documented against it. Remaining cost: no object versioning, no server-side
encryption or object locking, partial lifecycle — **unmeasured against the ten buckets, and the
one thing that could still disqualify it**.
- **SeaweedFS** — longest field record and erasure coding. Its console sign-on is a paid feature,
which is now beside the point. What weighs against it is narrower: its S3 surface is a gateway
translating onto its own file-system API, with no first-party support for the opaque consumer.
- **RustFS** — closest in shape to the incumbent, so the least porting. But it reached general
availability eight days before this was written, and carries an open defect in the credential
path. Two earlier claims about it are corrected in 01: it is **not** a drop-in that retains
existing data.
- **Ceph RGW** — remains rejected as disproportionate where the object store is an ordinary
module rather than a platform.
**This is now two decisions, not one:** whether to repoint to the fork or migrate, and — if
migrating — to which. Repointing does not foreclose migrating, which is the argument for taking it
first. On the corrected requirement the migration ranking is Garage, then SeaweedFS, and not yet
RustFS. Two measurements gate any graduation: **which S3 endpoints the consumers actually call**
(Garage cannot be ranked fairly until counted), and **whether the fork can read the incumbent's
on-disk format in place** — tested on a copy, because the migration between them is one-way. Both
are in [01](01-candidate-comparison.md#what-is-still-unmeasured).
## The migration track, in outline
Data movement is the easy half, and deliberately reversible.
1.**Stand the replacement up beside the incumbent**, on its own ports, its own provision type and
**its own data directory**. Nothing removed. The data directory matters: reusing one the
incumbent already holds would put a fresh single-drive store on top of a live erasure set.
2.**Copy bucket by bucket with a neutral tool.**`rclone` rather than the incumbent's own client
— the client has been withdrawn upstream too, so building the migration on it would inherit
the same dependency this effort exists to remove.
3.**Verify per bucket** — object counts and checksums, not a transfer exit code.
4.**Repoint consumers through the connection the module already publishes.** Consumers read an
API URL from the module's declared connections rather than addressing the store directly, so
the cutover surface is that value plus the provisioning and tool handlers.
5.**Freeze writes, final incremental sync, flip**, and keep the incumbent read-only as the
rollback until confidence is earned. For the opaque consumer this is **not optional and not
instant**: it stores objects by internal id with metadata in its own database, so a copy taken
while it runs will drift. It needs a maintenance window for the final sync, and the window is
proportional to 82,496 objects rather than to 230 GiB.
6.**Retire**, and only then remove the module.
The genuinely new work is not the copy. It is the **provisioning handler** and the **tool
handlers**, both written against the incumbent's admin API. *The OIDC wiring was previously listed
here and is struck: it is not a requirement and it never worked.*
## Open questions
- ~~How much of the OIDC requirement survives, and in which build?~~ **Answered, and it was the
wrong question.** The console requirement does not exist, and the login it referred to never
worked. What replaced it: which S3 endpoints consumers actually call, and whether the fork reads
the incumbent's format in place.
- Does the mesh's bucket provision translate to the candidate's identity model without weakening
what ADR 0049 says about a consumer's identity fitting the tightest backend?
- Should this effort also answer issue 113's general question — mirroring third-party images into
the mesh's own registry — or is that a separate decision? Replacing one withdrawn product with
another unmirrored upstream leaves the same one-way door in place, just further from the hinge.
- Is the four-node erasure-coded topology still warranted, or was it inherited? Worth re-asking
while the product is being chosen, rather than reproducing a shape by default.
**What.** Which rotation mechanisms the mesh's providers can actually support, measured against
their code rather than assumed. Every provider in the catalogue was read, found by listing every definition that provides something: how it names what it
makes for a consumer, what its remove destroys, whether it re-applies a password, whether its
backend can hold two secrets for one login or two logins on one resource, and how its own
administrative credential is set. The consumer side was read too: when a module reads a secret, and
what makes it read a new one.
**Why.** [ADR 0113](../../02-DECISIONS/0113-the-vault-makes-every-secret.md), as first drafted,
chose *overlap*: add a second login beside the first, move every reader, then remove the old one,
"through the adapter's existing create and remove", with "no consumer changes". A review showed that
claim false. In most providers the consumer's data is named after its login, and remove drops the data
with the login. Overlap as written would have deleted every consumer's database on its first
rotation. The mechanism has to be chosen on what the providers do.
**What it touches.** Rotation in 0113 and [to-be 27](../../03-DESIGN/01-to-be/27-a-module-requires-the-mesh-resolves.md),
which [ADR 0114](../../02-DECISIONS/0114-a-shared-credential-rotates-over-two-credentials.md) decided on
these findings. The identity budget in
[ADR 0049](../../02-DECISIONS/0049-a-consumers-identity-fits-the-tightest-backend.md), if a consumer
gets two logins. The rotation already implemented, which [to-be 13](../../03-DESIGN/01-to-be/13-credentials-and-their-rotation.md)
describes.
**Documents.**
- [01 — The providers](01-the-providers.md): the survey, one row per provider, and what it shows.
- [02 — The readers](02-the-readers.md): how a secret reaches a running process, and what already
recreates it.
- [03 — The options](03-the-options.md): each rotation mechanism against those facts, and a
recommendation.
**Finding, in one paragraph.** All nine credential providers already re-apply a consumer's password
in place on every create, and the controller's `rotate` command relies on that. It is a working
rotation with a stated window. Eight of the nine name the consumer's resource after its login, and five
destroy the consumer's data when they remove the login. The harness, keyed by login, would do the same
on any change of login. Only one backend holds two passwords on one login, and two more hold several
tokens. Eight backends can grant two logins the same rights over one resource; the ninth can give one
login a second token. So every provider can hold **two credentials** over one resource, but only after
each adapter separates *the consumer's resource* from *the credential that reaches it*. In postgres
that also means the resource belongs to a role no login owns. Administrative credentials are a
different case. They have one party and a fixed name, and five backends take them only at first
initialisation, so changing one needs the old and the new value at once.
| provider | the consumer's resource is named | remove destroys | create re-applies the password | two secrets on one login | two logins on one resource |
|---|---|---|---|---|---|
| postgres | a database named `login`, owned by the role `login` | the database and the role | yes, `ALTER ROLE … PASSWORD` when the role exists | no: a role has one password | yes, but only through a role that cannot log in owning the database, with each login working as it. Otherwise whatever one login creates is its own, and dropping that login means handing its objects over first. Not done today |
| mssql | a database named `login`, with the login mapped into it | the database and the login | yes, `ALTER LOGIN … WITH PASSWORD` | no: a login has one password | yes, two logins mapped to users in `db_owner`. A user owning a schema cannot be dropped, and a login with an open session cannot. Not done today |
| mongodb | a database named `login`, with a user holding `dbOwner` | the database and the user | yes, `updateUser` with the new password | no: a user has one credential | yes, two users with `dbOwner` on one database. Not done today |
| redis | the key prefix `login:` on an ACL user named `login` | the user, **not** its keys | yes: `ACL SETUSER … reset … >password` replaces all of them | **yes**: an ACL user holds several passwords, added with `>` and removed with `<`. Today's `reset` discards all but the new one | yes, two users on one key prefix, once the prefix is not the login |
| minio | a bucket derived from `login`, and a service account whose access key is `login` | the access key; the bucket **only if empty**. A bucket holding objects is left, and the failure logged | yes, by removing the access key and adding it again, which leaves a moment with no key | no, but an access key *is* the login: a second key is a second login | yes, two service accounts with one bucket policy. The access key is capped at 20 characters |
| lavinmq | a virtual host named `login`, and a user named `login` with permissions on it | the virtual host, with any queued messages, and the user | yes, the user is written again with the password | no: a user has one password | yes, permissions for two users on one virtual host |
| mosquitto | a client named `login`, with a role named for it on the topic prefix `login/#` | the client and its role | yes, the password is set when the client exists | no: a client has one password | yes, two clients holding one role, once the prefix is not the login. The MQTT client identifier is chosen by the consumer, not tied to the login; a duplicate one takes the older session over |
| mailu | a mailbox `login@domain`, unless the consumer contributes its own account name | the mailbox with its mail, for a login-named one; a contributed name is left for an operator | yes, the password is set when the user exists | no for the password; a user can hold several authentication tokens, per the backend's documentation | **no**: a mail user *is* its mailbox |
| gitea (npm) | a user named `login` on a team of an organisation that owns every package | the user; **packages survive**, because the organisation owns them | yes, the user's password is set on every run | no for the password; a user can hold several access tokens | yes, trivially: a second member of the same team |
## The other providers
| provider | answers with | credential |
|---|---|---|
| umami | a website, found by its public name | none. The site id it makes has no way back to the consumer today |
| cloudflare-dns | a public name derived from `login` | none handed to the consumer; its own API token is an operator value |
| showcase | a route | none |
| mesh-vault | custody: it records and withdraws sealed values in a ledger | it holds secrets; it makes none today |
verdaccio provides the npm registry too, and has no provisioner.
## Each provider's own administrative credential
| provider | identity | how the backend takes it |
|---|---|---|
| postgres | a fixed superuser | from a file **only at first initialisation** |
| mssql | `sa` | from the environment at first setup. The image documents no file form, and the definition records that as a declared exception |
| mongodb | a fixed `root` | from a file **only at first initialisation**, when the data directory is empty |
| mosquitto | a fixed admin client | seeded into the broker's dynamic-security file **once**; the seeding step skips when the file exists |
| lavinmq | a fixed admin name | per its own bootstrap code, **only on a first boot** with an empty data directory. No resource in the definition runs that bootstrap; what sets it on a running mesh is outside the catalogue |
| redis | the default user | from `requirepass` in a configuration the mesh renders, read when the server starts |
| minio | a fixed root user | from a file, read when the server starts |
**In five of seven, a new administrative value takes effect only through a command run with the old
one.** The credential file is mounted directly into both the server and the provisioner. So replacing
it recreates the provisioner, which then holds only the new value while the backend still expects the
old one, and the provisioner is locked out. That is worse than changing nothing.
Every provider module also has its own bus account, an own secret, read at start.
## What the tables show
1.**Every credential provider already rotates in place.** All nine re-apply the password on the
same login each time `create` runs. The controller's `rotate` command relies on that: it replaces
the credential in the inventory and sends both ends in one push. Its own comments state the window,
between the provider applying and the consumer restarting, in which the consumer cannot
authenticate.
2.**Eight of nine name the consumer's resource after its login.** Only gitea separates them,
because an organisation owns the packages. A second login therefore has no resource of its own to
reach, and cannot share the first one's without the adapter granting it.
3.**Five of nine destroy the consumer's data when they remove the login**: postgres, mssql and
mongodb drop the database, lavinmq drops the virtual host with its queued messages, and mailu
deletes the mailbox with its mail. minio drops only an empty bucket, and redis leaves the keys. In
those five, *retire a login* and *delete the consumer's data* are one call. With the harness keyed
by login, a changed login triggers it too.
4.**One backend holds two passwords on one login** (redis). Two hold several tokens beside one
password (gitea and mailu). A rotation built on two secrets per login would work for three
providers out of nine.
5.**Eight of nine can give two logins the same rights over one resource.** Group roles in postgres,
database roles in mssql and mongodb, permissions in lavinmq, a shared role in mosquitto, a shared
policy in minio, a shared key prefix in redis, a shared team in gitea. mailu cannot, because its
user is its mailbox, but it can give one user a second token. So every provider can hold **two
credentials** over one resource, though not every one as two logins. No adapter does either today.
6.**Ownership is a trap in two backends.** In postgres whatever a login creates is that login's, so a
second login cannot alter the first one's tables, and the first cannot be dropped while it owns
them. The one-step way out deletes them. In mssql, a login cannot be dropped with a session open,
nor its user while it owns a schema.
7.**The administrative credentials have one party and a fixed name**, and five backends take them
only at first initialisation. The provisioner needs the old and the new value at once to change
them. Today nothing can give it both.
8.**A consumer's identity is already the resource's name.** The login is derived from the
assignment, which is a module on a node, so the current login and "the consumer" are the same
string today. A second login would need a new name. The resource can keep the one it has.
## Seen on the way
The redis configuration names no ACL file, so a consumer's ACL user exists only in memory. A restart
of the redis server erases every consumer's user. The provisioner does not create them again until it
restarts itself, because its in-memory record says they are done. That is not a rotation finding, but
it is a live fault, and it is recorded here so it is not lost.
The operator's wish, written as behaviour: what a person or an agent sees the mesh doing. This is a
target to design toward, not a design. Every part of it is to be decided through a record before it
is built.
## Principles
**1. Every loop compares what should be with what is, never with what it did.** Desired state is the
mesh's: assignments, requirements, seats. Observed state is read from the thing itself: the container,
the backend, the node. A loop that compares against its own memory of what it applied is blind to
anything that changed behind its back. That is issue 120, and it is the pattern this whole effort is
written against.
**2. Healing is the ordinary path run again, never a second path.** Repairing a lost login is
provisioning it. Repairing a dead container is converging the node. Repairing a stale declaration is
delivering it. A repair that needs its own code is a second way of doing something, which is exactly
what the mesh is removing everywhere else.
**3. A repair never destroys.** Healing may recreate, re-provision, re-deliver and restart. It may
never delete a consumer's data, retire a credential someone still uses, or pick a winner between two
contradictory states. Where the only repair is destructive, it is escalated.
**4. Nothing fails silently.** Every condition the mesh cannot repair within its budget becomes
visible. It is named, it says since when, why, and who can resolve it. It is visible until it is
resolved, and resolved by observation, not by someone clicking it away.
**5. What the mesh cannot fix goes to an agent.** Per the mission, an agent may be human or not. A
condition that needs judgement is handed to one, as work, with what the mesh knows. It is not handed
over as a notification that someone may or may not read.
**6. Correctness, not only liveness.** A running process that authenticates with a dead credential,
serves an old version, or routes nowhere is not healthy. What a provision's contract promises is what
is checked: the credential authenticates, the route answers, the version is the declared one.
## The loop, everywhere
Every part of the mesh that owns something runs the same loop:
1.**know** what should be true: from assignments, requirements and seats;
2.**observe** what is true: from the thing itself, on its own cadence;
3.**repair** the difference by running the ordinary path again, within a budget of attempts and
time;
4.**raise** a *condition* when the budget is spent or the only repair is destructive;
5.**clear** the condition when observation shows it resolved.
A **condition** is a durable fact about something the mesh owns, such as a node, an assignment, a
provision, a seat or a rotation: what is wrong, since when, the evidence, what was tried, and who can
resolve it. Conditions are the one thing a person or an agent looks at to know whether the mesh is
right. `status` is the list of open conditions. When it is empty, the mesh is right, not just up.
## On NATS
[ADR 0106](../../02-DECISIONS/0106-the-bus-is-nats.md) moves the bus to NATS, and NATS makes most of
this cheaper, because observation becomes something every component publishes rather than something
a central process polls.
| the wish needs | on NATS |
|---|---|
| every component says it is alive | a heartbeat on a subject per node and assignment; silence past its interval is a condition, and nobody polls |
| every component says what it observed | observations published on subjects (`mesh.observed.<node>.<assignment>`, for instance), consumed by whoever owns the comparison |
| the last known state survives restarts | a JetStream key-value bucket of observed state per owner; the provisioner's "what I applied" and a rotation's step live there, not in memory |
| conditions are durable and watchable | conditions as entries in a key-value bucket, watched by anyone who cares: a surface, an agent, the controller |
| the bus itself is observed | the server's advisories (a consumer exceeding its deliveries, a slow consumer, a client disconnecting) and its monitoring endpoint become observations like any other |
| a repair is retried, not lost | JetStream redelivery with delay, which is the same mechanism [ADR 0083](../../02-DECISIONS/0083-one-push-leaves-the-mesh-consistent.md)'s guarantee moves to |
| work handed to an agent | a condition that needs judgement published as a task on a subject an agent's queue group consumes |
**Who compares.** Each owner compares its own: the host for its node's containers and files, a
provisioner for its backend, the vault for rotations, the controller for delivery and seats. The
controller's observability context does not repair anything. It holds conditions, their history,
and the view across the mesh. It notices what no owner can see about itself: an owner gone silent.
## What stays human
Some repairs need the operator's key, and the mesh says so rather than pretending otherwise:
re-raising the vault or the broker, and recovering a node's identity. These are conditions too, with
the procedure named, and they are the only ones that can never clear themselves.
## Open
- The budgets: how many attempts, over how long, per kind of repair.
- How a condition that needs judgement reaches an agent, and how the agent's action is recorded.
- Where the observability context stores history (to-be 06 left it open; volume argues against the
relational store).
- Which correctness probe each provision's contract offers, and how often it runs.
_Reconciliation note (2026-09-05): supersedes the earlier "repository structure" decision, which the consolidation folded; no standalone record remains to point at, so body references to it now point at the nearest surviving record, [ADR 0015](0015-applications-live-in-their-own-repository.md)._
## Context
The earlier "repository structure" decision (folded in consolidation; see the reconciliation note
above, and [ADR 0015](0015-applications-live-in-their-own-repository.md) as the nearest survivor)
named `mesh-sdk` "contracts shared across tiers:
types, not behaviour." That line is superseded here, because it draws the boundary in the wrong
place. The boundary that matters is not *types versus behaviour* — it is **how often the thing
changes**.
The current SDK is the cautionary tale, and its failure is precise. `hal/sdk` holds all the
code, including a per-module API client for every service (`clients/plex.ts`, `clients/gitea.ts`,
…) and a per-module tool implementation for each (`tools/plex.ts`, …). Every module depends on
the SDK, so **every edit to any of that per-module code rebuilds every module** — the cascade.
The SDK is under constant maintenance precisely because it became the place all the volatile
per-module logic accumulated.
The root cause is worth stating exactly, because the fix follows from it: the pressure was never
to share a client *between* modules. It was to share a client between one module's *own features*
— plex's tools, its health check and its hooks all wanted the same `PlexClient` — and the only
place to share code across a module's features was the global SDK. So **intra-module sharing
leaked out as inter-module coupling.**
## Decision
The SDK holds the **stable spine** that modules build against, and earns its place by rarely
changing. The test for membership is change-frequency, not kind.
### What it holds
- The **tool-serving harness** — the worker and registration mechanism, and the tool-definition
type. *How* a tool is declared and served is settled; it does not change when an individual
tool does.
- The **messaging and event framework** — the broker client, the event consumer, the envelope.
- The **contracts** — the manifest, declaration, provision and link shapes.
- **Core primitives** — sealing and crypto, semver, the shared resolution helpers.
These change rarely and deliberately. When one of them does change, a rebuild of everything is
the *correct* outcome, because the contract every module shares has genuinely changed.
### What it must not hold — the more important half
- **A module's API client.** A Plex client, a Gitea client, a MinIO client belong in their
module. They change when that service's API or the module's use of it changes, which is often,
and which has nothing to do with any other module.
- **A module's tool implementations.** Same reason, same place: in the module.
- **Anything volatile** — anything that changes when one service's features change.
The rule, stated so it can be applied without re-deriving it:
> If editing a thing recompiles unrelated modules **and** it changes often, it does not belong
> in the SDK.
Both conditions are load-bearing. A rare change that cascades is fine — that is a contract, and
the cascade is correct. A frequent change that stays local is fine — that is a module minding its
own business. Only **frequent *and* cascading** is the disease, and per-module clients and tools
are its carriers.
### Where per-module shared code lives instead
Code shared among a module's *own* features lives **in the module**. The default is the plainest
thing that works: an ordinary shared file the features import — `plex/client.ts`, imported by
`plex/tools/`. Within one module, features are files importing sibling files; no package
boundary, no ceremony.
A **module-local SDK** (a sub-package with its own version) is warranted only for the few modules
whose shared surface is large enough to version on its own. It is the exception, not the shape.
Either form gives the property the global SDK could not: editing a module's shared code rebuilds
**that module and nothing else**.
## Consequences
- The cascade becomes **structurally impossible for module logic**. There is no longer an edge
from one module's internals to another, so the only thing that can rebuild everything is a real
change to a shared contract in the SDK — which is rare, and when it happens, is right.
- The SDK is small and stable **by construction**, not by discipline. Its size is no longer a
thing anyone has to police.
- **Converting a module from the current system is partly a de-coupling, not just a move.** Its
client and its tools are pulled *out* of the shared SDK and *into* the module. A conversion
that copied `clients/plex.ts` into the SDK's replacement would rebuild the exact mistake.
- The host still does not import the SDK. It depends on nothing
([ADR 0005](0005-the-node-host.md)) and **mirrors** the contracts rather than
importing them, exactly as its apply-shapes table already does deliberately. The SDK is shared
by the tiers that *can* share code; the host is not one of them.
## References
- The earlier "repository structure" decision — named the repositories; its `mesh-sdk`
description ("types, not behaviour") is superseded by this record (folded in consolidation;
_Reconciliation note (2026-09-05): supersedes the earlier "grouped by domain" decision, which the consolidation folded into how-we-build.md; no standalone record remains to point at._
## Context
[ADR 0009](0009-modules-and-the-graph.md) settled that everything is a module, but never said what a
module *is* beyond "a directory the mesh processes." That gap let the catalogue's breadth read as a
smell: a module can carry a container, a built image, tools, a provisioner, migrations, health,
config, seat claims, requires and provides — so much that the unit seemed ill-defined.
The earlier "grouped by domain" decision (folded in consolidation; see the reconciliation note
above) tried to organise modules by domain, which is the wrong axis. This record states what a module is, drawn from the cases that
stress-tested it: the shell, i3-vs-sway, umami, and "database."
## Decision
**A module is one self-contained piece of software the mesh installs and manages** — everything
needed to make that one thing real and integrable: what runs, the seats it claims, what it provides
to other modules, what it requires from them, and what operates it.
The **software is the module's identity.** Capabilities, seats and provisioned resources are the
**relationships *between* modules**, not what a module is — and that is what binds a module into one
thing. umami is bound by *being umami*: its container runs umami, its provisioner creates umami sites,
its tools query umami, its `requires` gets umami a database. Every feature serves the one software.
### The three relationships
1.**Shared seat** — several modules fulfil a capability and coexist; one may be default. bash, zsh
and fish all join `shell`.
2.**Exclusive seat** — modules contend for a single slot; one holds it. i3 (needs x11) and sway
(needs wayland) contend for `display-session`.
3.**Provide / require** — a provider ships the **provisioner** that creates instances of the
resource it offers and returns sealed credentials; a consumer requires it and the mesh wires the
credential in. Symmetric: umami requires a database *and* provides analytics.
### Interfaces are mesh-owned; providers adapt to them
The mesh **defines the interface** for a capability — the provider-neutral contract of what a
consumer receives and how it integrates. Both sides conform: a provider's provisioner **adapts** its
software's real API to the mesh contract; a consumer depends on the **interface**, never on a
provider. Swap one provider for another and the consumer does not change.
### The naming rule — draw the interface at the consumer's real coupling
Name a `provides`/`requires` at the **widest boundary across which the consumer genuinely does not
care which implementation serves it**:
- Where the consumer's coupling is thin — an analytics embed snippet and dashboard, opaque to it —
the mesh defines a neutral interface (`analytics`) and providers (umami, amumi) adapt. Swappable
across vendors.
- Where the consumer **speaks a protocol** — a database's wire protocol and query dialect — the
interface *is* the protocol: `postgres-database`, `mssql-database`, `mongodb-database`. Swappable
only among protocol-compatible implementations, **never across**, because the application cannot
cross it either. "database" is not a capability; the protocol is.
- **Never false genericity.** A name must not promise a swap the contract cannot deliver
# 68. The lab takes requests, one at a time, and runs each from its own copy
## Context
**The lab is exclusive hardware, and today a person holds it.** Raising a scenario takes over
addresses and names on the workstation for as long as it stands, and only one scenario can stand
at a time. So a run is not merely slow — it occupies the machine and the person who started it,
who then waits rather than works.
**Running it in the background against the working copy is worse than waiting.** The obvious fix
is to start a run and carry on editing. But a run reads the working copy as it goes: binaries are
rebuilt from it, manifests are read out of it, and the bed's own code is loaded from it. Edit
while it runs and the result describes a state that never existed — a mixture of what was there
when each file happened to be read. A green result obtained that way is not evidence, and a red
one costs a day to disbelieve.
**Nothing today records what was asked for.** A run is a command line in somebody's terminal. What
commit it exercised, what it was trying to find out, and what it answered all live in scrollback,
which is why the same question gets re-run rather than looked up.
**Most of the parts already exist.** The lab writes a receipt of its last run. The mesh already
carries messages between nodes and can notify a person. The machine already runs work on a
schedule. What is missing is the thing in the middle.
## Considered Options
**1. Leave it as it is — a person drives the lab and waits.** Rejected. It is the loop
[ADR 0010](0010-delivery.md) removed everywhere else, kept here by habit rather than by argument,
and the cost compounds: because a run is expensive to start and blocks the person, fewer are run,
so faults are found later and in larger batches.
**2. Run in the background against the working copy.** Rejected on the reasoning above. The
failure is silent, which is the kind this repository exists to refuse.
**3. Put the lab behind the ordinary build pipeline.** Rejected for now. The pipeline builds
artifacts and does not own a machine that can raise virtual machines; giving it one makes the
pipeline's slowest job the lab's, and couples every push to hardware only one machine has. This
may become right later; it is not the smallest thing that works.
**4. A queue in front of the lab, and an isolated copy behind it.** Chosen.
## Decision
**The lab accepts requests rather than commands.** A request is recorded, queued, and answered.
The person who made it is told when it is answered and does not wait.
**A request names a bed and a commit, and nothing else.** This is the load-bearing restriction. A
request may say *run this bed, at this version of these repositories*. It may not say what to
install, on which machine, or with which settings — because a request that could say those things
would be a second way of installing a mesh, and the whole reason the installer exists is that the
lab already was one ([ADR 0067](0067-genesis-is-a-pivot.md)). The bed decides what is installed;
the request only decides which bed and which version.
**Requests are released one at a time.** The hardware admits one standing scenario, so the queue
enforces what the hardware already requires, rather than leaving it to whoever remembers.
**Every run happens in a copy the lab owns.** The lab checks the requested commit out into its own
path and builds and runs from there. A working copy is never read by a run. This is what makes the
queue safe to use while work continues, and without it the rest of this record is not worth
having.
**The lab is reached through tools, not only a command line.** A command line is available only
to whoever is sitting at the machine, which is the constraint this record exists to remove. The
lab answers three questions to anything that can reach the mesh — *what is standing now*, *what is
queued or running*, and *what did this request answer* — and accepts a request and a cancellation.
An agent can therefore start a run, stop attending to it, and come back; and somebody who did not
start a run can still see it, which is the difference between a shared lab and a private one.
**The restriction holds at every door.** A tool submits a bed and a commit, exactly as a command
line does. A tool that could name a module, a node or a setting would reintroduce the second
installer through a different entrance, and the entrance is not what made it dangerous.
**Every run leaves a record that outlives the terminal**: what was asked, which commit, when it
ran, what it answered, and where its output went. A question already answered is looked up rather
than re-run.
## Consequences
Work continues while the lab runs, which is the point. A second session may edit freely, because
nothing it edits is what the lab is reading.
A request is reproducible by construction: it names a commit, so the same request can be asked
again and compared. Today two runs of "the same thing" are only as alike as the tree happened to be.
The lab gains a second copy of every repository it exercises, costing disk and needing to be kept
from drifting into a place people edit by hand.
Anything that can reach the mesh can now see what the lab is doing, including an agent working on
something else. That is the intended gain and also the obvious hazard: a thing that is easy to ask
is easy to ask too often, and the hardware still admits one scenario at a time.
The queue becomes a thing that can fail — stuck, backed up, or lost — and a queue nobody watches
is worse than no queue, because it absorbs requests silently.
## How this is checked
| Rule | Checked by |
|---|---|
| A run never reads a working copy | The runner is given a path it owns and no other; a run started while a working copy is deliberately dirtied produces a result matching the commit, not the edits. |
| One scenario stands at a time | A second request submitted while one runs is observed to wait, not to raise. |
| A request cannot say what to install | The request format admits a bed and a commit only. A request naming a module, a node or a setting is refused, and the refusal is exercised. |
| A request is answered | Every queued request reaches a terminal state with a record. A request that vanishes is a failure of the queue, not a quiet nothing. |
| The lab can be asked from elsewhere | What is standing is asked from a session that did not raise it, and the answer matches the machine. A lab that only answers its own caller has not left the terminal. |
# 69. A module is a repository and a path within it
## Context
**The builder clones one repository and reads `module.json` at its root.** The to-be design says
so in as many words — *"one file at the root"* — and the code implements it: clone, read the root
manifest, build what it declares.
**Nothing that exists is shaped that way.** The catalogue holds sixty-seven modules, each in its
own directory, and has no manifest at its root. None of the five code repositories has one either.
So today the builder cannot be asked to build any module that exists: pointed at the catalogue it
finds no manifest, and pointed at a module's source it finds no manifest.
**The system being replaced already works the other way**, and has for years: a monorepo with one
directory per piece of software, and the coordinator builds a module from a repository and a path
inside it. The root-only assumption is not a simplification of that — it is a different model that
was never reconciled with it.
**And it splits what a build needs into two places.** The control plane's manifest sits in the
catalogue; the source it describes sits in the control plane's own repository. A build must read
one tree, so under the root-only model neither location can be built from.
## Considered Options
**1. One repository per module.** Rejected. Sixty-seven repositories for sixty-seven modules, most
of which are a single manifest naming a public image, and every one needing its own creation,
permissions and lifecycle. It also contradicts [ADR 0015](0015-applications-live-in-their-own-repository.md),
which put *applications* in their own repositories precisely because modules do not need one.
**2. Keep manifests in the catalogue and source elsewhere, and have a build fetch both.** Rejected.
A build would clone two trees whose versions can disagree, so "what commit is this module?" stops
having one answer — and that question is the whole basis of knowing when to rebuild.
**3. A module is a repository and a path within it.** Chosen. It is what the current system does,
what the catalogue already looks like, and it keeps a module's description beside the thing it
describes.
## Decision
**A module is named by a repository and a path within it.** The path holds `module.json`, and
everything that manifest declares is produced from that path. A module whose path is the root is
the ordinary case of this, not a separate one.
**A module's manifest lives beside its source.** Where a module has code, its directory holds both,
so one commit answers "what is this module, and what is it made of". Where a module has no source —
a manifest naming a public image — the directory holds only the manifest, and there is nothing to
build.
**This moves the core modules.** The control plane and the builder are built from the control
plane's repository, so their manifests belong in that repository at their own paths, not in the
catalogue. The catalogue keeps the modules whose source it holds, and the modules that are only a
manifest.
**One commit, one module version.** Because a module is one path in one repository, the commit that
built it identifies it exactly, and "the source has moved ahead of what the mesh holds" stays a
question with a yes or no answer.
## Consequences
The builder gains a path alongside the repository and the ref. A build is `repository, path, ref`,
and the manifest it returns is the module the mesh records.
The catalogue stops being the place every manifest lives, and becomes the place manifests live
*when their module has no other home*. That is a smaller claim than it sounds: most of the
sixty-seven stay exactly where they are.
Two repositories change shape — the control plane's gains manifests for the modules built from it.
Nothing else moves.
A repository can hold modules that are built and modules that are not, and no rule distinguishes
them beyond whether their manifest declares anything to build.
## How this is checked
| Rule | Checked by |
|---|---|
| A module is buildable from its repository and path | The builder is asked for a module by repository and path, and returns a manifest whose artifacts are pinned to digests the mesh's registry assigned. |
| A manifest sits beside what it describes | A module declaring something to build, whose path holds no source to build it from, is refused at build time rather than producing an empty result. |
| One commit identifies one module | Two builds of the same repository, path and commit produce the same digests. |
| The core modules are built like any other | The control plane is rebuilt from its own repository and path, and the running mesh is upgraded to it — the same path an ordinary module takes. |
# 70. The catalogue owns the module graph, and genesis builds rather than carries
## Context
**The module graph has no owner.** What modules exist, what each requires and provides, what each
claims, what each is made of — all of it lives inside the control plane because that is where it
was first written, not because anything decided it belonged there.
**The control plane's own test says it does not belong there.**
[`06-the-control-plane`](../03-DESIGN/01-to-be/06-the-controller.md) defines the tier as
*everything that needs to know about more than one node*, and states the corollary plainly:
anything a single machine could answer alone is not the control plane's. What a module is, and
what it needs, requires no knowledge of any node whatsoever.
**And nothing can query it.** The graph is the thing that answers *what must be rebuilt when this
changes*, *what would break if this were removed*, and *what can be installed here* — and today it
is reachable only as control-plane internals. [ADR 0009](0009-modules-and-the-graph.md) already
decided that a build edge is derived, that an artifact is stale when anything it was built against
moved, and that the rebuild set is therefore computable. None of that has anywhere to live.
**Genesis currently carries an image, and cannot produce a builder at all.**
[ADR 0067](0067-genesis-is-a-pivot.md) has the installer carry the control plane's image. That
works, but it leaves the builder with no route onto a fresh mesh — it cannot be fetched from the
public internet, because the mesh builds it, and the installer carries one image only. So a raised
mesh cannot build anything, including the modules it is made of.
## Decision
**The catalogue is a core module, beside the control plane and the builder, and it owns the module
graph.** Modules, what they require and provide, what they claim, their dependencies and their
build edges, assignments and configuration — the catalogue holds them and serves tools over them:
install a module, query what is available and what it needs, query the graph, query and update
settings.
**It is one per mesh**, expressed the way the control plane already expresses it — a claim scoped
to the mesh, not a new mechanism.
**The control plane consumes it.** Resolving what a module requires into an actual binding, and
composing what a machine should be, both need to know what modules are. So the dependency runs
from the control plane to the catalogue, which is the opposite of what the tiers suggest and is
therefore written down here rather than left to be inferred.
**Genesis builds the core modules rather than carrying them.** The installer ships an *init
builder* — the one thing carried — which is started, clones the source, and builds the control
plane, the catalogue and the builder. The installer then raises a temporary control plane, which
installs the catalogue, registers the permanent control plane and the builder, and assigns all
three to the first machine. The temporary control plane stops; the installer verifies that the
permanent one answers and can query the graph. The builder then sees a catalogue in its initial
state and builds the core modules into it.
**So exactly one thing is carried, and it is a builder rather than a result.** That is the
difference from [ADR 0067](0067-genesis-is-a-pivot.md), which carried the control plane's image:
carrying a builder produces every core module on the machine, including the builder itself, so
there is no component left without a route.
## Consequences
The catalogue joins the small set of things that cannot arrive through the ordinary path, because
it cannot be installed by something that needs it in order to install anything. It arrives the same
way everything else does under this record — built by the init builder before the mesh can install
anything — so the set is answered by one mechanism rather than three special cases.
The rebuild fan-out gains a home. *What was this built against* and *what must rebuild now* are
questions about the graph, and the graph now has an owner to hold the edges and answer them.
The catalogue holds state, so it owns a store in the substrate's database, the same way the control
plane's contexts do. That is the mesh's own store and not the `postgres` module, which is a
provider other modules consume.
A mesh without a catalogue cannot resolve anything, where previously it merely lacked an interface.
That is the cost of ownership over surfacing, and it is deliberate: one owner beats two copies.
## Open, and to be settled before this is built
**Where the init builder clones from.** [ADR 0067](0067-genesis-is-a-pivot.md) rejected building
from source at genesis partly because the source lives in a forge that runs on the mesh, so a
total rebuild would need the mesh it is rebuilding. Carrying a builder answers the toolchain half
of that objection and not this half. Genesis must therefore name a source that exists before the
mesh does.
**What the init builder publishes into.** Building produces artifacts that must be pinned by a
digest a registry assigned, and today the registry is installed after the machine has joined. If
the core modules are built first, the registry has to exist first, so the substrate's order needs
restating rather than assumed.
**Where the line falls between the catalogue and the control plane.** Claims and assignments need
to know about every node, which by the control plane's own test is its work. Whether the catalogue
holds them and asks, or the control plane holds them and the catalogue surfaces them, is not
settled here.
## How this is checked
| Rule | Checked by |
|---|---|
| The catalogue owns the graph | The control plane answers *what does this module require* by asking the catalogue, and a mesh whose catalogue is stopped cannot resolve — observed, not assumed. |
| One per mesh | A second catalogue assigned anywhere in the mesh is refused by the claim, and the refusal is exercised. |
| Genesis builds rather than carries | The installer carries exactly one artifact, and after installing, every core module is pinned to a digest the mesh's own registry assigned. |
| The builder has a route | A mesh raised by the installer, with no hand-placed image, can build a module. |
# 71. Genesis clones from a mesh, and checks what it got
## Context
**[ADR 0070](0070-the-catalogue-owns-the-module-graph.md) has the init builder clone the source,
and does not say from where.** [ADR 0067](0067-genesis-is-a-pivot.md) had already rejected building
at genesis partly for that reason: the forge holding the source runs *on* the mesh, so a total
rebuild would need the mesh it is rebuilding.
That objection is real but narrower than it reads. It only binds when the mesh being raised and the
mesh holding the source are the same one, which is true exactly once.
## Decision
**Genesis clones from a mesh's forge, reached by name.** Any mesh that holds the source can serve
it. The first mesh is not structurally special — it is simply the only one that existed when there
was nothing else to clone from.
**If the mesh serving the source is lost, the name moves to another mesh that holds a copy.**
Recovery is a name pointing somewhere else, not a backup being restored. This is what makes the
source's survival a property of there being more than one mesh, rather than a property of somebody
having remembered to take a copy. A mesh that has installed from that name holds the source
afterwards, so every installation adds a place the name could point.
**Genesis names a commit and checks what it got.** It does not clone whatever a branch happens to
point at. The forge a mesh installs from is the trust anchor for everything that mesh will ever
run, and a branch is a moving target that somebody else controls.
*This is not hypothetical.* On 2026-09-11 the forge that would serve this role was running a
cryptominer, and its git operations were being tampered with in flight — output injected into the
protocol stream by a hook that fired on every fetch. Nothing was altered: the repositories were
verified against local copies and found byte-identical. But a mesh installing from that name during
those hours had no way to establish that for itself, and would have had none.
## Consequences
The init builder needs a name it can resolve and a commit it can verify, and nothing else. It does
not need to know which mesh answers.
Whoever operates the mesh that name points at carries a responsibility to everyone installing from
it, and should know that. It is not merely a convenience host.
A mesh that cannot reach any forge cannot be raised. That is a real limit and it is accepted: the
alternative is carrying the whole source in the installer, which makes the installer a release
artifact that goes stale rather than a program that fetches what it was told to.
## Open — what relationship a mesh keeps afterwards
**Not decided, and named here so it is not decided by accident** by whoever writes the init
builder. Two shapes, and they are meaningfully different:
**A snapshot, and then independence.** A mesh installs once, mirrors the source into its own forge,
and has no upstream afterwards. It is fully self-hosted, in the sense that nothing it needs lives
anywhere else. Updates are then something an operator does deliberately, by pulling changes in —
tooling for which is possible and is not a priority.
**A continuing upstream for core modules**, the way a distribution serves packages and a separate
collection serves everything else. A mesh keeps looking at the origin for the modules that make a
mesh a mesh, and holds its own for the rest.
The first is more obviously aligned with the rest of this design, which is arranged so nothing a
mesh needs depends on somebody else continuing to host it. The second is more convenient and makes
a security problem in one forge everybody's problem. Neither is chosen here.
## How this is checked
| Rule | Checked by |
|---|---|
| Genesis needs only a name and a commit | A mesh is raised with the name pointed at a different mesh than the last time, and the result is identical. |
| What was cloned is what was asked for | Genesis is pointed at a commit and refuses a forge serving different content under it, rather than building what it received. |
| Losing the serving mesh is survivable | The name is repointed at a mesh that installed from it earlier, and a raise succeeds. |
# 72. Two graphs, and a build chain that orders itself
## Context
[ADR 0070](0070-the-catalogue-owns-the-module-graph.md) gave the catalogue the module graph and
said the control plane consumes it — that resolving a requirement and composing what a machine
should be *"both need to know what modules are, so the dependency runs from the control plane to
the catalogue."*
**That paragraph is wrong, and this record corrects it.** It was written before the two graphs had
been told apart, and it creates a dependency that does not need to exist: a catalogue that is down
would leave the control plane unable to compose the declaration that would repair it.
## Decision
**There are two graphs, with different owners, and they meet only when something is installed.**
| Graph | Owner | What it links | Answers |
|---|---|---|---|
| the module graph | the **catalogue** | module-versions to each other | what was this built against · what must rebuild now · what does this need |
| the runtime graph | the **control plane** | module-versions to nodes | what runs where · who consumes this provision · what breaks if this machine goes |
The catalogue does not know nodes exist. The control plane holds module-versions and nodes, along
with capabilities and claims, because deciding whether a machine qualifies needs every node.
**So the control plane never asks the catalogue anything.** It holds what it needs to compose a
declaration. A catalogue that is down stops new installs and stops the rebuild fan-out, and does not
touch anything already running or the ability to repair it.
**Everything between them travels as events over the broker**, like all module-to-module
communication. The builder finishes and announces that a module was built. The catalogue registers
it, places it among what it depends on, and announces that a module was upgraded. The control plane
reacts to *that*, not to build output — a semantic fact rather than an artifact.
**Nothing is lost if a receiver is down.** A consuming module's queue is durable with a dead-letter
exchange, declared by the mesh rather than by the module, so an event waits for a consumer that is
not there.
**Build order is not computed. It emerges from the chain.** The builder never consults the graph: it
builds what it is asked for, one at a time. The catalogue asks for the next build after the previous
registration, so *"do not start this until that is registered"* holds by construction rather than by
a schedule somebody maintains.
**The catalogue's rule is a condition, not an order:** ask for a module to be rebuilt once everything
it was built against is current. That handles a chain and a diamond with one rule, where an
order-based approach needs to know the shape in advance.
**Two refusals belong to the catalogue.** A cycle, because the chain would never settle. And a
rebuild whose artifacts are identical to what it replaced, which is not an upgrade and must not be
announced as one — or a single change ripples outward forever through modules that did not change.
**Genesis does none of this.** The init builder has a fixed, short list — control plane, catalogue,
builder — in a written order, because there is no catalogue yet to ask.
## Consequences
The builder stays simple, and independent of the catalogue. That is what makes genesis possible at
all: the thing that builds the catalogue cannot require the catalogue.
The two sides can be briefly out of step — the catalogue may hold a module-version a moment before
the control plane knows of it. Assigning in that instant fails, and should say why rather than
report that no such module exists.
A module's declared events stop being documentation and become its permissions: an account is
scoped from what a module emits and consumes, so a builder that announces what it built is granted
what it needs by the ordinary mechanism rather than by a special case.
## Open
**Whether an upgrade is applied or merely noticed.** Today the system this replaces deploys
automatically, and that is a defensible default for core modules on a mesh its operator runs. But
the design as it stands does the opposite: it records that the source moved ahead, makes it visible,
and waits to be told. This must become a setting with a chosen default rather than inherited
behaviour — and the choice matters most on the day a bad commit reaches something that carries mail.
**Whether a module assigned to several machines upgrades on all of them at once.** Doing so makes
one bad commit simultaneous everywhere. Doing one machine and pausing turns it into one casualty.
recorded for the shared runtime, and fixing it there while shipping it here on every new mesh would
be a strange place to stop.
## Decision
**The installer carries a builder, and nothing else.** One artifact, not a growing set. It clones
the source at a named commit, checks what it got ([ADR 0071](0071-where-genesis-gets-its-source.md)),
and produces the control plane from the same repository and path that any later rebuild of it would
use. What raises the mesh is therefore the same thing that will maintain it, and there is no second
mechanism kept in step with the first.
**The registry does not move, and the argument for moving it does not survive being made.**
It was put this way: a produced image has to be put somewhere before anything can fetch it, so the
registry must now precede the control plane, and
[ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md)'s answer to that question has to flip.
It does not, because the premise is false. **The thing that builds the image and the machine that
runs it are the same machine.** A built image is already in that machine's container runtime, and
the temporary control plane names it exactly as it names a carried one — by the digest of its own
configuration, a local identity that requires nothing to have served it. Building changes where the
bytes came from. It does not change where they are.
| | is it substrate? | must it precede the control plane? |
|---|---|---|
| the store | yes | yes — there is nowhere else to put the control plane's state |
| the broker | yes | yes — the control plane reaches a machine only over it |
| the image registry | yes — it cannot grant itself a repository | **still no** — the first machine neither fetches the control plane nor needs to, whether the image was carried in or made here |
So the registry stays where [ADR 0033](0033-the-substrate-is-a-store-and-a-broker.md) put it:
substrate by role, ordinary by delivery, installed by the temporary control plane as its first act.
The bundle carries two services and a control plane, as it did. **What publishes into the registry
is unchanged too** — the existing step that pushes the control plane's image into it, which is the
moment that image first receives a digest assigned by something other than itself. It now pushes
something this mesh built rather than something it was handed.
## Consequences
**Genesis gains one step and changes no others.** A build happens before the image is loaded. The
pivot described in [ADR 0067](0067-genesis-is-a-pivot.md) survives exactly as written, because the
step it pivots on never cared where the image came from.
**A fresh mesh can produce from the moment it exists.** The builder is present before the control
plane is, so the core modules, the catalogue and the builder's own module can be built in the
ordinary way rather than waiting for somebody to carry them in. The paragraphs in
[`17-raising-a-mesh`](../03-DESIGN/01-to-be/17-raising-a-mesh.md) that describe this were describing
something that could not start; they can start now.
**Genesis needs more of the outside world.** Carrying an image needed nothing but the installer.
Building one needs the source, and whatever the build itself reaches for. This is a real cost and
is not waved away: it makes genesis fail in more ways, all of them at a step that says what it was
doing. It is accepted because the alternative is a mesh that cannot rebuild its own control plane,
which fails in exactly one way, silently, later, and for ever.
**A pre-built bundle remains possible and is not this.** Nothing here forbids delivering artifacts
rather than building them; it fixes where they may come from. A bundle of pre-built core modules is
an **export of a mesh that built them**, carrying what the catalogue knows about each alongside the
artifact itself — so that loading one leaves the graph in the state building would have left it. A
bundle that carries images without that is the thing this decision rejects, whoever ships it.
## What this does not decide
**Whether the builder's own module is carried or built.** It builds everything else; what installs
*it* as an ordinary module afterwards, so that it too can be upgraded, is the same closed-list
# 74. The mesh defines a module protocol; an SDK is an implementation of it
## Context
[ADR 0039](0039-what-the-sdk-holds-and-refuses.md) says what the SDK holds: the tool-serving
harness, the messaging and event framework, the contracts, and core primitives. It settles what
belongs in *an* SDK. It does not say what happens when there is more than one.
There is already more than one. **The contracts are expressed twice** — as Go types in the control
plane and the host, and as TypeScript types in the SDK — and nobody has felt it because both live
in one repository and one head.
**A correction, made after inspecting the wire rather than the types** (2026-09-16). This record
first claimed the two implementations already disagreed — `resource` vs `Provision`, `consumer`
meaning the module in one and the node in the other, headers declared on one side and emitted by
neither. **On inspection the live wire agrees**, and the claim was wrong:
- The grant types that disagreed (`Grant`, `Interface`, `Credential` in the SDK's `contracts`)
were **dead** — exported and imported by nothing. The live provisioning wire is the contributions
file, whose shape (`as`, `secret`, `node`, `at`, `values`) is the same on both sides. Those dead
types have been removed.
- The envelope agrees too: Go emits all five required headers, and `x-causation-id`/`x-schema` are
**optional** — the SDK sets them when a handler has a causation or a schema, and a bare event
carrying neither is correct, not a drift.
So the danger was never live disagreement. It was **dead types that contradicted the live wire**,
which read as the contract and were not — and are exactly what led this record to assert a drift
that inspection did not find. That is a sharper reason for the decision below, not a weaker one: a
type is only as good as its being the wire, and the way to guarantee that is to specify the wire and
check implementations against it, rather than to trust a hand-kept type to still describe it.
A failure of this kind does not announce itself. Two implementations that disagree about an
envelope do not fail to compile — they ignore each other's messages, and a mesh where a module
stops reacting looks exactly like a mesh where nothing happened.
## The question this settles
A module may be written in any language the mesh can build
([`18-building-a-module`](../03-DESIGN/01-to-be/18-building-a-module.md)). Every language needs an
SDK. What is an SDK *of*?
Two answers were available, and the obvious one is wrong.
**Shared types, generated.** Write the shapes once — a schema, an IDL — and generate Go, TypeScript,
Rust. It is the familiar answer and it solves the smaller half of the problem. The shapes are not
where the difficulty is.
**A specified wire, with a conformance suite.** The shapes are a consequence; what an SDK must get
right is *behaviour*.
## Decision
**The mesh defines a module protocol. An SDK is an implementation of that protocol in one
language, and nothing more.**
That is the whole of what an SDK is. Not a library a language happens to have, not a convenience
layer, not a place for helpers to accumulate — an implementation of a specified protocol, finished
when it implements it and correct when it agrees with every other implementation.
### The protocol is split per capability
**A module does not use all of it, so an SDK need not implement all of it.** A module that only
consumes events uses the event capability. One that serves tools uses the tool capability. A
provider uses provisioning. Nothing about consuming an event requires knowing how a grant is
answered.
So the protocol is a floor plus capabilities:
| part | what it covers | who needs it |
|---|---|---|
| **connection** — the floor | reading the sealed credential, pinning the certificate fingerprint, taking identity from the credential rather than the environment | everything |
| **events** | the envelope and its headers, the durable per-consumer queue, binding, at-least-once with dedup on `x-event-id` | a module that emits or consumes |
| **tools** | registration, the shared durable `serve.<key>` queue, request and reply | a module with a surface |
| **provisioning** | a grant in, a credential out, and what each carries | a module that provides something |
**This is the same shape the host already has.** A host declares which resource kinds it can apply,
and a partial host — one that can write files and run things but not manage users or containers —
is a real thing rather than a broken one ([ADR 0005](0005-the-node-host.md)). An SDK that implements
the floor and events is exactly as legitimate, and a module written against it is a module that
does events.
**So a language arrives in pieces rather than all at once.** A Rust SDK implementing connection and
events is useful the day it exists; tools and provisioning follow when something needs them. The
alternative — a language is unsupported until it is entirely supported — is what makes adding one a
project rather than a contribution.
**And what a language can be used for is then a fact the mesh can state**, rather than something an
author discovers by writing a module that cannot be built: the toolchain list says which languages
exist, and the conformance results say what each can do.
### What the specification covers
Per capability, what two implementations can disagree about:
- **the exchanges and queues** — which exchanges exist, that a consumer's queue is durable and
named `<node>.<module>.events`, that a tool is served from a shared durable `serve.<key>`
- **the envelope** — every header, which are required, what an unknown `x-` header means, and that
ignoring one is correct rather than lax
- **identity** — that a module's node and module name come from its sealed credential and not from
its environment, so what it emits matches what the mesh authorised
- **delivery** — at-least-once, and that dedup is on `x-event-id`, which only the emitter can make
- **the credential** — the sealed document's fields, and that a connection pins a certificate
fingerprint rather than trusting an authority
- **provisioning** — a grant in, a credential out, and what each carries
- **the vocabulary** — that `consumer` is one thing, named once
### Conformance is per capability
**A suite per part, and an SDK claims the parts it passes.** A monolithic pass/fail would make a
partial implementation indistinguishable from a broken one, which is the distinction this is built
on.
**And the suite is executable, not prose.** A specification nobody can run is a document two
implementations drift from while both believe they conform. Conformance is a set of fixtures — an
emitted event, a served tool call, a grant and its answer — that every SDK must produce and consume
byte-for-byte.
**The existing two implementations are the first two to be made to pass it.** Not a future language:
the drift above is present, and a suite that only new SDKs must satisfy would leave the disagreement
that already exists in place while certifying everything added afterwards against it.
## Why not generated types
Generation makes the shapes agree and leaves everything that matters unspecified. Two SDKs
generated from one schema can still name their queues differently, take identity from the
environment, dedup on the wrong field, or omit a header the other requires — and every one of those
is a mesh that runs and quietly does not work.
It also makes the contract into whatever the generator supports, which is a decision nobody made
about a boundary everything else depends on.
**The shapes are worth generating once the wire is specified.** That is a convenience, and it comes
second.
## Consequences
**A language is a commitment, and now a divisible one.** Adding one means implementing the protocol
and passing the suites for the parts it claims. That is more work than transliterating types, and
it is the work that was always there — the difference is that it can be finished, and finished in
pieces, rather than believed.
**Versioning becomes possible.**`x-schema` exists for it and is never written. A specified envelope
with a version on the body is what lets a mesh hold a module built against an older SDK, which is
the ordinary state of any mesh that has been running for a while.
**The two current implementations agree on the live wire** — inspection showed it. What was wrong
was a set of dead types beside the wire, now removed. The suite's job here is therefore prevention:
to keep that agreement true as the wire changes, and to hold a new language's SDK to it, rather than
to repair a break that exists today.
**This does not make the mesh polyglot by itself**, and should not be reported as though it does. It
makes polyglot possible to do correctly. A Rust SDK is still a Rust SDK.
## How this is checked
| Rule | Checked by |
|---|---|
| One vocabulary | A word means one thing across implementations, checked by the fixtures using it. |
| The wire is what is specified | Both existing SDKs run the conformance suite in their own test suites, and a change to one that breaks a fixture fails there rather than in a mesh. |
| A new SDK is a passing SDK | A language is not listed as buildable for a capability until its SDK passes that capability's suite; the toolchain list and the conformance results name the same set. |
| A partial SDK is a real thing | An SDK implementing the floor and one capability passes, is listed for that capability, and a module using another is refused with the reason — rather than failing at runtime in a language nobody said was finished. |
| An unknown header is ignored | A fixture carries one, and every implementation accepts it. |
| Identity comes from the credential | A fixture sets an environment that disagrees with the credential, and the emitted event carries the credential's. |
**The framing that dissolves both:**`artifact-store` is already a provision, and `registry`
already provides it. So "should gitea be the registry" is not a question about replacing a
component. It is a question about **a second provider of an existing provision** — which this mesh
has a mechanism for, and uses for certificate authorities and VPNs already.
## Decision
**Two provisions, because they are two jobs.**
| provision | is | for |
|---|---|---|
| `artifact-store` | content-addressed blobs, pinned by digest, no versions, no ranges | what the **mesh** delivers to **machines** |
| `package-registry` | an ecosystem's own registry — npm, cargo, PyPI, Go | what **code** resolves when it is compiled |
They are not the same store with different clients. One is addressed by digest and immutable by
construction; the other is addressed by name and version, and resolves ranges. Conflating them is
how a mesh that pins everything ends up rebuilding one commit into two different things.
**`registry` remains the provider genesis installs.** Not because it is better, but because of what
it is: a directory and one container, no database, no control plane, installable at step 8 of an
install where neither exists yet. Gitea needs a store and provisioning, which means a control plane,
which means the pivot has already happened — and the pivot needs somewhere to publish to.
**Gitea also provides `artifact-store`, and a mesh may choose it.** Two providers of one provision
is a thing the mesh understands: it refuses, names both, and choosing is assigning the one you want.
**And it does not claim `the-artifact-store`.** That claim is node-scoped, so a module holding it
cannot share a machine with another that does — and a machine running gitea for git and packages
*alongside* a registry serving artifacts is an ordinary arrangement, not a conflict. They are
different ports doing different jobs.
The exclusivity that matters is mesh-wide and is already expressed: `provides` at mesh scope means
two providers are two answers, and the resolver refuses until one is assigned. Forbidding
co-residence adds nothing to that and forbids something reasonable. **Whether `registry` should
still hold that claim is left open here** — it may be protecting something about the port or the
data directory that is not written down, and removing a claim is not a thing to do from the outside
of a manifest.
A mesh that assigns gitea gets authentication and TLS for its artifacts — which is to say, **issues
042 and 048 are answered by choosing a provider that already solved them**, rather than by
reimplementing accounts and certificates in a registry that has none.
**Gitea provides `package-registry`.** That is ADR 0014's private registry, and it is one service
rather than one per ecosystem. `verdaccio` may provide it too, for npm alone, and is then a choice
somebody makes rather than the answer.
**The registry is not removed at the end of installing.** A mesh that never runs gitea still has an
artifact store. Retiring it is a migration a mesh performs, not a step an installation ends with.
## Why not simply gitea, from the start
Because genesis would need a control plane before the thing that stores the control plane's image,
and that is circular rather than merely awkward. It would also make one of the three things the
build loop cannot produce for itself into a stateful application with a database — the pivot is the
hardest part of this design already.
And it puts every artifact in the service that is also the trust anchor for everything the mesh will
ever run ([ADR 0071](0071-where-genesis-gets-its-source.md)), which records that forge serving a
cryptominer with tampered git operations. Two blast radii are better than one.
## Moving from one provider to the other is a designed act
**Not a removal.** Every image a machine runs is pinned to a digest at a named store, the control
plane's own included. Changing the provider means:
1. gitea installed, reachable, and holding an account the builder may publish with
2. every artifact mirrored
3. every declaration re-pinned, the control plane's **last**, because it is what performs the others
4.**every machine verified to have converged and to be able to pull from the new store**
5. only then the old provider unassigned, and its volume kept ([ADR 0030](0030-data-outlives-the-mesh-that-declared-it.md))
Step 4 is the one that is easy to skip and the only thing between this and a mesh that cannot
restart its own control plane. A machine that reboots mid-migration pulls from a store that no
longer exists, and a local image cache hides that until exactly the moment it matters.
## Consequences
**The bootstrap is unchanged**, which is the point of keeping the small provider.
**042 and 048 gain a second answer.** They can be fixed in the registry, or dissolved by choosing a
provider that already has accounts and TLS. The second is less work and more service.
**ADR 0014 becomes satisfiable.** There is a provision for the private registry, something that
provides it, and a module may depend on it — so the SDK can be published and consumed rather than
cloned, and issue 053 has somewhere to go.
**A mesh can be minimal or complete, and both are legitimate.** One with the small registry and no
gitea builds and runs modules and cannot serve packages. That is a real configuration, not a broken
one — the same way a partial host is real.
**And the bootstrap still has no package registry.** The first build of the shared base happens
before anything has installed one. That is the same pivot as everything else and it is **not solved
here**: it is named, so the next person does not discover it.
## How this is checked
| Rule | Checked by |
|---|---|
| Genesis needs no database | The installer raises a mesh of one on a machine with nothing, and the artifact store it installs has no store of its own. |
| Two providers are a choice, not a conflict | A mesh holding both is asked to resolve `artifact-store` and refuses, naming both, until one is assigned. |
| Providers may share a machine | A node is assigned both gitea and a registry, and both run — only one of them answers `artifact-store`. |
| The two stores are not interchangeable | A module depending on `package-registry` is not satisfied by `artifact-store`, and the refusal says why. |
| A migration is verified before it is finished | The old provider cannot be unassigned while any machine's declaration still names it. |
# 76. The SDK is a published package, and the toolchain resolves it by version
## Context
[ADR 0014](0014-no-npm-workspace.md) decided a module consumes its dependencies — the mesh's own
shared library included — from the private registry. [ADR 0075](0075-two-stores-and-which-provides-what.md)
decided the private registry is a `package-registry` provision, and that gitea provides it. What
neither settled, and what [issue 053](../04-ISSUES/053-the-sdk-is-pinned-twice-and-the-two-disagree/00-report.md)
left open, is the one build where the rule cannot simply be obeyed: **the first one.**
The TypeScript toolchain image is built *from* the SDK — it carries the SDK so that every module
compiled inside it resolves the shared library without each build fetching it. So the thing that
compiles TypeScript and the thing that contains the SDK were the same object, and that object
cannot be what builds the SDK. Stated as a question — "how does the SDK reach the registry before
the toolchain exists, when the toolchain is what builds it?" — it reads as a paradox.
It is not one. The paradox exists only because the toolchain *bakes a git-cloned copy* of the SDK.
The SDK itself is plain TypeScript: it needs `node` and `tsc` and nothing the mesh makes. A public
base image can build it. The circularity is a property of the workaround, not of the SDK.
## Decision
**The SDK is an ordinary published package in the mesh's `package-registry`, consumed by version.**
The git URL in the toolchain's manifest and the sibling-path lock beside it — the two halves of
issue 053 — are both removed. A build resolves the SDK the way it resolves any dependency, with a
lock that agrees with its manifest, so `npm ci` is the command and reproducibility is by
construction rather than by the machine the build ran on.
**The SDK is built with a public base image, not with the mesh's toolchain.** It is *not* one of
the components the loop cannot build — the control plane, the registry, the builder, the catalogue
([`12-a-module-repository`](../03-DESIGN/01-to-be/12-a-module-repository.md)), which arrive by
carrying an init builder because they are the loop's own machinery. The SDK is machinery for
nothing; it is an ordinary dependency the loop builds and publishes like any other. The only
constraint is narrow: it cannot be compiled *in the mesh toolchain*, because that toolchain is built
from it. So it is compiled on a public base image instead — which needs nothing the mesh makes — and
published before the toolchain that consumes it. It is not carried, because building it does not
wait on a mesh existing first.
**The toolchain base stays, thinned.** mesh-tools remains the image bundles are compiled in and the
one place the SDK is resolved — but it `npm ci`s the SDK by version from the registry instead of
baking a copy cloned from a git URL. Bundles keep borrowing its resolved dependencies; what changes
is that the version they borrow is named and honest. This was the shape chosen over dropping the
shared base entirely and having every bundle resolve the SDK itself: one resolution point, one
place to be right about the version.
**Genesis orders the publish before the first compile.** The package-registry provider is a public
image (gitea), so it comes up needing no toolchain; the SDK is published into it; only then is the
toolchain built, so the first `npm ci` has a registry to read from. Nothing in that chain is
circular, because the only thing that needed the toolchain — baking the SDK — is gone.
## Consequences
Each language's toolchain repeats the shape: its own SDK, built from that language's public base
image, published to the same registry, resolved by version with that ecosystem's lockfile-honest
install (`npm ci`, `cargo` against a vendored or registry source, `pip` against a pinned set). The
warning in issue 053 — that whatever the TypeScript repository does the others will copy — is
answered by making the copied thing the correct one.
A change to the SDK is publish-then-consume, exactly as [ADR 0014](0014-no-npm-workspace.md) already
priced it: publish the new SDK version, then bump the toolchain (and any module pinning it directly)
to consume it. There is no shortcut that resolves an unpublished SDK, which is the property that was
missing.
mesh-tools is no longer an SDK carrier in the sense that mattered — it does not contain a copy
whose provenance is a branch head somebody force-pushes. It contains a version.
A mesh with no package-registry cannot build TypeScript. This is accepted and is not new: it is the
same shape as a mesh that cannot reach a forge being unable to be raised
([ADR 0071](0071-where-genesis-gets-its-source.md)). Installing brings the registry up first.
## How this is checked
| Rule | Checked by |
|---|---|
| The SDK a build compiles against is named, not cloned from a branch | The toolchain manifest pins `@novox/mesh-sdk` to a version, and the build runs `npm ci`, which refuses a lock that disagrees with the manifest. Issue 053's two checks become this one. |
| The SDK builds without the mesh's own toolchain | The SDK's build recipe names a public base image. A recipe that named the mesh toolchain would reintroduce the cycle and is refused in review. |
| The registry is up before the first compile | The genesis bed asserts the package-registry answers, and the SDK is published, before the base build runs. A base build that ran first would fail its `npm ci` with no registry, which is the positive control. |
| A second language repeats the shape, not a new one | When a second SDK is added, its recipe is compared to this one: public base, publish by version, lockfile-honest install. |
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.