Verified on the machine: umami has zero restarts, applies all 26
migrations, reports the store up to date, and umami.novox.be answers 200
where it had answered 502 since 2026-09-25.
It was the same fault as 135, filed three days earlier and diagnosed
without either record noticing the other. 135's container IS umami — it
is named in 135's own evidence table, holding novox.internal:10.42.0.1
after the overlay range moved. The dial appeared to succeed and the first
real query timed out because the name pointed at an address that no
longer existed; the store, the path and the credential were all fine.
118's own reasoning is kept as a warning, because it is careful and
wrong: it argued path-MTU and conntrack, and it named the move that would
have found it — compare umami against a container created that week —
and did not make it. A stale name presents as a network fault, which is
why ADR 0148 stops copying names into containers rather than detecting
when the copies go bad.
Also: 118's located-in named mesh-catalog modules/umami, which was never
at fault; it names mesh-host's comparison now. And 135 gains the pointer
to 118, which is the back-reference I have now missed three times.
The certificate is genuine, from Mesh Internal CA, and nothing on the
workstation trusts it — verbatim the error the report gives. The public
name on the same proxy verifies cleanly, which puts the fault exactly
where the report puts it.
What is in the way is not an assignment. `ca-trust` is merged in the
catalogue and has never been registered with the mesh — 39 of 76
manifests are — so there is no module to assign. It dry-runs clean and
needs no artifact built.
Two findings from reproducing it, both their own issues:
157 — every routed name is published with an `.internal` alias that
nothing serves. The hosts file says keycloak.novox.be.internal; the proxy
serves keycloak.novox.internal and refuses the other by name. The first
three names I tried came from the hosts file and failed with a TLS alert
rather than a verification error, which pointed at a regression that had
not happened.
158 — the proxy re-logs all 52 routes every two seconds, 31 times a
minute. The one line that explained 157 sat between two of them.
Also recorded, because it was nearly filed as a defect and is not one:
step-ca publishes roots as /roots.pem, which is PEM, so ca-trust's fetch
and its refuse-a-non-certificate guard are both right. Its other endpoint
/roots returns JSON that contains the text the guard greps for, so the
guard is sound only because of which path is published.
Three records were left describing a mechanism a new record had moved,
and a reader arrives at them by following a citation: 0066 still said a
routed name is written into every container after 0148 replaced that with
resolution; 0016 still read as though the lab were the test bed after
0149; and issues 109 and 135 said nothing about 0148 ending the copying
that 135's own fix made comparable. Each was a citation leading to the
wrong answer in a record that was not wrong about anything it decided.
This is the second time in one session. The playbook rule I added last
round did not stop it, so the convention is now written where the record
conventions live, with the shape to use and three worked examples — and
with the honest note that it is NOT machine-checked and cannot be from
`extends:` alone: 102 records extend another, 87 have no back-reference,
and that is correct, because extending usually means building on a
context. Making it mechanical means a record declaring the relationship
in frontmatter, which is a schema change and is not mine to decide.
Also: designs 18 and 20 claimed `updated:` dates from before I edited
them, and 117's `fixed-by` gained the commit beside the record.
**0114 accepted.** The separation it draws — a consumer's resource and a
credential that reaches it are different things — is a data-loss rule,
and a record naming one should not sit unresolved while the code that
could hit it is written. Not built, and accepting it schedules nothing;
accepted-and-not-built is where 0141 and 0142 already are.
**0068 superseded by 0149, the live mesh is the test bed.** Never built,
and contradicted by practice that was written down nowhere but a handoff
note. The faults that cost the most are faults of a mesh that already
exists — bound consumers, containers made against an older roster, an
adopted machine — and a bed is by construction a mesh that does not.
Issue 156 settles it: that change was verified honestly, on the only path
where it cannot fail. The lab is not retired; raising a mesh from bare is
now its whole job.
**117 answered by 0150.** A module's own code runs as supervised
processes under the module's one account. 0047's argument was for a
runtime per module, not for a container — a unit satisfies it and runs as
an account. The invariant is the account, not the process count: its
worry was a second identity to scope and seal, and processes sharing one
account create none. Designs 18 and 20 now cite it and 0047 carries a
dated pointer, which closes the third disagreement — that neither design
knew the record existed.
Answers issue 151. Copying the roster into each container made the roster
part of each container's identity, so one name moving replaced every
container in the mesh — and it never stopped the staleness it was for,
since a copy taken at creation is stale the moment the roster moves (109,
135).
A container resolves through its machine's resolver instead, and nothing
is copied. Staleness stops being possible rather than detected, and a
name's blast radius becomes nothing.
Scoping each container to the names it binds was the close call and is
rejected: it contradicts anything-calls-anything, and leaves the roster
in the digest so the churn returns for a widely-bound name.
Gated on issue 110 — a container on the runtime's default network has no
DNS at all today. Removing the copy first reintroduces 109 and 135
silently on a live mesh. 151 stays open until the code lands; design 08's
file-not-resolver passage is narrowed to the machine's own roster.
Three issues named the branch that fixed them, and a branch is deleted
when it merges — so every `fixed-by:` was a pointer that resolved to
nothing by the time anyone followed it. They name commits and pull
requests now, and playbook 03 says to.
Two records were missing the thing a reader arrives for. 146 did not say
that one of its fixes crash-looped the control plane on a running mesh,
which is the whole reason the delivery subject carries the stream and the
raise path was the only one exercised. 151 did not say that 152 removed
the false reasons its roster moved, or that it stays open for the real
ones.
ADR 0080 enumerates what cycle.py enforces and named four things; it
enforces five. A progressive insight names the fifth — the decision
stands, the list had gone stale. The checks README and playbook 03 gained
the same rule, and 155 points at all three.
Measured on the broker's own config after the fixed controller composed
it: every node is now allowed _DELIVER.<node>.> as well as the bare
subject. That grant was missing when this was diagnosed, which is why
re-making the consumers would have silenced every machine.
Also records the five consumers it kept and reported, including the build
machine's worker, which is bound the same way and was not anticipated.
146's change is correct on a mesh being raised and fatal on one that is
running: the server will not move a push consumer's delivery subject
while a subscriber is bound, and a node is bound to its declaration
consumer the whole time it is up. Reproduced before fixing; an earlier
version of the test unsubscribed first and passed against the broken
code.
The consumer that works is kept and the assertion says so. Not re-made:
the nodes are granted the bare subject only, so re-making would have
silenced every machine.
Two machines filing issues in the same hour both read main correctly and
both took "the next free number". main lags every open pull request —
seven that evening — so they collided twice. The second collision
reached main with records, cycle and index all reporting success.
cycle.py now refuses a tree where two issue folders share a leading
number, and names both. Proven by adding a duplicate and watching it
fail. The colliding records become 153 and 154, renumbered in the branch
that lands last, because renumbering a branch whose author is still
pushing only moves the race.
The check catches a collision; it does not prevent one. Taking a number
still means reading the open pull requests as well as main — issue 155
says so.
It ran 22:46-23:11, five full replacements, and stopped when a pass
happened to read the store during a window it was up. The record said
nothing about it stops; that was wrong. Exiting by luck is the finding,
not a mitigation.
The three gatherers pass over a node whose set does not compose, and
raise anything else. 151 stays open: this removes the false reasons a
roster changes, not the fact that a real change still replaces every
container.
ace numbered its two records 147 and 148, which this repository already
uses; they become 150 and 151, as 127 became 149.
The new record is why the control node could not stop applying: the
roster alternates between two values because routeNamesInTheMesh
swallows a per-node plan failure, and the roster is part of every
container's identity. The loop closes through the control plane's own
store, which each pass replaces.
ADR 0138's reach is internal, public or both. ace's LAN devices (an IoT
switch on mosquitto, unifi access points, plex clients) are none of them;
the only reach that admits them is public, which states the wrong thing and
opens the port to the internet on any machine with a public address.
ADR 0112 decided that an assignment places a directory and says where an
access's data is; the control plane implements neither. ace's media stack
(a 40 TB ZFS library, plex state on a second disk) cannot be expressed
without writing ace's paths into manifests.
Both found migrating the first module on ace (searxng), adopted with route-adapter:
147 — assign is not a no-op for a routed module: its route reaches the predecessor's
proxy while its container is held, and the name serves a dead backend until take.
148 — every push that moved the mesh's names replaced every container on novox, twice,
the control plane's own store included.
It asked a question rather than reporting a defect, and the question was taken
five days later — the mesh's own components are binaries on the machine, and
third-party software stays a container because an image is the right way to
carry somebody else's build. The delivery of them is issue 142.
A number identifies a record, and two were given 127 on 2026-09-27. The one
three documents and three source files cite by number keeps it; the other
becomes 149, says so in its own heading, and its two inbound references are
repointed.
It is also resolved: an empty declaration is sent carrying owns_nothing rather
than skipped, and the host refuses an empty body that does not carry it, so
emptiness cannot be read as a truncated declaration.
0037 (where a module lives), 0113 (the vault makes every secret) and 0115 (one
assignment of a module per node) described arrangements the mesh has — a
catalogue of other people's software plus a table filled by module add, a vault
answering a secret requirement six modules make, and a rule the assignment
table's primary key already enforces. Each is accepted against what was built,
and says so in its own words.
0037's other half is not built: a manifest outside this catalogue has no check,
which is issue 148.
0068 (the lab takes requests) and 0114 (a shared credential rotates over two
credentials) stay proposed. Neither is built, and both are decisions rather
than records of something that happened.
The record catches README claiming this repository is indexed into a knowledge
base. Nothing indexed it, and since the cut-over there is nothing to index it
into — the surface that answered is on the transport the mesh removed (issue
147). Noted where the record is, so the next reader does not go looking for a
search that cannot exist.
088, 089, 120, 128 and 130 each name a commit that is on main and cites them —
the forge's address following a moved port, a route naming its endpoint, a
provisioner asking the backend what is there, the hosts file written into a
marked block, and undeclaring giving a unit back the state it was found in.
Each says it was closed by reading commits rather than by a run, so nobody
reads a green that was never measured.
129 stays located on purpose: ca-trust is merged and no machine holds it, so
the symptom it opened on is still true everywhere.
Written 2026-09-28, answered the same week by mesh-controller bdf965d and
c68d3a7, and never closed — so the mesh's own account said reach was declared
nowhere while three readers were reading it: the filter, the proxy's names, and
each authority's host policy. The resolution names them and what checks each.
What is left is retiring the older per-port keys the block replaces, which is
not a gap in what reach can say.
Diagnosed from the configuration: the tool server is HAL's brain, a local
process on the workstation with the predecessor's broker URL in the assistant's
own config. No manifest, no assignment, no seat, no account. Nothing regressed
— the mesh removed a transport this program still dials, and the program was
never part of the mesh. The mesh has a tool model and nothing publishes an
operator-facing surface onto it.
Every tool call fails with 'AMQP not connected', on every node including the
local one, because the tool surface still opens an AMQP connection and that
transport was deleted at the cut-over. The mesh reports healthy throughout —
what broke is the thing standing outside asking it questions, so nothing the
mesh checks is about it.
Not about enrolment. A push consumer delivers onto an ordinary subject and
everything subscribed to it gets a copy; the controller's two consumers were
both named after it, so both were given the same subject and the one process
acted on every message twice. Enrolment is where it drew blood because a second
enrolment mints a second credential.
Two more faults behind the three already fixed. The account a token is the
password of was never recorded, and the comment above the issuing code said
it was; issuing now records it. Placing the composed list at genesis is the
other half — the control plane says what it composed and whoever raises the
machine writes it beside the bus, because no declaration can reach a machine
that has not enrolled.
With that a first node enrols. It is then enrolled twice from one attempt,
each minting a credential, and it keeps the answer to the first while the mesh
keeps the second. The trail for that one stops at a duplicate that survived
message-id deduplication.
Raising a first node hits them in order: the bundle's bus image named for a
registry that is gone; the certificate made by openssl in an image that has
none; enrolment dialling TLS at a bus that speaks first, then refusing its own
token for the empty bus. Each was right until the bus changed and nothing has
raised a foundation since.
The fourth is not a patch: the installer carries the bus's first user list and
the controller composes the rest through the bus being a module, and at genesis
there is no module — so the first node cannot be let onto the bus it just
raised.
Found trying to run ADR 0147's bed. The older bundle raises the previous
broker and a control plane that refuses to start without MESH_BUS_NATS; the
newer one stops a step earlier, asking for a certificate from openssl in an
image that has none. No bed can run while this holds, so 0147's check section
now says what actually stands behind it — the rendering, not a machine.
Issue 129: every internal HTTPS name fails verification on every machine,
because nothing has ever written the mesh's root into a trust store. The
report proposed the controller inject it the way the private network writes
the registry's trust; this record rejects that — reachability and trust are
not the same fact, and where anchors live is the host's difference, not the
controller's. A module requiring internal-acme-ca does the whole of it, and
being unassigned undoes it.
0145's core stands — a module checks what the mesh claims, from where the callers
are — and everything it said about what to dial was wrong.
A raw port is not how anything in this mesh is reached. A real caller resolves a
name, the proxy answers, the proxy reaches the service; dialling a port tests the
last hop of a four-hop path and skips the three where most connectivity lives.
So each hosting form gets its own endpoint, route and name:
connect-docker.<node>.internal for a container, connect-process for the mesh's own
code in a unit it writes, connect-unit for a unit a package ships — and the same
under each public domain. Every machine checks every machine, by name, over TLS,
verifying the certificate against the authority that should have issued it. The
names are the instrument: a failure reads as connect-docker.g14.internal did not
answer.
No name is written anywhere: the machines come from the roster fact, the labels are
the module's, and a machine that joins is checked by the others on the next push.
Five things this needs that the mesh does not have, named rather than assumed —
a module every machine has (ScopeNode is exclusivity, not obligation), a container
running a bundle, a unit a package ships, the public domain in the roster fact, and
issue 129, until which every internal name will correctly fail verification.
The mesh asserts three things are callable and has never checked any of them. A
module on every machine serves an endpoint of its own and dials every other
machine's, from the position the callers are in — not the host and not the control
plane, both of which reach those addresses by paths no ordinary caller uses and
would have passed throughout the outage.
Its probe is its own endpoint, declared reachable over the private network, so it
is admitted by exactly the rule that governs every internally-exposed service. The
tempting target is a service every machine has, and those are the ones never closed
— ssh above all — which would have passed for all eleven hours.
Resolution and connection reported separately, because they have different owners.
One failure is not a fault, and the count travels with the result. It reports and
repairs nothing.
Built on a scheduled container, a rendered roster fact and a bundle: no new
vocabulary. Adopted on its own merits, which is what 0143 was not.
Everything should be able to call what runs on the same machine, another machine's
service exposed to the private network, and another machine's service exposed
publicly. Three cases; the filter had two.
The first was expressed as the machines' own addresses on the private network. A
caller on the machine carries such an address; a caller in one of its containers
carries a bridge address and matched nothing — measured, same destination and same
machine: src 10.10.0.1 against src 172.17.0.8. The second case worked by accident,
because the tunnel rewrites a caller's address to the sending machine's. Two of
three working is why it read as correct.
0143 answered the wrong question. It proposed verifying each grant from the
consumer's own network position and went to length about which position, because
whether a caller sat in a container changed the answer — and that difference was the
bug. Observing a configuration error is not its remedy. Superseded, and nothing
replaces it; whether the mesh should check a grant is still open in issue 145 and
must stand on its own.
And a module is not a container: 61 of 72 happen to use one, 11 do not, and a rule
reasoning about containers describes most of the mesh rather than the mesh.
A grant is four facts and a credential — the provision, the machine, the port, and
who the consumer is when it connects — and it is the whole mechanism by which
anything in the mesh reaches anything else. The mesh asserted it and never found
out whether it was true. Issue 145 is what that cost: eleven hours of 'every module
current with its source' while a module could not reach its database.
The consumer verifies it, from its own network position. Not the control plane,
which reaches the address by a path no consumer uses. Not the machine, whose own
packets carried a source address the filter admitted while every container's was
refused — so a host-side check would have passed throughout the outage it exists to
catch. That is inference from the rule that was loaded, stated as such.
A connection and nothing more; speaking each provision's protocol would be a second
implementation of every provision. One failure is not news, only consecutive ones,
and the count is reported rather than the last attempt. A consumer that is not
running reads unchecked, which is a different sentence from broken.
What it costs to be wrong is the constraint on all of it: a check that calls a
working provision broken trains a reader to ignore the report.
Converging the control node closed every path by which a module reached another by
the machine's own name, and it ran for eleven hours while the mesh answered 'all
doing what they were told, all heard from, running what the mesh would send them,
and every module current with its source'.
6,154 database failures in one module's log, beginning at the minute of the flip.
The service accepted TCP and never answered HTTP. Every check the mesh makes
passed, because every check the mesh makes is about the relationship between the
mesh and a machine — applied, current, containers running. None asks whether a
module can reach what it requires, though the mesh composes every grant and so
knows exactly who requires what.
The filter fault is fixed. The eleven hours are the measurement, not the bug.
Also records on 144 that 'more closed than the mesh believes' was harmless only
for what is reached from outside, and an outage for what is reached from within.
The record says internal and public each mean something to the filter. For an
endpoint the proxy serves, the second half is wrong, and ADR 0045 said so first:
a public service is exposed through the proxy, listening from the mesh, not by
opening its own port.
Found by trying to express one real module, not by review — routed name public
because browsers post to it, machine port private because it serves a dashboard in
cleartext. Under one value for both, saying public would have reopened a port an
operator had just closed. Measured the same evening: the routed name answered from
the internet over TLS while the port was refused from the same place.
Corrects a fact. One statement per endpoint with three things derived from it
stands; the filter column applies to an unrouted endpoint.
The first account said the declaration carries no resource that would disable the
found firewall and that nothing implemented the sentence the flip prints. Wrong:
the mechanism is a step in the host's own apply, retireFirewall, and it is careful
— it reads back that the mesh's table is loaded before retiring anything, records
the forward policies first, and verifies ufw reports inactive afterwards.
What is established: ufw was active and enabled two minutes after a flip that
reported the node converged with nothing failed; and the machine's record now
reads disabled_by_mesh: true, written by a reconcile fifty minutes later that
found ufw already inactive because an operator had disabled it by hand.
So the step did not take effect and the record says it did. The candidates are
named rather than chosen — the step is skipped silently when the apply has any
failure, the mesh's table is loaded by a service in the same apply so ordering is
open, and the host's detail lines do not reach the journal, which is why this has
candidates instead of a cause.
Both found by converging the control node — the first machine with a firewall to
flip, since the two before it had none.
143: the preview and the flip both say the found firewall is disabled. The node
reported converged, 372 resources applied, nothing failed, and ufw is still
enabled and active. A converged node's declaration carries no resource that would
disable it; the sentence is printed by the command and nothing implements it.
144: ufw was never what filtered the traffic that mattered there. Fifty forwarded
openings converged through it had matched zero packets, while a chain the
predecessor installed in the container runtime's pre-accept hook did the work —
in memory only, recreated by nothing. The mesh's filter now covers that path, so
the machine no longer depends on it, but the chain remains and is the only thing
refusing the bus and the registry, which the design requires reachable from
anywhere so a machine can enrol before it has a private address.
Measured: the host is a binary somebody copied to four machines, owned by no
package and built by nothing, while the controller, catalogue, builder and vault
are container images publishing no ports at all. Same language, same project,
same kind of work, delivered two ways — and the difference is not a judgement
about either, it is that images are the only delivery that works.
What it costs: genesis must raise a container runtime before the control plane
can exist; updating the control plane goes through a registry the control plane
runs; a host change cannot be rolled out at all; and compiling the language the
mesh is written in is not a capability of the builder, so the controller is built
from a hand-written Dockerfile — the incantation the bundle toolchain exists to
abolish.
Third-party software stays a container: the store, the registry, the broker are
somebody else's build. The container runtime stays on the machine for modules.
What changes is that the control plane no longer needs it to exist.
The receiving half is already built and tested (ADR 0141). Staged: compile Go, an
artifact names its target, deliver a binary, the host first, then the rest, genesis
last.
The record claimed a version reaches a machine as an ordinary archive with
nothing new needed. Two things it needs do not exist: no toolchain can compile
the host (the list is typescript and python, and the control plane, also Go, is
built as an image from a Dockerfile instead), and nothing interpolates a built
version into a resource path, so nothing can ask for .../versions/<version>/.
The decision, the options weighed and every consequence stand — the host half is
merged and tested. What was understated was the cost, so it is corrected in place
and dated rather than superseded.
The supervision was already right — a clean exit means the host stood aside and
the launcher runs what is on disk, failures are counted, and a rollback happens
at the limit. Two things made it dead code: nothing told the running host a
successor was waiting, and the rollback resolved its known-good version through
pacman, which no machine here uses and which two of three operating systems do
not have.
Keeping a version rather than a path was the clue. Versions live side by side in
directories named for them; the newest runs; the running one stands aside between
reconciles; a reconcile that completes records itself and retires what is older
than its predecessor; rollback starts that predecessor. No new resource kind and
nothing new on the bus — an archive already fetches by digest, and the path
written is never the path executing.
Answers issue 142.
A host change merged yesterday reached no machine without a person copying a
file. The host is not a build target, no declaration delivers it, and the half
that recovers from a bad host — noticing the executable changed, a known-good
record, a launcher that rolls back — is written, tested and called by nothing.
All four machines run a byte-identical hand-copied binary that no package owns
and no record names, so nothing can say a machine is behind.
Found because ADR 0140 needs the machine to report a new fact, and merging that
could not roll it out.
Reading a converged machine's rendered rules showed the cause: the chain blocks
everything passing through and then allows the machine's own containers back by
listing their address ranges. 0137 made that list typeable and 0139 tried to
generate it; both refined a list that should not exist, because the mesh has no
position on a container reaching outward. Constrain what arrives from outside,
allow what did not, and let the machine report which links face outside — one
fact instead of a list. Ports keep following the modules unchanged.
The records check now allows one record to supersede several, and stops
requiring a withdrawn record's own citations to be live.
Both follow from the same rule the mesh is built on — a node's configuration is
composed from the modules assigned to it. Reach was settled separately by the
filter, the proxy's names and the certificate authority, so "this must not be
public" could not be written; it becomes one value on the assignment that all
three read. And the forward chain consulted two constants plus a typed list
although modules already declare their networks; it now forwards what they
declared, with the host rendering the addresses it allocated.
Found preparing the control-node's convergence. Reach is settled independently by
the filter, the proxy's names and the certificate authority, so "this must not be
public" cannot be written and a public certificate is obtained regardless. And the
forward chain allows two hardcoded ranges plus a typed list, though the mesh
already knows which networks exist because its own modules declared them — a range
wide enough to keep four of them would have forwarded two predecessor leftovers too.
One container restarted 2286 times over five days while the mesh reported the machine as doing what it
was told. Its overlay address was five days out of date: the host compares a container by a digest of
its spec, and the mesh's names were not in it, so a container whose image and files never changed was
left alone holding a name that no longer resolved. Forty-eight others were current only because
something else had recreated them.
The same fault as issue 045, in the field that was left out. Resolved by putting the names in the
digest.
0120 was already accepted; what failed the check was that it rests on 0112, still marked proposed —
and so do four designs. The decision stands: a definition names no node, no mesh and no host path, and
everything a module needs is a requirement the mesh resolves.
Accepting it makes the gap visible rather than hiding it, so issue 134 states it. 0112 says how it is
checked — 'a catalogue test finds no domain name in any definition value' — and there is no such test.
Asked by hand: seven modules name this installation in a value the mesh acts on, and eight mention a
public name in prose nothing reads. The two are not the same fault and the fixes differ, which is why
the issue separates them rather than counting to fifteen.
Two of its statements are built — a version preparing its state, and the mesh saying what it applied —
so the document is in-progress rather than proposed, and names the code that owns them. Resting on a
live record rather than a superseded one: 0127 was replaced by 0131.
What this exposes is pre-existing: it also rests on ADR 0112, which is still proposed, and a document
that is not itself proposed may not. ADR 0120 has rested on it the same way for a while. Accepting or
superseding 0112 is a decision, not a cleanup, so it stays visible in the check rather than papered
over.
ADR 0135 made a step something the mesh derives for any module that prepares its state, which turned
ADR 0052's reach into a fault: a module whose database is briefly unreachable would stop every module
declared after it on that machine — the fault issue 011 already removed for every other shape, and
the reason the catalogue migrates itself at start rather than in a step.
A step now stops the rest of its own module and nothing else; an action still gates the machine,
because genesis is a row of them and they belong to no module. What was not attempted is reported as
skipped rather than left to be inferred from silence.
Two faults in 0133, both caught on review. It put the declaration on a container — one resource kind
the host applies — so every author would restate the machine's arrangement and a module's own
lifecycle would be tied to how its artifact happens to run. A module declares entrypoints for its
tools and its provisioner; preparing its state is the same vocabulary and nothing about a runtime.
And it derived the scope from the machine, which the facts already answer: a consumer is a module on
a machine (issue 022, migration 0015), so what the mesh provisions is per consumer. A module on three
machines has three databases, there is no shared state to race over, and the lock obligation 0133
invented was for a situation the mesh does not produce. The level question HAL answered with stages
dissolves — the scope of preparation is the scope of the state, and the mesh knows it.
0133 keeps its reasoning and gains a pointer; design 32 and issue 133 name the live record.
0133 — a module owns its migrations and the mesh owns when they run. A container declares what must
run before it; the mesh derives the gated step from the resource it precedes, so the image, the
environment and the credentials come from the one place they are described. The module owns the SQL,
the dialect and the lock; the mesh owns the moment and refuses to start a version whose step failed.
Per node, with no level: a step that ran once somewhere leaves every other machine ungated, and
'once, mesh-wide' is what holding a seat already means.
0134 — the pipeline is observable from a merge to an artifact and goes dark at the machine. What a
node now runs, and what it refused, become facts under the control plane's own seat, emitted when
what a machine runs changes rather than on every convergence pass.
Design 32's lifecycle carries both; issue 133 points at them as what ends the matter it opened.
The mesh replaced its own control plane with a build carrying a migration, applied none of it, and
then recorded no build for three quarters of an hour while saying everything was fine. ADR 0052
already prescribes the shape — a run-once step that gates the server — and the control plane was the
one module that did not use it.
A role's tools belong to the role, not to whichever module holds it today: the seat declares them
with their schemas, serving them is a condition of occupying the seat, and what the mesh can do
becomes a read of its own records rather than a question nothing answers. A module keeps its own
tools — the same module may run without the seat, and then only its own name is true.
Design 33 follows: the three families, addressing a node-scoped seat, discovery, and what serves
this to an agent.
Nine modules could not be rebuilt: their record named the repository and no directory, so every
build looked for a manifest at a repository root that has never had one. Resolved by mesh-controller
— `module add` takes the directory and the forge, and the rule is checked rather than described.
The AMQP transport is gone from the control plane and the hosts (mesh-controller #112,
mesh-host #39). On the way: no build had ever recorded its bases, so every bases-first order
walked an empty graph; the builder now reports what it was handed and the graph is read from
builds (mesh-controller #113/#114). Issue 131 is resolved by the forge module's merge event,
the control plane following it, and those edges.
Tasks 4.3, 5.2 and 5.4 are done as of 2026-09-28 02:25: every machine reports on
the new bus, the seat is held by the module that provides it, the old broker is
unassigned and forgotten, and every credential was minted afresh at the end.
5.2 records what it took, in the order it was found and each fixed on the trunk
before the next step, and how the bootstrap loop was broken once, by hand.
Until now the holder was derived — assigned and claiming — and a second eligible
assignment was refused, so a seat could not pass from one holder to the next
without a moment where nobody held it. The controller finds its own bus through
one of these seats, and that moment took the control plane down on 2026-09-27.
The holder is now a row the controller keeps, written by `seat <name> --to
<node>/<module>` in the same write that removes the previous one. No row means the
old rule, so nothing changes for a mesh that never hands a seat over; with a row,
another eligible assignment is silent rather than refused, which is what lets the
next holder run beside the current one until the switch. A holding is the
assignment's and goes when it does. Each rule names the test that checks it.
Under ADR 0131; design 28 task 5.3 is the work.
Taken during the outage of 2026-09-27, when the protocol leaked into the seat's
contract: to hold mesh-broker a module had to provide amqp, so the module that
will carry the bus could not hold the seat that names the bus, while the module
being retired could. Supersedes 0127. Modules depend on the seat and reach the
bus through the sdk; no manifest provides or requires amqp; the old broker's
module and the two modules that required it leave the catalogue; the AMQP
transport is deleted once every node reports on the new bus.
Design 28 step 5 rewritten under it: the seat handover becomes its own task and
is built first, because the seat the control plane dereferences cannot be empty
in between — that emptiness was the outage. The cost note now carries what was
measured rather than what was assumed.
0128 and 0130 extended 0127; each now rests on 0131 with a dated note and
changes nothing it decided. Every other citation of 0127 names its replacement.
records.py still fails on 0120/0112, which predates this branch.
Derived while merging and visible from no single repository, so it belongs
written down rather than re-derived later: the client library before the
catalogue, because it is where the subject is derived and a converted module
against the old one publishes the local name itself; the catalogue before the
controller, because the controller refuses an old-style name outright and would
make every unconverted module unregisterable.
Two tested properties are what make it safe and not merely ordered. An old-style
name passes through the derivation untouched, so an unconverted module keeps
working at every step. And a converted name derives to exactly the key the old bus
published, so nothing moves on the wire until 5.2 sets the variable.
The failure avoided is 127's own, which is why this is worth a table: a publisher
and a subscriber disagreeing about a subject log nothing anywhere.
Correcting design 26 to match what the merged code does, and fixing issue 112's
status, which used a word the vocabulary does not have.
A claim on a seat the manifest does not itself declare is refused at registration
rather than by the parser. A module may hold a seat another module declared —
that is why ADR 0126 has a caller name the seat and not its provider — so whether
the name exists is a fact about the whole catalogue.
And a declared seat may promise nothing. That is a marker seat, and most
node-scoped seats are markers: which module is this machine's packet filter. ADR
0126's "a declared seat carries a protocol" says what a holder must satisfy, not
that every seat offers something.
`records.py` still fails on ADR 0120 resting on a proposed ADR 0112, which is not
this branch's and not mine to decide.
Both lines of work numbered from the same point, so four decision records and one design
document existed twice with different content. The trunk keeps its numbers and this branch
yields — the only rule that scales, because the trunk's are already cited by what merged
before them.
0117 the bus is the only broker -> 0125
0118 a module declares its own seats -> 0126
0119 amqp is a provision, not the bus -> 0127
0120 the mesh bus is required -> 0128
0123 a seat carries its role's protocol -> 0129
0124 the predecessor is ending -> 0130
design 29, what a module declares -> design 32
Applied to the code repositories too, because a stale reference is worse when numbers
collide than when they dangle: the reader lands on a real record that decided something
else.
Two reconciliations the merge forced, both real:
**0110 was marked wholly superseded and was not.** Its successor says in as many words that
everything 0110 decided about what a seat *is* stands untouched — and two records that
landed on the trunk rest on exactly that part. So it is accepted again, extended rather than
replaced, with a note saying which of its claims moved and where.
**A seat's protocol becomes columns, not fields.** The trunk moved the seat set out of
compiled code into a table the controller owns. This branch had added what a role accepts,
emits and serves to the Go slice. The decision is unaffected and the mechanism is better for
it: giving a role a protocol is now a write rather than a rebuild, which is the trunk's own
argument applied to what this branch added.
One check still fails and it fails on main too: a record resting on ADR 0112 while that is
still 'proposed'. Left alone — it is not this merge's to answer.
ADR 0119 rejected giving the old broker a retirement condition and said why: "its
clients are not only the predecessor's, so the retirement condition describes a day
that will not come". The operator has said that day is coming — the predecessor is
deprecated, some of it still running, none of it being migrated, left to stop rather
than moved.
Recorded because three documents reason from the premise it overturns. Design 25 §5's
"no day anything is waiting for", §9's "the predecessor's clients never notice", and
design 28's closing note that the predecessor's world does not need to move.
**And it needs no new machinery, which is 0119 being paid off rather than revised.**
Because that record made the broker an ordinary provider rather than a compatibility
module, ending it is unassigning a provider whose provision nothing requires — something
the module system has done since it existed. So step 5.3 finishes instead of trailing
off, and the transitional double announcement of a build outcome has a date.
The consequence worth planning around: the predecessor's own mesh talks over that
broker, so shutting it down ends the tooling that reaches this installation's machines
from a workstation. The rollout is driven from the node, or driven before the broker
stops. That is a sequencing constraint on 5.2, not an afterthought.
What survives is `amqp` as a provision: a module that genuinely needs an AMQP broker can
still be given one. What retires is this broker's role as the predecessor's.
`rollout check` answers from records whether this mesh could move its bus, and names the
next step for each thing missing. The move itself waits on that check having been run
against a real mesh — writing the irreversible half before its question has ever been
asked of something real breaks the plan's own rule about beds by another route.
And the cost of being wrong is written down rather than assumed: the old broker stays for
its other clients, nothing in a served request's path goes over the mesh's own bus, and
what a failed move costs is the mesh's ability to change things rather than the services
its modules serve.
Three ways to do the replay were weighed and the right answer was that the bus being
moved to already does it. A queue on the old bus receives only what is published after
it is bound, so everything built before the catalogue existed was announced to nobody. A
stream is a log and a consumer is a position in it: a consumer created later starts at
the beginning and the builds are simply there. Checked against a running server, because
the decision rested on it.
So "who replays" has no answer because nothing replays. The mechanism was never about
builds — it was about a queue that could not remember, and carrying it across would have
carried a workaround for a limitation that no longer exists, with nothing looking wrong.
Retiring it belongs to step 5, with the rest of what only the old bus needs.
The account existed as a permission model and as nothing a person could be given; there
is a record and three commands now. The client is two surfaces over one thing, a command
line and an MCP server, both using the client a module's runtime uses — so what a person
may do is answered by the same permission list that answers it for a module.
Design 25 §7 says nothing of this is built before its bed passes, and this was built
before. Noted in the task rather than quietly ignored.
A node's intrusion filter should be composed from its assigned modules, like
its firewall (the Filtering mechanism): a service module (postgres, mssql,
mailu) declares its jail in its manifest (filter + stanza, no node/path per ADR
0112), and the mesh writes the jails of a node's modules into the fail2ban
holder's jail.d. The base (sshd, recidive, ignoreip=mesh-range) stays the
fail2ban module's. Records the model after novox's HAL per-module jails were
lost as dangling symlinks; the ignoreip is now safe on disk, the service jails
need this to be restored.
A foundation template that stands the server up, writes its settings and the mesh's
first user list beside them, and starts a controller on the new bus. The first user list
is the installer's because at genesis there is no mesh to compose one — a bootstrap
credential, rotated like the store's.
The carried list is checked against what the controller derives, since a mesh cannot be
raised twice to discover they disagreed. That check immediately found the composer
granting a role's whole event branch as well as the one event it follows.
What is left of 4.3 is running it, which is 4.1's bed.
Both sides behind a seam, one implementation per bus, and the outcome is the role's
own event so one publish reaches the asker, the controller and the catalogue. Checked
against a running server, including the part the decision rests on: a third party
hears the same outcome the asker does.
Reviews 0110/0121 after a session where renaming seats cost three freezes, a
builder deadlock, and hand-resolved manifests. The seat rules were right; the
set being a compiled Go slice referenced by name-string everywhere was the
mistake. Seats become a table keyed by a stable id; claims/held/production code
reference the id; a rename is one UPDATE, no rebuild, no re-registration, no
freeze. The build machine reads the set from the mesh instead of embedding it,
removing the controller/builder seat coupling. Closed set and scope naming
unchanged; only storage and reference change. Outstanding renames (registry
seats, private-network scope) wait for this — as data each is a write.
The mesh's own seats said who does a job and nothing about what may be said to
them or by them, and that gap showed up three times in one day looking like three
different problems: a build machine with three audiences for one outcome and no way
to derive a grant for any of them; an event genuinely about a role with nowhere to
live but the namespace of whichever module holds that role today; and a catalogue
catching up on builds, where every option needed a grant the design refuses.
One cause — the mesh has roles it cannot describe. So the `mesh-*` seats take the
same three fields a module's seat has, and the machinery that already derives
authority, queues and consumers from a declared seat does it for these too.
Builds become work submitted to a role, and `mesh.build.request`,
`mesh.control.built` and the BUILDS stream retire. A work queue shared by several
build machines is exactly what a seat's `accepts` is, so a second mechanism for it
was two places a permission could be wrong. The outcome is the seat's own event,
which means one publish still reaches whoever asked, the controller that records it
and the catalogue that places it — the fan-out a shared exchange gave for free,
written as a subject the mesh derived rather than a topology somebody configured.
That also avoids the grant that ruled out the alternatives: no holder needs
permission to publish into an asker's inbox.
The blocking gap is now named rather than incidental: the shared library has no way
for a module to publish on a seat. The build machine is Go and reaches the bus
directly, so it is unaffected; the artifact-store event waits.
Records the manual update process (module moved -> build -> reconcile; and the
breaking-change freeze/re-register recovery), and the two things that make
self-update more than a webhook: the build-on-push trigger is currently HAL's
(hal-gitea-tools on :9877), a retirement gap the mesh must replace with its own
forge-webhook trigger wired to every repo including mesh-controller; and the
builder validates manifests too, so a breaking change couples controller +
builder + manifests + hosts, and renaming the builder's own seat deadlocks its
rebuild. Names the transition discipline (accept old+new for one release) that
self-update needs so a push does not auto-freeze.
Every module named its events the way the old bus spelled a routing key, so on the
new bus every cross-module subscription pointed at a namespace nobody publishes to.
Nothing failed — the services started and none of them reacted. Converted, and the
rule now has checks at both scales: at registration for one manifest, and as a test
across the whole catalogue where a consumed event's emitter is present.
It was larger than the report said, in two directions nobody had looked. Forty-three
files of module code pass the event name at runtime, so the code mattered as much as
the manifests. And both clients had to learn the mapping — without that, converting
the modules would have broken the mesh that is actually running, which is the
opposite of what fixing this was for.
Design 29 gained three things it did not say: what a wildcard is (`*` for one name,
`**` for the rest, spelled the mesh's way and derived to each bus's own), that an
event about a role belongs on the seat and why that is not yet possible, and how the
rule is checked — because "a subscription that matches nothing is silence" is exactly
why nobody noticed thirty-seven manifests being wrong the same way.
4.2 and 4.3 are unblocked. The catch-up half of 4.5 is not: it is a decision, and it
narrowed rather than closed. It cannot be a reply to a module's inbox, because that
needs the blanket grant design 25 §4 refuses.