Verified on the machine: umami has zero restarts, applies all 26
migrations, reports the store up to date, and umami.novox.be answers 200
where it had answered 502 since 2026-09-25.
It was the same fault as 135, filed three days earlier and diagnosed
without either record noticing the other. 135's container IS umami — it
is named in 135's own evidence table, holding novox.internal:10.42.0.1
after the overlay range moved. The dial appeared to succeed and the first
real query timed out because the name pointed at an address that no
longer existed; the store, the path and the credential were all fine.
118's own reasoning is kept as a warning, because it is careful and
wrong: it argued path-MTU and conntrack, and it named the move that would
have found it — compare umami against a container created that week —
and did not make it. A stale name presents as a network fault, which is
why ADR 0148 stops copying names into containers rather than detecting
when the copies go bad.
Also: 118's located-in named mesh-catalog modules/umami, which was never
at fault; it names mesh-host's comparison now. And 135 gains the pointer
to 118, which is the back-reference I have now missed three times.
The certificate is genuine, from Mesh Internal CA, and nothing on the
workstation trusts it — verbatim the error the report gives. The public
name on the same proxy verifies cleanly, which puts the fault exactly
where the report puts it.
What is in the way is not an assignment. `ca-trust` is merged in the
catalogue and has never been registered with the mesh — 39 of 76
manifests are — so there is no module to assign. It dry-runs clean and
needs no artifact built.
Two findings from reproducing it, both their own issues:
157 — every routed name is published with an `.internal` alias that
nothing serves. The hosts file says keycloak.novox.be.internal; the proxy
serves keycloak.novox.internal and refuses the other by name. The first
three names I tried came from the hosts file and failed with a TLS alert
rather than a verification error, which pointed at a regression that had
not happened.
158 — the proxy re-logs all 52 routes every two seconds, 31 times a
minute. The one line that explained 157 sat between two of them.
Also recorded, because it was nearly filed as a defect and is not one:
step-ca publishes roots as /roots.pem, which is PEM, so ca-trust's fetch
and its refuse-a-non-certificate guard are both right. Its other endpoint
/roots returns JSON that contains the text the guard greps for, so the
guard is sound only because of which path is published.
Three records were left describing a mechanism a new record had moved,
and a reader arrives at them by following a citation: 0066 still said a
routed name is written into every container after 0148 replaced that with
resolution; 0016 still read as though the lab were the test bed after
0149; and issues 109 and 135 said nothing about 0148 ending the copying
that 135's own fix made comparable. Each was a citation leading to the
wrong answer in a record that was not wrong about anything it decided.
This is the second time in one session. The playbook rule I added last
round did not stop it, so the convention is now written where the record
conventions live, with the shape to use and three worked examples — and
with the honest note that it is NOT machine-checked and cannot be from
`extends:` alone: 102 records extend another, 87 have no back-reference,
and that is correct, because extending usually means building on a
context. Making it mechanical means a record declaring the relationship
in frontmatter, which is a schema change and is not mine to decide.
Also: designs 18 and 20 claimed `updated:` dates from before I edited
them, and 117's `fixed-by` gained the commit beside the record.
**0114 accepted.** The separation it draws — a consumer's resource and a
credential that reaches it are different things — is a data-loss rule,
and a record naming one should not sit unresolved while the code that
could hit it is written. Not built, and accepting it schedules nothing;
accepted-and-not-built is where 0141 and 0142 already are.
**0068 superseded by 0149, the live mesh is the test bed.** Never built,
and contradicted by practice that was written down nowhere but a handoff
note. The faults that cost the most are faults of a mesh that already
exists — bound consumers, containers made against an older roster, an
adopted machine — and a bed is by construction a mesh that does not.
Issue 156 settles it: that change was verified honestly, on the only path
where it cannot fail. The lab is not retired; raising a mesh from bare is
now its whole job.
**117 answered by 0150.** A module's own code runs as supervised
processes under the module's one account. 0047's argument was for a
runtime per module, not for a container — a unit satisfies it and runs as
an account. The invariant is the account, not the process count: its
worry was a second identity to scope and seal, and processes sharing one
account create none. Designs 18 and 20 now cite it and 0047 carries a
dated pointer, which closes the third disagreement — that neither design
knew the record existed.
Answers issue 151. Copying the roster into each container made the roster
part of each container's identity, so one name moving replaced every
container in the mesh — and it never stopped the staleness it was for,
since a copy taken at creation is stale the moment the roster moves (109,
135).
A container resolves through its machine's resolver instead, and nothing
is copied. Staleness stops being possible rather than detected, and a
name's blast radius becomes nothing.
Scoping each container to the names it binds was the close call and is
rejected: it contradicts anything-calls-anything, and leaves the roster
in the digest so the churn returns for a widely-bound name.
Gated on issue 110 — a container on the runtime's default network has no
DNS at all today. Removing the copy first reintroduces 109 and 135
silently on a live mesh. 151 stays open until the code lands; design 08's
file-not-resolver passage is narrowed to the machine's own roster.
Three issues named the branch that fixed them, and a branch is deleted
when it merges — so every `fixed-by:` was a pointer that resolved to
nothing by the time anyone followed it. They name commits and pull
requests now, and playbook 03 says to.
Two records were missing the thing a reader arrives for. 146 did not say
that one of its fixes crash-looped the control plane on a running mesh,
which is the whole reason the delivery subject carries the stream and the
raise path was the only one exercised. 151 did not say that 152 removed
the false reasons its roster moved, or that it stays open for the real
ones.
ADR 0080 enumerates what cycle.py enforces and named four things; it
enforces five. A progressive insight names the fifth — the decision
stands, the list had gone stale. The checks README and playbook 03 gained
the same rule, and 155 points at all three.
Measured on the broker's own config after the fixed controller composed
it: every node is now allowed _DELIVER.<node>.> as well as the bare
subject. That grant was missing when this was diagnosed, which is why
re-making the consumers would have silenced every machine.
Also records the five consumers it kept and reported, including the build
machine's worker, which is bound the same way and was not anticipated.
146's change is correct on a mesh being raised and fatal on one that is
running: the server will not move a push consumer's delivery subject
while a subscriber is bound, and a node is bound to its declaration
consumer the whole time it is up. Reproduced before fixing; an earlier
version of the test unsubscribed first and passed against the broken
code.
The consumer that works is kept and the assertion says so. Not re-made:
the nodes are granted the bare subject only, so re-making would have
silenced every machine.
Two machines filing issues in the same hour both read main correctly and
both took "the next free number". main lags every open pull request —
seven that evening — so they collided twice. The second collision
reached main with records, cycle and index all reporting success.
cycle.py now refuses a tree where two issue folders share a leading
number, and names both. Proven by adding a duplicate and watching it
fail. The colliding records become 153 and 154, renumbered in the branch
that lands last, because renumbering a branch whose author is still
pushing only moves the race.
The check catches a collision; it does not prevent one. Taking a number
still means reading the open pull requests as well as main — issue 155
says so.
It ran 22:46-23:11, five full replacements, and stopped when a pass
happened to read the store during a window it was up. The record said
nothing about it stops; that was wrong. Exiting by luck is the finding,
not a mitigation.
The three gatherers pass over a node whose set does not compose, and
raise anything else. 151 stays open: this removes the false reasons a
roster changes, not the fact that a real change still replaces every
container.
ace numbered its two records 147 and 148, which this repository already
uses; they become 150 and 151, as 127 became 149.
The new record is why the control node could not stop applying: the
roster alternates between two values because routeNamesInTheMesh
swallows a per-node plan failure, and the roster is part of every
container's identity. The loop closes through the control plane's own
store, which each pass replaces.
ADR 0138's reach is internal, public or both. ace's LAN devices (an IoT
switch on mosquitto, unifi access points, plex clients) are none of them;
the only reach that admits them is public, which states the wrong thing and
opens the port to the internet on any machine with a public address.
ADR 0112 decided that an assignment places a directory and says where an
access's data is; the control plane implements neither. ace's media stack
(a 40 TB ZFS library, plex state on a second disk) cannot be expressed
without writing ace's paths into manifests.
Both found migrating the first module on ace (searxng), adopted with route-adapter:
147 — assign is not a no-op for a routed module: its route reaches the predecessor's
proxy while its container is held, and the name serves a dead backend until take.
148 — every push that moved the mesh's names replaced every container on novox, twice,
the control plane's own store included.
It asked a question rather than reporting a defect, and the question was taken
five days later — the mesh's own components are binaries on the machine, and
third-party software stays a container because an image is the right way to
carry somebody else's build. The delivery of them is issue 142.
A number identifies a record, and two were given 127 on 2026-09-27. The one
three documents and three source files cite by number keeps it; the other
becomes 149, says so in its own heading, and its two inbound references are
repointed.
It is also resolved: an empty declaration is sent carrying owns_nothing rather
than skipped, and the host refuses an empty body that does not carry it, so
emptiness cannot be read as a truncated declaration.
0037 (where a module lives), 0113 (the vault makes every secret) and 0115 (one
assignment of a module per node) described arrangements the mesh has — a
catalogue of other people's software plus a table filled by module add, a vault
answering a secret requirement six modules make, and a rule the assignment
table's primary key already enforces. Each is accepted against what was built,
and says so in its own words.
0037's other half is not built: a manifest outside this catalogue has no check,
which is issue 148.
0068 (the lab takes requests) and 0114 (a shared credential rotates over two
credentials) stay proposed. Neither is built, and both are decisions rather
than records of something that happened.
The record catches README claiming this repository is indexed into a knowledge
base. Nothing indexed it, and since the cut-over there is nothing to index it
into — the surface that answered is on the transport the mesh removed (issue
147). Noted where the record is, so the next reader does not go looking for a
search that cannot exist.
088, 089, 120, 128 and 130 each name a commit that is on main and cites them —
the forge's address following a moved port, a route naming its endpoint, a
provisioner asking the backend what is there, the hosts file written into a
marked block, and undeclaring giving a unit back the state it was found in.
Each says it was closed by reading commits rather than by a run, so nobody
reads a green that was never measured.
129 stays located on purpose: ca-trust is merged and no machine holds it, so
the symptom it opened on is still true everywhere.
Written 2026-09-28, answered the same week by mesh-controller bdf965d and
c68d3a7, and never closed — so the mesh's own account said reach was declared
nowhere while three readers were reading it: the filter, the proxy's names, and
each authority's host policy. The resolution names them and what checks each.
What is left is retiring the older per-port keys the block replaces, which is
not a gap in what reach can say.
Diagnosed from the configuration: the tool server is HAL's brain, a local
process on the workstation with the predecessor's broker URL in the assistant's
own config. No manifest, no assignment, no seat, no account. Nothing regressed
— the mesh removed a transport this program still dials, and the program was
never part of the mesh. The mesh has a tool model and nothing publishes an
operator-facing surface onto it.
Every tool call fails with 'AMQP not connected', on every node including the
local one, because the tool surface still opens an AMQP connection and that
transport was deleted at the cut-over. The mesh reports healthy throughout —
what broke is the thing standing outside asking it questions, so nothing the
mesh checks is about it.
Not about enrolment. A push consumer delivers onto an ordinary subject and
everything subscribed to it gets a copy; the controller's two consumers were
both named after it, so both were given the same subject and the one process
acted on every message twice. Enrolment is where it drew blood because a second
enrolment mints a second credential.
Two more faults behind the three already fixed. The account a token is the
password of was never recorded, and the comment above the issuing code said
it was; issuing now records it. Placing the composed list at genesis is the
other half — the control plane says what it composed and whoever raises the
machine writes it beside the bus, because no declaration can reach a machine
that has not enrolled.
With that a first node enrols. It is then enrolled twice from one attempt,
each minting a credential, and it keeps the answer to the first while the mesh
keeps the second. The trail for that one stops at a duplicate that survived
message-id deduplication.
Raising a first node hits them in order: the bundle's bus image named for a
registry that is gone; the certificate made by openssl in an image that has
none; enrolment dialling TLS at a bus that speaks first, then refusing its own
token for the empty bus. Each was right until the bus changed and nothing has
raised a foundation since.
The fourth is not a patch: the installer carries the bus's first user list and
the controller composes the rest through the bus being a module, and at genesis
there is no module — so the first node cannot be let onto the bus it just
raised.
Found trying to run ADR 0147's bed. The older bundle raises the previous
broker and a control plane that refuses to start without MESH_BUS_NATS; the
newer one stops a step earlier, asking for a certificate from openssl in an
image that has none. No bed can run while this holds, so 0147's check section
now says what actually stands behind it — the rendering, not a machine.
Issue 129: every internal HTTPS name fails verification on every machine,
because nothing has ever written the mesh's root into a trust store. The
report proposed the controller inject it the way the private network writes
the registry's trust; this record rejects that — reachability and trust are
not the same fact, and where anchors live is the host's difference, not the
controller's. A module requiring internal-acme-ca does the whole of it, and
being unassigned undoes it.
0145's core stands — a module checks what the mesh claims, from where the callers
are — and everything it said about what to dial was wrong.
A raw port is not how anything in this mesh is reached. A real caller resolves a
name, the proxy answers, the proxy reaches the service; dialling a port tests the
last hop of a four-hop path and skips the three where most connectivity lives.
So each hosting form gets its own endpoint, route and name:
connect-docker.<node>.internal for a container, connect-process for the mesh's own
code in a unit it writes, connect-unit for a unit a package ships — and the same
under each public domain. Every machine checks every machine, by name, over TLS,
verifying the certificate against the authority that should have issued it. The
names are the instrument: a failure reads as connect-docker.g14.internal did not
answer.
No name is written anywhere: the machines come from the roster fact, the labels are
the module's, and a machine that joins is checked by the others on the next push.
Five things this needs that the mesh does not have, named rather than assumed —
a module every machine has (ScopeNode is exclusivity, not obligation), a container
running a bundle, a unit a package ships, the public domain in the roster fact, and
issue 129, until which every internal name will correctly fail verification.
The mesh asserts three things are callable and has never checked any of them. A
module on every machine serves an endpoint of its own and dials every other
machine's, from the position the callers are in — not the host and not the control
plane, both of which reach those addresses by paths no ordinary caller uses and
would have passed throughout the outage.
Its probe is its own endpoint, declared reachable over the private network, so it
is admitted by exactly the rule that governs every internally-exposed service. The
tempting target is a service every machine has, and those are the ones never closed
— ssh above all — which would have passed for all eleven hours.
Resolution and connection reported separately, because they have different owners.
One failure is not a fault, and the count travels with the result. It reports and
repairs nothing.
Built on a scheduled container, a rendered roster fact and a bundle: no new
vocabulary. Adopted on its own merits, which is what 0143 was not.