Measured on the broker's own config after the fixed controller composed
it: every node is now allowed _DELIVER.<node>.> as well as the bare
subject. That grant was missing when this was diagnosed, which is why
re-making the consumers would have silenced every machine.
Also records the five consumers it kept and reported, including the build
machine's worker, which is bound the same way and was not anticipated.
146's change is correct on a mesh being raised and fatal on one that is
running: the server will not move a push consumer's delivery subject
while a subscriber is bound, and a node is bound to its declaration
consumer the whole time it is up. Reproduced before fixing; an earlier
version of the test unsubscribed first and passed against the broken
code.
The consumer that works is kept and the assertion says so. Not re-made:
the nodes are granted the bare subject only, so re-making would have
silenced every machine.
Two machines filing issues in the same hour both read main correctly and
both took "the next free number". main lags every open pull request —
seven that evening — so they collided twice. The second collision
reached main with records, cycle and index all reporting success.
cycle.py now refuses a tree where two issue folders share a leading
number, and names both. Proven by adding a duplicate and watching it
fail. The colliding records become 153 and 154, renumbered in the branch
that lands last, because renumbering a branch whose author is still
pushing only moves the race.
The check catches a collision; it does not prevent one. Taking a number
still means reading the open pull requests as well as main — issue 155
says so.
It ran 22:46-23:11, five full replacements, and stopped when a pass
happened to read the store during a window it was up. The record said
nothing about it stops; that was wrong. Exiting by luck is the finding,
not a mitigation.
The three gatherers pass over a node whose set does not compose, and
raise anything else. 151 stays open: this removes the false reasons a
roster changes, not the fact that a real change still replaces every
container.
ace numbered its two records 147 and 148, which this repository already
uses; they become 150 and 151, as 127 became 149.
The new record is why the control node could not stop applying: the
roster alternates between two values because routeNamesInTheMesh
swallows a per-node plan failure, and the roster is part of every
container's identity. The loop closes through the control plane's own
store, which each pass replaces.
ADR 0138's reach is internal, public or both. ace's LAN devices (an IoT
switch on mosquitto, unifi access points, plex clients) are none of them;
the only reach that admits them is public, which states the wrong thing and
opens the port to the internet on any machine with a public address.
ADR 0112 decided that an assignment places a directory and says where an
access's data is; the control plane implements neither. ace's media stack
(a 40 TB ZFS library, plex state on a second disk) cannot be expressed
without writing ace's paths into manifests.
Both found migrating the first module on ace (searxng), adopted with route-adapter:
147 — assign is not a no-op for a routed module: its route reaches the predecessor's
proxy while its container is held, and the name serves a dead backend until take.
148 — every push that moved the mesh's names replaced every container on novox, twice,
the control plane's own store included.
It asked a question rather than reporting a defect, and the question was taken
five days later — the mesh's own components are binaries on the machine, and
third-party software stays a container because an image is the right way to
carry somebody else's build. The delivery of them is issue 142.
A number identifies a record, and two were given 127 on 2026-09-27. The one
three documents and three source files cite by number keeps it; the other
becomes 149, says so in its own heading, and its two inbound references are
repointed.
It is also resolved: an empty declaration is sent carrying owns_nothing rather
than skipped, and the host refuses an empty body that does not carry it, so
emptiness cannot be read as a truncated declaration.
0037 (where a module lives), 0113 (the vault makes every secret) and 0115 (one
assignment of a module per node) described arrangements the mesh has — a
catalogue of other people's software plus a table filled by module add, a vault
answering a secret requirement six modules make, and a rule the assignment
table's primary key already enforces. Each is accepted against what was built,
and says so in its own words.
0037's other half is not built: a manifest outside this catalogue has no check,
which is issue 148.
0068 (the lab takes requests) and 0114 (a shared credential rotates over two
credentials) stay proposed. Neither is built, and both are decisions rather
than records of something that happened.
The record catches README claiming this repository is indexed into a knowledge
base. Nothing indexed it, and since the cut-over there is nothing to index it
into — the surface that answered is on the transport the mesh removed (issue
147). Noted where the record is, so the next reader does not go looking for a
search that cannot exist.
088, 089, 120, 128 and 130 each name a commit that is on main and cites them —
the forge's address following a moved port, a route naming its endpoint, a
provisioner asking the backend what is there, the hosts file written into a
marked block, and undeclaring giving a unit back the state it was found in.
Each says it was closed by reading commits rather than by a run, so nobody
reads a green that was never measured.
129 stays located on purpose: ca-trust is merged and no machine holds it, so
the symptom it opened on is still true everywhere.
Written 2026-09-28, answered the same week by mesh-controller bdf965d and
c68d3a7, and never closed — so the mesh's own account said reach was declared
nowhere while three readers were reading it: the filter, the proxy's names, and
each authority's host policy. The resolution names them and what checks each.
What is left is retiring the older per-port keys the block replaces, which is
not a gap in what reach can say.
Diagnosed from the configuration: the tool server is HAL's brain, a local
process on the workstation with the predecessor's broker URL in the assistant's
own config. No manifest, no assignment, no seat, no account. Nothing regressed
— the mesh removed a transport this program still dials, and the program was
never part of the mesh. The mesh has a tool model and nothing publishes an
operator-facing surface onto it.
Every tool call fails with 'AMQP not connected', on every node including the
local one, because the tool surface still opens an AMQP connection and that
transport was deleted at the cut-over. The mesh reports healthy throughout —
what broke is the thing standing outside asking it questions, so nothing the
mesh checks is about it.
Not about enrolment. A push consumer delivers onto an ordinary subject and
everything subscribed to it gets a copy; the controller's two consumers were
both named after it, so both were given the same subject and the one process
acted on every message twice. Enrolment is where it drew blood because a second
enrolment mints a second credential.
Two more faults behind the three already fixed. The account a token is the
password of was never recorded, and the comment above the issuing code said
it was; issuing now records it. Placing the composed list at genesis is the
other half — the control plane says what it composed and whoever raises the
machine writes it beside the bus, because no declaration can reach a machine
that has not enrolled.
With that a first node enrols. It is then enrolled twice from one attempt,
each minting a credential, and it keeps the answer to the first while the mesh
keeps the second. The trail for that one stops at a duplicate that survived
message-id deduplication.
Raising a first node hits them in order: the bundle's bus image named for a
registry that is gone; the certificate made by openssl in an image that has
none; enrolment dialling TLS at a bus that speaks first, then refusing its own
token for the empty bus. Each was right until the bus changed and nothing has
raised a foundation since.
The fourth is not a patch: the installer carries the bus's first user list and
the controller composes the rest through the bus being a module, and at genesis
there is no module — so the first node cannot be let onto the bus it just
raised.
Found trying to run ADR 0147's bed. The older bundle raises the previous
broker and a control plane that refuses to start without MESH_BUS_NATS; the
newer one stops a step earlier, asking for a certificate from openssl in an
image that has none. No bed can run while this holds, so 0147's check section
now says what actually stands behind it — the rendering, not a machine.
Issue 129: every internal HTTPS name fails verification on every machine,
because nothing has ever written the mesh's root into a trust store. The
report proposed the controller inject it the way the private network writes
the registry's trust; this record rejects that — reachability and trust are
not the same fact, and where anchors live is the host's difference, not the
controller's. A module requiring internal-acme-ca does the whole of it, and
being unassigned undoes it.
0145's core stands — a module checks what the mesh claims, from where the callers
are — and everything it said about what to dial was wrong.
A raw port is not how anything in this mesh is reached. A real caller resolves a
name, the proxy answers, the proxy reaches the service; dialling a port tests the
last hop of a four-hop path and skips the three where most connectivity lives.
So each hosting form gets its own endpoint, route and name:
connect-docker.<node>.internal for a container, connect-process for the mesh's own
code in a unit it writes, connect-unit for a unit a package ships — and the same
under each public domain. Every machine checks every machine, by name, over TLS,
verifying the certificate against the authority that should have issued it. The
names are the instrument: a failure reads as connect-docker.g14.internal did not
answer.
No name is written anywhere: the machines come from the roster fact, the labels are
the module's, and a machine that joins is checked by the others on the next push.
Five things this needs that the mesh does not have, named rather than assumed —
a module every machine has (ScopeNode is exclusivity, not obligation), a container
running a bundle, a unit a package ships, the public domain in the roster fact, and
issue 129, until which every internal name will correctly fail verification.
The mesh asserts three things are callable and has never checked any of them. A
module on every machine serves an endpoint of its own and dials every other
machine's, from the position the callers are in — not the host and not the control
plane, both of which reach those addresses by paths no ordinary caller uses and
would have passed throughout the outage.
Its probe is its own endpoint, declared reachable over the private network, so it
is admitted by exactly the rule that governs every internally-exposed service. The
tempting target is a service every machine has, and those are the ones never closed
— ssh above all — which would have passed for all eleven hours.
Resolution and connection reported separately, because they have different owners.
One failure is not a fault, and the count travels with the result. It reports and
repairs nothing.
Built on a scheduled container, a rendered roster fact and a bundle: no new
vocabulary. Adopted on its own merits, which is what 0143 was not.
Everything should be able to call what runs on the same machine, another machine's
service exposed to the private network, and another machine's service exposed
publicly. Three cases; the filter had two.
The first was expressed as the machines' own addresses on the private network. A
caller on the machine carries such an address; a caller in one of its containers
carries a bridge address and matched nothing — measured, same destination and same
machine: src 10.10.0.1 against src 172.17.0.8. The second case worked by accident,
because the tunnel rewrites a caller's address to the sending machine's. Two of
three working is why it read as correct.
0143 answered the wrong question. It proposed verifying each grant from the
consumer's own network position and went to length about which position, because
whether a caller sat in a container changed the answer — and that difference was the
bug. Observing a configuration error is not its remedy. Superseded, and nothing
replaces it; whether the mesh should check a grant is still open in issue 145 and
must stand on its own.
And a module is not a container: 61 of 72 happen to use one, 11 do not, and a rule
reasoning about containers describes most of the mesh rather than the mesh.
A grant is four facts and a credential — the provision, the machine, the port, and
who the consumer is when it connects — and it is the whole mechanism by which
anything in the mesh reaches anything else. The mesh asserted it and never found
out whether it was true. Issue 145 is what that cost: eleven hours of 'every module
current with its source' while a module could not reach its database.
The consumer verifies it, from its own network position. Not the control plane,
which reaches the address by a path no consumer uses. Not the machine, whose own
packets carried a source address the filter admitted while every container's was
refused — so a host-side check would have passed throughout the outage it exists to
catch. That is inference from the rule that was loaded, stated as such.
A connection and nothing more; speaking each provision's protocol would be a second
implementation of every provision. One failure is not news, only consecutive ones,
and the count is reported rather than the last attempt. A consumer that is not
running reads unchecked, which is a different sentence from broken.
What it costs to be wrong is the constraint on all of it: a check that calls a
working provision broken trains a reader to ignore the report.
Converging the control node closed every path by which a module reached another by
the machine's own name, and it ran for eleven hours while the mesh answered 'all
doing what they were told, all heard from, running what the mesh would send them,
and every module current with its source'.
6,154 database failures in one module's log, beginning at the minute of the flip.
The service accepted TCP and never answered HTTP. Every check the mesh makes
passed, because every check the mesh makes is about the relationship between the
mesh and a machine — applied, current, containers running. None asks whether a
module can reach what it requires, though the mesh composes every grant and so
knows exactly who requires what.
The filter fault is fixed. The eleven hours are the measurement, not the bug.
Also records on 144 that 'more closed than the mesh believes' was harmless only
for what is reached from outside, and an outage for what is reached from within.
The record says internal and public each mean something to the filter. For an
endpoint the proxy serves, the second half is wrong, and ADR 0045 said so first:
a public service is exposed through the proxy, listening from the mesh, not by
opening its own port.
Found by trying to express one real module, not by review — routed name public
because browsers post to it, machine port private because it serves a dashboard in
cleartext. Under one value for both, saying public would have reopened a port an
operator had just closed. Measured the same evening: the routed name answered from
the internet over TLS while the port was refused from the same place.
Corrects a fact. One statement per endpoint with three things derived from it
stands; the filter column applies to an unrouted endpoint.
The first account said the declaration carries no resource that would disable the
found firewall and that nothing implemented the sentence the flip prints. Wrong:
the mechanism is a step in the host's own apply, retireFirewall, and it is careful
— it reads back that the mesh's table is loaded before retiring anything, records
the forward policies first, and verifies ufw reports inactive afterwards.
What is established: ufw was active and enabled two minutes after a flip that
reported the node converged with nothing failed; and the machine's record now
reads disabled_by_mesh: true, written by a reconcile fifty minutes later that
found ufw already inactive because an operator had disabled it by hand.
So the step did not take effect and the record says it did. The candidates are
named rather than chosen — the step is skipped silently when the apply has any
failure, the mesh's table is loaded by a service in the same apply so ordering is
open, and the host's detail lines do not reach the journal, which is why this has
candidates instead of a cause.