The work order's group-3 question answered: an ordinary module the mesh assigns to the machine a
person sits at, holding a minted credential, calling tools under a manifest grant (invokes), serving
MCP on loopback. Design 34; pointers in 33, 25 and 0095; module check designed into 12 (issue 148);
README stops claiming an indexing nothing provides (issue 006).
Reported by the operator: ssh by the bare name logs in, by the mesh name
is refused. The predecessor's generator writes the bare name only; the
mesh's ssh-client roster already matches both and is not yet shipped.
Found and fixed the afternoon ADR 0148 landed: mailu-admin lost its
database behind Mailu's own resolver. Two catalogue PRs; an insight on
0148 that a container's dns is a decision, not a preference.
Issue 110's cause was not the filter: the runtime had never been told,
and the resolver dropped a query arriving on a bridge. ADR 0148 step 3
landed once it did (109, 151 resolved). ADR 0151 composes a route's
internal name under the serving node and drops the suffixed alias
(139, 157 resolved). Design 08 amended; a fact in 0148 corrected.
Hosts first, then the controller — a build and a push each, now that the
mesh delivers the host. The host refuses a lower sequence than it kept
and drains a batch by sequence rather than arrival; the controller
numbers each send under the node's hold, inside the signed bytes.
Measured: two pushes, sequence 2 in the kept declaration, counters in
the store agree, no machine reads as behind. That last one is the
subtlety: the mesh compares the digest of what it would send against
what it did, and a number changes the bytes, so the read-only comparison
composes with the last number sent rather than a fresh one.
All four machines run a host the mesh built, published and delivered, the
last delivery unattended: each stood aside once for a genuinely newer
version and the delivered launcher started it. A following push that
delivered nothing new was applied and reported by every machine and stood
nobody aside.
The crossover needs one restart of the unit per machine, once, because
the running launcher executes from its own inode. Measured timing: three
seconds on the machine, 17-20 as the operator sees it, the difference
being the control plane composing before it sends.
107 is unblocked: a declaration field is now a build and a push.
Asked whether a newer host was delivered using the link-time stamp, which
every delivered host carries as 'development build' now that the version
comes from where the binary sits. Never matched, so it stood aside on
every push for ever; standing aside cancels the report, so the mesh never
heard from it. Read as healthy throughout.
The three-minute push wait made it invisible: a wait long enough to
absorb a whole apply is long enough to hide that the machine never
answered.
161 resolved and verified on a machine: the workstation runs a host the
mesh compiled, published, delivered and started, applying declarations
and reporting the version it was delivered as.
The system it was built for comes from the artifact — the one thing a
toolchain takes from a module, which 0142 already allowed because the
target is a property of the artifact. The version comes from where the
binary sits, which 0142 decided and nothing had implemented.
Two mistakes on the way, both caught by reading the output rather than
the line that claimed success. A second -ldflags does not merge with the
first: the binary gained its system and lost -s -w, 12.2MB against 8.5MB.
And the delivered binary was named after its package, so the first
delivery was correct, reported success and was invisible to the launcher.
A delivered host that cannot apply is a machine the mesh cannot repair,
because the declaration that would fix it is the one it cannot apply. The
launcher's fallback is what made that an inconvenience instead of an
expedition.
162 is new and not about the host: an archive has no removal, so a module
using one can never be unassigned, and the attempt takes the whole apply
with it — the machine applies nothing else either. It is how undoing the
first delivery froze the workstation.
The loop closed on the workstation: the version landed, the launcher was
replaced, the running host stood aside, and after one restart the launcher
started a binary the mesh had compiled, published and delivered.
The launcher goes as a file resource rather than inside the archive, and
that is the safety rather than a preference. A file is written atomically,
so the running launcher keeps the inode it started from; an archive writes
in place with truncate and would cut a script a shell is reading. The
manifest carries a second copy and a test refuses any drift from the one
in packaging.
Then it would have refused the first declaration it was asked to apply.
The Makefile links in two facts the mesh's toolchain does not, on purpose,
and one of them is the system the host was built for — read before
anything is applied, so the failure is safe and total. Nothing reports it:
the unit is active, the bus link is up, and the log says it is hearing
what the node should be.
Worse, the declaration that would fix it is the declaration it cannot
apply, so the mesh cannot repair such a machine. Restored by moving the
delivered versions aside and letting the launcher fall back, which is the
fallback working as designed.
0142 already settles the version — it comes from where the component sits,
not from its linker — and that is unimplemented. The system pin has no
answer, and the candidates are a decision rather than a fix: put it in the
path too, carry it in a file beside the binary, or stop pinning at link
time at all, which is 0005's to change.
Filed as a to-do. Nothing is broken by it: every machine here is amd64 and
reports so, and one architecture is enough for now.
When a machine joins, the mesh should collect what it reasonably can about
it and refresh that daily. It already asks what a machine can do; what it
is made of is the same question one level down.
More is already collected than it looks — eight capabilities, the links
that face outside, the host version, and on an adopted machine what it
holds, what is reachable and the firewall and tunnel it was found with.
The architecture and the kernel are in there too, and nothing reads
either: measured, all four machines report amd64 and linux, and node show
prints the capabilities beside them without printing them.
Missing: memory, disk, the processor beyond its architecture, the
distribution and its version, virtual or physical, cores, uptime. Several
are what somebody wants when deciding where a module goes, and the
placement code's own comment already imagines them.
Also missing: the refresh. A machine publishes after an apply, and the
five-minute reconcile publishes nothing, so the mesh's picture is as old
as the last push. Same mechanism 087 wanted.
This is the third thing in one day found to be collected and read nowhere,
after held resources and the host version. Whatever gets added should say
in the same breath which surface shows it, or it will be the fourth.
159 gains the note that the architecture is already reported, so matching
an artifact to a machine needs no new fact — only the comparison and a
compiler told what to target.
Asked whether the host is built for more than one architecture. It is
not, and the reason is worse than a missing feature.
A bundle in a compiled language must name a system, must name one of
alpine, android or arch, and is refused with a careful message if it gets
that wrong. The field is then read by nothing: it does not reach the
compiler, no machine is matched against it, and nothing chooses between
two artifacts by it. The compile runs with no target named and produces a
binary for whatever the build machine happens to be.
The host is x86-64 because the build machine is, not because the
declaration said so. Correct for this mesh by coincidence — four machines,
all x86-64 Arch.
A module declaring two systems would get two identical binaries, both
published and both pinned, and the one sent to the machine it was not
built for would fail at exec. android is the sharp end: not an x86-64
platform, and an artifact declared for it today would be an x86-64 binary
wearing the label.
A field that is checked and ignored is worse than one that does not
exist, because the check is what persuades you it works.
Also noted: the processor is a second dimension the manifest has no word
for, so even implementing the present field would not answer the question
that found this. And since the Go toolchain builds statically, one binary
would run on all three systems anyway — so the pin is a policy rather
than a necessity, which is a decision and not a fix.
Both of the things ADR 0141's insight named as remaining are built. A Go
toolchain based on a new mesh-tools-go module, so the compiler is named
and not pinned; and ${version} in any value of a resource that uses an
archive or a bundle.
Measured rather than asserted: the mesh built the host through its own
toolchain, published it to its own registry, and the bundle fetched back
out is a statically linked stripped binary that runs and says it is the
host.
The cost was larger again than 0141's note said. Three more things in the
path assumed one language or one shape — an entrypoint became a .ts file
whatever the language, the output directory was the compiler's to create,
and a bundle was refused if it named what it is built from — and a fourth
was in the base image, which is Alpine where the first Dockerfile ran
apt-get. That last one is issue 136 in an image, and the build refused
rather than a module failing later.
The version in a path is the digest, not the commit: two builds of one
commit are the same bytes, so a content-addressed version keeps the path
an unchanged build already had.
Still nothing delivers a version to a machine. The host module declares no
resources, so the bundle sits in the registry and no machine is asked to
take it. 0141 carries the insight and 142 the account.
It reported "N machines run an older host than another" by comparing
versions as strings. A host reports its version as a commit, and commits
have no order. On the live mesh it named the three machines running the
NEWER host as the ones behind — ced54d4 sorts above 04a27ca and that is
all it means.
The code even carried a caveat saying versions compare as strings and that
this "is enough for the timestamps and commits this mesh uses". That was
the error, written down and not noticed: enough for timestamps, meaningless
for commits, and the mesh reports commits.
It now reports the split and claims no ordering, which is more useful as
well as more honest — the reader sees who is on which side, and that is
what decides whether a field can be sent. Ordering is left with the host,
which would have to report something ordered for anybody to have it.
This is issue 145 arriving by my own door an hour after I closed it: a
report that confidently says the opposite of the truth is worse than one
that says less.
Measured live. All four machines run the identical host binary — same
digest, installed within eighteen seconds — and at first only one reported
a version, which read as a difference where there was none.
A machine publishes a report after an apply. The five-minute reconcile
publishes nothing, because it is the machine keeping itself as declared
rather than answering anything. So a current, idle machine never says, and
the mesh cannot tell that from a machine running something ancient.
Confirmed by pushing: not reported, then 04a27ca.
Enough for the purpose, not enough for the claim. For deciding whether a
new declaration field is safe it is sufficient — pushing is what the mesh
is about to do, and the answer arrives with the act. For knowing what the
mesh runs it is not, and "not reported" is worded as "nobody has asked
recently" for that reason.
Making a heartbeat carry it would close the gap and would change what a
heartbeat is — a bare word that the node is there, deliberately carrying
nothing else. Left alone rather than widened in passing.
145, partly resolved. The sentence that was true for eleven hours of a
mesh in which no module could reach another now says what it is not a
claim about: that is the mesh and the machines agreeing, and nothing here
dials a provision. It checks nothing and does not pretend to — ADR 0146
decides the check and is deliberately not built. What changed is that the
report no longer implies otherwise. Stays open for that reason.
Carried forward: 0146's check needs an internal name fetched over TLS with
the certificate verified, and until today no machine trusted the mesh's
authority. Three of four do now, so whoever builds it does not have to
solve that first.
107, diagnosed and deliberately not built. The premise is confirmed in the
host's own words — unknown fields are refused because "a field the host
does not know is a thing the control plane believes it asked for" — so the
fix is a flag day, not an addition. 087 now makes the cost measurable, and
the measurement is why it waits: one machine of four runs an older host,
it is parked, and nothing delivers a host at all (142). Shipping the field
means hand-placing binaries and unparking a machine, and one missed in
that sequence is unreachable, not degraded. The fault it prevents has
never been observed.
142 gains the note that it is 107's gate, and that it is what makes a
declaration field cost a rollout instead of an expedition.
A judgement about order, not a refusal, and cheap to overrule.
The machine has reported its host version since ADR 0141, whose own
comment says why it must: without it nothing can say a machine is behind.
The controller's copy of the report did not have the field, so it
unmarshalled into nothing and was thrown away on arrival. Two structs
describe one message and only the sending side had it.
node show names it per machine, "not reported" where the mesh has not been
told. status names every machine running an older host than another does,
and which is newest.
Disagreement rather than staleness, deliberately: nothing delivers a host
version yet, so the mesh holds no canonical current one and "behind" has
no fixed point. What it can say is that the oldest host in the mesh is
what the mesh may send.
Two refusals to guess: a machine that reported nothing is not called
behind, and versions compare as strings — right for the timestamps this
mesh uses, wrong for a scheme where 10 sorts before 9, said at the place
that would have to learn.
107 gains the note that this is what makes its new field safe to consider,
and that one machine of four is behind today, so it is not free yet.
Two of the four surfaces the report named already carried it — the host
has reported Held since ADR 0100, and node show reads the machine's own
list with an `as of` beside it. Recorded as checked rather than assumed.
Two did not. The apply line counted what it applied and said nothing
about the difference; status read the mesh's take-time listing, so a
module assigned after it showed nothing at all.
Both now say it, and the part that carries the weight: a hold suppresses
"all doing what they were told, all heard from, running what the mesh
would send them". That sentence was true for the whole outage, and acting
on it is what stopped the predecessor's proxy. Being adopted still does
not suppress it — a mode somebody chose is not a half-finished action.
Status does not call a hold a fault, deliberately. It is correct
behaviour, and a reader trained to see red for something the mesh did
right stops reading.
Extended after the first was proven. Every converged machine now holds
the anchor and verifies an internal name with a plain client; before,
the two unassigned ones answered 'unable to get local issuer
certificate' and held no entry for the mesh.
ace is excluded on purpose: it is adopted, so a module assigned there is
held rather than run, which is right and is not trust.
Both the resolution and 0147's insight said one machine of four, which
was true for about twenty minutes.
Registered ca-trust from the catalogue — it was merged and had never been
registered, which is why "assign it to one machine" had no module to name
— assigned it to the workstation, and verified.
Verified in the form ADR 0147 prescribes, against the authority's own API
so the handshake needs nothing else in the mesh to be right: 200, issuer
Mesh Internal CA, Verify return code 0. Four routed internal names verify
too, and `git ls-remote https://…` works, which is the consequence the
report named.
Removal exercised for the first time. Unassign and push removes the
anchor, empties the trust store of the mesh's authority, and returns the
plain client to the original error; assigning again restores it. That is
the half 0147 claimed and nothing had shown.
One thing it found that is not in the module: removal works only because
the host removes the service before the script. Stopping the unit is what
deletes the certificate and refreshes the bundles, and it needs the script
to still exist. The symmetry rests on an ordering nothing states.
0147's "written, and not yet run" now carries a progressive insight
saying it has run, where, and that it ran on one machine of four.
Verified on the machine: umami has zero restarts, applies all 26
migrations, reports the store up to date, and umami.novox.be answers 200
where it had answered 502 since 2026-09-25.
It was the same fault as 135, filed three days earlier and diagnosed
without either record noticing the other. 135's container IS umami — it
is named in 135's own evidence table, holding novox.internal:10.42.0.1
after the overlay range moved. The dial appeared to succeed and the first
real query timed out because the name pointed at an address that no
longer existed; the store, the path and the credential were all fine.
118's own reasoning is kept as a warning, because it is careful and
wrong: it argued path-MTU and conntrack, and it named the move that would
have found it — compare umami against a container created that week —
and did not make it. A stale name presents as a network fault, which is
why ADR 0148 stops copying names into containers rather than detecting
when the copies go bad.
Also: 118's located-in named mesh-catalog modules/umami, which was never
at fault; it names mesh-host's comparison now. And 135 gains the pointer
to 118, which is the back-reference I have now missed three times.
The certificate is genuine, from Mesh Internal CA, and nothing on the
workstation trusts it — verbatim the error the report gives. The public
name on the same proxy verifies cleanly, which puts the fault exactly
where the report puts it.
What is in the way is not an assignment. `ca-trust` is merged in the
catalogue and has never been registered with the mesh — 39 of 76
manifests are — so there is no module to assign. It dry-runs clean and
needs no artifact built.
Two findings from reproducing it, both their own issues:
157 — every routed name is published with an `.internal` alias that
nothing serves. The hosts file says keycloak.novox.be.internal; the proxy
serves keycloak.novox.internal and refuses the other by name. The first
three names I tried came from the hosts file and failed with a TLS alert
rather than a verification error, which pointed at a regression that had
not happened.
158 — the proxy re-logs all 52 routes every two seconds, 31 times a
minute. The one line that explained 157 sat between two of them.
Also recorded, because it was nearly filed as a defect and is not one:
step-ca publishes roots as /roots.pem, which is PEM, so ca-trust's fetch
and its refuse-a-non-certificate guard are both right. Its other endpoint
/roots returns JSON that contains the text the guard greps for, so the
guard is sound only because of which path is published.
Three records were left describing a mechanism a new record had moved,
and a reader arrives at them by following a citation: 0066 still said a
routed name is written into every container after 0148 replaced that with
resolution; 0016 still read as though the lab were the test bed after
0149; and issues 109 and 135 said nothing about 0148 ending the copying
that 135's own fix made comparable. Each was a citation leading to the
wrong answer in a record that was not wrong about anything it decided.
This is the second time in one session. The playbook rule I added last
round did not stop it, so the convention is now written where the record
conventions live, with the shape to use and three worked examples — and
with the honest note that it is NOT machine-checked and cannot be from
`extends:` alone: 102 records extend another, 87 have no back-reference,
and that is correct, because extending usually means building on a
context. Making it mechanical means a record declaring the relationship
in frontmatter, which is a schema change and is not mine to decide.
Also: designs 18 and 20 claimed `updated:` dates from before I edited
them, and 117's `fixed-by` gained the commit beside the record.
**0114 accepted.** The separation it draws — a consumer's resource and a
credential that reaches it are different things — is a data-loss rule,
and a record naming one should not sit unresolved while the code that
could hit it is written. Not built, and accepting it schedules nothing;
accepted-and-not-built is where 0141 and 0142 already are.
**0068 superseded by 0149, the live mesh is the test bed.** Never built,
and contradicted by practice that was written down nowhere but a handoff
note. The faults that cost the most are faults of a mesh that already
exists — bound consumers, containers made against an older roster, an
adopted machine — and a bed is by construction a mesh that does not.
Issue 156 settles it: that change was verified honestly, on the only path
where it cannot fail. The lab is not retired; raising a mesh from bare is
now its whole job.
**117 answered by 0150.** A module's own code runs as supervised
processes under the module's one account. 0047's argument was for a
runtime per module, not for a container — a unit satisfies it and runs as
an account. The invariant is the account, not the process count: its
worry was a second identity to scope and seal, and processes sharing one
account create none. Designs 18 and 20 now cite it and 0047 carries a
dated pointer, which closes the third disagreement — that neither design
knew the record existed.
Answers issue 151. Copying the roster into each container made the roster
part of each container's identity, so one name moving replaced every
container in the mesh — and it never stopped the staleness it was for,
since a copy taken at creation is stale the moment the roster moves (109,
135).
A container resolves through its machine's resolver instead, and nothing
is copied. Staleness stops being possible rather than detected, and a
name's blast radius becomes nothing.
Scoping each container to the names it binds was the close call and is
rejected: it contradicts anything-calls-anything, and leaves the roster
in the digest so the churn returns for a widely-bound name.
Gated on issue 110 — a container on the runtime's default network has no
DNS at all today. Removing the copy first reintroduces 109 and 135
silently on a live mesh. 151 stays open until the code lands; design 08's
file-not-resolver passage is narrowed to the machine's own roster.
Three issues named the branch that fixed them, and a branch is deleted
when it merges — so every `fixed-by:` was a pointer that resolved to
nothing by the time anyone followed it. They name commits and pull
requests now, and playbook 03 says to.
Two records were missing the thing a reader arrives for. 146 did not say
that one of its fixes crash-looped the control plane on a running mesh,
which is the whole reason the delivery subject carries the stream and the
raise path was the only one exercised. 151 did not say that 152 removed
the false reasons its roster moved, or that it stays open for the real
ones.
ADR 0080 enumerates what cycle.py enforces and named four things; it
enforces five. A progressive insight names the fifth — the decision
stands, the list had gone stale. The checks README and playbook 03 gained
the same rule, and 155 points at all three.
Measured on the broker's own config after the fixed controller composed
it: every node is now allowed _DELIVER.<node>.> as well as the bare
subject. That grant was missing when this was diagnosed, which is why
re-making the consumers would have silenced every machine.
Also records the five consumers it kept and reported, including the build
machine's worker, which is bound the same way and was not anticipated.
146's change is correct on a mesh being raised and fatal on one that is
running: the server will not move a push consumer's delivery subject
while a subscriber is bound, and a node is bound to its declaration
consumer the whole time it is up. Reproduced before fixing; an earlier
version of the test unsubscribed first and passed against the broken
code.
The consumer that works is kept and the assertion says so. Not re-made:
the nodes are granted the bare subject only, so re-making would have
silenced every machine.
Two machines filing issues in the same hour both read main correctly and
both took "the next free number". main lags every open pull request —
seven that evening — so they collided twice. The second collision
reached main with records, cycle and index all reporting success.
cycle.py now refuses a tree where two issue folders share a leading
number, and names both. Proven by adding a duplicate and watching it
fail. The colliding records become 153 and 154, renumbered in the branch
that lands last, because renumbering a branch whose author is still
pushing only moves the race.
The check catches a collision; it does not prevent one. Taking a number
still means reading the open pull requests as well as main — issue 155
says so.
It ran 22:46-23:11, five full replacements, and stopped when a pass
happened to read the store during a window it was up. The record said
nothing about it stops; that was wrong. Exiting by luck is the finding,
not a mitigation.
The three gatherers pass over a node whose set does not compose, and
raise anything else. 151 stays open: this removes the false reasons a
roster changes, not the fact that a real change still replaces every
container.
ace numbered its two records 147 and 148, which this repository already
uses; they become 150 and 151, as 127 became 149.
The new record is why the control node could not stop applying: the
roster alternates between two values because routeNamesInTheMesh
swallows a per-node plan failure, and the roster is part of every
container's identity. The loop closes through the control plane's own
store, which each pass replaces.
ADR 0138's reach is internal, public or both. ace's LAN devices (an IoT
switch on mosquitto, unifi access points, plex clients) are none of them;
the only reach that admits them is public, which states the wrong thing and
opens the port to the internet on any machine with a public address.
ADR 0112 decided that an assignment places a directory and says where an
access's data is; the control plane implements neither. ace's media stack
(a 40 TB ZFS library, plex state on a second disk) cannot be expressed
without writing ace's paths into manifests.
Both found migrating the first module on ace (searxng), adopted with route-adapter:
147 — assign is not a no-op for a routed module: its route reaches the predecessor's
proxy while its container is held, and the name serves a dead backend until take.
148 — every push that moved the mesh's names replaced every container on novox, twice,
the control plane's own store included.
It asked a question rather than reporting a defect, and the question was taken
five days later — the mesh's own components are binaries on the machine, and
third-party software stays a container because an image is the right way to
carry somebody else's build. The delivery of them is issue 142.
A number identifies a record, and two were given 127 on 2026-09-27. The one
three documents and three source files cite by number keeps it; the other
becomes 149, says so in its own heading, and its two inbound references are
repointed.
It is also resolved: an empty declaration is sent carrying owns_nothing rather
than skipped, and the host refuses an empty body that does not carry it, so
emptiness cannot be read as a truncated declaration.
0037 (where a module lives), 0113 (the vault makes every secret) and 0115 (one
assignment of a module per node) described arrangements the mesh has — a
catalogue of other people's software plus a table filled by module add, a vault
answering a secret requirement six modules make, and a rule the assignment
table's primary key already enforces. Each is accepted against what was built,
and says so in its own words.
0037's other half is not built: a manifest outside this catalogue has no check,
which is issue 148.
0068 (the lab takes requests) and 0114 (a shared credential rotates over two
credentials) stay proposed. Neither is built, and both are decisions rather
than records of something that happened.
The record catches README claiming this repository is indexed into a knowledge
base. Nothing indexed it, and since the cut-over there is nothing to index it
into — the surface that answered is on the transport the mesh removed (issue
147). Noted where the record is, so the next reader does not go looking for a
search that cannot exist.
088, 089, 120, 128 and 130 each name a commit that is on main and cites them —
the forge's address following a moved port, a route naming its endpoint, a
provisioner asking the backend what is there, the hosts file written into a
marked block, and undeclaring giving a unit back the state it was found in.
Each says it was closed by reading commits rather than by a run, so nobody
reads a green that was never measured.
129 stays located on purpose: ca-trust is merged and no machine holds it, so
the symptom it opened on is still true everywhere.
Written 2026-09-28, answered the same week by mesh-controller bdf965d and
c68d3a7, and never closed — so the mesh's own account said reach was declared
nowhere while three readers were reading it: the filter, the proxy's names, and
each authority's host policy. The resolution names them and what checks each.
What is left is retiring the older per-port keys the block replaces, which is
not a gap in what reach can say.
Diagnosed from the configuration: the tool server is HAL's brain, a local
process on the workstation with the predecessor's broker URL in the assistant's
own config. No manifest, no assignment, no seat, no account. Nothing regressed
— the mesh removed a transport this program still dials, and the program was
never part of the mesh. The mesh has a tool model and nothing publishes an
operator-facing surface onto it.
Every tool call fails with 'AMQP not connected', on every node including the
local one, because the tool surface still opens an AMQP connection and that
transport was deleted at the cut-over. The mesh reports healthy throughout —
what broke is the thing standing outside asking it questions, so nothing the
mesh checks is about it.
Not about enrolment. A push consumer delivers onto an ordinary subject and
everything subscribed to it gets a copy; the controller's two consumers were
both named after it, so both were given the same subject and the one process
acted on every message twice. Enrolment is where it drew blood because a second
enrolment mints a second credential.
Two more faults behind the three already fixed. The account a token is the
password of was never recorded, and the comment above the issuing code said
it was; issuing now records it. Placing the composed list at genesis is the
other half — the control plane says what it composed and whoever raises the
machine writes it beside the bus, because no declaration can reach a machine
that has not enrolled.
With that a first node enrols. It is then enrolled twice from one attempt,
each minting a credential, and it keeps the answer to the first while the mesh
keeps the second. The trail for that one stops at a duplicate that survived
message-id deduplication.
Raising a first node hits them in order: the bundle's bus image named for a
registry that is gone; the certificate made by openssl in an image that has
none; enrolment dialling TLS at a bus that speaks first, then refusing its own
token for the empty bus. Each was right until the bus changed and nothing has
raised a foundation since.
The fourth is not a patch: the installer carries the bus's first user list and
the controller composes the rest through the bus being a module, and at genesis
there is no module — so the first node cannot be let onto the bus it just
raised.
Found trying to run ADR 0147's bed. The older bundle raises the previous
broker and a control plane that refuses to start without MESH_BUS_NATS; the
newer one stops a step earlier, asking for a certificate from openssl in an
image that has none. No bed can run while this holds, so 0147's check section
now says what actually stands behind it — the rendering, not a machine.
Issue 129: every internal HTTPS name fails verification on every machine,
because nothing has ever written the mesh's root into a trust store. The
report proposed the controller inject it the way the private network writes
the registry's trust; this record rejects that — reachability and trust are
not the same fact, and where anchors live is the host's difference, not the
controller's. A module requiring internal-acme-ca does the whole of it, and
being unassigned undoes it.
0145's core stands — a module checks what the mesh claims, from where the callers
are — and everything it said about what to dial was wrong.
A raw port is not how anything in this mesh is reached. A real caller resolves a
name, the proxy answers, the proxy reaches the service; dialling a port tests the
last hop of a four-hop path and skips the three where most connectivity lives.
So each hosting form gets its own endpoint, route and name:
connect-docker.<node>.internal for a container, connect-process for the mesh's own
code in a unit it writes, connect-unit for a unit a package ships — and the same
under each public domain. Every machine checks every machine, by name, over TLS,
verifying the certificate against the authority that should have issued it. The
names are the instrument: a failure reads as connect-docker.g14.internal did not
answer.
No name is written anywhere: the machines come from the roster fact, the labels are
the module's, and a machine that joins is checked by the others on the next push.
Five things this needs that the mesh does not have, named rather than assumed —
a module every machine has (ScopeNode is exclusivity, not obligation), a container
running a bundle, a unit a package ships, the public domain in the roster fact, and
issue 129, until which every internal name will correctly fail verification.
The mesh asserts three things are callable and has never checked any of them. A
module on every machine serves an endpoint of its own and dials every other
machine's, from the position the callers are in — not the host and not the control
plane, both of which reach those addresses by paths no ordinary caller uses and
would have passed throughout the outage.
Its probe is its own endpoint, declared reachable over the private network, so it
is admitted by exactly the rule that governs every internally-exposed service. The
tempting target is a service every machine has, and those are the ones never closed
— ssh above all — which would have passed for all eleven hours.
Resolution and connection reported separately, because they have different owners.
One failure is not a fault, and the count travels with the result. It reports and
repairs nothing.
Built on a scheduled container, a rendered roster fact and a bundle: no new
vocabulary. Adopted on its own merits, which is what 0143 was not.
Everything should be able to call what runs on the same machine, another machine's
service exposed to the private network, and another machine's service exposed
publicly. Three cases; the filter had two.
The first was expressed as the machines' own addresses on the private network. A
caller on the machine carries such an address; a caller in one of its containers
carries a bridge address and matched nothing — measured, same destination and same
machine: src 10.10.0.1 against src 172.17.0.8. The second case worked by accident,
because the tunnel rewrites a caller's address to the sending machine's. Two of
three working is why it read as correct.
0143 answered the wrong question. It proposed verifying each grant from the
consumer's own network position and went to length about which position, because
whether a caller sat in a container changed the answer — and that difference was the
bug. Observing a configuration error is not its remedy. Superseded, and nothing
replaces it; whether the mesh should check a grant is still open in issue 145 and
must stand on its own.
And a module is not a container: 61 of 72 happen to use one, 11 do not, and a rule
reasoning about containers describes most of the mesh rather than the mesh.
A grant is four facts and a credential — the provision, the machine, the port, and
who the consumer is when it connects — and it is the whole mechanism by which
anything in the mesh reaches anything else. The mesh asserted it and never found
out whether it was true. Issue 145 is what that cost: eleven hours of 'every module
current with its source' while a module could not reach its database.
The consumer verifies it, from its own network position. Not the control plane,
which reaches the address by a path no consumer uses. Not the machine, whose own
packets carried a source address the filter admitted while every container's was
refused — so a host-side check would have passed throughout the outage it exists to
catch. That is inference from the rule that was loaded, stated as such.
A connection and nothing more; speaking each provision's protocol would be a second
implementation of every provision. One failure is not news, only consecutive ones,
and the count is reported rather than the last attempt. A consumer that is not
running reads unchecked, which is a different sentence from broken.
What it costs to be wrong is the constraint on all of it: a check that calls a
working provision broken trains a reader to ignore the report.
Converging the control node closed every path by which a module reached another by
the machine's own name, and it ran for eleven hours while the mesh answered 'all
doing what they were told, all heard from, running what the mesh would send them,
and every module current with its source'.
6,154 database failures in one module's log, beginning at the minute of the flip.
The service accepted TCP and never answered HTTP. Every check the mesh makes
passed, because every check the mesh makes is about the relationship between the
mesh and a machine — applied, current, containers running. None asks whether a
module can reach what it requires, though the mesh composes every grant and so
knows exactly who requires what.
The filter fault is fixed. The eleven hours are the measurement, not the bug.
Also records on 144 that 'more closed than the mesh believes' was harmless only
for what is reached from outside, and an outage for what is reached from within.
The record says internal and public each mean something to the filter. For an
endpoint the proxy serves, the second half is wrong, and ADR 0045 said so first:
a public service is exposed through the proxy, listening from the mesh, not by
opening its own port.
Found by trying to express one real module, not by review — routed name public
because browsers post to it, machine port private because it serves a dashboard in
cleartext. Under one value for both, saying public would have reopened a port an
operator had just closed. Measured the same evening: the routed name answered from
the internet over TLS while the port was refused from the same place.
Corrects a fact. One statement per endpoint with three things derived from it
stands; the filter column applies to an unrouted endpoint.
The first account said the declaration carries no resource that would disable the
found firewall and that nothing implemented the sentence the flip prints. Wrong:
the mechanism is a step in the host's own apply, retireFirewall, and it is careful
— it reads back that the mesh's table is loaded before retiring anything, records
the forward policies first, and verifies ufw reports inactive afterwards.
What is established: ufw was active and enabled two minutes after a flip that
reported the node converged with nothing failed; and the machine's record now
reads disabled_by_mesh: true, written by a reconcile fifty minutes later that
found ufw already inactive because an operator had disabled it by hand.
So the step did not take effect and the record says it did. The candidates are
named rather than chosen — the step is skipped silently when the apply has any
failure, the mesh's table is loaded by a service in the same apply so ordering is
open, and the host's detail lines do not reach the journal, which is why this has
candidates instead of a cause.
Both found by converging the control node — the first machine with a firewall to
flip, since the two before it had none.
143: the preview and the flip both say the found firewall is disabled. The node
reported converged, 372 resources applied, nothing failed, and ufw is still
enabled and active. A converged node's declaration carries no resource that would
disable it; the sentence is printed by the command and nothing implements it.
144: ufw was never what filtered the traffic that mattered there. Fifty forwarded
openings converged through it had matched zero packets, while a chain the
predecessor installed in the container runtime's pre-accept hook did the work —
in memory only, recreated by nothing. The mesh's filter now covers that path, so
the machine no longer depends on it, but the chain remains and is the only thing
refusing the bus and the registry, which the design requires reachable from
anywhere so a machine can enrol before it has a private address.
Measured: the host is a binary somebody copied to four machines, owned by no
package and built by nothing, while the controller, catalogue, builder and vault
are container images publishing no ports at all. Same language, same project,
same kind of work, delivered two ways — and the difference is not a judgement
about either, it is that images are the only delivery that works.
What it costs: genesis must raise a container runtime before the control plane
can exist; updating the control plane goes through a registry the control plane
runs; a host change cannot be rolled out at all; and compiling the language the
mesh is written in is not a capability of the builder, so the controller is built
from a hand-written Dockerfile — the incantation the bundle toolchain exists to
abolish.
Third-party software stays a container: the store, the registry, the broker are
somebody else's build. The container runtime stays on the machine for modules.
What changes is that the control plane no longer needs it to exist.
The receiving half is already built and tested (ADR 0141). Staged: compile Go, an
artifact names its target, deliver a binary, the host first, then the rest, genesis
last.
The record claimed a version reaches a machine as an ordinary archive with
nothing new needed. Two things it needs do not exist: no toolchain can compile
the host (the list is typescript and python, and the control plane, also Go, is
built as an image from a Dockerfile instead), and nothing interpolates a built
version into a resource path, so nothing can ask for .../versions/<version>/.
The decision, the options weighed and every consequence stand — the host half is
merged and tested. What was understated was the cost, so it is corrected in place
and dated rather than superseded.
The supervision was already right — a clean exit means the host stood aside and
the launcher runs what is on disk, failures are counted, and a rollback happens
at the limit. Two things made it dead code: nothing told the running host a
successor was waiting, and the rollback resolved its known-good version through
pacman, which no machine here uses and which two of three operating systems do
not have.
Keeping a version rather than a path was the clue. Versions live side by side in
directories named for them; the newest runs; the running one stands aside between
reconciles; a reconcile that completes records itself and retires what is older
than its predecessor; rollback starts that predecessor. No new resource kind and
nothing new on the bus — an archive already fetches by digest, and the path
written is never the path executing.
Answers issue 142.
A host change merged yesterday reached no machine without a person copying a
file. The host is not a build target, no declaration delivers it, and the half
that recovers from a bad host — noticing the executable changed, a known-good
record, a launcher that rolls back — is written, tested and called by nothing.
All four machines run a byte-identical hand-copied binary that no package owns
and no record names, so nothing can say a machine is behind.
Found because ADR 0140 needs the machine to report a new fact, and merging that
could not roll it out.
Reading a converged machine's rendered rules showed the cause: the chain blocks
everything passing through and then allows the machine's own containers back by
listing their address ranges. 0137 made that list typeable and 0139 tried to
generate it; both refined a list that should not exist, because the mesh has no
position on a container reaching outward. Constrain what arrives from outside,
allow what did not, and let the machine report which links face outside — one
fact instead of a list. Ports keep following the modules unchanged.
The records check now allows one record to supersede several, and stops
requiring a withdrawn record's own citations to be live.
Both follow from the same rule the mesh is built on — a node's configuration is
composed from the modules assigned to it. Reach was settled separately by the
filter, the proxy's names and the certificate authority, so "this must not be
public" could not be written; it becomes one value on the assignment that all
three read. And the forward chain consulted two constants plus a typed list
although modules already declare their networks; it now forwards what they
declared, with the host rendering the addresses it allocated.